What the simulator says: SRAM size against key traffic, the NTT-bound and memory-bound regimes, power and energy per bootstrap under a TDP, the algorithmic acceleration techniques, a design-space sweep, silicon area as a third axis (a three-way power, performance and area front), the model's limitations, a reading list (F1, CraterLake, BTS, ARK, SHARP, GPU work, HEIR, OpenFHE, Lattigo) and practice questions.
SRAM sizingRegimesEnergy / bootstrapMin-KSAreaParetoReading list
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Cryptography decks or LLM Inference Simulators.
How much on-chip SRAM stops key traffic dominating?
When is a design NTT-bound, memory-bound or power-bound?
Which algorithmic techniques matter, and on which designs?
What does a bootstrap cost in energy under a power limit?
And in silicon area: which designs are worth their mm²?
Ground rules
Every number in this deck comes from examples/results.py in FHE_Accelerator_Sim, which regenerates them all; re-run it after any model change. The workload is one full-slot ark-set bootstrap (N = 216, L = 23, dnum = 4) unless stated. Hardware coefficients are illustrative: read ratios, trends and crossovers, not absolute milliseconds.
02
SRAM Against Key Traffic and Area
SRAM
Baseline: bootstrap
key GB
all HBM GB
Min-KS + seeded + OTF: bootstrap
key GB
all HBM GB
128 MiB
26.98 ms
9.29
26.53
21.22 ms
4.35
18.22
256 MiB
17.10 ms
6.99
16.37
8.71 ms
0.70
4.33
512 MiB
13.94 ms
6.74
12.44
7.19 ms
0.58
0.78
1 GiB
11.69 ms
6.74
10.26
7.19 ms
0.58
0.78
4 GiB
11.41 ms
6.74
10.02
7.19 ms
0.58
0.78
With the baseline algorithm, no SRAM size stops key traffic. Each of the 85 rotation keys is used once per bootstrap, so 6.74 GB is compulsory from 512 MiB upwards. More SRAM only removes ciphertext spills.
Real programs agree. LoLa as compiled by HEIR, with a forced bootstrap at the N = 217, 47-prime parameters HEIR chooses, takes 203 ms with 512 MiB of scratchpad and 101 ms with 2 GiB (deck 03).
Across bootstraps keys could be reused, but only if the whole key set fits: two back-to-back bootstraps first share keys at 16 GiB (3.37 GB of keys per bootstrap, 7.99 ms).
With key reuse built into the algorithm, 512 MiB is enough: total HBM traffic falls to 0.78 GB per bootstrap. The answer to "how much SRAM?" depends on the algorithm. That is why ARK co-designed the two, and why the simulator models both.
Monotonicity (more SRAM never raises traffic) is a tested property, so search.min_sram can bisect for the smallest adequate size.
The third axis: area. SRAM is not free. The simulator's area model (7 nm; units from ARK's published breakdown, SRAM from a CACTI sweep; illustrative) prices the same sweep in mm², with all three key-traffic techniques. "Pareto" marks designs no other design beats on latency, energy and area:
Past 512 MiB, SRAM costs area and buys nothing: the latency is flat at 7.19 ms while perf/mm² falls from 0.354 (256 MiB) to 0.200 (1 GiB). Beyond about 2 GiB no reticle could print the die at all.
Compute or SRAM? From a 256 MiB design running the baseline algorithm (memory-bound), +128 MiB of SRAM saves 3.74 ms per 100 mm² added, and 2× NTT + MAC only 0.21. With all three techniques it is MAC-bound and the two are even (1.71 against 1.68). The general treatment is SimEng 13: Power, Performance and Area.
03
NTT-Bound Against Memory-Bound
Bootstrap latency and verdict as NTT throughput and HBM bandwidth vary (MAC lanes = 2 × NTT butterflies; 512 MiB; 250 W TDP under the dynamic power manager):
Baseline algorithm NTT bfly/cycle \ HBM
500 GB/s
1 TB/s
2 TB/s
4 TB/s
512
34.2 ms memory
21.9 ms NTT
17.9 ms NTT
16.7 ms NTT
1,024
28.8 ms memory
17.2 ms memory
11.5 ms NTT
9.8 ms NTT
2,048
27.0 ms memory
14.8 ms memory
8.8 ms memory
6.7 ms MAC
4,096
26.4 ms memory
13.9 ms memory
7.9 ms memory
5.6 ms MAC
8,192
26.2 ms memory
13.8 ms memory
7.7 ms memory
5.5 ms MAC
Min-KS + seeded + OTF NTT bfly/cycle \ HBM
500 GB/s
1 TB/s
2 TB/s
4 TB/s
512
28.7 ms NTT
28.6 ms NTT
28.6 ms NTT
28.6 ms NTT
2,048
9.6 ms MAC
9.5 ms MAC
9.4 ms MAC
9.4 ms MAC
8,192
6.7 ms MAC
6.3 ms MAC
6.3 ms MAC
6.2 ms MAC
The ridge: with the baseline algorithm, 1 TB/s supports only about 512–1,024 butterflies per cycle; beyond that, NTT hardware waits on keys. Reducing key traffic moves the ridge, after which the MAC lanes bind.
Power binds only when the real draw does. With one worst-case clock (every unit and HBM flat out must fit the TDP), the 4 TB/s column was slower than 2 TB/s, because HBM alone reserved 120 W, and the fast designs were power-bound. The dynamic power manager (slide 05) removes that artefact: no cell here reaches the 250 W limit for long enough to be power-bound.
04
Acceleration Techniques
Algorithm
ARK-class: bootstrap
key GB
verdict
Small digital: bootstrap
verdict
Baseline (hoisted BSGS)
13.94 ms
6.74
memory
17.46 ms
NTT
No hoisting
14.02 ms
6.74
memory
20.14 ms
NTT
OpenFHE's BSGS (lazy ModDown)
19.77 ms
8.60
memory
18.99 ms
NTT
+ Min-KS
8.51 ms
1.15
memory
21.54 ms
NTT
+ Min-KS + seeded keys
8.86 ms
0.58
MAC
22.07 ms
NTT
+ Min-KS + seeded keys + OTF plaintexts
7.19 ms
0.58
MAC
26.65 ms
NTT
SlotToCoeff first
11.51 ms
5.73
memory
12.74 ms
NTT
+ all three + SlotToCoeff first
5.42 ms
0.50
MAC
20.01 ms
NTT
On the memory-bound design key reuse (Min-KS) is worth 1.6× on its own and the full set 1.9×. Seeded keys on their own lose a little once the design is no longer memory-bound: their PRNG work lands on the MAC lanes, which are now the bottleneck.
SlotToCoeff first is the exception: it helps every design. It runs the second DFT on the nearly exhausted input (cheap keys), needs one EvalMod instead of two for real-valued data, and leaves 10 levels instead of 7, so the cost per useful level falls from 1.99 to 1.15 ms on the ARK-class design and from 2.49 to 1.27 ms on the NTT-starved one. With the key-traffic techniques it reaches 0.54 ms per useful level.
A CPU library's choice can be wrong for an accelerator. OpenFHE's BSGS split, recorded from a real OpenFHE bootstrap (deck 02), saves a third of the NTTs, which is right for a CPU. It needs 109 rotations instead of 85 and 28% more key traffic, so on the memory-bound ARK-class design it is 42% slower. It doesn't even win on the NTT-starved design, because its extra traffic outweighs the saved transforms.
On the NTT-starved design every technique hurts: each trades arithmetic for bytes, and there are no spare cycles to trade. Hoisting, which saves NTTs, is the only one that helps (no hoisting is 15% slower).
Lesson: techniques are not universally good. The simulator's job is to say which ones pay on this design.
05
Power and Energy per Bootstrap
The simulator enforces the TDP with a dynamic power manager. Each kernel gets the highest clock, and each HBM chunk the highest bandwidth, that fits the headroom left by everything running at that moment; if even the lowest setting does not fit, it waits. The alternative, one worst-case clock at which every unit and HBM at full rate fit the TDP, is kept for comparison (250 W TDP throughout):
Configuration
Power mode
Bootstrap
Mean clock
Avg / peak W
mJ
Verdict
ARK-class, baseline
either
13.94 ms
100%
82 / 176
1,139
memory
… with DVFS
either
15.83 ms
50%
72 / 103
1,134
memory
ARK-class, all techniques
either
7.19 ms
100%
84 / 150
607
MAC
2× NTT and MAC, all techniques
worst-case
6.95 ms
91%
82 / 179
569
power (MAC)
2× NTT and MAC, all techniques
dynamic
6.34 ms
100%
90 / 218
573
MAC
4× NTT and MAC, all techniques
worst-case
8.28 ms
74%
70 / 135
577
power (MAC)
4× NTT and MAC, all techniques
dynamic
6.20 ms
100%
92 / 239
567
MAC
4× design, all techniques
Worst-case clock
Dynamic manager
TDP 250 W
8.28 ms at 74%, peak 135 W
6.20 ms at 100%, peak 239 W
TDP 150 W
11.62 ms at 53%, peak 95 W
6.41 ms, mean clock 99%, peak 150 W
TDP 100 W
cannot run (worst case at the lowest clock exceeds the TDP)
6.68 ms, mean clock 92%, power-bound
TDP 80 W
cannot run
7.73 ms, mean clock 89%, power-bound
Energy follows bytes and time. Cutting HBM traffic halves energy per bootstrap (1,139 → 607 mJ). Static power is about half the total, so finishing sooner is itself an energy optimisation.
Worst-case clocking wastes the budget: it reserves power for a coincidence (every unit and HBM flat out) that this workload never produces. The manager runs the 4× design at full clock under 250 W and 1.8× faster than worst-case clocking at 150 W, and it can run at TDPs where worst-case clocking cannot run at all.
DVFS barely helps a memory-bound design: halving the clock saves dynamic energy but stretches the run 14% and pays static power for longer, so the net saving is 0.4%.
The guarantee is kept: peak power never exceeds the TDP under either mode (tested, including at 100 W), and the JavaScript port reproduces the manager's decisions bit for bit.
06
Interactive: Design-Space Map
Runs a 5 × 4 grid of full bootstraps in the browser (the bit-exact JavaScript port) and colours each cell by its verdict. Change the algorithm, SRAM, TDP or power control and watch the regions move. Try 150 W with the worst-case clock, then with the dynamic manager.
250
07
Sweeps, Pareto Fronts and Simulator Speed
A 36-point sweep (NTT 1,024–8,192 butterflies/cycle × SRAM 256 MiB–1 GiB × HBM 0.5–2 TB/s, all techniques) ran in 0.4 s on 8 processes. Its latency/energy Pareto front is a single point:
NTT bfly/cycle
SRAM
HBM
Bootstrap
Energy
Verdict
8,192
512 MiB
2 TB/s
6.34 ms
573 mJ
MAC-bound
Once key traffic is under control, the MAC lanes bind (base conversion and multiply-adds), so the best design is the widest one with enough bandwidth, and the fastest point is also the most energy-efficient. Under worst-case clocking the same sweep's best points were power-bound; that verdict was an artefact of reserving power the workload never draws.
Simulator speed: one ark-set bootstrap (203 HE ops, 1,153 kernels) takes about 70 ms in Python/SimPy and about 12 ms in the JavaScript port (70–74 ms and 11–13 ms across recorded runs). A one-event-per-kernel abstraction is what makes interactive design maps possible (sister series, deck 08).
Cheaper still:search.analytic_bound bounds latency from the trace with no simulation (10.01 ms for the default design, against 13.94 ms simulated). Use it to prune a search before simulating, never to replace simulation; the 28% gap is the overlap structure it cannot see.
08
Against Published Numbers
Reference
What it reports
Use here
OpenFHE on this machine (i7-3770, 8 threads)
HMult 352.6 ms, HRotate 324.8 ms (N = 216, 24 limbs, dnum 4); sparse bootstrap 12.1 s
Calibration: one parameter fitted to HMult; HRotate predicted within 10%; the bootstrap is predicted 30% low by the scheme model and 11% low by replaying OpenFHE's own recorded kernel stream
ASIC papers mostly report amortised metrics (TA.S., application time) at their own parameters, so their numbers are references for scale, not targets to fit. A like-for-like ASIC comparison would need their exact algorithms, which is what the trace interface is for.
09
Limitations and Extensions
Each limitation is a self-contained exercise, roughly in order of difficulty:
Complex-valued data with SlotToCoeff first: the model assumes real-valued data (one EvalMod); add the second EvalMod for complex slots, and measure other level budgets for the two DFTs.
Smarter power policies. The power manager is greedy and first-come: the first kernel to ask gets the highest clock that fits. Try fair-share or critical-path-aware allocation, and DVFS transition latency.
Banked scratchpad and on-chip network contention, replacing "every unit gets full port bandwidth".
Interleaving bootstraps (or many ciphertexts) to reuse keys across operations, ARK's "inter-operation key reuse".
More from the compiler. The HEIR front end (deck 03) reads HEIR's ckks IR. Next steps: support several bootstraps per program with per-segment stage spans; let the simulator feed costs back into HEIR's bootstrap placement (its ILP placer accepts a cost model); and cross-check the Lattigo backend as well as OpenFHE.
Optical extensions: chained optical stages, redundancy-based error correction, optical base conversion; each with its own precision analysis in precision.py.
Correlation against RTL or an FPGA prototype of one kernel (say an NTT unit), to replace an illustrative coefficient with a measured one.
10
Reading List
F1 (Feldmann et al., MICRO 2021)The first programmable FHE accelerator: wide-vector units for NTT, modular arithmetic and permutations, and a compiler-managed memory hierarchy, because "data movement becomes the bottleneck".
CraterLake (Samardzic et al., ISCA 2022)F1's successor for unbounded (bootstrapped) computation; 28-bit words, 256 MB on chip.
BTS (Kim et al., ISCA 2022)Bootstrapping as a first-class citizen; analyses the off-chip bandwidth bootstrapping needs and the parameter sets that suit accelerators.
ARK (Kim et al., MICRO 2022)Runtime data generation and inter-operation key reuse (Min-KS, OF-Limb): the algorithm-architecture co-design modelled in this series.
SHARP (Kim et al., ISCA 2023)A short-word (36-bit) hierarchical accelerator: word length as a design parameter.
Over 100x Faster Bootstrapping (Jung et al., TCHES 2021) and TensorFHE (Fan et al., HPCA 2023)GPU bootstrapping: memory-centric kernel fusion, dnum choice, tensor cores for NTTs.
FAB and SoK: FHE AcceleratorsAn FPGA bootstrapping accelerator, and a survey that benchmarks the field side by side.
Why is CKKS bootstrapping memory-bound on a large accelerator, and what would you change first?
The homomorphic DFT stages use dozens of distinct rotation keys, each 100+ MiB at high levels and each used once per bootstrap, plus a plaintext per diagonal: gigabytes of compulsory traffic against modest arithmetic. More SRAM does not help (each key is used once). Change the algorithm first: key reuse (Min-KS), seeded keys and on-the-fly plaintexts cut key traffic by an order of magnitude; then size SRAM to the new working set and re-check the bound, which moves to compute or power. (Decks 02, 05.)
How would you validate an FHE accelerator simulator before RTL exists?
A ladder: sizes and operation counts against hand formulas and published tables; invariants (dependencies, conservation of work, determinism); analytic checks of the engine (serial chains, flow-shop makespans, roofline bounds); behavioural expectations (more SRAM never raises traffic); property-based tests over random configurations; calibration of throughput against a real library such as OpenFHE with out-of-sample predictions; and differential tests between independent implementations. Then plan correlation against RTL kernels as they arrive. (Deck 03.)
An optical engine computes FFTs very fast. What do you need to know before believing FHE gets faster?
Its effective precision (ENOB) and how many digit planes exact rounding then needs at your limb width.
Conversions per transform point against butterflies offloaded per point.
Converter energy per sample (a Walden figure of merit) and its throughput under the power budget.
Static laser and tuning power.
Whether the target design is NTT-bound at all: if it is waiting on keys, faster transforms change nothing. (Deck 04.)
How would you get real FHE programs into the simulator?
Define a kernel-level trace format (operations, objects, sizes, keys, dependencies) as the contract. Produce it from a compiler such as HEIR at a polynomial or modular-arithmetic level of its IR, or by instrumenting a library's key-switching and NTT routines. This repo does the library route for OpenFHE: an 82-line patch logs the stream, and replaying it predicts the measured bootstrap time to within 11%. Use the same program compiled to a CPU backend as a functional and timing reference, and track coverage (which operations and parameter sets the front end supports) as an engineering metric. (Decks 03, 05.)
12
What to Take Away
SRAM alone does not fix key traffic for a baseline bootstrap; algorithms that reuse keys do, after which 512 MiB is enough here.
The ridge: at 1 TB/s a baseline bootstrap saturates HBM beyond about 1,000 butterflies per cycle; with key reuse, the MAC lanes bind; power binds only under a much lower TDP.
Techniques are design-dependent: a 1.9× win on a memory-bound design, a loss on an NTT-starved one. SlotToCoeff first is the exception and helps both, cutting the cost per useful level by 42–49%.
Energy follows bytes and time: 1,139 → 607 mJ per bootstrap; static power is half of it; a dynamic power manager keeps over-provisioned compute useful under a TDP that worst-case clocking would waste.
Optics needs both precision and a design that is NTT-bound (deck 04).
Every number here is reproducible from one script, and every claim about the model is a test.