FHE Accelerator Simulators — Presentation 03

Simulating an FHE Accelerator

The modelled machine (NTT units, modular multiply-add lanes, an automorphism network, a scratchpad, HBM and key streaming), the operation-trace format and where traces come from (a scheme model, HEIR, OpenFHE), the SimPy engine, its metrics and its validation ladder, and the whole simulator running live in the browser.

SimPy Trace format Scratchpad HBM Hot-spots Live simulator
Trace → Issue → Scratchpad → HBM → Units → Metrics
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Cryptography decks or LLM Inference Simulators.

01

What the Simulator Must Answer

LatencyHow long does one bootstrap take, and how is it split across ModRaise, CoeffToSlot, EvalMod and SlotToCoeff?ms, per stage
BoundIs the design NTT-bound, memory-bound (keys) or power-bound, and which resource is the hot-spot in each stage?utilisation
MemoryHow much on-chip SRAM stops evaluation-key traffic dominating, and under which algorithm?GB by class
EnergyJoules per bootstrap and peak power under a TDP; does an optical NTT engine change either?mJ, W

These are architecture-exploration questions, so the model sits at the transaction / discrete-event rung of the fidelity ladder (LLM Inference Simulators deck 01): one event per kernel and per memory chunk, not per cycle. A full bootstrap simulates in tens of milliseconds, so a design sweep of hundreds of points takes seconds.

02

The Modelled Architecture

NTT units4,096 butterflies/cycle MAC lanes8,192 mul-adds/cycle Automorphism4,096 words/cycle optical engine (opt.) Scratchpad 512 MiB LRU: ciphertexts, keys,plaintexts; pinned while used 90 MiB reserved for onekey switch's temporaries 20 TB/s per unit port HBM, 1 TB/s evaluation keysDFT plaintextsspilled ciphertexts one shared resource,4 MiB chunks, FCFS spillload

The defaults are an ARK-class digital design at 1 GHz, sized like the published ASICs (ARK: 512 MB scratchpad, 1 TB/s HBM). The units are aggregates: four NTT pipelines become one resource with their combined throughput. All coefficients are illustrative.

03

The Trace: Contract Between Scheme and Hardware

The scheme model emits HE operations in program order. Each names its inputs and output, the evaluation key and plaintexts it needs, and its primitive kernels as [kind, amount, on-chip words]. Two real lines from fhe-sim --params small --dump-trace t.json (two hoisted baby-step rotations):

fhe-sim-trace/1 (excerpt)
{"id": 2, "op": "hrot", "stage": "cts", "level": 11, "boot": 0, "inputs": ["b0.c1", "b0.c2"],
 "output": "b0.c3", "key": ["cts0.b1", 3145728], "pts": [],
 "kernels": [["auto", 245760, 491520], ["mac", 393216, 720896], ["intt", 8, 65536],
             ["bconv", 524288, 262144], ["ntt", 24, 196608]]}
{"id": 3, "op": "hrot", "stage": "cts", "level": 11, "boot": 0, "inputs": ["b0.c1", "b0.c2"],
 "output": "b0.c4", "key": ["cts0.b2", 3145728], ...}
04

Where Traces Come From

SourceHowStatus here
Scheme model (workload.py)Builds bootstrapping from the algorithm: key-switch recipes, BSGS DFT levels, Chebyshev EvalMod, level trackingImplemented; operation counts are tested against closed forms
HEIR (heir.dev; Ali et al., arXiv:2508.11095)An MLIR-based FHE compiler. After --torch-linalg-to-ckks with unrolled kernel loops, the server function is straight-line code in HEIR's ckks dialect, with every value's level in its type. HEIR has already chosen the packing, rotations, rescales and bootstrap placementDone (heir_frontend.py, release v2026.10.01): one HE operation per ckks op, dependencies from SSA, levels checked against HEIR's types. Three of HEIR's example programs are in the repo (slide 12). HEIR's dialects change quickly, so the front end names the release it was validated on
OpenFHE (ePrint 2022/915)Instrument the C++ library's NTT, base-conversion, key-switching, automorphism and rescale routines; log each call with its shape and the identity of its key or plaintext; replay the streamDone for two bootstraps of OpenFHE v1.5.1 (an 82-line patch, in calibration/openfhe_trace). The model is checked against them in deck 02; the replay is the best timing predictor (slide 11)
Lattigo (Go)The same instrumentation approachNot done
Why a compiler front end matters

A hand-built scheme model reproduces the algorithm someone chose. A trace from the compiler that will actually target the hardware captures its real level management, packing and fusion decisions. The OpenFHE traces show the catch: a CPU library's trace carries CPU-tuned choices. OpenFHE's BSGS split saves a third of the NTTs but needs more rotation keys, which makes a memory-bound accelerator 42% slower (deck 05). That is why the compiler targeting the hardware, not a CPU library, should produce the trace; deck 09 of the sister series discusses the same point for PyTorch and ONNX.

05

The SimPy Engine

Every HE operation is a SimPy process; every functional unit and HBM is a simpy.Resource. The core of sim.py, lightly trimmed:

sim.py: Simulation.run_op (bookkeeping lines elided)
def run_op(self, o, pl: OpPlan):
    env = self.env
    pf = env.process(self.prefetch(o, pl)) if (pl.key_load or pl.pt_load) else None  # keys don't wait for data
    for name, size in pl.writebacks:                      # spills decided at issue
        ev = env.process(self.writeback(name, size))
        self.wb_ev[name] = ev
    for x in o.inputs:                                    # dependencies
        p = self.producer.get(x)
        if p is not None:
            yield from self.wait_done(p)
    for x, size in pl.ct_load:                            # reload spilled ciphertexts
        w = self.wb_ev.get(x)
        if w is not None and not w.processed:
            yield w
        yield from self.xfer(size, "ct_read", o.stage)
    if pf is not None and not pf.processed:
        yield pf
    for k in o.kernels:
        for seg in self.cost.segments(k):                   # cost model: unit, time, joules
            req = self.units[seg.unit].request()
            yield req
            ...                                              # power and statistics bookkeeping
            yield env.timeout(seg.time)
            self.units[seg.unit].release(req)
06

Scratchpad and Traffic

Tested property

For every algorithm tried, more SRAM never increases key traffic or total HBM traffic (sweeps from 128 MiB to 2 GiB). That justifies using bisection to find the smallest adequate scratchpad (deck 05).

07

Cost and Power Model

Time per kernel

  • NTT: limbs × (N/2) log2N butterflies ÷ butterfly rate.
  • MAC and base conversion: multiply-adds ÷ lane rate; automorphism: words ÷ word rate.
  • Each is floored by on-chip bandwidth: words × 8 B ÷ port bandwidth.
  • All rates scale with the clock fraction s.

Energy and power

  • Static 40 W, plus 10 pJ per butterfly, 5 pJ per multiply-add, 1 pJ per permuted word, 1 pJ per SRAM byte and 30 pJ per HBM byte (illustrative); logic energy scales with s2.
  • TDP enforced by a dynamic power manager (default): when a kernel starts, it gets the highest clock (down to smin) whose power fits the headroom left by everything running at that moment; an HBM chunk gets the highest bandwidth that fits (down to 25%). If nothing fits, the request waits for power to be released.
  • Optional DVFS caps the clock of a memory-bound run at the level that just keeps the busiest compute unit as busy as HBM.

The manager guarantees the peak never exceeds the TDP, and a test checks it on every configuration, including Hypothesis-generated ones. It is greedy and first-come, with the clock chosen by bisection (no transcendentals, so the JavaScript port matches it bit for bit). The older alternative, one worst-case clock at which every unit and HBM flat out fit the TDP, is kept as power_mode="worst-case". It reserves power for a coincidence the workload never produces, so it can be 1.8× slower, or unable to run at all at a low TDP (deck 05). Methodology for all of this is in LLM Inference Simulators deck 07.

08

Metrics and Hot-Spots

MetricDefinition
Latency, per stageSimulated time; span of each stage from its first operation's start to its last operation's end
UtilisationBusy time ÷ total time for NTT, MAC, automorphism, optical and HBM
Bound attributionThe most-utilised resource: HBM → memory-bound; NTT (or optical) → NTT-bound; MAC → MAC-bound; plus "power-bound" when a compute-bound design lost more than 10% of the run to the power manager (or, with a worst-case clock, ran below 100%)
Hot-spot per stageFor each stage, the resource with the most busy time attributed to that stage's operations
Traffic by classHBM bytes for keys, plaintexts, ciphertext reads and writes; key share of the total
Energy, powerStatic × time + dynamic per unit + HBM; average power; peak instantaneous power from active units
Lower boundsBusiest resource's busy time (a schedule bound) and an analytic roofline bound from the trace alone
Perfetto trace--trace boot.json: one row per unit and for HBM, one slice per kernel and per chunk

Default run (ark set, ARK-class design): 13.94 ms, memory-bound (HBM 89% busy, NTT 13%, MAC 29%), CoeffToSlot 72% of the time with HBM as its hot-spot, EvalMod's hot-spot the MAC lanes. Keys are 54% of 12.44 GB of HBM traffic; 1,139 mJ per bootstrap at 82 W average and 176 W peak against a 250 W TDP.

09

Interactive: Run the Simulator

The JavaScript port of the SimPy model, with a minimal SimPy core that reproduces SimPy's event ordering. It matches the Python package bit for bit, including clocks, bytes, energy and per-operation end times; the repo's tests check this. Choose a configuration and press Run, SRAM sweep or Compare algorithms.

4096
8192
512
1000
250
4
1
where the bootstrap's time goes (stage spans)
timeline: busy intervals per resource, coloured by stage
sweep / comparison
10

Experiments to Try

  1. Find the bound. Run the defaults: memory-bound, with CoeffToSlot's hot-spot on HBM. The timeline shows HBM saturated while NTT and MAC idle in the orange stretches, then compute busy in EvalMod.
  2. Buy SRAM. Press SRAM sweep. Key traffic does not move from 512 MiB upwards (each key is used once); only ciphertext spills disappear.
  3. Change the algorithm instead. Tick Min-KS and seeded keys, or press Compare algorithms: key traffic falls about tenfold and the verdict moves to the MAC lanes.
  4. Starve the NTTs. Choose small digital and compare algorithms again: on an NTT-bound design the same techniques make bootstrapping slower.
  5. Hit the power wall. Raise NTT to 16,384 and MAC to 32,768, tick the three key techniques and lower the TDP to 150 W. The dynamic manager still runs at a mean clock near 100%; switch Power control to worst-case clock and the same design slows to about half speed. At 100 W the worst-case clock cannot run at all, and the manager becomes power-bound.
  6. Back-to-back bootstraps. Set 2 bootstraps with a large SRAM: keys are reused across bootstraps only when the whole key set fits (about 16 GiB, beyond the slider).
11

Validation

89 tests, organised as a verification ladder (the method is in LLM Inference Simulators deck 06):

RungChecks
Unit: sizesCiphertext, plaintext and key sizes against hand formulas and ARK's Table III (24/12/120 MiB; 25/12.5/150 MiB)
Unit: operation countsKey-switch transforms and base-conversion counts; rotations, PMults, keys, HMults and levels per bootstrap stage against closed forms, for four parameter sets
InvariantsDependencies respected; work conserved on every unit; traffic at least compulsory; reads only after writes; determinism; probes passive
AnalyticA dependent chain sums exactly; independent load-then-compute operations match the two-machine flow-shop makespan a + b + (n−1)·max(a, b) to 1e-12; simulated time ≥ the roofline bound
BehaviourMore SRAM never raises traffic; the baseline is memory-bound; key-reuse techniques shift the bound; they hurt NTT-bound designs; optics helps NTT-bound designs and not memory-bound ones
PowerEnergy accounts add up; peak ≤ TDP under both power modes, down to a 100 W TDP; the dynamic manager beats worst-case clocking where that wastes the budget and throttles when the real draw binds; DVFS saves energy when memory-bound
Property-basedHypothesis: 60 random parameter sets, algorithms, hardware and power modes; all invariants, the TDP and determinism
Real libraryTwo recorded OpenFHE bootstraps: identical key-switch traffic where the algorithms agree, identical bootstrap depth (16 and 20), DFT transforms within 20% under OpenFHE's BSGS, and EvalMod's rescale-on-use difference pinned down (deck 02)
CompilerThree HEIR-compiled programs: parameters, one HE op per ckks op, every level matching HEIR's types, bootstraps and polynomial activations expanded, and the same LoLa matching OpenFHE's execution on rotations, keys and relinearisations
ExternalOptical precision rule against a functional model (deck 04); timing calibration against OpenFHE; JavaScript against Python, bit-exact on 26 configurations, in both power modes
Calibration (openfhe-python, i7-3770, 8 threads)OpenFHESimulatedError
HMult, N = 216, 24 limbs, dnum 4 (fitted)352.6 ms352.6 ms0%
HRotate, same parameters (predicted)324.8 ms292.5 ms−10%
Bootstrap, N = 216, 8 slots, dnum 3: scheme model (predicted)12,113 ms8,473 ms−30%
Same bootstrap: OpenFHE's own recorded kernel stream, replayed (predicted)12,113 ms10,759 ms−11%

One parameter was fitted (0.18 modular operations per cycle per unit). The scheme model under-predicts the bootstrap mainly because OpenFHE rescales copies of each input before every multiplication, doubling EvalMod's transforms. Replaying the recorded stream removes that gap; most of the remaining −11% is element-wise work the tracer does not log. A full-slot N = 216 OpenFHE bootstrap could not be measured or recorded: its keys and precomputed plaintexts exceeded the memory available on the 15 GB test machine.

12

A HEIR Front End: Real Programs

Three of HEIR's own example programs, compiled with the HEIR v2026.10.01 release binaries and simulated on the ARK-class design (fhe-sim --heir). HEIR picked every parameter, rotation, rescale and bootstrap; the front end expands each op into kernels. LoLa is an MNIST CNN with square activations (55 rotations); the MLP has a polynomial ReLU; the third row is LoLa compiled with a level budget of 2, so HEIR must place a bootstrap. All four runs are memory-bound.

ProgramN, limbsHE opsScratchpadLatencyKeys / plaintexts / ciphertexts
LoLa (MNIST CNN)215, 11409512 MiB0.61 ms0.35 / 0.17 / 0.00 GB
MNIST MLP215, 131,655512 MiB2.25 ms0.73 / 0.95 / 0.27 GB
LoLa + HEIR bootstrap217, 47625512 MiB202.75 ms61.43 / 24.58 / 112.46 GB
… same217, 476252 GiB101.45 ms47.92 / 24.58 / 21.20 GB
13

Reading the Code

FHE_Accelerator_Sim follows the structure of the sister repo, Disaggregated_Inference_Sim: scheme model, cost model, engine and probes in separate files.

ModuleRole
params.pyCKKS parameter sets and sizes
workload.pyScheme model: bootstrap and single-op traces, JSON dump and replay
hardware.pyAccelerator, optical engine, cost and power model, TDP clock
sim.pySimPy engine: units, HBM, scratchpad, issue window, prefetch, write-backs, DVFS
metrics.py, trace.pyMetrics, bound attribution, hot-spots; Perfetto export
precision.pyFunctional model of an analogue NTT (deck 04)
search.pyAnalytic bound, sweeps, bisection, Pareto front
heir_frontend.pyReads HEIR's ckks-dialect output: parameters, levels, SSA dependencies, bootstraps and Chebyshev activations, as a trace
openfhe_trace.pyReads a kernel stream recorded from instrumented OpenFHE; per-stage counts; converts it to a replayable trace
web/sim_engine.jsThe port running on this page, with its minimal SimPy core
Quick start
pip install -e .[dev] && pytest                        # 89 tests
fhe-sim                                                # the default run on slide 08
fhe-sim --counts                                       # deck 02's per-stage table
fhe-sim --sweep-sram 128 256 512 1024 2048             # traffic against SRAM
fhe-sim --min-ks --seeded-keys --otf-pt --trace b.json # open in ui.perfetto.dev
fhe-sim --openfhe-log calibration/openfhe_trace/sparse16.log.gz --hw cpu   # a real OpenFHE bootstrap
fhe-sim --heir calibration/heir/lola.ckks.mlir.gz         # LoLa, compiled by HEIR
python examples/results.py                             # every number in this series
14

What to Take Away

Next

Deck 04 adds an optical transform engine and asks what precision and conversion cost.