The modelled machine (NTT units, modular multiply-add lanes, an automorphism network, a scratchpad, HBM and key streaming), the operation-trace format and where traces come from (a scheme model, HEIR, OpenFHE), the SimPy engine, its metrics and its validation ladder, and the whole simulator running live in the browser.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Cryptography decks or LLM Inference Simulators.
These are architecture-exploration questions, so the model sits at the transaction / discrete-event rung of the fidelity ladder (LLM Inference Simulators deck 01): one event per kernel and per memory chunk, not per cycle. A full bootstrap simulates in tens of milliseconds, so a design sweep of hundreds of points takes seconds.
The defaults are an ARK-class digital design at 1 GHz, sized like the published ASICs (ARK: 512 MB scratchpad, 1 TB/s HBM). The units are aggregates: four NTT pipelines become one resource with their combined throughput. All coefficients are illustrative.
The scheme model emits HE operations in program order. Each names its inputs and output, the evaluation key and plaintexts it needs, and its primitive kernels as [kind, amount, on-chip words]. Two real lines from fhe-sim --params small --dump-trace t.json (two hoisted baby-step rotations):
{"id": 2, "op": "hrot", "stage": "cts", "level": 11, "boot": 0, "inputs": ["b0.c1", "b0.c2"],
"output": "b0.c3", "key": ["cts0.b1", 3145728], "pts": [],
"kernels": [["auto", 245760, 491520], ["mac", 393216, 720896], ["intt", 8, 65536],
["bconv", 524288, 262144], ["ntt", 24, 196608]]}
{"id": 3, "op": "hrot", "stage": "cts", "level": 11, "boot": 0, "inputs": ["b0.c1", "b0.c2"],
"output": "b0.c4", "key": ["cts0.b2", 3145728], ...}
b0.c2 is the shared ModUp output (hoisting), so both rotations depend on it. Their keys differ, which is the whole data-movement story of deck 02.| Source | How | Status here |
|---|---|---|
Scheme model (workload.py) | Builds bootstrapping from the algorithm: key-switch recipes, BSGS DFT levels, Chebyshev EvalMod, level tracking | Implemented; operation counts are tested against closed forms |
| HEIR (heir.dev; Ali et al., arXiv:2508.11095) | An MLIR-based FHE compiler. After --torch-linalg-to-ckks with unrolled kernel loops, the server function is straight-line code in HEIR's ckks dialect, with every value's level in its type. HEIR has already chosen the packing, rotations, rescales and bootstrap placement | Done (heir_frontend.py, release v2026.10.01): one HE operation per ckks op, dependencies from SSA, levels checked against HEIR's types. Three of HEIR's example programs are in the repo (slide 12). HEIR's dialects change quickly, so the front end names the release it was validated on |
| OpenFHE (ePrint 2022/915) | Instrument the C++ library's NTT, base-conversion, key-switching, automorphism and rescale routines; log each call with its shape and the identity of its key or plaintext; replay the stream | Done for two bootstraps of OpenFHE v1.5.1 (an 82-line patch, in calibration/openfhe_trace). The model is checked against them in deck 02; the replay is the best timing predictor (slide 11) |
| Lattigo (Go) | The same instrumentation approach | Not done |
A hand-built scheme model reproduces the algorithm someone chose. A trace from the compiler that will actually target the hardware captures its real level management, packing and fusion decisions. The OpenFHE traces show the catch: a CPU library's trace carries CPU-tuned choices. OpenFHE's BSGS split saves a third of the NTTs but needs more rotation keys, which makes a memory-bound accelerator 42% slower (deck 05). That is why the compiler targeting the hardware, not a CPU library, should produce the trace; deck 09 of the sister series discusses the same point for PyTorch and ONNX.
Every HE operation is a SimPy process; every functional unit and HBM is a simpy.Resource. The core of sim.py, lightly trimmed:
def run_op(self, o, pl: OpPlan):
env = self.env
pf = env.process(self.prefetch(o, pl)) if (pl.key_load or pl.pt_load) else None # keys don't wait for data
for name, size in pl.writebacks: # spills decided at issue
ev = env.process(self.writeback(name, size))
self.wb_ev[name] = ev
for x in o.inputs: # dependencies
p = self.producer.get(x)
if p is not None:
yield from self.wait_done(p)
for x, size in pl.ct_load: # reload spilled ciphertexts
w = self.wb_ev.get(x)
if w is not None and not w.processed:
yield w
yield from self.xfer(size, "ct_read", o.stage)
if pf is not None and not pf.processed:
yield pf
for k in o.kernels:
for seg in self.cost.segments(k): # cost model: unit, time, joules
req = self.units[seg.unit].request()
yield req
... # power and statistics bookkeeping
yield env.timeout(seg.time)
self.units[seg.unit].release(req)
For every algorithm tried, more SRAM never increases key traffic or total HBM traffic (sweeps from 128 MiB to 2 GiB). That justifies using bisection to find the smallest adequate scratchpad (deck 05).
The manager guarantees the peak never exceeds the TDP, and a test checks it on every configuration, including Hypothesis-generated ones. It is greedy and first-come, with the clock chosen by bisection (no transcendentals, so the JavaScript port matches it bit for bit). The older alternative, one worst-case clock at which every unit and HBM flat out fit the TDP, is kept as power_mode="worst-case". It reserves power for a coincidence the workload never produces, so it can be 1.8× slower, or unable to run at all at a low TDP (deck 05). Methodology for all of this is in LLM Inference Simulators deck 07.
| Metric | Definition |
|---|---|
| Latency, per stage | Simulated time; span of each stage from its first operation's start to its last operation's end |
| Utilisation | Busy time ÷ total time for NTT, MAC, automorphism, optical and HBM |
| Bound attribution | The most-utilised resource: HBM → memory-bound; NTT (or optical) → NTT-bound; MAC → MAC-bound; plus "power-bound" when a compute-bound design lost more than 10% of the run to the power manager (or, with a worst-case clock, ran below 100%) |
| Hot-spot per stage | For each stage, the resource with the most busy time attributed to that stage's operations |
| Traffic by class | HBM bytes for keys, plaintexts, ciphertext reads and writes; key share of the total |
| Energy, power | Static × time + dynamic per unit + HBM; average power; peak instantaneous power from active units |
| Lower bounds | Busiest resource's busy time (a schedule bound) and an analytic roofline bound from the trace alone |
| Perfetto trace | --trace boot.json: one row per unit and for HBM, one slice per kernel and per chunk |
Default run (ark set, ARK-class design): 13.94 ms, memory-bound (HBM 89% busy, NTT 13%, MAC 29%), CoeffToSlot 72% of the time with HBM as its hot-spot, EvalMod's hot-spot the MAC lanes. Keys are 54% of 12.44 GB of HBM traffic; 1,139 mJ per bootstrap at 82 W average and 176 W peak against a 250 W TDP.
The JavaScript port of the SimPy model, with a minimal SimPy core that reproduces SimPy's event ordering. It matches the Python package bit for bit, including clocks, bytes, energy and per-operation end times; the repo's tests check this. Choose a configuration and press Run, SRAM sweep or Compare algorithms.
89 tests, organised as a verification ladder (the method is in LLM Inference Simulators deck 06):
| Rung | Checks |
|---|---|
| Unit: sizes | Ciphertext, plaintext and key sizes against hand formulas and ARK's Table III (24/12/120 MiB; 25/12.5/150 MiB) |
| Unit: operation counts | Key-switch transforms and base-conversion counts; rotations, PMults, keys, HMults and levels per bootstrap stage against closed forms, for four parameter sets |
| Invariants | Dependencies respected; work conserved on every unit; traffic at least compulsory; reads only after writes; determinism; probes passive |
| Analytic | A dependent chain sums exactly; independent load-then-compute operations match the two-machine flow-shop makespan a + b + (n−1)·max(a, b) to 1e-12; simulated time ≥ the roofline bound |
| Behaviour | More SRAM never raises traffic; the baseline is memory-bound; key-reuse techniques shift the bound; they hurt NTT-bound designs; optics helps NTT-bound designs and not memory-bound ones |
| Power | Energy accounts add up; peak ≤ TDP under both power modes, down to a 100 W TDP; the dynamic manager beats worst-case clocking where that wastes the budget and throttles when the real draw binds; DVFS saves energy when memory-bound |
| Property-based | Hypothesis: 60 random parameter sets, algorithms, hardware and power modes; all invariants, the TDP and determinism |
| Real library | Two recorded OpenFHE bootstraps: identical key-switch traffic where the algorithms agree, identical bootstrap depth (16 and 20), DFT transforms within 20% under OpenFHE's BSGS, and EvalMod's rescale-on-use difference pinned down (deck 02) |
| Compiler | Three HEIR-compiled programs: parameters, one HE op per ckks op, every level matching HEIR's types, bootstraps and polynomial activations expanded, and the same LoLa matching OpenFHE's execution on rotations, keys and relinearisations |
| External | Optical precision rule against a functional model (deck 04); timing calibration against OpenFHE; JavaScript against Python, bit-exact on 26 configurations, in both power modes |
| Calibration (openfhe-python, i7-3770, 8 threads) | OpenFHE | Simulated | Error |
|---|---|---|---|
| HMult, N = 216, 24 limbs, dnum 4 (fitted) | 352.6 ms | 352.6 ms | 0% |
| HRotate, same parameters (predicted) | 324.8 ms | 292.5 ms | −10% |
| Bootstrap, N = 216, 8 slots, dnum 3: scheme model (predicted) | 12,113 ms | 8,473 ms | −30% |
| Same bootstrap: OpenFHE's own recorded kernel stream, replayed (predicted) | 12,113 ms | 10,759 ms | −11% |
One parameter was fitted (0.18 modular operations per cycle per unit). The scheme model under-predicts the bootstrap mainly because OpenFHE rescales copies of each input before every multiplication, doubling EvalMod's transforms. Replaying the recorded stream removes that gap; most of the remaining −11% is element-wise work the tracer does not log. A full-slot N = 216 OpenFHE bootstrap could not be measured or recorded: its keys and precomputed plaintexts exceeded the memory available on the 15 GB test machine.
Three of HEIR's own example programs, compiled with the HEIR v2026.10.01 release binaries and simulated on the ARK-class design (fhe-sim --heir). HEIR picked every parameter, rotation, rescale and bootstrap; the front end expands each op into kernels. LoLa is an MNIST CNN with square activations (55 rotations); the MLP has a polynomial ReLU; the third row is LoLa compiled with a level budget of 2, so HEIR must place a bootstrap. All four runs are memory-bound.
| Program | N, limbs | HE ops | Scratchpad | Latency | Keys / plaintexts / ciphertexts |
|---|---|---|---|---|---|
| LoLa (MNIST CNN) | 215, 11 | 409 | 512 MiB | 0.61 ms | 0.35 / 0.17 / 0.00 GB |
| MNIST MLP | 215, 13 | 1,655 | 512 MiB | 2.25 ms | 0.73 / 0.95 / 0.27 GB |
| LoLa + HEIR bootstrap | 217, 47 | 625 | 512 MiB | 202.75 ms | 61.43 / 24.58 / 112.46 GB |
| … same | 217, 47 | 625 | 2 GiB | 101.45 ms | 47.92 / 24.58 / 21.20 GB |
FHE_Accelerator_Sim follows the structure of the sister repo, Disaggregated_Inference_Sim: scheme model, cost model, engine and probes in separate files.
| Module | Role |
|---|---|
params.py | CKKS parameter sets and sizes |
workload.py | Scheme model: bootstrap and single-op traces, JSON dump and replay |
hardware.py | Accelerator, optical engine, cost and power model, TDP clock |
sim.py | SimPy engine: units, HBM, scratchpad, issue window, prefetch, write-backs, DVFS |
metrics.py, trace.py | Metrics, bound attribution, hot-spots; Perfetto export |
precision.py | Functional model of an analogue NTT (deck 04) |
search.py | Analytic bound, sweeps, bisection, Pareto front |
heir_frontend.py | Reads HEIR's ckks-dialect output: parameters, levels, SSA dependencies, bootstraps and Chebyshev activations, as a trace |
openfhe_trace.py | Reads a kernel stream recorded from instrumented OpenFHE; per-stage counts; converts it to a replayable trace |
web/sim_engine.js | The port running on this page, with its minimal SimPy core |
pip install -e .[dev] && pytest # 89 tests
fhe-sim # the default run on slide 08
fhe-sim --counts # deck 02's per-stage table
fhe-sim --sweep-sram 128 256 512 1024 2048 # traffic against SRAM
fhe-sim --min-ks --seeded-keys --otf-pt --trace b.json # open in ui.perfetto.dev
fhe-sim --openfhe-log calibration/openfhe_trace/sparse16.log.gz --hw cpu # a real OpenFHE bootstrap
fhe-sim --heir calibration/heir/lola.ckks.mlir.gz # LoLa, compiled by HEIR
python examples/results.py # every number in this series
Deck 04 adds an optical transform engine and asks what precision and conversion cost.