The three axes an architect trades: performance (latency and throughput), power (dynamic, static, DVFS, dark silicon) and area (SRAM, logic, PHYs and wires, estimated before layout), and why area is cost (dies per wafer, Poisson and Murphy yield, the reticle). Composite metrics (perf/W, perf/mm², EDP, TCO), Pareto fronts, a live explorer on FHE_Accelerator_Sim's new area model, a worked three-way trade-off, and the other trade-offs the simulation series make, named.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.
PPA is power, performance and area: the three quantities every chip design is judged on. The simulation series in this GitHub measure performance and power carefully; this deck adds the third axis, and names the trade-offs between all three.
| Who | Decides | Trades, for example |
|---|---|---|
| Architecture | Which units, how many, how much on-chip memory, which memory technology, which algorithms the hardware is built for | More SRAM cuts off-chip traffic (performance, energy) and costs area |
| Micro-architecture | Pipelining, banking, datapath widths, the clock target, how units share ports and wires | A deeper pipeline raises the clock and costs registers (area, power) |
| Physical design | Floorplan, cell libraries and threshold voltages, SRAM bitcell choice, wire layers, power grid | Low-leakage cells cut static power and are slower |
Performance is two different quantities, and a design can win one while losing the other:
The metrics, their statistics and how to attribute them to hot-spots are covered in InfSim 06; this deck does not repeat them. Two points matter here:
Dynamic: Pdyn = α C V² f, for activity factor α (the fraction of the capacitance C switched each cycle), supply voltage V and clock f. Static (leakage): Pstatic = V · Ileak, drawn whenever the circuit is powered, busy or not.
Four things take the silicon: SRAM bitcells (with their decoders and sense amplifiers), logic (here modular multipliers, which dominate FHE datapaths), I/O and PHYs (the analogue interfaces to HBM, PCIe, die-to-die links) and wiring, with the repeaters and pipeline registers long wires need. A published breakdown shows the proportions: ARK, an FHE accelerator modelled at 7 nm (Kim et al., MICRO 2022, Table IV), reproduced by the area model in FHE_Accelerator_Sim:
| design | NTT | MAC | permute | SRAM | uncore | HBM PHY | die mm² |
|---|---|---|---|---|---|---|---|
| ARK as published (8,192 bfly, 1,024 perm. words/cycle) | 57.2 | 18.2 | 20.6 | 229.2 | 63.4 | 29.6 | 418.2 |
| ARK-class (this model's default) | 28.6 | 18.2 | 82.4 | 229.2 | 69.9 | 29.6 | 457.9 |
| small digital | 3.6 | 4.5 | 20.6 | 229.2 | 50.3 | 59.2 | 367.4 |
| small + hybrid optical* | 3.6 | 4.5 | 20.6 | 229.2 | 50.3 | 59.2 | 368.9 |
| small + ideal optical* | 3.6 | 4.5 | 20.6 | 229.2 | 50.3 | 59.2 | 368.9 |
Source: examples/results.md in FHE_Accelerator_Sim
Before RTL, area comes from analytical models; after RTL, from synthesis; only layout gives the real number. Three pre-RTL estimators are explained, with their accuracy, on SimEng 12's cards: CACTI (SRAM and caches), McPAT (processors) and Accelergy and Timeloop (accelerators, from component tables). FHE_Accelerator_Sim's area model (ppa.py) is built the same way, from three sources:
| MiB | 22 nm mm²/MiB (CACTI) | array efficiency | 7 nm mm²/MiB (model) | 7 nm mm² (model) |
|---|---|---|---|---|
| 64 | 0.9189 | 72.6% | 0.4688 | 30.0 |
| 128 | 0.8901 | 74.9% | 0.4541 | 58.1 |
| 256 | 0.9020 | 73.9% | 0.4602 | 117.8 |
| 512 | 0.8774 | 76.0% | 0.4477 | 229.2 |
| 1024 | 0.8189 | 81.4% | 0.4178 | 427.8 |
| 2048 | 0.7995 | 83.4% | 0.4079 | 835.4 |
Source: examples/results.md in FHE_Accelerator_Sim
Dies per wafer (wafer diameter d, die area A): π(d/2)²/A − πd/√(2A); the second term is the partial dies lost round the edge. Poisson yield: Y = e−A·D0 for defect density D0, if defects land independently and uniformly. Murphy yield (Murphy, 1964): Y = ((1 − e−A·D0) / (A·D0))², for a defect density that varies across the wafer. Cost per good die = wafer price / (dies per wafer × Y).
Combining axes into one number makes designs comparable, and each combination hides something. With perf = tasks per second, t = time per task and E = energy per task:
| Metric | Formula | Rewards | Hides |
|---|---|---|---|
| perf/W | (1/t) / (E/t) = 1/E | Energy per task, and nothing else | Speed: a slow design with the same energy scores the same |
| perf/mm² | 1 / (t · A) | Throughput per unit of silicon (cost) | Energy, and yield (a mm² on a big die costs more) |
| perf/$ | (1/t) / cost per good die | Throughput per unit of manufacturing cost | Running cost, and everything but silicon |
| EDP | E · t | Energy and speed equally | Area |
| ED²P | E · t² | Speed more than energy; roughly independent of supply voltage where f rises with V (E ∝ V², t ∝ 1/V), so it compares designs, not operating points | Area |
| TCO | purchase cost amortised over the service life + energy × electricity price × datacentre overhead | The money an operator actually spends | Nothing, but it needs prices you rarely have pre-silicon |
Design A dominates design B if A is no worse on every axis and better on at least one. The Pareto front is the set of designs nothing dominates. Every design off the front can be improved for free; every design on it is a different compromise, and choosing among them needs the product's priorities, not more simulation.
search.pareto_nd compares every pair (O(n²), fine for hundreds of designs). Finding the points is the expensive part, which is why smarter experiments (bounds, bisection, surrogates) matter.Each design is simulated in your browser by FHE_Accelerator_Sim's JavaScript port (bit-exact with the Python package) running one full-slot ark-set bootstrap, then priced by the ported area model. Grey points are a grid of 80 designs; green ones are on the three-way (latency, energy, area) Pareto front; the ring is the design the controls select. The plot shows two of the three axes, so a green point can look dominated here and still win on energy (hover for its numbers). Area coefficients are 7 nm and illustrative; the node controls scale them.
FHE_Accelerator_Sim's scratchpad sweep, priced in area (results.md §22). With the baseline algorithm every rotation key is used once per bootstrap, so the design stays memory-bound and SRAM keeps buying latency:
| SRAM MiB | bootstrap | mJ | die mm² | perf/mm² (1/s/mm²) | perf/W (1/J) | EDP (mJ·s) | Pareto |
|---|---|---|---|---|---|---|---|
| 128 | 26.98 ms | 2084 | 253.5 | 0.146 | 0.48 | 56.22 | yes |
| 256 | 17.10 ms | 1383 | 324.8 | 0.180 | 0.72 | 23.65 | yes |
| 384 | 14.57 ms | 1193 | 392.3 | 0.175 | 0.84 | 17.39 | yes |
| 512 | 13.94 ms | 1139 | 457.9 | 0.157 | 0.88 | 15.87 | yes |
| 768 | 13.33 ms | 1096 | 581.1 | 0.129 | 0.91 | 14.60 | yes |
| 1024 | 11.69 ms | 984 | 695.2 | 0.123 | 1.02 | 11.50 | yes |
| 2048 | 11.41 ms | 966 | 1182.3 (> reticle) | 0.074 | 1.04 | 11.02 | yes |
| 4096 | 11.41 ms | 966 | 2180.5 (> reticle) | 0.040 | 1.04 | 11.02 |
Source: examples/results.md in FHE_Accelerator_Sim
With Min-KS, seeded keys and on-the-fly plaintexts the key traffic is gone by 512 MiB and the design turns MAC-bound:
| SRAM MiB | bootstrap | mJ | die mm² | perf/mm² (1/s/mm²) | perf/W (1/J) | EDP (mJ·s) | Pareto |
|---|---|---|---|---|---|---|---|
| 128 | 21.22 ms | 1691 | 253.5 | 0.186 | 0.59 | 35.88 | yes |
| 256 | 8.71 ms | 774 | 324.8 | 0.354 | 1.29 | 6.74 | yes |
| 384 | 7.56 ms | 646 | 392.3 | 0.337 | 1.55 | 4.88 | yes |
| 512 | 7.19 ms | 607 | 457.9 | 0.304 | 1.65 | 4.37 | yes |
| 768 | 7.19 ms | 607 | 581.1 | 0.239 | 1.65 | 4.37 | |
| 1024 | 7.19 ms | 607 | 695.2 | 0.200 | 1.65 | 4.37 | |
| 2048 | 7.19 ms | 607 | 1182.3 (> reticle) | 0.118 | 1.65 | 4.37 | |
| 4096 | 7.19 ms | 607 | 2180.5 (> reticle) | 0.064 | 1.65 | 4.37 |
Source: examples/results.md in FHE_Accelerator_Sim
From a 256 MiB design, one change at a time; "ms saved per 100 mm²" is the return on the area spent (§23). Baseline algorithm (memory-bound):
| design | die mm² | area added | bootstrap | mJ | perf/mm² | ms saved per 100 mm² |
|---|---|---|---|---|---|---|
| start: 256 MiB, 4,096 bfly, 8,192 MAC | 324.8 | +0.0 | 17.10 ms, memory-bound | 1383 | 0.180 | - |
| 2x NTT + MAC | 380.7 | +55.9 | 16.98 ms, memory-bound | 1379 | 0.155 | 0.21 |
| 4x NTT + MAC | 492.5 | +167.8 | 16.96 ms, memory-bound | 1378 | 0.120 | 0.08 |
| 2x MAC only | 346.5 | +21.7 | 17.14 ms, memory-bound | 1385 | 0.168 | -0.20 |
| +128 MiB SRAM (384) | 392.3 | +67.5 | 14.57 ms, memory-bound | 1193 | 0.175 | 3.74 |
| +256 MiB SRAM (512) | 457.9 | +133.1 | 13.94 ms, memory-bound | 1139 | 0.157 | 2.38 |
| permutation network cut to 1,024 words/cycle | 250.9 | -73.8 | 17.09 ms, memory-bound | 1383 | 0.233 | - |
Source: examples/results.md in FHE_Accelerator_Sim
All three techniques (MAC-bound):
| design | die mm² | area added | bootstrap | mJ | perf/mm² | ms saved per 100 mm² |
|---|---|---|---|---|---|---|
| start: 256 MiB, 4,096 bfly, 8,192 MAC | 324.8 | +0.0 | 8.71 ms, MAC-bound | 774 | 0.354 | - |
| 2x NTT + MAC | 380.7 | +55.9 | 7.77 ms, MAC-bound | 736 | 0.338 | 1.68 |
| 4x NTT + MAC | 492.5 | +167.8 | 7.63 ms, MAC-bound | 731 | 0.266 | 0.64 |
| 2x MAC only | 346.5 | +21.7 | 8.64 ms, MAC-bound | 771 | 0.334 | 0.34 |
| +128 MiB SRAM (384) | 392.3 | +67.5 | 7.56 ms, MAC-bound | 646 | 0.337 | 1.71 |
| +256 MiB SRAM (512) | 457.9 | +133.1 | 7.19 ms, MAC-bound | 607 | 0.304 | 1.14 |
| permutation network cut to 1,024 words/cycle | 250.9 | -73.8 | 8.74 ms, MAC-bound | 775 | 0.456 | - |
Source: examples/results.md in FHE_Accelerator_Sim
Splitting a die that is too big (§24): the 2 GiB design as equal chiplets, silicon only.
| chiplets | mm² each | fits the reticle | Murphy yield each | silicon $ for the set |
|---|---|---|---|---|
| 1 | 1182.3 | no | 34.4% | $727 |
| 2 | 591.1 | yes | 57.0% | $381 |
| 4 | 295.6 | yes | 75.0% | $267 |
| 8 | 147.8 | yes | 86.4% | $219 |
Source: examples/results.md in FHE_Accelerator_Sim
The simulation series make several trade-offs without calling them that. Each is a front with two axes and no single best point:
| Trade-off | The lever | What it shows | Where |
|---|---|---|---|
| Latency against throughput | Batch size | Bigger batches fill the hardware and raise throughput; every request in the batch waits longer | InfSim 03 · batching, InfSim 05 · results |
| Time to first token against time per output token | Disaggregating prefill and decode | Separate pools stop prefill bursts stalling decode, at the cost of moving the KV cache and of hardware that cannot be shared | InfSim 05 · the interference problem, the idea |
| Precision against energy | ENOB of an optical engine's converters | Each extra bit doubles converter energy; too few bits forces narrower digits and more passes | FHESim 04 · the precision tax, conversion energy |
| Memory against compute | Min-KS, seeded keys, on-the-fly plaintexts | Regenerating data on chip trades HBM traffic for multiplications: on the memory-rich ARK-class design a bootstrap falls from 13.94 ms to 7.19 ms; on the NTT-starved design it rises from 17.46 ms to 26.65 ms | FHESim 02 · algorithmic levers, FHESim 05 · acceleration techniques |
| Fidelity against simulation speed | The model's level of abstraction | Each rung down the fidelity ladder is slower and more exact; loosely timed SystemC models trade timing error for speed through the quantum | InfSim 01 · the fidelity ladder, InfSim 08 · choose the abstraction, SimEng 03 · the quantum trade-off |
Every number comes from examples/results.py in FHE_Accelerator_Sim (§21–25), whose area model is src/fhe_sim/ppa.py with its CACTI calibration in calibration/cacti/; the deck build inlines them from results.md. fhe-sim --area prints the breakdown for any design.