Simulation Engineering Toolkit — Presentation 13

Power, Performance and Area: the Architect's Trade-offs

The three axes an architect trades: performance (latency and throughput), power (dynamic, static, DVFS, dark silicon) and area (SRAM, logic, PHYs and wires, estimated before layout), and why area is cost (dies per wafer, Poisson and Murphy yield, the reticle). Composite metrics (perf/W, perf/mm², EDP, TCO), Pareto fronts, a live explorer on FHE_Accelerator_Sim's new area model, a worked three-way trade-off, and the other trade-offs the simulation series make, named.

PPA Dynamic and static power Area and yield perf/W, perf/mm², EDP Pareto fronts Trade-offs
Design → Simulate → Price the area → Combine → Pareto → Choose
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

01

What PPA Is, and Who Trades It

PPA is power, performance and area: the three quantities every chip design is judged on. The simulation series in this GitHub measure performance and power carefully; this deck adds the third axis, and names the trade-offs between all three.

WhoDecidesTrades, for example
ArchitectureWhich units, how many, how much on-chip memory, which memory technology, which algorithms the hardware is built forMore SRAM cuts off-chip traffic (performance, energy) and costs area
Micro-architecturePipelining, banking, datapath widths, the clock target, how units share ports and wiresA deeper pipeline raises the clock and costs registers (area, power)
Physical designFloorplan, cell libraries and threshold voltages, SRAM bitcell choice, wire layers, power gridLow-leakage cells cut static power and are slower
02

Performance: Latency and Throughput

Performance is two different quantities, and a design can win one while losing the other:

The metrics, their statistics and how to attribute them to hot-spots are covered in InfSim 06; this deck does not repeat them. Two points matter here:

03

Power: Dynamic, Static, DVFS and Dark Silicon

The two kinds of power

Dynamic: Pdyn = α C V² f, for activity factor α (the fraction of the capacitance C switched each cycle), supply voltage V and clock f. Static (leakage): Pstatic = V · Ileak, drawn whenever the circuit is powered, busy or not.

04

Area: What Sets It

Four things take the silicon: SRAM bitcells (with their decoders and sense amplifiers), logic (here modular multipliers, which dominate FHE datapaths), I/O and PHYs (the analogue interfaces to HBM, PCIe, die-to-die links) and wiring, with the repeaters and pipeline registers long wires need. A published breakdown shows the proportions: ARK, an FHE accelerator modelled at 7 nm (Kim et al., MICRO 2022, Table IV), reproduced by the area model in FHE_Accelerator_Sim:

designNTTMACpermuteSRAMuncoreHBM PHYdie mm²
ARK as published (8,192 bfly, 1,024 perm. words/cycle)57.218.220.6229.263.429.6418.2
ARK-class (this model's default)28.618.282.4229.269.929.6457.9
small digital3.64.520.6229.250.359.2367.4
small + hybrid optical*3.64.520.6229.250.359.2368.9
small + ideal optical*3.64.520.6229.250.359.2368.9

Source: examples/results.md in FHE_Accelerator_Sim

05

Estimating Area Before Layout

Before RTL, area comes from analytical models; after RTL, from synthesis; only layout gives the real number. Three pre-RTL estimators are explained, with their accuracy, on SimEng 12's cards: CACTI (SRAM and caches), McPAT (processors) and Accelergy and Timeloop (accelerators, from component tables). FHE_Accelerator_Sim's area model (ppa.py) is built the same way, from three sources:

  1. Units: ARK's per-unit areas (above) divided by its unit counts: about 0.007 mm² per NTT butterfly and 0.002 mm² per multiply-add lane, at 7 nm, wiring included. Scaling linearly to other counts is the illustrative part: wiring grows faster than the units it joins.
  2. SRAM: CACTI 7, built and run here, gives the shape of area against capacity for banked scratchpads (22 nm is its smallest node). One factor scales it to 7 nm so that 512 MiB matches ARK's scratchpad; BTS's 512 MB agrees within 3%.
  3. Everything else: HBM PHYs per stack, and register files plus on-chip network as ARK's fraction of the rest.
MiB22 nm mm²/MiB (CACTI)array efficiency7 nm mm²/MiB (model)7 nm mm² (model)
640.918972.6%0.468830.0
1280.890174.9%0.454158.1
2560.902073.9%0.4602117.8
5120.877476.0%0.4477229.2
10240.818981.4%0.4178427.8
20480.799583.4%0.4079835.4

Source: examples/results.md in FHE_Accelerator_Sim

06

Why Area Is Cost: Dies per Wafer and Yield

Three formulas

Dies per wafer (wafer diameter d, die area A): π(d/2)²/A − πd/√(2A); the second term is the partial dies lost round the edge. Poisson yield: Y = e−A·D0 for defect density D0, if defects land independently and uniformly. Murphy yield (Murphy, 1964): Y = ((1 − e−A·D0) / (A·D0))², for a defect density that varies across the wafer. Cost per good die = wafer price / (dies per wafer × Y).

07

Composite Metrics: perf/W, perf/mm², EDP and TCO

Combining axes into one number makes designs comparable, and each combination hides something. With perf = tasks per second, t = time per task and E = energy per task:

MetricFormulaRewardsHides
perf/W(1/t) / (E/t) = 1/EEnergy per task, and nothing elseSpeed: a slow design with the same energy scores the same
perf/mm²1 / (t · A)Throughput per unit of silicon (cost)Energy, and yield (a mm² on a big die costs more)
perf/$(1/t) / cost per good dieThroughput per unit of manufacturing costRunning cost, and everything but silicon
EDPE · tEnergy and speed equallyArea
ED²PE · t²Speed more than energy; roughly independent of supply voltage where f rises with V (E ∝ V², t ∝ 1/V), so it compares designs, not operating pointsArea
TCOpurchase cost amortised over the service life + energy × electricity price × datacentre overheadThe money an operator actually spendsNothing, but it needs prices you rarely have pre-silicon
08

Pareto Fronts: Why There Is No Single Best Design

Design A dominates design B if A is no worse on every axis and better on at least one. The Pareto front is the set of designs nothing dominates. Every design off the front can be improved for free; every design on it is a different compromise, and choosing among them needs the product's priorities, not more simulation.

area (lower is better) latency (lower is better) on the front: nothing beats it dominated arrows: front designs that beat the dominated one on both axes schematic, not data: the live version is the next slide
09

Interactive: PPA Explorer

Each design is simulated in your browser by FHE_Accelerator_Sim's JavaScript port (bit-exact with the Python package) running one full-slot ark-set bootstrap, then priced by the ported area model. Grey points are a grid of 80 designs; green ones are on the three-way (latency, energy, area) Pareto front; the ring is the design the controls select. The plot shows two of the three axes, so a green point can look dominated here and still win on energy (hover for its numbers). Area coefficients are 7 nm and illustrative; the node controls scale them.

10

Worked Example: The Scratchpad Sweep

FHE_Accelerator_Sim's scratchpad sweep, priced in area (results.md §22). With the baseline algorithm every rotation key is used once per bootstrap, so the design stays memory-bound and SRAM keeps buying latency:

SRAM MiBbootstrapmJdie mm²perf/mm² (1/s/mm²)perf/W (1/J)EDP (mJ·s)Pareto
12826.98 ms2084253.50.1460.4856.22yes
25617.10 ms1383324.80.1800.7223.65yes
38414.57 ms1193392.30.1750.8417.39yes
51213.94 ms1139457.90.1570.8815.87yes
76813.33 ms1096581.10.1290.9114.60yes
102411.69 ms984695.20.1231.0211.50yes
204811.41 ms9661182.3 (> reticle)0.0741.0411.02yes
409611.41 ms9662180.5 (> reticle)0.0401.0411.02

Source: examples/results.md in FHE_Accelerator_Sim

With Min-KS, seeded keys and on-the-fly plaintexts the key traffic is gone by 512 MiB and the design turns MAC-bound:

SRAM MiBbootstrapmJdie mm²perf/mm² (1/s/mm²)perf/W (1/J)EDP (mJ·s)Pareto
12821.22 ms1691253.50.1860.5935.88yes
2568.71 ms774324.80.3541.296.74yes
3847.56 ms646392.30.3371.554.88yes
5127.19 ms607457.90.3041.654.37yes
7687.19 ms607581.10.2391.654.37
10247.19 ms607695.20.2001.654.37
20487.19 ms6071182.3 (> reticle)0.1181.654.37
40967.19 ms6072180.5 (> reticle)0.0641.654.37

Source: examples/results.md in FHE_Accelerator_Sim

11

Compute or SRAM, and When to Split the Die

From a 256 MiB design, one change at a time; "ms saved per 100 mm²" is the return on the area spent (§23). Baseline algorithm (memory-bound):

designdie mm²area addedbootstrapmJperf/mm²ms saved per 100 mm²
start: 256 MiB, 4,096 bfly, 8,192 MAC324.8+0.017.10 ms, memory-bound13830.180-
2x NTT + MAC380.7+55.916.98 ms, memory-bound13790.1550.21
4x NTT + MAC492.5+167.816.96 ms, memory-bound13780.1200.08
2x MAC only346.5+21.717.14 ms, memory-bound13850.168-0.20
+128 MiB SRAM (384)392.3+67.514.57 ms, memory-bound11930.1753.74
+256 MiB SRAM (512)457.9+133.113.94 ms, memory-bound11390.1572.38
permutation network cut to 1,024 words/cycle250.9-73.817.09 ms, memory-bound13830.233-

Source: examples/results.md in FHE_Accelerator_Sim

All three techniques (MAC-bound):

designdie mm²area addedbootstrapmJperf/mm²ms saved per 100 mm²
start: 256 MiB, 4,096 bfly, 8,192 MAC324.8+0.08.71 ms, MAC-bound7740.354-
2x NTT + MAC380.7+55.97.77 ms, MAC-bound7360.3381.68
4x NTT + MAC492.5+167.87.63 ms, MAC-bound7310.2660.64
2x MAC only346.5+21.78.64 ms, MAC-bound7710.3340.34
+128 MiB SRAM (384)392.3+67.57.56 ms, MAC-bound6460.3371.71
+256 MiB SRAM (512)457.9+133.17.19 ms, MAC-bound6070.3041.14
permutation network cut to 1,024 words/cycle250.9-73.88.74 ms, MAC-bound7750.456-

Source: examples/results.md in FHE_Accelerator_Sim

Splitting a die that is too big (§24): the 2 GiB design as equal chiplets, silicon only.

chipletsmm² eachfits the reticleMurphy yield eachsilicon $ for the set
11182.3no34.4%$727
2591.1yes57.0%$381
4295.6yes75.0%$267
8147.8yes86.4%$219

Source: examples/results.md in FHE_Accelerator_Sim

12

The Other Trade-offs, Named

The simulation series make several trade-offs without calling them that. Each is a front with two axes and no single best point:

Trade-offThe leverWhat it showsWhere
Latency against throughputBatch sizeBigger batches fill the hardware and raise throughput; every request in the batch waits longerInfSim 03 · batching, InfSim 05 · results
Time to first token against time per output tokenDisaggregating prefill and decodeSeparate pools stop prefill bursts stalling decode, at the cost of moving the KV cache and of hardware that cannot be sharedInfSim 05 · the interference problem, the idea
Precision against energyENOB of an optical engine's convertersEach extra bit doubles converter energy; too few bits forces narrower digits and more passesFHESim 04 · the precision tax, conversion energy
Memory against computeMin-KS, seeded keys, on-the-fly plaintextsRegenerating data on chip trades HBM traffic for multiplications: on the memory-rich ARK-class design a bootstrap falls from 13.94 ms to 7.19 ms; on the NTT-starved design it rises from 17.46 ms to 26.65 msFHESim 02 · algorithmic levers, FHESim 05 · acceleration techniques
Fidelity against simulation speedThe model's level of abstractionEach rung down the fidelity ladder is slower and more exact; loosely timed SystemC models trade timing error for speed through the quantumInfSim 01 · the fidelity ladder, InfSim 08 · choose the abstraction, SimEng 03 · the quantum trade-off
13

Pitfalls

14

What to Take Away

Re-running this deck's numbers

Every number comes from examples/results.py in FHE_Accelerator_Sim (§21–25), whose area model is src/fhe_sim/ppa.py with its CACTI calibration in calibration/cacti/; the deck build inlines them from results.md. fhe-sim --area prints the breakdown for any design.