Speed is half of performance. Static and dynamic power, why data movement dominates energy, the simulator's calibrated power model, prefill running hot and decode cool, DVFS and power caps as a third roof, energy proportionality, the link to RTL power and thermal analysis, and what photonic systems must pay for — all measured on the companion simulator.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.
Speed is only half of performance. For AI hardware the other half is power and energy, and it is often the half that decides the product.
P = P_static + P_dynamic
P_static ≈ V · I_leak # leakage, clocks, always-on logic: paid whenever powered
P_dynamic ≈ α · C · V² · f # switching: activity x capacitance x voltage^2 x frequency
E_op ≈ α · C · V² # energy per operation: why lowering V (with f) pays quadratically
Horowitz's widely cited energy table (ISSCC 2014, 45 nm) makes the key point: moving data costs far more than computing on it.
| Operation (45 nm) | Energy | Relative to an 8-bit add |
|---|---|---|
| 8-bit integer add | 0.03 pJ | 1× |
| 32-bit float multiply | 3.7 pJ | ~120× |
| 64-bit read from an 8 KB SRAM | 10 pJ | ~330× |
| 64-bit read from DRAM | 1.3–2.6 nJ | ~40,000–90,000× |
Absolute values have fallen at newer nodes, but the ratios still hold. This is why the simulator charges separately for FLOPs and for HBM bytes, and why decode, which streams almost every weight on every step (all but the input-embedding table, which is a lookup), is an energy problem as well as a bandwidth problem.
Four coefficients per device and one per link, attached to the same FLOP and byte counts the roofline already computes. Nothing new is counted; it is only priced.
E_total = Σinstances idle_W · T_run # static: charged busy or not
+ Σsteps (FLOPs · pJ/FLOP · s² + HBM bytes · pJ/byte) # dynamic, with DVFS factor s
+ Σtransfers KV bits · pJ/bit # interconnect
P_step = idle_W + E_step / t_step # tracked for the peak
| Device (illustrative) | TDP | Idle | pJ / FLOP | pJ / HBM byte |
|---|---|---|---|---|
| H100-SXM | 700 W | 100 W | 1.0 | 60 |
| A100-SXM | 400 W | 60 W | 1.6 | 70 |
| Hypothetical optical MAC | 700 W | 180 W | 0.1 | 60 |
| Links (pJ/bit) | NVLink 5 · PCIe 6 · InfiniBand / Ethernet 15 (SerDes, NIC and switch together) | |||
These coefficients are chosen to give realistic totals (prefill near TDP, decode around 300 W per H100). To calibrate for real hardware, run micro-benchmarks that sweep FLOP rate and byte rate independently, read power with DCGM or nvidia-smi, and fit P − idle = a·FLOP/s + b·byte/s by least squares. a and b are your pJ/FLOP and pJ/byte. For pre-silicon hardware the same coefficients come from RTL power analysis of the blocks (slide 09).
Per step, Llama-3-70B on 4×H100, from the simulator's cost model (as corrected on 2026-10-03; see slide 07):
| Step | Bound | Time | Power per device | Dynamic energy |
|---|---|---|---|---|
| Prefill, 2,048-token prompt | compute | 133.9 ms | 658 W (near the 700 W TDP) | 299 J |
| Decode, batch 16, ~2.3k context | memory | 14.6 ms | 295 W | 11.4 J |
| Decode, same, with DVFS | memory | 14.6 ms | 262 W | 9.4 J |
Dynamic voltage and frequency scaling lowers the clock (and with it the voltage). In the model the compute clock runs at a fraction s of nominal; HBM is unaffected.
t(s) = max(tc / s, tm) # compute slows; memory doesn't
E(s) = Ec · s² + Em # V tracks f, so energy per op ~ s^2
P(s) = idle + E(s) / t(s) # increases with s: lower clocks always reduce power
dvfs on: memory-bound step → s = max(s_min, tc/tm) # same time, less energy
power cap: choose the largest s with P(s) ≤ cap # closed form: sqrt, or a cubic (Cardano)
still too hot at s_min → stretch the step to fit the budget
Not enough FLOP/s. Prefill.
Not enough bytes/s. Decode.
Not enough watts. A capped prefill: the work could go faster, but not without exceeding the cap.
The simulator labels each step with whichever roof bound it and reports the fraction of busy time spent power-bound. Real DVFS has discrete operating points, voltage floors and transition latency; the continuous model is a first-order stand-in to refine when the hardware's operating points are known.
This uses the simulator's own cost model (the browser engine from deck 05). Pick a step, then move the per-device cap and toggle DVFS. The chart sweeps the cap from TDP down to just above idle and plots time against energy: the energy–delay trade-off.
Things to notice: for prefill, tighter caps trade time for energy until static power (charged for the longer step) starts to dominate and total energy rises again. That minimum is the energy-optimal operating point: for a 2,048-token prefill on H100s it sits near a 160 W cap, at 186 J against 352 J uncapped, but at 2.2× the step time. For decode, the curve is flat in time over a wide range of caps: free savings.
Full runs: Llama-3-70B, 4×H100 per instance, 4 req/s, 2,048/256-token mean prompts/outputs, SLOs TTFT ≤ 1 s and TPOT ≤ 25 ms, 800 requests. Every row is disagg-sim with the flags shown.
| Configuration | TTFT p99 | TPOT p99 | SLO met | Avg power | J / token |
|---|---|---|---|---|---|
| Colocated, 2 instances | 630 ms | 25.9 ms | 97.8% | 2,993 W | 3.09 |
| Disaggregated 1P1D | 854 ms | 15.2 ms | 99.7% | 2,687 W | 2.77 |
… + --dvfs | 854 ms | 15.2 ms | 99.7% | 2,570 W | 2.65 |
… + --decode-power-cap 250 | 854 ms | 17.2 ms | 99.7% | 2,503 W | 2.59 |
… + --decode-power-cap 200 | 854 ms | 27.2 ms | 12.6% | 2,284 W | 2.43 |
1P1D, --power-cap 400 --dvfs | 1,240 ms | 15.2 ms | 95.3% | 2,190 W | 2.26 |
1P1D, --power-cap 300 --dvfs | 1,651 ms | 15.2 ms | 83.3% | 2,013 W | 2.08 |
2P1D, --power-cap 400 --dvfs | 689 ms | 15.1 ms | 100% | 2,593 W | 2.68 |
Every step had been charged the whole input-embedding table, where a lookup reads only the rows it uses. And decode attention had left out each new token's attention to itself. Tracing the real model found both (Toolkit deck 10). All the tables in this deck were rerun with the corrected model. Decode TPOT fell slightly. The J/token ranking, the power-bound flags and the hot-spots did not change. The 200 W cliff remains, though 12.6% of requests now meet the SLOs, against 1.7% before. Before and after: examples/results.md.
An idle accelerator still burns its static power. At low load that is spread over few tokens.
| Offered load (1P1D) | Avg power | J / output token | Static share of energy |
|---|---|---|---|
| 0.5 req/s | 1,472 W | 11.8 | 54% |
| 1 req/s | 1,729 W | 7.0 | 46% |
| 2 req/s | 2,066 W | 4.2 | 39% |
| 4 req/s | 2,687 W | 2.8 | 30% |
| 6 req/s | 3,254 W | 2.3 | 25% |
Power modelling follows the same fidelity ladder as timing (deck 01), and the levels feed each other the same way.
| Level | Power method | Tools |
|---|---|---|
| Analytical / DES | Energy per event × event counts (this simulator) | Custom models; Accelergy for per-access energy tables |
| Architecture | Component power from structure: caches, NoCs, cores | McPAT, CACTI (SRAM/cache energy), Accelergy + Timeloop |
| RTL | Switching activity from simulation (SAIF, VCD, FSDB) applied to the netlist or RTL | Commercial RTL and gate-level power analysis tools |
| Emulation | Activity over real software runs, millions of cycles | Emulator-based dynamic power analysis |
| Silicon | Measurement | Board telemetry, DCGM, power analysers |
An optical multiply can cost very little energy. A photonic system pays for everything around it, and a simulator must account for all of it.
| Consumer | Behaviour | How to model it |
|---|---|---|
| Lasers | Continuous optical power, limited wall-plug efficiency; on whether or not work arrives | Static power, with gating policies and their wake-up latency |
| Thermal tuning | Heaters hold resonant devices on wavelength; depends on ambient and neighbouring heat | Static plus thermal-state dependent |
| DACs and modulators | Energy per conversion grows with sample rate and resolution | Per-sample energy × samples; a Walden-style figure of merit (energy ∝ 2ENOB) |
| Photodetectors, TIAs and ADCs | As above, on the way back to electronics | Per-sample energy; precision against power trade-off |
| Digital pre- and post-processing | Normalisation, accumulation, modular reduction (FHE), error correction | Ordinary pJ/op, as for any digital block |
| Memory and data movement | Same HBM and SRAM costs as everyone else | pJ/byte, as in this simulator |
The energy win of optics is realised only when each conversion is amortised over many optical operations (high reuse per sample) and the core is kept busy enough to amortise static power. Both are workload- and mapping-dependent, so they can only be judged by simulating real workloads through the full toolchain. For a worked example, an optical transform engine in a serving simulator's prefill pool, with conversions, averaging passes, mask rewrites and static power charged and the break-even points found: Fourier Optics for Inference 03.
--prefill-power-cap, --decode-power-cap).| Metric | Question | In the simulator |
|---|---|---|
| Average power (W) | What does the deployment draw? | energy.avg_power_W |
| Peak step power (W) | Will it trip the board or rack limit? | per_instance.peak_step_w |
| Energy per output token (J) | What does a token cost? | J_per_output_token |
| Tokens per joule | Efficiency headline (equivalently, tokens/s per W) | output_tokens_per_J |
| Energy per SLO-meeting request | What does useful work cost? | J_per_slo_met_request |
| Energy breakdown | Static, compute, memory or link: where to optimise | energy.breakdown |
| Power-bound time fraction | Is the cap binding? | per_instance.power_bound_frac |
| Energy–delay product | A single figure balancing speed and energy | E × latency, from the above |
Treat these like latency metrics: report them per configuration, with confidence intervals, under common random numbers, and track the energy-model correlation error alongside the timing one.
Deck 08 turns to the simulator's own speed: how to make it faster without changing a single answer.