LLM Inference Simulators — Presentation 07

Power & Energy in Inference Simulators

Speed is half of performance. Static and dynamic power, why data movement dominates energy, the simulator's calibrated power model, prefill running hot and decode cool, DVFS and power caps as a third roof, energy proportionality, the link to RTL power and thermal analysis, and what photonic systems must pay for — all measured on the companion simulator.

Joules per token DVFS Power caps Static power RTL power Photonics
Static → Dynamic → Data movement → DVFS → Caps → J/token
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.

01

Why Power Is a First-Class Simulator Output

Speed is only half of performance. For AI hardware the other half is power and energy, and it is often the half that decides the product.

Why it matters

  • Datacentres are power-limited. New AI capacity is constrained by available megawatts and cooling at least as much as by chips.
  • Energy is running cost. Over a deployment's life, electricity and cooling are a large share of total cost of ownership.
  • Power sets the package. Board limits (TDP), power delivery, current density and cooling are fixed early and are expensive to change.
  • Efficiency is the pitch. Novel compute, photonics included, usually competes on performance per watt rather than raw speed.

Questions the simulator must answer

  • What are the average and peak power for this workload?
  • How many joules per output token, and where do they go?
  • What happens to latency under a power cap?
  • Which phase can be throttled for free?
  • How does energy per token change with load?
  • Is the design limited by compute, memory, or power?
02

Where the Joules Go: Static, Dynamic, and Data Movement

CMOS power, first order
P = P_static + P_dynamic
P_static  ≈ V · I_leak                 # leakage, clocks, always-on logic: paid whenever powered
P_dynamic ≈ α · C · V² · f               # switching: activity x capacitance x voltage^2 x frequency
E_op      ≈ α · C · V²                   # energy per operation: why lowering V (with f) pays quadratically

Horowitz's widely cited energy table (ISSCC 2014, 45 nm) makes the key point: moving data costs far more than computing on it.

Operation (45 nm)EnergyRelative to an 8-bit add
8-bit integer add0.03 pJ1×
32-bit float multiply3.7 pJ~120×
64-bit read from an 8 KB SRAM10 pJ~330×
64-bit read from DRAM1.3–2.6 nJ~40,000–90,000×

Absolute values have fallen at newer nodes, but the ratios still hold. This is why the simulator charges separately for FLOPs and for HBM bytes, and why decode, which streams almost every weight on every step (all but the input-embedding table, which is a lookup), is an energy problem as well as a bandwidth problem.

03

The Simulator's Power Model

Four coefficients per device and one per link, attached to the same FLOP and byte counts the roofline already computes. Nothing new is counted; it is only priced.

Energy of a run
E_total = Σinstances idle_W · T_run                    # static: charged busy or not
        + Σsteps (FLOPs · pJ/FLOP · s² + HBM bytes · pJ/byte)   # dynamic, with DVFS factor s
        + Σtransfers KV bits · pJ/bit                     # interconnect
P_step  = idle_W + E_step / t_step                       # tracked for the peak
Device (illustrative)TDPIdlepJ / FLOPpJ / HBM byte
H100-SXM700 W100 W1.060
A100-SXM400 W60 W1.670
Hypothetical optical MAC700 W180 W0.160
Links (pJ/bit)NVLink 5 · PCIe 6 · InfiniBand / Ethernet 15 (SerDes, NIC and switch together)
Calibrate, don't trust

These coefficients are chosen to give realistic totals (prefill near TDP, decode around 300 W per H100). To calibrate for real hardware, run micro-benchmarks that sweep FLOP rate and byte rate independently, read power with DCGM or nvidia-smi, and fit P − idle = a·FLOP/s + b·byte/s by least squares. a and b are your pJ/FLOP and pJ/byte. For pre-silicon hardware the same coefficients come from RTL power analysis of the blocks (slide 09).

04

Prefill Runs Hot, Decode Runs Cool

Per step, Llama-3-70B on 4×H100, from the simulator's cost model (as corrected on 2026-10-03; see slide 07):

StepBoundTimePower per deviceDynamic energy
Prefill, 2,048-token promptcompute133.9 ms658 W (near the 700 W TDP)299 J
Decode, batch 16, ~2.3k contextmemory14.6 ms295 W11.4 J
Decode, same, with DVFSmemory14.6 ms262 W9.4 J
05

DVFS and Power Caps: The Third Roof

Dynamic voltage and frequency scaling lowers the clock (and with it the voltage). In the model the compute clock runs at a fraction s of nominal; HBM is unaffected.

hardware.py, CostModel.step_time
t(s) = max(tc / s, tm)                # compute slows; memory doesn't
E(s) = Ec · s² + Em                    # V tracks f, so energy per op ~ s^2
P(s) = idle + E(s) / t(s)             # increases with s: lower clocks always reduce power

dvfs on:     memory-bound step  →  s = max(s_min, tc/tm)        # same time, less energy
power cap:   choose the largest s with P(s) ≤ cap                   # closed form: sqrt, or a cubic (Cardano)
             still too hot at s_min → stretch the step to fit the budget

Compute roof

Not enough FLOP/s. Prefill.

Memory roof

Not enough bytes/s. Decode.

Power roof

Not enough watts. A capped prefill: the work could go faster, but not without exceeding the cap.

The simulator labels each step with whichever roof bound it and reports the fraction of busy time spent power-bound. Real DVFS has discrete operating points, voltage floors and transition latency; the continuous model is a first-order stand-in to refine when the hardware's operating points are known.

06

Interactive: One Step Under a Power Cap

This uses the simulator's own cost model (the browser engine from deck 05). Pick a step, then move the per-device cap and toggle DVFS. The chart sweeps the cap from TDP down to just above idle and plots time against energy: the energy–delay trade-off.

2048
700
Step time
—
Power / device
—
Dynamic energy
—
Bound
—

Things to notice: for prefill, tighter caps trade time for energy until static power (charged for the longer step) starts to dominate and total energy rises again. That minimum is the energy-optimal operating point: for a 2,048-token prefill on H100s it sits near a 160 W cap, at 186 J against 352 J uncapped, but at 2.2× the step time. For decode, the curve is flat in time over a wide range of caps: free savings.

07

System Results: Latency Against Energy

Full runs: Llama-3-70B, 4×H100 per instance, 4 req/s, 2,048/256-token mean prompts/outputs, SLOs TTFT ≤ 1 s and TPOT ≤ 25 ms, 800 requests. Every row is disagg-sim with the flags shown.

ConfigurationTTFT p99TPOT p99SLO metAvg powerJ / token
Colocated, 2 instances630 ms25.9 ms97.8%2,993 W3.09
Disaggregated 1P1D854 ms15.2 ms99.7%2,687 W2.77
… + --dvfs854 ms15.2 ms99.7%2,570 W2.65
… + --decode-power-cap 250854 ms17.2 ms99.7%2,503 W2.59
… + --decode-power-cap 200854 ms27.2 ms12.6%2,284 W2.43
1P1D, --power-cap 400 --dvfs1,240 ms15.2 ms95.3%2,190 W2.26
1P1D, --power-cap 300 --dvfs1,651 ms15.2 ms83.3%2,013 W2.08
2P1D, --power-cap 400 --dvfs689 ms15.1 ms100%2,593 W2.68
Cost model corrected, 2026-10-03

Every step had been charged the whole input-embedding table, where a lookup reads only the rows it uses. And decode attention had left out each new token's attention to itself. Tracing the real model found both (Toolkit deck 10). All the tables in this deck were rerun with the corrected model. Decode TPOT fell slightly. The J/token ranking, the power-bound flags and the hot-spots did not change. The 200 W cliff remains, though 12.6% of requests now meet the SLOs, against 1.7% before. Before and after: examples/results.md.

08

Energy Proportionality: Static Power and Load

An idle accelerator still burns its static power. At low load that is spread over few tokens.

Offered load (1P1D)Avg powerJ / output tokenStatic share of energy
0.5 req/s1,472 W11.854%
1 req/s1,729 W7.046%
2 req/s2,066 W4.239%
4 req/s2,687 W2.830%
6 req/s3,254 W2.325%
09

From the Simulator to RTL Power and Thermal Analysis

Power modelling follows the same fidelity ladder as timing (deck 01), and the levels feed each other the same way.

LevelPower methodTools
Analytical / DESEnergy per event × event counts (this simulator)Custom models; Accelergy for per-access energy tables
ArchitectureComponent power from structure: caches, NoCs, coresMcPAT, CACTI (SRAM/cache energy), Accelergy + Timeloop
RTLSwitching activity from simulation (SAIF, VCD, FSDB) applied to the netlist or RTLCommercial RTL and gate-level power analysis tools
EmulationActivity over real software runs, millions of cyclesEmulator-based dynamic power analysis
SiliconMeasurementBoard telemetry, DCGM, power analysers
10

Power in Photonic and Novel Compute

An optical multiply can cost very little energy. A photonic system pays for everything around it, and a simulator must account for all of it.

ConsumerBehaviourHow to model it
LasersContinuous optical power, limited wall-plug efficiency; on whether or not work arrivesStatic power, with gating policies and their wake-up latency
Thermal tuningHeaters hold resonant devices on wavelength; depends on ambient and neighbouring heatStatic plus thermal-state dependent
DACs and modulatorsEnergy per conversion grows with sample rate and resolutionPer-sample energy × samples; a Walden-style figure of merit (energy ∝ 2ENOB)
Photodetectors, TIAs and ADCsAs above, on the way back to electronicsPer-sample energy; precision against power trade-off
Digital pre- and post-processingNormalisation, accumulation, modular reduction (FHE), error correctionOrdinary pJ/op, as for any digital block
Memory and data movementSame HBM and SRAM costs as everyone elsepJ/byte, as in this simulator
The architectural consequence

The energy win of optics is realised only when each conversion is amortised over many optical operations (high reuse per sample) and the core is kept busy enough to amortise static power. Both are workload- and mapping-dependent, so they can only be judged by simulating real workloads through the full toolchain. For a worked example, an optical transform engine in a serving simulator's prefill pool, with conversions, averaging passes, mask rewrites and static power charged and the break-even points found: Fourier Optics for Inference 03.

11

The Datacentre View

12

Power Metrics Checklist

MetricQuestionIn the simulator
Average power (W)What does the deployment draw?energy.avg_power_W
Peak step power (W)Will it trip the board or rack limit?per_instance.peak_step_w
Energy per output token (J)What does a token cost?J_per_output_token
Tokens per jouleEfficiency headline (equivalently, tokens/s per W)output_tokens_per_J
Energy per SLO-meeting requestWhat does useful work cost?J_per_slo_met_request
Energy breakdownStatic, compute, memory or link: where to optimiseenergy.breakdown
Power-bound time fractionIs the cap binding?per_instance.power_bound_frac
Energy–delay productA single figure balancing speed and energyE × latency, from the above

Treat these like latency metrics: report them per configuration, with confidence intervals, under common random numbers, and track the energy-model correlation error alongside the timing one.

13

What to Take Away

Next

Deck 08 turns to the simulator's own speed: how to make it faster without changing a single answer.