Prefill versus decode, the roofline, KV-cache arithmetic, batching and parallelism — the first-order physics every LLM inference simulator must get right, with the formulas to put in a cost model.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.
Every LLM request runs in two phases with opposite hardware characters. Getting this distinction into the simulator is the single most important modelling step.
A machine tuned for one phase is mistuned for the other. Running both on the same device makes them interfere: a long prefill stalls every decode in the batch. That tension motivates disaggregation (deck 05).
For a dense decoder with P matmul parameters, L layers and model width d:
linear layers 2 · P # one multiply + one add per weight
attention at pos c 4 · L · d · c # QKᵀ and AV, each 2·d·c per layer
prefill(s) = 2·P·s + 2·L·d·s(s+1) # causal: sum of c over 1..s
decode(B,c) = 2·P·B + 4·L·d·(Σcᵢ + B) # B sequences, cᵢ cached tokens each, + itself
| Model | Total params | Linear FLOPs / token | Attention FLOPs / token at 8k context |
|---|---|---|---|
| Llama-3-8B (32 L, d=4096) | 8.03 B | 15.0 GFLOP | 4.3 GFLOP |
| Llama-3-70B (80 L, d=8192) | 70.6 B | 139 GFLOP | 21.5 GFLOP |
The parameter counts come from the shapes alone (GQA attention, SwiGLU MLP, untied embeddings) and match the published model sizes; the simulator's unit tests check this. Note that attention FLOPs grow with context and become significant at long context even though the linear layers dominate at short context.
Time is set by whichever is slower: doing the arithmetic or moving the data. Bytes moved from HBM per forward pass:
weights = (params − V·d) × bytes_per_weight # layers + LM head, once per step
emb(n) = n · d · bytes_per_weight # embedding lookup: rows used only
kv_per_token = 2 · L · n_kv_heads · head_dim · bytes # K and V
prefill(s) = weights + emb(s) + s · kv_per_token # write the new KV
decode(B,c) = weights + emb(B) + (Σcᵢ + B) · kv_per_token # read all KV, write one more token each
For Llama-3-70B in BF16 the weights occupy 141 GB. A decode step must stream 139 GB of them from memory however small the batch: everything except the 2.1 GB input-embedding table, of which it reads only the rows its tokens look up. On four H100s (4 × 3.35 TB/s at 80% achievable) that floor is about 13 ms per step before a single KV byte is read. Batching amortises this one read across many sequences, which is why serving systems batch hard.
Corrected on 2026-10-03. Until then the simulator charged every step all 141 GB, embedding table included, and left out each decoded token's attention to itself (slide 02). Tracing the real model found both (Toolkit deck 10). The floor moved from 13.16 to 12.97 ms.
FP8 or INT4 weights halve or quarter the dominant term in decode bytes. They do little for prefill, which is compute-bound, unless the arithmetic also runs at lower precision.
Fetching a byte from HBM costs tens of times the energy of a FLOP on it. Decode is therefore an energy problem as well as a bandwidth problem, and the same byte counts drive the simulator's power model (deck 07).
The roofline model (Williams, Waterman & Patterson, 2009) plots attainable performance against arithmetic intensity I = FLOPs / bytes. Below the ridge point the kernel is memory-bound; above it, compute-bound.
t = max( FLOPs / (peak_flops · MFU_eff), bytes / (mem_bw · MBU_eff) ) + overhead
ridge = (peak_flops · MFU_eff) / (mem_bw · MBU_eff) # H100: 295 FLOP/B at peak, ~203 derated
| Llama-3-70B on 4×H100 | Intensity (FLOP/B) | Bound | Step time |
|---|---|---|---|
| Decode, batch 1, 2k context | 1.0 | memory | 13.7 ms |
| Decode, batch 32 | 28 | memory | 15.7 ms |
| Decode, batch 256 | 118 | memory | 29.7 ms |
| Prefill, 128-token prompt | 126 | memory | 13.7 ms |
| Prefill, 2,048-token prompt | 2,047 | compute | 134 ms |
Two lessons. In BF16 decode, intensity is roughly the batch size, so you need hundreds of concurrent sequences to approach the ridge, and the KV reads pull it further down. Prefill crosses the ridge after a few hundred tokens. These numbers come from the same cost model used in the simulator.
Choose a model and device, then move the decode batch, the context length and the prompt length. Both phases are plotted on a log–log roofline.
Try the hypothetical optical MAC: four times the matmul throughput, same memory. Prefill speeds up; decode hardly moves, because it is bandwidth-bound. Any novel compute fabric has to answer the memory question as well.
The KV cache holds K and V for every layer and every token of every live sequence. With grouped-query attention (GQA) only the n_kv_heads are stored.
| Llama-3-8B | Llama-3-70B | |
|---|---|---|
| KV per token (BF16) | 2·32·8·128·2 = 128 KiB | 2·80·8·128·2 = 320 KiB |
| One 2,048-token prompt | 268 MB | 671 MB |
| Capacity at 90% HBM, after weights | ~430k tokens on 1×H100 | ~448k tokens on 4×H100 |
| Technique | Idea | What the simulator must model |
|---|---|---|
| Static batching | Form a batch, run it until every member finishes | Batch formation; the slowest request holds the batch |
| Continuous batching (Orca, OSDI 2022) | Re-form the batch at every iteration; finished sequences leave, new ones join | An iteration-level scheduler loop: the core of every serving simulator |
| Paged KV (vLLM PagedAttention, SOSP 2023) | KV in fixed-size blocks with a block table, allocated on demand | KV capacity in blocks; admission, pre-emption and swapping when full |
| Chunked prefill (Sarathi-Serve, OSDI 2024) | Split long prefills into chunks and piggy-back decodes on each step | A per-step token budget mixing prefill chunks and decode tokens |
| Prefix caching | Reuse KV for shared prompt prefixes | Hit rates from the workload; cache eviction |
| Speculative decoding | A draft model proposes k tokens; the target verifies them in one pass | Acceptance-rate distribution; extra compute per step |
Chunked prefill and disaggregation are two answers to the same interference problem. The first bounds the stall; the second removes it by putting the phases on different hardware.
Large models span many devices. Each form of parallelism adds communication that the simulator must cost on the interconnect.
| Form | Split | Communication | Typical fabric |
|---|---|---|---|
| Tensor (TP) | Each weight matrix across devices | Two all-reduces of the activations per layer | NVLink / scale-up |
| Pipeline (PP) | Layers across devices | Point-to-point activations between stages; bubbles | Any |
| Expert (EP) | MoE experts across devices | All-to-all dispatch and combine per MoE layer | Scale-up or fast scale-out |
| Data / replica | Whole model copies | None at inference, apart from routing | — |
| Disaggregated | Prefill and decode on separate pools | KV-cache transfer per request | RDMA / NVLink |
t_allreduce ≈ 2 · (n − 1) / n · bytes / link_bw + 2 · (n − 1) · latency
The companion simulator treats a TP group as one larger device and ignores the all-reduce, a simplification its README states plainly. Adding it is a good first extension, and ASTRA-sim (deck 04) is the tool for when collectives are the question.
| Metric | Definition | Set mainly by |
|---|---|---|
| TTFT | Arrival to first output token | Prefill queueing and prefill compute |
| TPOT | (finish − first token) / (output tokens − 1) | Decode step time, batch size, interference |
| ITL | Gap between consecutive tokens of one request | As TPOT, but shows the stalls that the mean hides |
| E2E latency | Arrival to last token | Everything |
| Throughput | Output tokens/s, or requests/s | Batch size and utilisation |
| Goodput | Requests/s that meet both the TTFT and TPOT SLOs | The balance of all the above: the metric DistServe optimises |
| MFU / MBU | Achieved FLOP/s or bytes/s as a fraction of peak | How well the hardware is used |
| Energy per token | Joules per output token (or tokens per joule) | Static power, FLOP and byte energy, utilisation (deck 07) |
Report p50, p90 and p99, never the mean alone. Interference from colocated prefill barely moves the mean ITL but can multiply the p99 by ten, as the simulator in deck 05 shows.
The roofline gets the first-order answer. Before trusting a simulator's absolute numbers, check what it does about each of these:
The professional fix is calibration: profile real kernels (or RTL), fit per-operator tables, and let the simulator interpolate. This is Vidur's approach, and the one deck 06 builds into a validation flow.
Deck 04 surveys the existing LLM simulators: what each models, at which level, and when to reach for which.