LLM Inference Simulators — Presentation 05

Disaggregated Inference, Simulated

Why splitting prefill and decode onto separate hardware raises goodput (Splitwise, DistServe, Mooncake, NVIDIA Dynamo), what the KV-cache transfer costs, and a live discrete-event simulator you can drive in the browser — a JavaScript port of the SimPy model in Disaggregated_Inference_Sim.

Splitwise DistServe Mooncake Dynamo / NIXL KV transfer Live simulator
Arrivals → Prefill pool → KV link → Decode pool → Metrics
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.

01

The Interference Problem

On a colocated server, prefill and decode share each GPU. A vLLM-v0-style scheduler runs a waiting prefill before the next decode step, so every sequence mid-generation stalls for the whole prefill.

One colocated GPU, Llama-3-70B on 4×H100 prefill 2k-token prompt ~134 ms another prefill green = decode step (~15 ms, every running sequence gets a token) orange = prefill: every running sequence waits; its inter-token latency jumps ~10x Mean TPOT moves a little. p99 inter-token latency explodes. Users see the stutter.

Chunked prefill (Sarathi-Serve) bounds the stall by slicing prompts into pieces. Disaggregation removes it by putting the two phases on different machines.

02

The Idea: Split the Phases

Two papers from 2024 made the case. Splitwise (Patel et al., Microsoft, ISCA 2024) characterised production traces and argued for separate prompt and token machine pools, even different GPU generations for each. DistServe (Zhong et al., OSDI 2024) framed the goal as goodput, the request rate served within both TTFT and TPOT SLOs, and searched placements by simulation.

What you gain

  • No interference: decode steps never wait for prefills.
  • Independent scaling: buy prefill capacity for TTFT, decode capacity for TPOT.
  • Independent parallelism: prefill might want more TP for latency; decode wants large batches.
  • Heterogeneous hardware: compute-rich parts for prefill, bandwidth-rich or cheaper parts for decode.

What you pay

  • KV transfer: every prompt's KV cache crosses the network.
  • A new hot-spot: the link and its queue.
  • Provisioning risk: a fixed prefill:decode ratio is wrong when the workload mix shifts.
  • Complexity: two schedulers, a router, a transfer engine.
Why a simulator is essential here

The right prefill:decode ratio, link speed and batch limits depend on the workload's prompt and output length distribution and the SLOs. There are too many combinations to try on real clusters. That is exactly the job DistServe gives its simulator, and the job of the one in this deck.

03

Disaggregation in Production

SystemFromNotable ideas
MooncakeMoonshot AI (Kimi); FAST 2025KV-cache-centric architecture: a disaggregated KV store spanning CPU memory and SSD across the cluster, prefix reuse, and SLO-aware early rejection under overload
NVIDIA DynamoNVIDIA, open-sourced 2025Disaggregated serving across engines (TensorRT-LLM, vLLM, SGLang), a KV-aware router, and the NIXL library for low-latency KV movement over NVLink, RDMA and storage
vLLMOpen sourceDisaggregated prefill through pluggable KV connectors (for example LMCache and NIXL-based ones)
SGLangOpen sourcePrefill–decode disaggregation with RDMA KV transfer, used for large MoE deployments
llm-dRed Hat, Google, IBM and others; 2025Kubernetes-native distributed inference built on vLLM, with disaggregation and cache-aware routing

These systems move quickly; treat the table as a map for further reading, not a feature matrix. The common thread is that the KV cache has become a first-class object that is stored, moved, routed on and reused, and that is what makes data movement the core modelling problem.

04

What the KV Transfer Costs

Bytes to move = prompt tokens × KV bytes per token. For Llama-3-70B in BF16 that is 320 KiB per token, so a 2,048-token prompt carries 671 MB.

Link (one direction, per channel)Bandwidth2k-token 70B KVvs a 15 ms decode step
NVLink 4450 GB/s1.5 msnegligible
PCIe Gen5 x1664 GB/s10.5 msunder one step
InfiniBand NDR 400G50 GB/s13.4 msabout one step
100 GbE12.5 GB/s54 msabout four steps
25 GbE3.1 GB/s215 msa link that saturates at ~4.6 prompts/s
05

The Simulator: Model and Assumptions

Poissonarrivals prefill-0batch ≤ 8192 tok prefill-1TTFT stamped KV linkResource(channels) decode-0continuous batching decode-1KV capacity admit metrics +hot-spots routers: least queued prompt tokens (prefill) / fewest sequences (decode)

Modelled

  • Roofline step time for each prefill batch and decode iteration (deck 03's formulas).
  • Prefill batching up to a token budget; continuous decode batching up to 256 sequences.
  • KV capacity per decode instance; full sequence reserved at admission.
  • A shared KV fabric with N FCFS channels and per-transfer latency.
  • Colocated mode with prefill priority, for comparison on the same GPU count.
  • Power and energy: static power, pJ per FLOP, per HBM byte and per link bit; DVFS and per-pool power caps (deck 07).

Simplified (deliberately)

  • A TP group is one big device; no all-reduce cost.
  • KV is transferred after the whole prefill, not layer by layer.
  • No pre-emption, swapping or prefix caching.
  • No chunked prefill in colocated mode.
  • Efficiency factors are constants, not shape-dependent.
06

Interactive: Run the Simulator

This is a line-for-line JavaScript port of the SimPy model. On identical request streams it reproduces the Python simulator's per-request timestamps exactly; the repo's test suite checks this. Power cap, decode-pool cap and DVFS controls drive the power model from deck 07. Set a configuration and press Run, Compare (colocated on the same number of instances) or Sweep (goodput against load, both modes).

4
1
1
1
4
800
2048
256
0.5
1000
25
off
off
1
where the average request's time goes
latency CDFs
queues over time
07

Six Experiments to Try

  1. Find the interference. Defaults (70B, 4×H100 per instance, 1 prefill + 1 decode, 4 req/s), then Compare. Colocated runs two instances that each do both jobs; its dashed inter-token curve grows a long tail of ~130–200 ms stalls while disaggregated stays flat near 15 ms.
  2. Push the load. Sweep from the defaults. Colocated loses goodput first, because the TPOT SLO fails; disaggregated holds until the single prefill instance saturates, and then TTFT fails instead. Add a second prefill instance (2P1D) and sweep again.
  3. Starve the link. Set 2 prefill instances, 6 req/s and 25 GbE. The hot-spot moves to kv_wait → kv-link and most of each request's life is spent queueing for the network. Add channels or a faster link and watch it move back.
  4. Make compute nearly free. Switch to the hypothetical optical MAC. Prefill (compute-bound) gets several times faster, TTFT drops, and decode barely changes, because it is bandwidth-bound. Fast compute shifts the bottleneck to memory and data movement.
  5. Spend fewer watts. Turn DVFS on: energy per token drops and latency does not move. Then lower the decode-pool cap: down to about 250 W per device TPOT barely changes; near 200 W it collapses. Now cap everything at 400 W: TTFT suffers because prefill is compute-bound and runs near TDP. Add a second prefill instance and the latency comes back at lower power per device.
  6. Change the workload. 8k-token prompts with 64-token outputs (summarisation) against 256-token prompts with 1k-token outputs (generation). The best prefill:decode ratio flips; no fixed ratio suits both, which is the provisioning risk of disaggregation.
08

Results from the Python CLI

The same experiments from the command line (disagg-sim --compare --sweep-rate 2 3 4 5 6 8): Llama-3-70B, 4×H100 per instance, 2,048-token mean prompts, 256-token mean outputs, SLOs TTFT ≤ 1 s and TPOT ≤ 25 ms, 800 requests.

Offered load (req/s)Colocated TPOT p99Colocated SLO metDisagg 1P1D TPOT p99Disagg 1P1D TTFT p99Disagg 1P1D SLO met
219.3 ms100%14.5 ms640 ms100%
425.9 ms97.8%15.2 ms854 ms99.7%
529.1 ms89.7%15.5 ms1,025 ms98.5%
634.8 ms64.7%15.8 ms1,339 ms89.9%
846.8 ms14.0%16.2 ms6,085 ms21.0%

The same runs report energy: colocated uses 3.09 J per output token at 2,993 W; disaggregated 2.77 J at 2,687 W, because decode steps never idle behind prefills. Deck 07 builds on this.

With two prefill instances (2P1D, three instances in total) the disaggregated system meets both SLOs for 100% of requests up to 10 req/s with TPOT p99 of 17 ms. At 4 req/s the decode instance runs at a mean batch of about 14 with an MBU of about 77%: memory-bound, as deck 03 predicted.

Cost model corrected, 2026-10-03

Decode steps had been charged the whole input-embedding table and had left out each new token's attention to itself (found by tracing the real model, Toolkit deck 10). This table and the live simulator above use the corrected model. Decode steps got 1–6% shorter: colocated TPOT p99 at 4 req/s fell from 27.6 to 25.9 ms, and no conclusion on this slide changed. Every figure is in examples/results.md, with before and after.

Reading the result like an architect

Disaggregation does not make anything faster; it makes the two bottlenecks separable. TPOT becomes a function of the decode pool alone and TTFT of the prefill pool alone, so each can be provisioned against its own SLO. The simulator turns that qualitative claim into the numbers needed to size a deployment.

09

Reading the Code

The Python package Disaggregated_Inference_Sim is about 2,400 lines of Python (it was 1,150 before the 2026-10-04 and 2026-10-05 extensions), written to be read top to bottom.

ModuleRole (deck 02's layers)
hardware.pyArchitecture config and cost model: ModelSpec, Accelerator, Link, CostModel (roofline)
workload.pyWorkload: Request with all its timestamps, lognormal lengths, Poisson arrivals, JSON replay
sim.pyEngine and behaviour: Prefill, Decode and Colocated instances as SimPy processes; KV link as a Resource; routers
metrics.pyMetrics: percentiles, goodput, utilisation, MFU/MBU, stage breakdown, hot-spot attribution, Little's law
trace.pyProbes: Chrome trace-event export for Perfetto
sim.py: FastDecodeInstanceThe exact accelerated decode path (incremental state, lazy bookkeeping, macro-steps); deck 08
search.pyAnalytic capacity bounds, parallel sweeps, bisection for the maximum sustainable load
web/sim_engine.jsThe JavaScript port that runs on this page
tests/154 tests: cost-model units (pinned against an operator trace of the real model since the 2026-10-03 correction), invariants, the M/D/1 analytic check, behaviour, property-based (Hypothesis), fast path against baseline, and this page's JavaScript against Python
Quick start
pip install -e .[dev] && pytest
disagg-sim --compare                                  # colocated vs disaggregated
disagg-sim --prefill 2 --sweep-rate 4 6 8 10          # where does goodput collapse?
disagg-sim --prefill 2 --rate 6 --link eth-25g        # watch the hot-spot move to the link
disagg-sim --trace run.json                           # open in https://ui.perfetto.dev
10

Limitations and Extensions

Each of these is a self-contained exercise, roughly in order of difficulty, and each is a real feature of production serving simulators.

  1. Layer-wise KV streaming: overlap the transfer with prefill by sending one layer's KV at a time.
  2. Tensor-parallel communication: add the all-reduce term from deck 03 and a scale-up link per instance. Now built: --tp, --pp and --ep, measured in Inference Trade-offs Explained, chapter 8.
  3. Chunked prefill in colocated mode: a per-step token budget mixing prefill chunks with decodes. Is it as good as disaggregation? Now built: --batch-policy chunked; the answer depends on the workload (Inference Trade-offs Explained, chapter 1).
  4. Heterogeneous pools: H100 for prefill, A100 or an optical part for decode (Splitwise's proposal). Now built: --prefill-device and --decode-device, with an optical transform engine in the prefill pool, in Fourier Optics for Inference 03 (optical prefill pools, simulated).
  5. Asymmetric models: a model whose prefill runs only part of the network (a causal encoder-decoder). Now built: --ced-encoder-layers, --ced-replay and --ced-replay-on, on the next slide.
  6. Pre-emption and paged KV: reserve KV incrementally rather than up front; evict or swap when full. Now built: --kv-policy paged and --preemption recompute|swap, in both pools (chapter 2).
  7. Prefix caching and a KV store (Mooncake): cache hits skip prefill and transfer. Prefix caching is now built (--prefix-caching, chapter 3); a KV store shared between pools is not.
  8. Calibrated cost model: replace the roofline with a lookup table of measured or RTL-derived step times.
  9. Per-request energy attribution: split each step's joules across the sequences in its batch, so energy can be reported per request class (the pool-level power model is already in place, deck 07).
11

Asymmetric Models: Encoder-Only Prefill

Everything so far assumed prefill and decode run the same model. A 2026 design breaks that. DeepSeek-V4.1-Flash’s Causal Encoder-Decoder (DeepSeek-AI, arXiv:2609.19969, section 2.2) makes the bottom half of its 40 layers a causal encoder and projects every decoder layer’s KV from the last encoder state, so a prompt token activates 8B parameters and a generated token 16B. The architecture is explained in Modern Architectures 06; this slide is about what it does to the pools.

Replay on prefill (the paper)

  • The prefill instance runs the encoder over the prompt, the decoder’s KV projections, and the last 128 prompt tokens through the decoder (to rebuild its sliding-window state).
  • Both pools hold the whole model; the first token leaves the prefill pool as before.

Replay on decode (an SGLang RFC)

  • sgl-project/sglang #39963 (open, September 2026): prefill workers hold only the encoder and the first decoder layer’s KV projection, about half the weights.
  • The decode worker runs the final 128-token window through all layers as its first step, mixed into its decode batch, and emits the first token. TTFT now includes the hand-off.

Both placements are in the simulator (--model llama3-70b-ced, --ced-replay-on prefill|decode; Python and the JavaScript above agree exactly). On a dense Llama-3-70B shape split 40 + 40 (illustrative), one 8,192-token prefill takes 288.3 ms instead of 564.3 ms. Highest request rate with 90% of requests inside both SLOs, best split of six 4×H100 instances:

Prompt : output (mean)Decoder-onlyCED, replay on prefillCED, replay on decode
2,048 : 51225.72 req/s (4P2D)39.20 (3P3D), 1.52×41.07 (3P3D), 1.60×
4,096 : 25613.10 (4P2D)27.01 (4P2D), 2.06×27.90 (4P2D), 2.13×
8,192 : 1287.89 (5P1D)13.23 (4P2D), 1.68×13.23 (4P2D), 1.68×
16,384 : 643.81 (5P1D)6.89 (5P1D), 1.81×8.17 (5P1D), 2.14×

Every figure is in sections 16–18 of examples/results.md. The prompt:output ratios stand in for agent loops and are illustrative; the proxy has no MoE, sparse attention or prefix-cache hits. Animated, with the pools at work: Inference Trade-offs Explained, chapter 7.

12

What to Take Away

Next

Deck 06 looks at the metrics, hot-spot analysis and validation behind these numbers, and how to make them hold up in review.