Why splitting prefill and decode onto separate hardware raises goodput (Splitwise, DistServe, Mooncake, NVIDIA Dynamo), what the KV-cache transfer costs, and a live discrete-event simulator you can drive in the browser — a JavaScript port of the SimPy model in Disaggregated_Inference_Sim.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.
On a colocated server, prefill and decode share each GPU. A vLLM-v0-style scheduler runs a waiting prefill before the next decode step, so every sequence mid-generation stalls for the whole prefill.
Chunked prefill (Sarathi-Serve) bounds the stall by slicing prompts into pieces. Disaggregation removes it by putting the two phases on different machines.
Two papers from 2024 made the case. Splitwise (Patel et al., Microsoft, ISCA 2024) characterised production traces and argued for separate prompt and token machine pools, even different GPU generations for each. DistServe (Zhong et al., OSDI 2024) framed the goal as goodput, the request rate served within both TTFT and TPOT SLOs, and searched placements by simulation.
The right prefill:decode ratio, link speed and batch limits depend on the workload's prompt and output length distribution and the SLOs. There are too many combinations to try on real clusters. That is exactly the job DistServe gives its simulator, and the job of the one in this deck.
| System | From | Notable ideas |
|---|---|---|
| Mooncake | Moonshot AI (Kimi); FAST 2025 | KV-cache-centric architecture: a disaggregated KV store spanning CPU memory and SSD across the cluster, prefix reuse, and SLO-aware early rejection under overload |
| NVIDIA Dynamo | NVIDIA, open-sourced 2025 | Disaggregated serving across engines (TensorRT-LLM, vLLM, SGLang), a KV-aware router, and the NIXL library for low-latency KV movement over NVLink, RDMA and storage |
| vLLM | Open source | Disaggregated prefill through pluggable KV connectors (for example LMCache and NIXL-based ones) |
| SGLang | Open source | Prefill–decode disaggregation with RDMA KV transfer, used for large MoE deployments |
| llm-d | Red Hat, Google, IBM and others; 2025 | Kubernetes-native distributed inference built on vLLM, with disaggregation and cache-aware routing |
These systems move quickly; treat the table as a map for further reading, not a feature matrix. The common thread is that the KV cache has become a first-class object that is stored, moved, routed on and reused, and that is what makes data movement the core modelling problem.
Bytes to move = prompt tokens × KV bytes per token. For Llama-3-70B in BF16 that is 320 KiB per token, so a 2,048-token prompt carries 671 MB.
| Link (one direction, per channel) | Bandwidth | 2k-token 70B KV | vs a 15 ms decode step |
|---|---|---|---|
| NVLink 4 | 450 GB/s | 1.5 ms | negligible |
| PCIe Gen5 x16 | 64 GB/s | 10.5 ms | under one step |
| InfiniBand NDR 400G | 50 GB/s | 13.4 ms | about one step |
| 100 GbE | 12.5 GB/s | 54 ms | about four steps |
| 25 GbE | 3.1 GB/s | 215 ms | a link that saturates at ~4.6 prompts/s |
This is a line-for-line JavaScript port of the SimPy model. On identical request streams it reproduces the Python simulator's per-request timestamps exactly; the repo's test suite checks this. Power cap, decode-pool cap and DVFS controls drive the power model from deck 07. Set a configuration and press Run, Compare (colocated on the same number of instances) or Sweep (goodput against load, both modes).
kv_wait → kv-link and most of each request's life is spent queueing for the network. Add channels or a faster link and watch it move back.The same experiments from the command line (disagg-sim --compare --sweep-rate 2 3 4 5 6 8): Llama-3-70B, 4×H100 per instance, 2,048-token mean prompts, 256-token mean outputs, SLOs TTFT ≤ 1 s and TPOT ≤ 25 ms, 800 requests.
| Offered load (req/s) | Colocated TPOT p99 | Colocated SLO met | Disagg 1P1D TPOT p99 | Disagg 1P1D TTFT p99 | Disagg 1P1D SLO met |
|---|---|---|---|---|---|
| 2 | 19.3 ms | 100% | 14.5 ms | 640 ms | 100% |
| 4 | 25.9 ms | 97.8% | 15.2 ms | 854 ms | 99.7% |
| 5 | 29.1 ms | 89.7% | 15.5 ms | 1,025 ms | 98.5% |
| 6 | 34.8 ms | 64.7% | 15.8 ms | 1,339 ms | 89.9% |
| 8 | 46.8 ms | 14.0% | 16.2 ms | 6,085 ms | 21.0% |
The same runs report energy: colocated uses 3.09 J per output token at 2,993 W; disaggregated 2.77 J at 2,687 W, because decode steps never idle behind prefills. Deck 07 builds on this.
With two prefill instances (2P1D, three instances in total) the disaggregated system meets both SLOs for 100% of requests up to 10 req/s with TPOT p99 of 17 ms. At 4 req/s the decode instance runs at a mean batch of about 14 with an MBU of about 77%: memory-bound, as deck 03 predicted.
Decode steps had been charged the whole input-embedding table and had left out each new token's attention to itself (found by tracing the real model, Toolkit deck 10). This table and the live simulator above use the corrected model. Decode steps got 1–6% shorter: colocated TPOT p99 at 4 req/s fell from 27.6 to 25.9 ms, and no conclusion on this slide changed. Every figure is in examples/results.md, with before and after.
Disaggregation does not make anything faster; it makes the two bottlenecks separable. TPOT becomes a function of the decode pool alone and TTFT of the prefill pool alone, so each can be provisioned against its own SLO. The simulator turns that qualitative claim into the numbers needed to size a deployment.
The Python package Disaggregated_Inference_Sim is about 2,400 lines of Python (it was 1,150 before the 2026-10-04 and 2026-10-05 extensions), written to be read top to bottom.
| Module | Role (deck 02's layers) |
|---|---|
hardware.py | Architecture config and cost model: ModelSpec, Accelerator, Link, CostModel (roofline) |
workload.py | Workload: Request with all its timestamps, lognormal lengths, Poisson arrivals, JSON replay |
sim.py | Engine and behaviour: Prefill, Decode and Colocated instances as SimPy processes; KV link as a Resource; routers |
metrics.py | Metrics: percentiles, goodput, utilisation, MFU/MBU, stage breakdown, hot-spot attribution, Little's law |
trace.py | Probes: Chrome trace-event export for Perfetto |
sim.py: FastDecodeInstance | The exact accelerated decode path (incremental state, lazy bookkeeping, macro-steps); deck 08 |
search.py | Analytic capacity bounds, parallel sweeps, bisection for the maximum sustainable load |
web/sim_engine.js | The JavaScript port that runs on this page |
tests/ | 154 tests: cost-model units (pinned against an operator trace of the real model since the 2026-10-03 correction), invariants, the M/D/1 analytic check, behaviour, property-based (Hypothesis), fast path against baseline, and this page's JavaScript against Python |
pip install -e .[dev] && pytest
disagg-sim --compare # colocated vs disaggregated
disagg-sim --prefill 2 --sweep-rate 4 6 8 10 # where does goodput collapse?
disagg-sim --prefill 2 --rate 6 --link eth-25g # watch the hot-spot move to the link
disagg-sim --trace run.json # open in https://ui.perfetto.dev
Each of these is a self-contained exercise, roughly in order of difficulty, and each is a real feature of production serving simulators.
--tp, --pp and --ep, measured in Inference Trade-offs Explained, chapter 8.--batch-policy chunked; the answer depends on the workload (Inference Trade-offs Explained, chapter 1).--prefill-device and --decode-device, with an optical transform engine in the prefill pool, in Fourier Optics for Inference 03 (optical prefill pools, simulated).--ced-encoder-layers, --ced-replay and --ced-replay-on, on the next slide.--kv-policy paged and --preemption recompute|swap, in both pools (chapter 2).--prefix-caching, chapter 3); a KV store shared between pools is not.Everything so far assumed prefill and decode run the same model. A 2026 design breaks that. DeepSeek-V4.1-Flash’s Causal Encoder-Decoder (DeepSeek-AI, arXiv:2609.19969, section 2.2) makes the bottom half of its 40 layers a causal encoder and projects every decoder layer’s KV from the last encoder state, so a prompt token activates 8B parameters and a generated token 16B. The architecture is explained in Modern Architectures 06; this slide is about what it does to the pools.
Both placements are in the simulator (--model llama3-70b-ced, --ced-replay-on prefill|decode; Python and the JavaScript above agree exactly). On a dense Llama-3-70B shape split 40 + 40 (illustrative), one 8,192-token prefill takes 288.3 ms instead of 564.3 ms. Highest request rate with 90% of requests inside both SLOs, best split of six 4×H100 instances:
| Prompt : output (mean) | Decoder-only | CED, replay on prefill | CED, replay on decode |
|---|---|---|---|
| 2,048 : 512 | 25.72 req/s (4P2D) | 39.20 (3P3D), 1.52× | 41.07 (3P3D), 1.60× |
| 4,096 : 256 | 13.10 (4P2D) | 27.01 (4P2D), 2.06× | 27.90 (4P2D), 2.13× |
| 8,192 : 128 | 7.89 (5P1D) | 13.23 (4P2D), 1.68× | 13.23 (4P2D), 1.68× |
| 16,384 : 64 | 3.81 (5P1D) | 6.89 (5P1D), 1.81× | 8.17 (5P1D), 2.14× |
Every figure is in sections 16–18 of examples/results.md. The prompt:output ratios stand in for agent loops and are illustrative; the proxy has no MoE, sparse attention or prefix-cache hits. Animated, with the pools at work: Inference Trade-offs Explained, chapter 7.
Deck 06 looks at the metrics, hot-spot analysis and validation behind these numbers, and how to make them hold up in review.