Prefill runs half the model, decode runs all of it: DeepSeek-V4.1-Flash’s Causal Encoder-Decoder (CED), where every decoder layer projects its KV from the last encoder state. What a causal encoder is, where the decoder’s KV comes from, the parameter, FLOP and KV-cache arithmetic, how to serve it with asymmetric pools, and what it costs.
Model facts come from the DeepSeek-V4.1-Flash technical report, DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression (DeepSeek-AI, arXiv:2609.19969, September 2026), cited as “p. N” or by section. Serving numbers come from a discrete-event simulator run on a dense proxy model, recorded in results.md; they are illustrative, not measurements of DeepSeek’s model.
A chat turn reads a prompt once and writes a long answer. An agent loop is the other way round: each step appends a tool result or a file to an ever-growing context and then writes a short tool call. The report opens with exactly this: “the widespread adoption of long-horizon agents has made model workloads increasingly input-heavy” (abstract, p. 1).
The prefix seen on earlier turns can be reused from a KV cache; DeepSeek keeps global KV on SSD for at least 72 hours (section 3.2.1, p. 19). The new suffix, and every prompt whose cache was evicted, must still be prefilled: “frequent tool calls generate extensive prefill requests, imposing severe computational overhead when KV caches miss” (section 2.2, p. 9).
Generating output needs the whole network; reading input, the bet goes, needs less of it. CED makes prefill cost about half of a full forward pass, “lowering the cost of processing new or uncached inputs in agentic workloads with growing contexts” (section 2.1, p. 8), while decode keeps every layer.
Prefill processes many tokens per step and is compute-bound; decode produces one token per sequence per step and is memory-bandwidth-bound. Halving prefill FLOPs therefore helps exactly where an input-heavy workload spends its compute. The background is in LLM Inference Simulators 03.
One stack of L causal layers. Every layer computes its own keys and values from its own hidden state, so prefill runs all L layers on every prompt token. Prompt and output are one token stream.
A bidirectional encoder reads the input; a separate decoder writes the output and cross-attends to the encoder’s states (Arch 05). Input and output are two streams, and appending to the input changes every encoder state, so a growing agent context cannot be encoded incrementally.
Still one token stream and causal attention everywhere. The bottom L/2 layers are the causal encoder: position i sees only positions ≤ i, so earlier states never change when tokens are appended and the KV cache works as usual. The top L/2 layers, the decoder, take their global keys and values from the encoder’s last layer.
“Causal” is what separates CED from T5: it is a decoder-only model in everything that matters for serving (one stream, prefix caching, autoregressive decode), but with the property that a prompt token’s contribution to every layer’s global KV is fixed once it has passed through the encoder. DeepSeek-V4.1-Flash has 40 layers, split 20 + 20 (Figure 3, p. 7; section 4.2.1, p. 21).
A later token’s decoder layers need, from each earlier prompt token, only that token’s decoder KV. Under CED those KV entries are a linear projection of the encoder output, so they can be produced without running the decoder on the prompt at all (section 2.2, p. 9).
For global attention, CED treats the bottom L/2 layers as the causal encoder. For a decoder layer l > L/2, the KV entries are “not derived from their respective hidden states Hl” but “projected directly from the hidden state of the (L/2)-th layer, HL/2, using layer-dependent projection weights” (section 2.2, equation 1, p. 9):
C_l = H_(L/2) · W_l^KV Z_l = H_(L/2) · W_l^Z for l > L/2
# C: the decoder layer's global KV entries; Z: their compression weights (the CSA2 compressor)
The sliding-window (SWA) branch is different: every layer, decoder layers included, still computes its local keys and values from its own hidden state Hl (section 2.2, p. 9). That is what forces the replay on the next slide. In the shipped model the decoder’s global KV is also shared across layers by CSA2: only the decoder’s first layer (Full mode) computes it from the encoder state, and later decoder layers reuse it (sections 2.3.1 and 4.2.1, pp. 11 and 22).
Prefill computes the encoder for every prompt token, plus the decoder’s global-KV projections from the encoder output. The decoder’s per-layer computation is skipped for all but the last nwin prompt tokens.
The first decode steps need each decoder layer’s SWA KV for the most recent nwin = 128 positions (window size, section 4.2.1, p. 22), and those come from the decoder’s own hidden states. Rebuilding them exactly would mean running the decoder over the last L/2 × nwin tokens, because sliding-window dependencies accumulate layer by layer (sections 2.2 and 3.2.2, pp. 9 and 20).
Instead, at every prefill the last nwin prompt tokens’ encoder outputs are fed through the decoder with SWA truncated to that segment. The result “is not mathematically equivalent” to a full decoder pass; the report finds a “negligible impact on response quality” and simulates the same replay in post-training (section 3.2.2, p. 20).
decoder-only: O(N · L)
CED: O(N · L/2 + n_win · L/2) ≈ O(N · L/2) # for N >> n_win: "effectively halving"
The same trick lets DeepSeek drop SWA KV from the persistent (SSD) cache. If a request hits the cached global KV but its encoder SWA KV has been evicted, Encoder SWA Bounded Replay recomputes only the last nwin tokens of the cached prefix instead of an L × nwin-token forward pass. SWA KV lives in a pool of 10% of each machine’s host DRAM with a TTL of minutes (sections 3.2.1–3.2.2, pp. 19–20).
Two kinds of sparsity now stack. Mixture-of-experts routing activates a few experts per token (Arch 01: total versus active parameters); CED then decides which layers a token runs at all.
| Base model | Backbone parameters | Activated per token | Source |
|---|---|---|---|
| DeepSeek-V4-Flash-Base | 284B | 13B | Table 1, p. 24 |
| DeepSeek-V4-Pro-Base | 1.6T | 49B | Table 1, p. 24 |
| DeepSeek-V4.1-Flash-Base | 552B (plus 196B Engram memory parameters) | 8B per prompt token, 16B per generated token | Table 1, p. 24; section 2.1, p. 7 |
DeepSeek-V4.1-Flash: 40 layers, hidden size 5,120; every layer is DeepSeekMoE with 1 shared and 384 routed experts, 6 routed experts active per token, expert width 2,304 (section 4.2.1, pp. 21–22). Section 4.3.2 and Table 1 report the base model comparable to DeepSeek-V4-Pro-Base on world knowledge with a third of the backbone parameters (p. 23–24).
On a dense Llama-3-70B-shape model split 40 + 40 (the simulator’s illustrative proxy, 4× H100, roofline cost model), a prompt token touches 34.9B matmul parameters against 69.5B for a generated token: a ratio of 0.502, the extra above one half being the decoder’s K/V projections (0.67B). One prefill step for one prompt:
| Prompt tokens | Decoder-only | CED, replay on prefill | ratio | CED, encoder only | ratio |
|---|---|---|---|---|---|
| 2,048 | 133.9 ms | 71.8 ms | 0.54× | 67.5 ms | 0.50× |
| 8,192 | 564.3 ms | 288.3 ms | 0.51× | 283.5 ms | 0.50× |
| 32,768 | 2,741 ms | 1,382 ms | 0.50× | 1,375 ms | 0.50× |
The ratio approaches one half as the prompt grows, because the 128-token replay and the LM head are fixed costs and causal attention grows with the square of the prompt in both halves alike. “Encoder only” is the prefill instance of the asymmetric deployment on slide 07, where the replay moves to the decode instance.
CED saves compute; the KV-cache saving in DeepSeek-V4.1-Flash comes mostly from cross-layer sharing (CSA2), FP4 storage and the bounded replay. The headline numbers, all relative to DeepSeek-V4-Flash at the same sequence length:
| Cache | Where it lives | V4.1-Flash | Source |
|---|---|---|---|
| Global KV (main KV + indexer K) | HBM, always | 890 bytes per token, about 1/4 of V4-Flash | abstract, p. 1; section 6, p. 37 |
| Persistent KV (prefix cache) | SSD or host memory | about 1/8 of V4-Flash | abstract, p. 1; section 3.2.1, p. 19 |
| Versus DeepSeek-V1 | global KV per token | about 437× smaller | Figure 1(b), p. 1 |
So the “about 8×” that circulates is the persistent cache; the cache held in HBM shrank about 4×. The 1/8 is the product of two factors (section 3.2.1, p. 19):
no SWA KV in the persistent cache ≈ 1/2 # SWA KV was "nearly half" of V4's persistent cache
global KV, CSA2 + FP4 ≈ 1/4
≈ 1/8
FP4 entry: 512 channels × (4 bits E2M1) + 512/16 E4M3 scales × 8 bits
= 256 B + 32 B = 288 B # 4.5 bits per value vs FP8's 8: "nearly halves"
KV-producing (Full-mode) layers:
encoder: 3 groups of 6 CSA2 layers, 1 Full each, compression m = 2 → 3 × 1/2 entry per token
decoder: 5 groups of 4, only group 1 starts in Full mode, m = 1 → 1 entry per token
main KV = (3 × 1/2 + 1) × 288 B = 2.5 × 288 = 720 B per token
the rest of the 890 B (≈ 170 B) is indexer K (128-dim, FP4 in V4), 2.5 entries per token
The paper states only the 890-byte total; the split above is a consistency check, not a quoted figure. It also shows how small CED’s own share is: the decoder contributes one KV-producing layer, projected from the encoder state, instead of twenty.
DeepSeek deploys with Encoder–Prefill–Decode (EPD) disaggregation: vision encoding, prefill and decoding scale independently (section 3.2, p. 19). CED makes the prefill and decode pools unequal in a new way: they no longer run the same model.
The prefill instance runs the encoder over the whole prompt, the decoder’s KV projections, and the 128-token decoder replay; the decode instance runs the whole model per generated token. Both hold all the weights.
SGLang issue #39963 (an open RFC, September 2026) proposes that the prefill worker hold only the embedding, the causal encoder and the first decoder layer’s KV/indexer projection, and that the decode worker run the final 128-token window through all 40 layers as a bounded prefill, then decode. Its stated goal: remove close to half the model weights from each prefill replica, and fit one production prefill replica on a single B300 GPU.
What a prefill instance must hold, for the dense 70B-shape proxy (BF16 weights; KV room = 90% of HBM minus weights, in tokens of KV for all 80 layers):
| Prefill instance | Whole model: weights | KV room (tokens) | Encoder only: weights | KV room (tokens) |
|---|---|---|---|---|
| 1× H100 | 141.1 GB | does not fit | 71.9 GB | 321 |
| 2× H100 | 141.1 GB | 8,835 | 71.9 GB | 220,048 |
| 4× H100 | 141.1 GB | 448,288 | 71.9 GB | 659,501 |
The decode instance now runs two step shapes, a bounded prefill of up to 128 tokens per newly admitted request and one-token decode steps, often mixed in one batch, and the first token is produced on the decode pool after the KV hand-off, so time-to-first-token includes the transfer. The RFC notes the need for piecewise CUDA graphs for the mixed batches. The simulator in LLM Inference Simulators 05 models both placements.
Six instances of 4× H100 (24 GPUs), split every way between a prefill and a decode pool; for each split, the highest Poisson request rate at which 90% of requests meet both SLOs (TPOT ≤ 25 ms; TTFT ≤ 5× the decoder-only model’s unloaded prefill time for the mean prompt), found by bisection on full simulator runs. Prompt and output lengths are lognormal (cv 0.5); the ratios stand in for agent loops and are illustrative.
| Prompt : output (mean tokens) | Model | Best split of 6 | Max req/s in SLO | vs decoder-only |
|---|---|---|---|---|
| 2,048 : 512 (4:1) TTFT SLO 669.3 ms | Decoder-only | 4P + 2D | 25.72 | 1.00× |
| CED, replay on prefill | 3P + 3D | 39.20 | 1.52× | |
| CED, replay on decode | 3P + 3D | 41.07 | 1.60× | |
| 4,096 : 256 (16:1) TTFT SLO 1,361 ms | Decoder-only | 4P + 2D | 13.10 | 1.00× |
| CED, replay on prefill | 4P + 2D | 27.01 | 2.06× | |
| CED, replay on decode | 4P + 2D | 27.90 | 2.13× | |
| 8,192 : 128 (64:1) TTFT SLO 2,821 ms | Decoder-only | 5P + 1D | 7.89 | 1.00× |
| CED, replay on prefill | 4P + 2D | 13.23 | 1.68× | |
| CED, replay on decode | 4P + 2D | 13.23 | 1.68× | |
| 16,384 : 64 (256:1) TTFT SLO 6,045 ms | Decoder-only | 5P + 1D | 3.81 | 1.00× |
| CED, replay on prefill | 5P + 1D | 6.89 | 1.81× | |
| CED, replay on decode | 5P + 1D | 8.17 | 2.14× |
Halving prefill cost raised capacity 1.52–2.06× with the paper’s placement. The best split moved one instance towards decode at 2,048 : 512 (from 4P + 2D to 3P + 3D) and 8,192 : 128 (from 5P + 1D to 4P + 2D): the prefill pool needs fewer instances for the same TTFT. Moving the replay to decode helped most at 16,384 : 64 (2.14× against 1.81×). Gains above 2× are queueing, not FLOPs: with a fixed TTFT SLO, halving the prefill service time cuts the queueing delay by more than half.
With CED’s prefill instances on 2 H100s each (still 24 GPUs; 8,192 : 128 workload), the best splits sustain:
| Model | Best split of 24 GPUs | Max req/s in SLO |
|---|---|---|
| CED, replay on prefill | 8 prefill × 2 GPUs + 2 decode × 4 GPUs | 12.95 |
| CED, replay on decode | 8 prefill × 2 GPUs + 2 decode × 4 GPUs | 13.09 |
About the same as with 4-GPU instances: in a roofline, smaller instances add no capacity. Their value is that the encoder fits with room to spare (slide 07). The decoder-only model on 2-GPU instances is left out: its result there is limited by the simulator’s router, not by compute (results.md, section 18).
Source: sections 16–18 of results.md (--ced-encoder-layers, --ced-replay, --ced-replay-on). A dense proxy on a roofline: no MoE, no sparse attention, no prefix-cache hits, no KV compression, so the absolute numbers say nothing about DeepSeek’s model; the direction and rough size of the effect are the point. The simulator’s prefill router ignores work in flight, which slightly favours the model with shorter prefill steps.
The report says CED cuts “nearly half of the prefill computation while maintaining performance comparable to the baseline” (section 2.2, p. 9) but publishes no CED-only ablation. Table 1 (p. 24) compares whole base models that also differ in data, CSA2, Engram and size, so CED’s own quality cost cannot be read from it.
A prompt token’s global KV is computed at depth L/2 and never revised by the decoder; only the last 128 prompt tokens ever run the decoder layers. What a decoder-only model would have computed about the prompt in layers 21–40 is gone. CED’s per-layer projections are the paper’s answer to YOCO’s single shared cache, to “enhance … the computational depth of KV generation” (p. 9); how much is lost is an open question.
Both bounded replays reconstruct SWA state approximately; the suffix’s KV then depends on where the cache hit happened (section 3.2.2, p. 20). The limitations section names “approximate state reconstruction in SWA Bounded Replay” as a possible cause of degradation in untested boundary cases (section 6, p. 37).
Our reading: training still runs every layer on every token (a loss is needed at every position), so CED saves inference prefill, not pre-training FLOPs; the paper claims no training saving. Decode still activates 16B per token, so output-heavy workloads gain nothing, and post-training must simulate the replay (p. 20).
Is half the right split, or would a deeper encoder cost little quality and save more? How does CED interact with speculative decoding, where the drafter (DSpark, section 2.4.3) must also be served? And at what prompt:output ratio does a deployment stop benefiting: the simulator says the gain shrinks as outputs grow, but the crossover depends on SLOs and hardware.
CED is “inspired by YOCO” (section 2.2, p. 9) and sits in a family of designs that stop every layer from keeping its own KV cache.
| Design | What is shared | Prefill | Source |
|---|---|---|---|
| YOCO (You Only Cache Once) | A self-decoder produces one global KV cache; a cross-decoder stacked on it reuses that cache via cross-attention | Can “early exit” after the self-decoder “without changing the final output” | Sun et al., 2024, arXiv:2405.05254 |
| CED | One source state (the encoder output); each decoder layer has its own KV projection of it; SWA stays per layer | Encoder only, plus a 128-token decoder replay | DeepSeek-AI, 2026, arXiv:2609.19969 |
| CLA (Cross-Layer Attention) | Adjacent layers share key/value heads, on top of MQA/GQA: about 2× less KV “while maintaining nearly the same accuracy” (1B and 3B models trained from scratch) | Unchanged | Brandon et al., 2024, arXiv:2405.12981 |
| CSA2 | Main KV and indexer keys shared across layers, top-k indices reused (Full, Reindex and Reuse modes) | Unchanged; less indexing | Same report, section 2.3 (pp. 9–12) |
| Cross-layer sparse attention | Built on YOCO-style KV sharing; one indexer’s top-k selection reused across the cross-decoder layers | Shares YOCO’s cheap prefill | Sun et al., 2026, arXiv:2606.06467 |
The difference that matters for serving: CLA and CSA2 shrink the cache but leave prefill compute alone; YOCO and CED shrink prefill compute by making the upper layers’ KV a function of a lower layer. CED keeps YOCO’s prefill saving while giving each decoder layer its own projection, and pays for keeping layer-wise sliding windows with the replay. Zamba’s single shared attention layer (Arch 05) is a cousin from the hybrid-SSM side.
The serving side is simulated in LLM Inference Simulators 05; mixture-of-experts sparsity is Arch 01; the classic encoder-decoder is Arch 05. The series index links all six decks.