Modern Architectures Series — Presentation 06

Asymmetric Causal Encoder-Decoder

Prefill runs half the model, decode runs all of it: DeepSeek-V4.1-Flash’s Causal Encoder-Decoder (CED), where every decoder layer projects its KV from the last encoder state. What a causal encoder is, where the decoder’s KV comes from, the parameter, FLOP and KV-cache arithmetic, how to serve it with asymmetric pools, and what it costs.

CED DeepSeek-V4.1-Flash Encoder-only prefill Cross-layer KV YOCO Agent workloads P/D disaggregation
Prompt → Causal encoder → KV projections → Decoder (last 128) → Decode: whole model
00

Topics We’ll Cover

Sources

Model facts come from the DeepSeek-V4.1-Flash technical report, DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression (DeepSeek-AI, arXiv:2609.19969, September 2026), cited as “p. N” or by section. Serving numbers come from a discrete-event simulator run on a dense proxy model, recorded in results.md; they are illustrative, not measurements of DeepSeek’s model.

01

Why Prefill Is the Bottleneck in Agent Loops

A chat turn reads a prompt once and writes a long answer. An agent loop is the other way round: each step appends a tool result or a file to an ever-growing context and then writes a short tool call. The report opens with exactly this: “the widespread adoption of long-horizon agents has made model workloads increasingly input-heavy” (abstract, p. 1).

context
+
tool output (prefill)
→
short tool call (decode)
→
run tool
↻

Prefix caching helps, but not enough

The prefix seen on earlier turns can be reused from a KV cache; DeepSeek keeps global KV on SSD for at least 72 hours (section 3.2.1, p. 19). The new suffix, and every prompt whose cache was evicted, must still be prefilled: “frequent tool calls generate extensive prefill requests, imposing severe computational overhead when KV caches miss” (section 2.2, p. 9).

The design bet

Generating output needs the whole network; reading input, the bet goes, needs less of it. CED makes prefill cost about half of a full forward pass, “lowering the cost of processing new or uncached inputs in agentic workloads with growing contexts” (section 2.1, p. 8), while decode keeps every layer.

Prefill versus decode, in one line

Prefill processes many tokens per step and is compute-bound; decode produces one token per sequence per step and is memory-bandwidth-bound. Halving prefill FLOPs therefore helps exactly where an input-heavy workload spends its compute. The background is in LLM Inference Simulators 03.

02

Three Shapes: Decoder-Only, Encoder-Decoder, Causal Encoder-Decoder

Decoder-only

One stack of L causal layers. Every layer computes its own keys and values from its own hidden state, so prefill runs all L layers on every prompt token. Prompt and output are one token stream.

Classic encoder-decoder (T5)

A bidirectional encoder reads the input; a separate decoder writes the output and cross-attends to the encoder’s states (Arch 05). Input and output are two streams, and appending to the input changes every encoder state, so a growing agent context cannot be encoded incrementally.

Causal encoder-decoder (CED)

Still one token stream and causal attention everywhere. The bottom L/2 layers are the causal encoder: position i sees only positions ≤ i, so earlier states never change when tokens are appended and the KV cache works as usual. The top L/2 layers, the decoder, take their global keys and values from the encoder’s last layer.

“Causal” is what separates CED from T5: it is a decoder-only model in everything that matters for serving (one stream, prefix caching, autoregressive decode), but with the property that a prompt token’s contribution to every layer’s global KV is fixed once it has passed through the encoder. DeepSeek-V4.1-Flash has 40 layers, split 20 + 20 (Figure 3, p. 7; section 4.2.1, p. 21).

Why the decoder can be skipped during prefill

A later token’s decoder layers need, from each earlier prompt token, only that token’s decoder KV. Under CED those KV entries are a linear projection of the encoder output, so they can be produced without running the decoder on the prompt at all (section 2.2, p. 9).

03

Where the Decoder’s KV Comes From

For global attention, CED treats the bottom L/2 layers as the causal encoder. For a decoder layer l > L/2, the KV entries are “not derived from their respective hidden states Hl” but “projected directly from the hidden state of the (L/2)-th layer, HL/2, using layer-dependent projection weights” (section 2.2, equation 1, p. 9):

Equation (1), arXiv:2609.19969
C_l = H_(L/2) · W_l^KV        Z_l = H_(L/2) · W_l^Z        for l > L/2
# C: the decoder layer's global KV entries; Z: their compression weights (the CSA2 compressor)
Causal encoder (layers 1 … L/2): every prompt token tokens → embedding encoder layer 1own KV from H_1 …own KV from H_l encoder layer L/2output H_(L/2) Decoder (layers L/2+1 … L) decoder layer L/2+1global KV = H_(L/2) W^KV_(L/2+1) decoder layer …global KV = H_(L/2) W^KV_l decoder layer Lglobal KV = H_(L/2) W^KV_L LM headonly for tokens that run the decoder dashed: one projection per decoder layer, from the same H_(L/2)

The sliding-window (SWA) branch is different: every layer, decoder layers included, still computes its local keys and values from its own hidden state Hl (section 2.2, p. 9). That is what forces the replay on the next slide. In the shipped model the decoder’s global KV is also shared across layers by CSA2: only the decoder’s first layer (Full mode) computes it from the encoder state, and later decoder layers reuse it (sections 2.3.1 and 4.2.1, pp. 11 and 22).

04

What Prefill Skips, and the Replay Window

Prefill computes the encoder for every prompt token, plus the decoder’s global-KV projections from the encoder output. The decoder’s per-layer computation is skipped for all but the last nwin prompt tokens.

Why anything is replayed

The first decode steps need each decoder layer’s SWA KV for the most recent nwin = 128 positions (window size, section 4.2.1, p. 22), and those come from the decoder’s own hidden states. Rebuilding them exactly would mean running the decoder over the last L/2 × nwin tokens, because sliding-window dependencies accumulate layer by layer (sections 2.2 and 3.2.2, pp. 9 and 20).

Decoder SWA Bounded Replay

Instead, at every prefill the last nwin prompt tokens’ encoder outputs are fed through the decoder with SWA truncated to that segment. The result “is not mathematically equivalent” to a full decoder pass; the report finds a “negligible impact on response quality” and simulates the same replay in post-training (section 3.2.2, p. 20).

Prefill work, sequence length N, L layers (section 2.2, p. 9)
decoder-only:  O(N · L)
CED:           O(N · L/2  +  n_win · L/2)  ≈  O(N · L/2)      # for N >> n_win: "effectively halving"
A second replay, for cache misses

The same trick lets DeepSeek drop SWA KV from the persistent (SSD) cache. If a request hits the cached global KV but its encoder SWA KV has been evicted, Encoder SWA Bounded Replay recomputes only the last nwin tokens of the cached prefix instead of an L × nwin-token forward pass. SWA KV lives in a pool of 10% of each machine’s host DRAM with a TTL of minutes (sections 3.2.1–3.2.2, pp. 19–20).

05

Activated Parameters and FLOPs per Token

Two kinds of sparsity now stack. Mixture-of-experts routing activates a few experts per token (Arch 01: total versus active parameters); CED then decides which layers a token runs at all.

Base modelBackbone parametersActivated per tokenSource
DeepSeek-V4-Flash-Base284B13BTable 1, p. 24
DeepSeek-V4-Pro-Base1.6T49BTable 1, p. 24
DeepSeek-V4.1-Flash-Base552B (plus 196B Engram memory parameters)8B per prompt token, 16B per generated tokenTable 1, p. 24; section 2.1, p. 7

DeepSeek-V4.1-Flash: 40 layers, hidden size 5,120; every layer is DeepSeekMoE with 1 shared and 384 routed experts, 6 routed experts active per token, expert width 2,304 (section 4.2.1, pp. 21–22). Section 4.3.2 and Table 1 report the base model comparable to DeepSeek-V4-Pro-Base on world knowledge with a third of the backbone parameters (p. 23–24).

The same split on a dense model

On a dense Llama-3-70B-shape model split 40 + 40 (the simulator’s illustrative proxy, 4× H100, roofline cost model), a prompt token touches 34.9B matmul parameters against 69.5B for a generated token: a ratio of 0.502, the extra above one half being the decoder’s K/V projections (0.67B). One prefill step for one prompt:

Prompt tokensDecoder-onlyCED, replay on prefillratioCED, encoder onlyratio
2,048133.9 ms71.8 ms0.54×67.5 ms0.50×
8,192564.3 ms288.3 ms0.51×283.5 ms0.50×
32,7682,741 ms1,382 ms0.50×1,375 ms0.50×

The ratio approaches one half as the prompt grows, because the 128-token replay and the LM head are fixed costs and causal attention grows with the square of the prompt in both halves alike. “Encoder only” is the prefill instance of the asymmetric deployment on slide 07, where the replay moves to the decode instance.

06

The KV-Cache Saving, Arithmetic Shown

CED saves compute; the KV-cache saving in DeepSeek-V4.1-Flash comes mostly from cross-layer sharing (CSA2), FP4 storage and the bounded replay. The headline numbers, all relative to DeepSeek-V4-Flash at the same sequence length:

CacheWhere it livesV4.1-FlashSource
Global KV (main KV + indexer K)HBM, always890 bytes per token, about 1/4 of V4-Flashabstract, p. 1; section 6, p. 37
Persistent KV (prefix cache)SSD or host memoryabout 1/8 of V4-Flashabstract, p. 1; section 3.2.1, p. 19
Versus DeepSeek-V1global KV per tokenabout 437× smallerFigure 1(b), p. 1

So the “about 8×” that circulates is the persistent cache; the cache held in HBM shrank about 4×. The 1/8 is the product of two factors (section 3.2.1, p. 19):

Persistent KV, V4.1-Flash / V4-Flash
no SWA KV in the persistent cache   ≈ 1/2     # SWA KV was "nearly half" of V4's persistent cache
global KV, CSA2 + FP4               ≈ 1/4
                                    ≈ 1/8

Reconstructing the 890 bytes (our arithmetic, from the report’s configuration)

Main KV, from sections 2.4.4 and 4.2.1 (pp. 14, 22)
FP4 entry:  512 channels × (4 bits E2M1) + 512/16 E4M3 scales × 8 bits
         =  256 B + 32 B = 288 B                       # 4.5 bits per value vs FP8's 8: "nearly halves"
KV-producing (Full-mode) layers:
  encoder: 3 groups of 6 CSA2 layers, 1 Full each, compression m = 2  → 3 × 1/2 entry per token
  decoder: 5 groups of 4, only group 1 starts in Full mode,      m = 1  → 1 entry per token
main KV  =  (3 × 1/2 + 1) × 288 B  =  2.5 × 288  =  720 B per token
the rest of the 890 B (≈ 170 B) is indexer K (128-dim, FP4 in V4), 2.5 entries per token

The paper states only the 890-byte total; the split above is a consistency check, not a quoted figure. It also shows how small CED’s own share is: the decoder contributes one KV-producing layer, projected from the encoder state, instead of twenty.

07

Serving It: Asymmetric Prefill and Decode Pools

DeepSeek deploys with Encoder–Prefill–Decode (EPD) disaggregation: vision encoding, prefill and decoding scale independently (section 3.2, p. 19). CED makes the prefill and decode pools unequal in a new way: they no longer run the same model.

The paper’s placement

The prefill instance runs the encoder over the whole prompt, the decoder’s KV projections, and the 128-token decoder replay; the decode instance runs the whole model per generated token. Both hold all the weights.

The SGLang RFC’s placement

SGLang issue #39963 (an open RFC, September 2026) proposes that the prefill worker hold only the embedding, the causal encoder and the first decoder layer’s KV/indexer projection, and that the decode worker run the final 128-token window through all 40 layers as a bounded prefill, then decode. Its stated goal: remove close to half the model weights from each prefill replica, and fit one production prefill replica on a single B300 GPU.

What a prefill instance must hold, for the dense 70B-shape proxy (BF16 weights; KV room = 90% of HBM minus weights, in tokens of KV for all 80 layers):

Prefill instanceWhole model: weightsKV room (tokens)Encoder only: weightsKV room (tokens)
1× H100141.1 GBdoes not fit71.9 GB321
2× H100141.1 GB8,83571.9 GB220,048
4× H100141.1 GB448,28871.9 GB659,501
The cost of moving the replay

The decode instance now runs two step shapes, a bounded prefill of up to 128 tokens per newly admitted request and one-token decode steps, often mixed in one batch, and the first token is produced on the decode pool after the KV hand-off, so time-to-first-token includes the transfer. The RFC notes the need for piecewise CUDA graphs for the mixed batches. The simulator in LLM Inference Simulators 05 models both placements.

08

Measured: Pool Split and Capacity in a Simulator

Six instances of 4× H100 (24 GPUs), split every way between a prefill and a decode pool; for each split, the highest Poisson request rate at which 90% of requests meet both SLOs (TPOT ≤ 25 ms; TTFT ≤ 5× the decoder-only model’s unloaded prefill time for the mean prompt), found by bisection on full simulator runs. Prompt and output lengths are lognormal (cv 0.5); the ratios stand in for agent loops and are illustrative.

Prompt : output (mean tokens)ModelBest split of 6Max req/s in SLOvs decoder-only
2,048 : 512 (4:1)
TTFT SLO 669.3 ms
Decoder-only4P + 2D25.721.00×
CED, replay on prefill3P + 3D39.201.52×
CED, replay on decode3P + 3D41.071.60×
4,096 : 256 (16:1)
TTFT SLO 1,361 ms
Decoder-only4P + 2D13.101.00×
CED, replay on prefill4P + 2D27.012.06×
CED, replay on decode4P + 2D27.902.13×
8,192 : 128 (64:1)
TTFT SLO 2,821 ms
Decoder-only5P + 1D7.891.00×
CED, replay on prefill4P + 2D13.231.68×
CED, replay on decode4P + 2D13.231.68×
16,384 : 64 (256:1)
TTFT SLO 6,045 ms
Decoder-only5P + 1D3.811.00×
CED, replay on prefill5P + 1D6.891.81×
CED, replay on decode5P + 1D8.172.14×
What moves

Halving prefill cost raised capacity 1.52–2.06× with the paper’s placement. The best split moved one instance towards decode at 2,048 : 512 (from 4P + 2D to 3P + 3D) and 8,192 : 128 (from 5P + 1D to 4P + 2D): the prefill pool needs fewer instances for the same TTFT. Moving the replay to decode helped most at 16,384 : 64 (2.14× against 1.81×). Gains above 2× are queueing, not FLOPs: with a fixed TTFT SLO, halving the prefill service time cuts the queueing delay by more than half.

Smaller prefill instances

With CED’s prefill instances on 2 H100s each (still 24 GPUs; 8,192 : 128 workload), the best splits sustain:

ModelBest split of 24 GPUsMax req/s in SLO
CED, replay on prefill8 prefill × 2 GPUs + 2 decode × 4 GPUs12.95
CED, replay on decode8 prefill × 2 GPUs + 2 decode × 4 GPUs13.09

About the same as with 4-GPU instances: in a roofline, smaller instances add no capacity. Their value is that the encoder fits with room to spare (slide 07). The decoder-only model on 2-GPU instances is left out: its result there is limited by the simulator’s router, not by compute (results.md, section 18).

Source: sections 16–18 of results.md (--ced-encoder-layers, --ced-replay, --ced-replay-on). A dense proxy on a roofline: no MoE, no sparse attention, no prefix-cache hits, no KV compression, so the absolute numbers say nothing about DeepSeek’s model; the direction and rough size of the effect are the point. The simulator’s prefill router ignores work in flight, which slightly favours the model with shorter prefill steps.

09

Trade-offs and Open Questions

Quality: claimed, not isolated

The report says CED cuts “nearly half of the prefill computation while maintaining performance comparable to the baseline” (section 2.2, p. 9) but publishes no CED-only ablation. Table 1 (p. 24) compares whole base models that also differ in data, CSA2, Engram and size, so CED’s own quality cost cannot be read from it.

Depth of the prompt’s KV

A prompt token’s global KV is computed at depth L/2 and never revised by the decoder; only the last 128 prompt tokens ever run the decoder layers. What a decoder-only model would have computed about the prompt in layers 21–40 is gone. CED’s per-layer projections are the paper’s answer to YOCO’s single shared cache, to “enhance … the computational depth of KV generation” (p. 9); how much is lost is an open question.

Approximate states

Both bounded replays reconstruct SWA state approximately; the suffix’s KV then depends on where the cache hit happened (section 3.2.2, p. 20). The limitations section names “approximate state reconstruction in SWA Bounded Replay” as a possible cause of degradation in untested boundary cases (section 6, p. 37).

Training and decode are not cheaper

Our reading: training still runs every layer on every token (a loss is needed at every position), so CED saves inference prefill, not pre-training FLOPs; the paper claims no training saving. Decode still activates 16B per token, so output-heavy workloads gain nothing, and post-training must simulate the replay (p. 20).

Open questions

Is half the right split, or would a deeper encoder cost little quality and save more? How does CED interact with speculative decoding, where the drafter (DSpark, section 2.4.3) must also be served? And at what prompt:output ratio does a deployment stop benefiting: the simulator says the gain shrinks as outputs grow, but the crossover depends on SLOs and hardware.

10

Relatives: YOCO, CLA and Cross-Layer Sharing

CED is “inspired by YOCO” (section 2.2, p. 9) and sits in a family of designs that stop every layer from keeping its own KV cache.

DesignWhat is sharedPrefillSource
YOCO (You Only Cache Once)A self-decoder produces one global KV cache; a cross-decoder stacked on it reuses that cache via cross-attentionCan “early exit” after the self-decoder “without changing the final output”Sun et al., 2024, arXiv:2405.05254
CEDOne source state (the encoder output); each decoder layer has its own KV projection of it; SWA stays per layerEncoder only, plus a 128-token decoder replayDeepSeek-AI, 2026, arXiv:2609.19969
CLA (Cross-Layer Attention)Adjacent layers share key/value heads, on top of MQA/GQA: about 2× less KV “while maintaining nearly the same accuracy” (1B and 3B models trained from scratch)UnchangedBrandon et al., 2024, arXiv:2405.12981
CSA2Main KV and indexer keys shared across layers, top-k indices reused (Full, Reindex and Reuse modes)Unchanged; less indexingSame report, section 2.3 (pp. 9–12)
Cross-layer sparse attentionBuilt on YOCO-style KV sharing; one indexer’s top-k selection reused across the cross-decoder layersShares YOCO’s cheap prefillSun et al., 2026, arXiv:2606.06467

The difference that matters for serving: CLA and CSA2 shrink the cache but leave prefill compute alone; YOCO and CED shrink prefill compute by making the upper layers’ KV a function of a lower layer. CED keeps YOCO’s prefill saving while giving each decoder layer its own projection, and pays for keeping layer-wise sliding windows with the replay. Zamba’s single shared attention layer (Arch 05) is a cousin from the hybrid-SSM side.

11

What to Take Away

Where to go next

The serving side is simulated in LLM Inference Simulators 05; mixture-of-experts sparsity is Arch 01; the classic encoder-decoder is Arch 05. The series index links all six decks.