Modern Architectures Series — Presentation 05

Hybrid & Encoder-Decoder Architectures

Beyond decoder-only: Jamba and Zamba’s Mamba-attention interleaving, T5/FLAN-T5 for structured tasks, when bidirectional encoding wins, byte-level models (ByT5, Charformer), and a practical decision tree for architecture selection.

Jamba Zamba T5 FLAN-T5 ByT5 Charformer Encoder-Decoder
Decoder-only → Hybrid (Jamba) → Encoder-Decoder → Byte-level → Choose
00

Topics We’ll Cover

01

Hybrid Architectures — Jamba, Zamba

The hybrid Mamba-attention family emerged in 2024 as a pragmatic synthesis: use Mamba layers for cheap O(T) bulk computation while retaining periodic full-attention layers for in-context recall and associative retrieval. Two models define the space.

ModelTotal paramsArchitectureContextFits on
Jamba (AI21, Mar 2024)52B (12B active MoE)Transformer-Mamba blocks + MoE FFN; 1 attn per 7 Mamba256K1× 80GB GPU
Jamba-1.5 Mini (Aug 2024)12BSame hybrid, smaller256K1× 24GB GPU
Zamba (Zyphra, Jun 2024)7BShared single attn layer + 6 Mamba per block; 35B training tokens32K1× 24GB GPU
Zamba2-7B (Nov 2024)7BShared & routed attn layers + LoRA adapters between blocks128K1× 24GB GPU

Jamba layer structure

Attention
→
Mamba
→
Mamba
→
Mamba
→
MoE FFN
→
Mamba
→
Mamba
→
Mamba
→
Attention
→
…
Why MoE + Mamba hybrid is synergistic

Jamba combines MoE (deck 01) with Mamba (deck 02) in the same model. The MoE FFN provides large total parameter count at low active-param cost; Mamba provides O(1) inference state; the sparse attention layers provide recall. Each of the three components contributes orthogonally to the quality-cost trade-off.

02

Why Mix Mamba and Attention

The mixing ratio is not arbitrary — it reflects a deliberate allocation of KV cache budget and compute. Understanding the engineering trade-off requires comparing what each layer type contributes per FLOP.

What attention gives you

Exact in-context retrieval — any query can attend to any key, O(T2). Perfect for: copy tasks, multi-hop reasoning over recently-read facts, associative recall (MQAR). Cost: KV cache grows with context; quadratic FLOPs at training.

What Mamba gives you

Constant-state recurrence — O(1) per token at inference. Good for: pattern accumulation, compression of long context into a summary state, local syntactic processing. Cost: cannot exactly recall a token from 50K steps ago if the state compressed it away.

Empirical ratio findings

Ablation studies in the Jamba and Zamba papers:

Shared-attention variant (Zamba)

Zamba uses a single attention layer whose KV cache is shared across multiple Mamba blocks. This reduces KV cache to a constant regardless of depth (only one set of KVs is retained, not one per attention layer). The shared attention layer effectively acts as a global information bus that all Mamba layers can query.

03

T5 / FLAN-T5 — Encoder-Decoder for Structured Tasks

T5 (Raffel et al. 2020, Google) reframed every NLP task as a text-to-text problem, unifying classification, translation, summarisation, and QA under a single encoder-decoder transformer. FLAN-T5 (Wei et al. 2022) fine-tuned T5 on >1800 instruction tasks, producing the strongest encoder-decoder instruction follower available in open weights.

ModelParamsEncoder layersDecoder layersdmodel
T5-Small60M66512
T5-Base220M1212768
T5-Large770M24241024
T5-XL3B24242048
T5-XXL / FLAN-T5-XXL11B24244096

T5 architectural choices

Relative position biases

T5 uses a simple relative position encoding: a learned scalar bias per (head, relative-distance-bucket) pair, added to attention logits. 32 distance buckets, log-spaced. No absolute positional encodings at all — which is why T5 generalises to longer sequences at inference without modification.

Layer norm placement

T5 uses pre-LN (normalise before the sublayer), a simplified RMSNorm variant without learned bias or affine parameters. This stabilises training at larger scales and was later adopted by LLaMA. The original T5 released in 2020 predates its widespread use in decoder-only models.

04

When Encoder-Decoder Beats Decoder-Only

The encoder-decoder architecture has a structural advantage for tasks where the input should be fully processed before the output begins. Decoder-only models process input and output in the same left-to-right pass — each output token can attend to input tokens, but input tokens cannot attend to future input tokens that appear after them in a long prompt.

Encoder: bidirectional context over input

Every encoder token attends to every other encoder token. For a 512-token source document, the encoder builds a representation where token 1 has seen tokens 1–512 before the decoder generates its first output token. Bidirectional attention captures both left and right context.

Decoder: causal attention over output

The decoder is a standard causal transformer. It cross-attends to the encoder output at every layer. Cross-attention keys and values are from the encoder; queries from the decoder. This is the information pathway from fully-contextualised input to generated output.

Tasks where encoder-decoder excels

Translation: The encoder can see the full source sentence before generating the first target word. Critical for long-distance word-order differences (German subordinate clauses, Japanese SOV).
Classification over long inputs: A [CLS]-like pooled encoder representation over a 4K-token document outperforms a decoder-only model’s final-token embedding.
Summarisation (extractive): Selecting spans from the input benefits from full bidirectional context.
Structured prediction with schema: Schema-constrained output (XML, SQL, JSON) where the encoder has processed the schema and the decoder generates a valid instantiation.

A causal encoder-decoder (2026)

DeepSeek-V4.1-Flash keeps a single causal token stream but splits its 40 layers into a 20-layer causal encoder and a 20-layer decoder whose keys and values are projected from the last encoder state, so prefill runs only the encoder (8B activated parameters per token) while decode runs the whole model (16B). Arch 06 — Asymmetric Causal Encoder-Decoder covers it.

05

Modern Uses — Translation, Classification, Structured Extraction

Encoder-decoder models did not vanish when GPT-3 appeared — they continue to dominate specific production use-cases in 2025.

Machine translation

NLLB-200 (Meta, 2022): 3.3B encoder-decoder supporting 200 languages. Still the standard for low-resource language pairs. OPUS-MT (Helsinki NLP): 1000+ T5-style translation models under 300M params each, deployed at Google Translate and DeepL as quality-check components.

Document classification

FLAN-T5-Large (780M) outperforms GPT-3 (175B) on zero-shot classification benchmarks (SuperGLUE). Reason: the encoder’s bidirectional representation is a better input to a classification head than a decoder-only model’s causal embedding. Used in many enterprise document-routing pipelines.

Structured information extraction

Extract JSON records from unstructured text: product attributes from catalogue descriptions, medical entities from clinical notes. FLAN-T5 with few-shot prompting generates schema-valid JSON more reliably than decoder-only models of similar size due to the bidirectional encoder’s full source comprehension before generation begins.

Semantic similarity & reranking

Encoder-only models (BERT, DeBERTa, sentence-transformers) are still the gold standard for embedding-based retrieval. Cross-encoder rerankers (MonoT5, RankT5) use the encoder-decoder architecture for query-document relevance scoring — covered in the RAG series.

FLAN-T5 as a distillation target

FLAN-T5-XXL (11B) is frequently used as the teacher model for distilling smaller task-specific models. Its instruction-tuned representations are highly transferable. The student is typically an encoder-only or encoder-decoder model of 100–300M parameters, making FLAN-T5-based distillation a cost-effective way to build production classifiers and extractors.

06

Char/Byte Models (ByT5, Charformer) — Tokenisation-Free

BPE and SentencePiece tokenisation introduce a systematic blind spot: languages with complex morphology (Finnish, Turkish, Arabic), code-switching, noisy user-generated text, and adversarial inputs (character-level attacks). Byte-level models eliminate the tokeniser entirely.

ModelGranularityVocab sizeSequence overheadKey paper
ByT5 (Google, 2022)UTF-8 bytes256 + 3 special~3–4× longer vs T5Xue et al., arXiv 2105.13626
Charformer (Google, 2022)Char n-grams → subword via GBSTUnicode chars~2× longerTay et al., arXiv 2106.12672
MegaByte (Meta, 2023)Raw bytes, patch hierarchy256Patch-level amortisationYu et al., arXiv 2305.07185

ByT5 architecture

ByT5 is a vanilla T5 operating on UTF-8 byte sequences. The encoder is made wider and deeper than the decoder (6× encoder layers per decoder layer) to handle the longer input sequences without proportional compute increase in the decoder. The byte-level encoder is the expensive part; the decoder remains compact.

ByT5-Large: encoder/decoder dimension split
Encoder: 36 layers, d_model=1024   # handles long byte sequences
Decoder: 6  layers, d_model=1024   # generates subword-level output
Total params: 1.2B vs T5-Large 770M  # ~57% overhead for byte granularity
Where ByT5 wins

ByT5 matches T5 on standard English NLP (GLUE, SuperGLUE) and substantially outperforms it on: noisy text (>+5 F1 on GermEval, TweetNLP), multilingual morphological tasks, character-level robustness (adversarial character swaps), and low-resource languages with OOV tokens. The 3–4× sequence-length overhead is the main engineering cost — mitigated by MegaByte-style patching in later work.

07

Rotation Between Paradigms — What’s Likely to Dominate

The LLM field has oscillated between paradigms every few years. Understanding the trajectory helps make architectural bets.

EraDominant paradigmWhy it wonWhy the next era superseded it
2017–2019Encoder-only (BERT)Bidirectional pretraining on MLMGPT-3: decoder-only scales better for generation
2019–2021Encoder-decoder (T5)Text-to-text unificationFew-shot GPT-3 removed need for fine-tuning
2020–presentDecoder-only (GPT, LLaMA)Simpler pretraining, better scaling, instruction tuningEfficiency pressure → MoE, SSM hybrids
2024–?Hybrid decoder + MoE/SSMBetter param/FLOP trade-off, long contextDiffusion LMs if quality gap closes?
The pragmatic 2025 view

Decoder-only transformers will remain dominant for general-purpose LLMs through 2026 — the ecosystem (frameworks, RLHF pipelines, tooling) is too mature to displace quickly. MoE will be standard at 100B+ scale. SSM hybrids (Jamba-class) will capture the cost-sensitive long-context niche. Encoder-decoder will persist in structured prediction and translation. Diffusion LMs are likely 2–3 years from production dominance but will appear in code generation pipelines earlier (Mercury is already deployed).

08

Decision Tree for Architecture Selection

Given a new task, which architecture should you reach for? The following decision tree encodes the empirical wisdom from decks 01–05.

Need generation (text output)? START HERE YES Requires full bidirectional input comprehension before output? NO YES Encoder-Decoder T5, FLAN-T5, NLLB Context > 128K tokens? or memory-constrained? YES NO Hybrid SSM-Attn Jamba, Zamba, or SSM 100B+ param budget? train from scratch? MAYBE YES MoE decoder-only Mixtral / DeepSeek-V3 NO Dense decoder-only LLaMA-3, Qwen-2.5 Noisy/multilingual/OOV text? Morphology heavy? NO-GEN YES Byte-level model ByT5, MegaByte NO Encoder-only BERT, DeBERTa
09

What to Take Away

Next

You have now covered MoE sparse gating (Arch 01), Mamba state-space models (Arch 02), long-context positional encoding and attention strategies (Arch 03), discrete diffusion language models (Arch 04), and the hybrid/encoder-decoder landscape (Arch 05). Arch 06 takes the encoder-decoder idea back inside a causal model: an encoder-only prefill with asymmetric prefill and decode costs. The series index links all six decks alongside the parent LLMs hub.