Hardware · AI Accelerators

The Hardware Matrix Multiplier Unit

From Kung & Leiserson 1978 to Blackwell FP4 — what an AI accelerator's MMUL really is, why every modern chip has one, and how it shapes the rest of the system.

Systolic arrays Tensor Cores FP8 · FP4 · MXFP TPU · H100 · B200 Memory hierarchy Power & thermals
01

What this deck covers

Twenty-three slides, broadly grouped:

  1. Why a dedicated MMUL exists — energy and the roofline
  2. A short history: from MAC to TPU
  3. The systolic principle
  4. Output-stationary array — TPU style
  5. Weight-stationary array — NVDLA / edge style
  6. Input-stationary & row-stationary (Eyeriss)
  7. Dot-product trees — NVIDIA Tensor Cores
  8. What runs on an MMUL? Mapping to the Transformer
  9. Number formats — FP32 to FP4
  10. Block-scaled microformats — MXFP, NVFP4
  11. Stochastic rounding & mixed accumulation
  12. Real systems — TPU lineage
  13. Real systems — NVIDIA Tensor Cores
  14. Real systems — Cerebras / Groq / Dojo / Trainium / ANE / SME / Sohu
  15. Memory hierarchy — what feeds the MMUL
  16. Are caches useful?
  17. Performance: peak vs sustained
  18. Power: every joule comes from somewhere
  19. Thermals & power delivery
  20. Tradeoffs: array size, format, fill bubbles
  21. Interactive: dataflow comparator
  22. Interactive: number-format explorer
  23. Where this is going (2025 → 2027)
02

Why a dedicated MMUL exists

A modern LLM forward pass is more than 99% multiply-accumulate (MAC). On a general-purpose CPU, almost all the energy spent on that workload is not the math — it's fetching operands from DRAM and shuttling them through register files.

Horowitz, ISSCC 2014 — the numbers that built an industry

Operation (45 nm)EnergyRelative to MAC
32-bit FP MUL3.7 pJ1×
32-bit FP ADD0.9 pJ0.25×
32-bit FP MAC (effective)4.6 pJ1×
32-bit register-file read0.1 pJ0.02×
32-bit on-chip SRAM read5 pJ~1×
Off-chip DRAM access640 pJ~140×

An MMUL only wins if every operand fetched from DRAM is reused O(N) times across many MACs. That reuse — temporal in registers and scratchpads, spatial across thousands of PEs — is the entire reason the architecture exists.

The roofline argument

Williams et al. (2009): performance is bounded by min(peak FLOP/s, AI · peak BW), where arithmetic intensity AI = FLOPs per byte of DRAM traffic.

Bottom line

Matmul is the rare workload where you can spend 1 byte of DRAM bandwidth and get hundreds of FLOPs back. The MMUL is the structure that turns that arithmetic intensity into actual joules saved.

03

A short history

YearEvent
1976TRW MPY-016 single-chip MAC — the immediate ancestor of every "tensor" instruction
1978Kung & Leiserson, Systolic Arrays (for VLSI) — Sparse Matrix Proceedings, SIAM
1982H. T. Kung, Why Systolic Architectures? — IEEE Computer. Often misdated as 1979 — the canonical citation
1984Warp (CMU + GE/Honeywell) — 10-cell linear systolic, 100 MFLOPS/cell @ 10 MHz
1990iWarp (CMU + Intel) — single-chip 700K-transistor LIW with on-die mesh; ancestor of "compute + on-chip routing"
2015Google TPU v1 deployed internally — 256×256 INT8 systolic, 92 TOPS @ 700 MHz
2017NVIDIA Volta V100 ships first Tensor Core (4×4×4, FP16 mul / FP32 acc)
2018Turing T4: INT8/INT4 in TC; NVDLA RTL released open-source
2020A100: TF32, BF16, FP64 TC, 2:4 structured sparsity; 312 TFLOPS BF16
2022H100: FP8 (E4M3, E5M2), TMA, Transformer Engine; 1979 TFLOPS FP8 dense
2023OCP MX v1.0 specification — block-scaled FP8/6/4 with shared E8M0 scale
2024Blackwell B200: FP4 (E2M1) + NVFP4; 20 PF FP4 dense per package
2024TPU v6e Trillium: MXU enlarged from 128×128 to 256×256 — first resize since 2017
2025TPU v7 Ironwood; AMD MI355X (FP4); Microsoft Maia 100 — block-scaled formats become default

The throughline: spatial reuse via locally-connected PEs. Forty-eight years separate Kung's 1978 paper from the Blackwell datasheet — the slide deck has changed, the principle hasn't.

04

The systolic principle

Kung's metaphor: blood through a heart. Data pumps through a regular array of locally-connected processing elements. Operands enter at the edges, results emerge after a fixed number of cycles, no element ever stalls waiting on global memory.

Three properties that fall out for free

A[0,*]→ A[1,*]→ A[2,*]→ A[3,*]→ B[*,0]↓ B[*,1]↓ B[*,2]↓ B[*,3]↓ PE 0,0PE 0,1PE 0,2PE 0,3 PE 1,0PE 1,1PE 1,2PE 1,3 PE 2,0PE 2,1PE 2,2PE 2,3 PE 3,0PE 3,1PE 3,2PE 3,3 acc[r,c]

A 4×4 output-stationary mesh. A streams left→right (blue), B streams top→bottom (purple), accumulators stay put (green).

05

Output-stationary array — the TPU pattern

Used by every Google TPU (v1 onwards), AWS Trainium / Inferentia, and most "MXU-class" datacentre accelerators.

Cycle-by-cycle

This repo's mmul_pe.sv is the canonical PE: one signed multiply, one signed add, three flops (a-skew, w-skew, accumulator). mmul_systolic_os.sv instances ROWS·COLS of them with edge wiring.

rtl/mmul_pe.sv — single PE
always_ff @(posedge clk or negedge rst_n) begin
  if (!rst_n)         { a_out, w_out, acc_out } <= '0;
  else if (clear)     { a_out, w_out, acc_out } <= '0;
  else if (enable) begin
    a_out   <= a_in;                     // skew right
    w_out   <= w_in;                     // skew down
    acc_out <= acc_out + acc_t'(a_in * w_in);
  end
end

Latency

The "fill bubble" is the cost. For 256×256 with K=64 (typical attention head), 511 cycles of fill+drain wrap 64 productive cycles — 87% bubble. Real TPU code amortises this by chaining many K-tiles back-to-back without draining.

06

Weight-stationary array — the edge / inference pattern

Used by NVDLA, Apple Neural Engine, ARM Ethos-N, Hailo-8, Qualcomm Hexagon NPU, and most mobile / edge accelerators.

Idea

For inference, weights are fixed and reused across many activations (especially across a batch). So load weights once, hold them, stream activations. Partial sums spill to a separate buffer.

Output-stationary

acc lives in PE; both operands stream. Best for training (activations and weights are equally hot). Used by TPU, Trainium.

Weight-stationary

B held in PE registers; A streams; partial sums move out per cycle. Best for inference where weights >> activations × batch. Used by edge NPUs.

The cost ledger

This repo's mmul_weight_stationary.sv implements an M × N register-file held over K compute steps; a separate load_w handshake fills it from the off-chip weight stream.

07

Input-stationary & row-stationary (Eyeriss)

Two more dataflows for completeness — both more specialised, both interesting.

Input-stationary

Activations pinned, weights stream, partial sums flow. Less common in dense matmul, but appears in some training accelerators where the activation reuse pattern of ∂L/∂W = A^T · ∂L/∂Y dominates the gradient pass.

Row-stationary — Eyeriss (MIT, ISCA 2016)

The most thoughtful of the four. Maps a 1-D row of CNN filter weights onto a single PE; ifmap rows stream horizontally; partial sums flow vertically. Sze and Chen showed this maximises all three reuse patterns simultaneously — weights, ifmaps, and psums.

Practical note

The dataflow choice is a workload-dependent decision. Modern compilers (XLA for TPU, CUTLASS for NVIDIA, TT-Metalium for Tenstorrent) re-tile across multiple physical arrays so the user-facing kernel is dataflow-agnostic — but underneath, every implementation makes a choice.

08

Dot-product trees — NVIDIA Tensor Cores

NVIDIA's Tensor Core is conceptually not a Kung-style systolic. It's a parallel reduction tree: K multipliers feeding an adder tree, with operands fetched per warp from the register file.

1st-gen V100 Tensor Core (2017)

Why a tree, not an array?

This repo's mmul_dotproduct.sv is a single-cycle K-wide combinational dot product — the simplest model of one Tensor-Core lane. With K = 8 it computes a Q8.8 inner product into a 32-bit accumulator combinationally.

Generation by generation

GenGPUNew formatsPeak (per GPU, dense)
1stV100 (2017)FP16125 TFLOPS
2ndT4 / RTX 20 (2018)INT8, INT4130 TOPS INT8
3rdA100 (2020)TF32, BF16, FP64, 2:4 sparsity312 TFLOPS BF16
4thH100 (2022)FP8 (E4M3, E5M2), TMA1979 TFLOPS FP8
5thB200 (2024)FP4, NVFP4, MXFP20 PFLOPS FP4
09

What does an MMUL run, in a Transformer?

Almost all of it. A typical decoder-only LLM forward pass spends ≥99% of its FLOPs in dense matrix multiplications — so the MMUL is the part of the chip that decides how fast a Llama or GPT runs.

Per-block FLOP map (decoder layer, sequence length S, model dim D, head dim d, FFN dim 4D)

OperationShapeFLOPsOn the MMUL?Comment
Token embeddingS × V → S × D—No (gather)Indexed table lookup; some chips fuse as one-hot · W
RMSNorm / LayerNormS × DO(S·D)NoReductions + elementwise; vector lane
Q projectionS × D · D × D2·S·D²YesDense GEMM
K projectionS × D · D × D2·S·D²YesDense GEMM (smaller for GQA / MQA)
V projectionS × D · D × D2·S·D²YesDense GEMM
Q · KT (scores)S × d · d × S (per head)2·S²·DYes"Batched" GEMM, one per head; quadratic in S
scale + softmaxS × SO(S²)NoSubtract max, exp, normalise — vector lane / SFU
scores · VS × S · S × d2·S²·DYesBatched GEMM
Output projection WOS × D · D × D2·S·D²YesDense GEMM
FFN up + gate (SwiGLU)S × D · D × 4D (×2)16·S·D²YesTwo GEMMs; the largest in the layer
SiLU / GeLUS × 4DO(S·D)NoElementwise; SFU
FFN downS × 4D · 4D × D8·S·D²YesDense GEMM
Residual addS × DO(S·D)NoElementwise add
RoPE rotationS × DO(S·D)No2-element rotations on Q,K (vector lane)
LM head logitsS × D · D × V2·S·D·VYesFinal dense GEMM (often largest matrix in the model)

What this looks like on H100 / B200 silicon

Two regimes that matter for hardware

Prefill / training (compute-bound)

S = 1024+ tokens at once. GEMM shapes (M=N=K) are large; AI is high; MMUL utilisation 50–85%. This is what sales decks measure.

Decode / single-token (memory-bound)

S = 1. Every projection is a tall-skinny GEMM (1 × D · D × D); arithmetic intensity drops to ~D / 2 ≈ 4096 / 2 ≈ 2000, but the actual constraint is HBM bandwidth for weights and KV. MMUL utilisation often < 5%. This is why "tokens / sec" benchmarks look so different from peak TFLOPS.

Practical takeaway

For training, the MMUL is genuinely the bottleneck — and chasing FP4 / FP8 / NVFP4 is the right thing. For LLM inference at batch=1 you can put a 20-PFLOP B200 next to a 1-PFLOP MI300X and they will both be limited by HBM bandwidth — which is why memory bandwidth (not peak TFLOPS) drives every modern inference card's spec sheet.

10

Number formats: from FP32 to FP4

Halving the bit-width of the multiplier is approximately 4× fewer transistors (multiplier area is quadratic in mantissa width) and ~2× lower energy per MAC. Every generation has chased this curve.

FormatBitsS/E/MRange (≈)Used in
FP32 (IEEE 754)321/8/23±3.4e38Baseline, accumulators everywhere
TF32 (NVIDIA)19 stored as 321/8/10same as FP32A100+ TC, FP32 in/out
FP16 (IEEE binary16)161/5/10±65,504V100 onwards
BF16161/8/7±3.4e38TPU v2+, all modern; "FP32 truncated"
FP8 E4M381/4/3±448H100, MI300X — fwd weights/activations
FP8 E5M281/5/2±57,344H100, MI300X — gradients
FP6 E3M261/3/2±28OCP MX, Blackwell
FP4 E2M141/2/1±6OCP MX, Blackwell, MI355X
INT8 / INT48 / 4—±127 / ±7TPU v1, NPUs, quantised inference

The two design choices that matter

11

Block-scaled microformats — MXFP & NVFP4

Below FP8, the dynamic range of any single element is too small to cover a full tensor. The industry's answer: group elements into a small block, give each block a shared exponent.

OCP MX (Microscaling) — September 2023

NVFP4 — Blackwell-specific

Hardware impact

An MXFP MAC is not just a small multiplier. The scale must be fanned out across the block on every cycle and re-applied to the partial sum. Blackwell adds a dedicated scale broadcast network running in parallel with the main dot-product fabric — silicon area you would not see in a pre-2023 design.

12

Stochastic rounding & mixed accumulation

The systematic-bias problem

Round-to-nearest applied many times is biased. Adding 1.0 + 0.49 a million times in FP4 gives 1.0, not 491,490. In low-precision training, this kills convergence.

Stochastic rounding

Two-stage accumulation

Keep the multiplier narrow, keep the accumulator wide:

Without this, every halving of bit-width would silently halve model quality. With it, the chip behaves numerically like a much wider design at a fraction of the area.

13

Real systems — Google TPU lineage

DeviceYearProcessMXU shapeDatatypePeak / chip
TPU v1201528 nm256×256INT8 mul / INT32 acc92 TOPS @ 700 MHz
TPU v22017—128×128 (×2 cores)BF16 mul / FP32 acc~46 TFLOPS BF16
TPU v32018—128×128 (×2)BF16~123 TFLOPS
TPU v420217 nm4× 128×128BF16, INT8275 TFLOPS BF16, OCS interconnect
TPU v5p20235 nm128×128BF16, INT8~459 TFLOPS BF16
TPU v6e Trillium20244 nm256×256BF16, INT8~926 TFLOPS BF16
TPU v7 Ironwood2025——FP8~4.6 PFLOPS FP8

Things to notice

14

Real systems — NVIDIA Tensor Cores

H100 (Hopper, 2022)

  • 4N TSMC, 80 GB HBM3
  • 4th-gen TC, FP8 (E4M3 + E5M2)
  • 1979 TFLOPS FP8 dense / 3958 sparse
  • Transformer Engine: per-tile auto-scaling, FP8 ↔ BF16 paths
  • TMA: async multi-dim DMA into shared mem

B200 (Blackwell, 2024)

  • 4NP TSMC, dual-die, 192 GB HBM3e, 8 TB/s
  • 5th-gen TC + TMEM (256 KB/SM tensor-private mem)
  • 20 PFLOPS FP4 dense / 40 sparse
  • NVFP4: 16-element block, FP8 E4M3 scale
  • FP6 + MXFP added in addition to FP4

What changed across generations

Common confusion

"H100 does 989 TFLOPS" is the FP16/BF16 figure. The FP8 dense figure is 1979 TFLOPS. Many tech press articles cite the BF16 figure when comparing to TPU/AMD FP8 numbers — apples vs oranges.

15

Real systems — the rest of the field

DeviceTopologyDatatypePeakOn-chip mem
Cerebras WSE-3900K cores, full 300 mm waferFP16, BF16, FP8125 PF (sparse FP16)44 GB SRAM, 21 PB/s
Groq LPUFunctional slices, deterministicFP16, INT8188 TF FP16 / 750 TOPS INT8230 MB SRAM, no DRAM
Tesla Dojo D1354 nodes/chip × 25-chip tileFP32, BF16, CFloat89 PF BF16 / training tile11 GB SRAM/tile
AWS Trainium2128×128 systolic Tensor EngineBF16, FP8~667 TFLOPS BF1696 GB HBM3e
Apple M4 ANE16-core, weight-stationaryFP16 native38 TOPS (= FP16 × 2 by convention)shared with SoC
Intel AMX8 tile-regs per coreBF16, INT8 (+FP16 on Granite Rapids)~1024 INT8 MACs/cyc/coreper-core
ARM SME / SME2Streaming-SVE ZA tileFP32, BF16, FP16, INT8 (+2:4 in 2024)VL-dependentper-core
Tenstorrent Blackhole140 Tensix++ coresFP8/16, BFP2/4/8774 TF FP8210 MB SRAM
Etched SohuHardwired transformer dataflowFP8 (vendor-claimed)500K tok/s Llama-70B (claim)144 GB HBM3e

Three philosophies on display

16

Memory hierarchy — what feeds the MMUL

An MMUL is a giant operand-consumer. A 256×256 BF16 array at 1 GHz needs 256 KB/cycle of A operands and 256 KB/cycle of B. That's ~256 TB/s of pure operand bandwidth. No HBM stack delivers that — it has to come from on-die SRAM.

The hierarchy you actually find on a modern AI chip

LevelCapacityBWLatencyEnergy / 32-bit access
PE register / accumulatortens of bytes1 PE/cyc0 cyc~0.1 pJ
Scratchpad SRAM (TMEM, CMEM, GLB)0.1 – 1 MB1 op/cyc/bank1–4 cyc~1 pJ
L2 / Unified Buffer10s of MB (50 MB on H100)~5–10 TB/s20–60 cyc~5 pJ
HBM3 / HBM3etens to hundreds of GB2 – 8 TB/s200–500 cyc~7 pJ/bit (~28 pJ/word)
NVLink / OCS / NeuronLinkoff-package~900 GB/s (NVL5) / per direction~µsmuch higher
Off-rack network (Ethernet/IB)cluster-wide200–800 Gbps/linktens of µshighest

The structures you see on every modern AI chip

17

Are caches useful?

Short answer: not in the CPU sense — and modern accelerators have largely abandoned them in favour of explicitly-managed scratchpads.

Why CPU-style caches don't fit an MMUL

Where caches DO appear on modern AI chips

ChipCache?What it actually is
NVIDIA H10050 MB L2 ✓True coherent cache — needed because GPU runs general CUDA code, not only matmul
NVIDIA B200L2 + TMEM ✓L2 stays; TMEM is a scratchpad, software-managed
Google TPU v1–v6No cacheUnified Buffer + CMEM, both software-managed scratchpads
Cerebras WSE-3No cache44 GB on-die SRAM, all explicitly addressed
Groq LPUNo cache230 MB SRAM, deterministic compiler-scheduled access
Apple ANE / ARM Ethos-NNo cacheTile buffers + DMA

Why GPUs keep an L2

The H100's 50 MB L2 isn't really there to serve the Tensor Core MMUL — it's there to serve the rest of the SM (CUDA cores, texture units, atomics, indirect memory accesses). On dedicated AI chips that don't run general code, the L2 budget gets reallocated as larger user-visible scratchpad.

The system-level implication

Putting an MMUL into a CPU socket (Intel AMX, ARM SME) inherits a CPU's cache hierarchy whether or not it's optimal. Intel AMX tile loads go through L1d/L2/L3 — convenient for programmers, less efficient than a dedicated tile RF would be. AMX is roughly half as energy-efficient per FLOP as a dedicated NPU at the same node, and the cache hierarchy is the main reason.

18

Performance: peak vs sustained

The peak number on the datasheet is the multiplier count × clock frequency × ops per multiplier × 2 (for FMA). Real workloads achieve 30–70% of that.

What eats the gap

Indicative numbers from public benchmarks

WorkloadHardwarePeakSustained%
GPT-3 175B training (BF16)A100312 TFLOPS~150–180 TFLOPS50–58%
Llama-70B training (BF16)H100989 TFLOPS~400–500 TFLOPS40–50%
Llama-70B training (FP8)H1001979 TFLOPS~700–900 TFLOPS35–45%
Llama-70B inference (single-batch)H1001979 TFLOPS~10–30 TFLOPSmemory-bound
GEMM (M=N=K=8192) BF16H100989 TFLOPS~830 TFLOPS~85%

The 5× gap between "MoEFLOPS marketing slide" and "actual throughput on your model" is not bad design — it's the cost of running real workloads on real memory hierarchies. Compiler engineering closes most of it.

19

Power: every joule comes from somewhere

Where the watts go on an H100

Energy per useful operation — order of magnitude

FormatMultiplier energy (45 nm baseline)Roughly at TSMC 4N
FP32 MAC~4.6 pJ~0.9 pJ
BF16 MAC~1.5 pJ~0.3 pJ
FP8 MAC~0.5 pJ~0.1 pJ
FP4 MAC~0.15 pJ~0.03 pJ
HBM3 access (per 32-bit word)~28 pJ~28 pJ (PHY-dominated, weak process scaling)

Two things to notice: (1) HBM access does not scale with process node — it's PHY/wire-dominated. So as compute energy per MAC drops, the relative cost of operand movement becomes higher. (2) The whole point of MX/FP4 is to get the multiplier energy below the operand-movement energy that flanks it.

Package power is the hard limit

Above ~700 W per package, air cooling gives up. Hence the industry-wide pivot to liquid-cooled cabinets in 2024–2026.

20

Thermals & power delivery

A modern MMUL package consumes more current than an early-2000s consumer CPU's whole motherboard. The interesting engineering isn't in the array — it's in getting the joules in and the heat out.

Power delivery network (PDN)

Thermal envelope

CoolingCeilingUsed by
Forced air~400 W packageRTX 6000 Ada workstation, edge boxes
Direct-to-chip liquid~1500 W packageH100/H200 SXM, B200, TPU v4/v5/v6
Two-phase liquid (immersion-style)2 kW+ packagesome hyperscale 2025 deployments
Full immersion3 kW+, rack-density gainexperimental, Microsoft Azure pilots

Hotspot density

A 256×256 array at FP8 dissipates roughly 100 W in ~50 mm² of silicon — that's a flux of ~200 W/cm², several times higher than a kitchen hotplate. Most of the silicon area in a Tensor-Core SM is not multipliers; it's the operand-fan-in network, the result-spill paths, and the local SRAM banks that surround them — partly because this spreads the heat out so hotspot temperatures stay below the 105 °C electromigration limit.

Why every paper has a "thermal" slide now

At 4 nm and below, instantaneous power density (W/cm²) — not average TDP — is the design constraint. A 5th-gen Tensor Core's clock-gating scheme is as carefully tuned as its multiplier topology, because letting the array run at 100% utilisation for too long simply melts the metal stack above it.

21

Tradeoffs & design space

Array size

Format vs accuracy

Specialisation vs generality

Reconfiguration cost

22

Interactive — dataflow comparator

Choose an array geometry and see the resulting cycle count and operand bandwidth for each dataflow.

32
32
128
OS cycles
—
WS cycles
—
DPT cycles
—
Total MACs
—
OS utilisation
—
Operand BW (GB/s @ 1 GHz)
—

Notice how the OS array's utilisation collapses for small K (fill bubble dominates) and the dot-product tree wins on latency but loses on per-cycle bandwidth. Real chips combine both — TPUs use OS-systolic for the MXU, NVIDIA TCs use trees, and both schedule many K-tiles back-to-back to hide the bubble.

23

Interactive — number-format explorer

Pick a format and a real number; see what gets stored.

BF16 FP16 FP8 E4M3 FP8 E5M2 FP4 E2M1 INT8
Stored bits
—
Decoded value
—
Abs error
—
Range
—

Try 1e-4 in FP4 (you'll round to zero) versus BF16 (full precision). Try 50000 in FP8 E4M3 (saturates to ±448) versus FP8 E5M2 (representable). This is exactly why H100 pairs E4M3 (small range, more precision — for forward) with E5M2 (large range — for gradients).

24

Where this is going (2025 → 2027)

Closing thought

The MMUL is the most architecturally settled part of an AI chip. The interesting work in 2025–2027 is not "what shape is the array" — that's been answered — but everything around it: how operands arrive, how partial sums leave, how the joules get in, and how the heat gets out.