NVIDIA GPU Architectures Series — Presentation 25

Silicon & Substrates — Process, Reticle Limits, and Multi-Die Packaging

Behind every NVIDIA GPU is a TSMC fab, an EUV scanner, a CoWoS line, and an absolute reticle limit of about 830 mm² per die. Walk through the process node history, why Blackwell broke through reticle limits with two dies, NV-HBI die-to-die links, CoWoS-S vs CoWoS-L, and what's coming next.

TSMC 4N4NP3nm A16EUVReticle NV-HBICoWoS-SCoWoS-L ChipletBackside power
Wafer → Litho → Reticle → Dies → Interposer → Substrate → Package
00

Topics We'll Cover

The fab-and-package layer of the NVIDIA stack — not because everyone needs to know the litho stepper sequence, but because a 2026 AI engineer who doesn't know what CoWoS-L is can't reason about why H200 was a year-long shortage.

01

NVIDIA's Process Lineage — Pascal to Blackwell

Eight years, four real process node generations, two reticle-limit walls, and one chiplet pivot. The progression of NVIDIA's flagship datacenter dies tracks the limits of TSMC's bleeding-edge logic process more closely than any other product family in the industry.

FamilyYearFoundryNodeTop-die transistorsTop-die area
Pascal GP1002016TSMC16FF15.3 B610 mm²
Volta GV1002017TSMC12FFN21.1 B815 mm²
Turing TU1022018TSMC12FFN18.6 B754 mm²
Ampere GA1002020TSMC7N54.2 B826 mm²
Ada AD1022022TSMC4N76.3 B608 mm²
Hopper GH1002022TSMC4N80 B814 mm²
Blackwell B2002024TSMC4NP~104 B × 2 = 208 B~800 mm² × 2
Read between the lines

Notice GV100, GA100, GH100 all sit at 815, 826, 814 mm² — the reticle ceiling. NVIDIA's flagship strategy for a decade was simply "build the largest die TSMC can print." That approach hit the wall: Blackwell B200 is two ~800 mm² dies stitched together because one die could not be made bigger. The transistor count doubling between Hopper and Blackwell came almost entirely from going dual-die, not from process shrink.

A note on labels

"TSMC 4N" (used by Hopper, Ada, RTX 40-series) is a custom NVIDIA-tuned variant of TSMC's N4P. It is not a different node from N4P — the same transistor density, the same back-end-of-line metal pitches. The differences live in the cell library and physical-implementation flow. Slide 02 unpacks why this label is misleading.

02

What "TSMC 4N" Actually Means

NVIDIA's process labels have become deliberately misleading marketing artefacts. Knowing the real foundry-node alignment matters because it tells you the actual transistor density, power envelope, and (more practically) which AI accelerators are competing for the same wafer slots.

The TSMC node ladder for this era

The headline nodes

  • N5 — HVM 2020. First TSMC EUV-heavy node. Used by Apple A14, M1.
  • N5P — 2021 refresh of N5. ~5% perf or ~10% power.
  • N4 — 2022. Effectively a polished N5; same density.
  • N4P — 2022 refresh of N4. ~6% perf, ~22% efficiency vs N5.
  • N3 (N3B) — HVM 2022/23. First true new node since N5.
  • N3E — 2024. Relaxed mask cost, better yield than N3B.
  • N3P — 2025. Performance refresh of N3E.

What "NVIDIA-customised" means

  • NVIDIA engages TSMC for physical-implementation tweaks: cell library variants, custom metal stack, restricted Vt mix.
  • No actual feature-size shrink — the lithography and BEOL pitches are stock.
  • Yields the gain you'd get from a careful in-house P&R team plus a tuned PDK; not from a node shrink.
  • Mostly enables higher achievable Fmax at thermal envelopes NVIDIA cares about.

The label-vs-reality map

Marketing labelReal TSMC alignmentWho uses it
NVIDIA 4NN5 family, NVIDIA-customised; closer to N4 than to stock N5Hopper GH100, Ada AD102, RTX 40-series
NVIDIA 4NPN4P with NVIDIA-specific physical tweaksBlackwell B100/B200, RTX 50-series
AMD "5 nm"Stock N5MI300A/X compute dies (with N6 IO chiplets)
Apple "3 nm"N3B (M3, A17 Pro), N3E (M3 refresh, M4)Apple silicon
Intel 18AIntel internal node, ~TSMC N2 class with backside power (PowerVia)Lunar Lake successors, Clearwater Forest
Why "4N"?

The label was chosen carefully. 4N sounds adjacent to N4, which sounds newer than N5. In reality NVIDIA Hopper shipped on a customised N5-class process while Apple's contemporaneous M2 Max also ran on N5 — same generation, different marketing. AMD MI300X uses N5 + N6 chiplets; Apple M3/M4 uses N3B/N3E. Naming aside, what matters is who's competing for the same wafer slot at TSMC: in 2024, it was NVIDIA, AMD, Apple, and several hyperscaler ASICs all on N5/N4-class capacity.

03

The Reticle Limit — ~830 mm²

EUV scanners print one reticle field at a time. The reticle is a stencil mask; the scanner steps it across the wafer, exposing one rectangular field per shot. The maximum field size is set by the optics — specifically the lens diameter and the scanner's slit dimensions — and it is roughly 26 mm × 33 mm = 858 mm². Practical usable die area is closer to ~830 mm² once scribe lanes and alignment marks are subtracted.

EUV reticle field — 26 mm × 33 mm = 858 mm² (theoretical max) reticle field outline 26 mm 33 mm GH100 (Hopper) 814 mm² ~95% of usable field scribe lanes + alignment marks consume edges

Why the limit exists

How NVIDIA stacked up against the wall

GP100
610 mm² — 71% of reticle
2016
GV100
815 mm² — at the wall
2017
GA100
826 mm² — at the wall
2020
GH100
814 mm² — at the wall
2022
B200 (one die)
~800 mm² × 2 dies
2024

Volta, Ampere, Hopper sat permanently against the ceiling. Blackwell broke the wall not by building a bigger die but by building two and connecting them — the subject of slide 05.

04

Defect Density & Binning

Wafers carry random defects: lithography particles, mask flaws, etch micro-cracks, contamination. Each defect that lands in a critical area kills the die. Yield versus die area follows a negative-binomial model:

Murphy / negative-binomial yield

Y ≈ (1 + A·D / k)^−k — where A = die area (cm²), D = defect density (defects/cm²), k = clustering parameter (typically 2–5). At TSMC 4N maturity in 2024, D ≈ 0.08/cm². For an 800 mm² (8 cm²) die, base yield without redundancy is ~50%.

Why GH100 ships as multiple SKUs

GH100 has 144 SMs physical on the silicon. NVIDIA never ships all 144 enabled — they bin parts based on which SMs failed test, then disable failures and clusters of failures to meet target SM counts. The result is a SKU ladder out of identical silicon:

SKUSMs enabledHBM stacksNotes
GH100 die (full)144 (physical)6 (physical)Never shipped at full count
H100 SXM51325 active (80 GB HBM3)Top-bin shipped product; ~92% SM yield
H100 PCIe1145 active (80 GB)Lower-bin parts; lower clocks; PCIe form factor
H8001325 activeExport-restricted China SKU; capped NVLink BW
H20785 active (96 GB HBM3)China-market further-cut SKU; FP64 throttled

Where the redundancy hides

Fictional yield-arithmetic illustration (units: cm², defects/cm²)
# Murphy model with k=3 (typical TSMC mature node)
def yield_pct(area_cm2, D, k=3):
    return (1 + area_cm2 * D / k) ** (-k) * 100

# GH100 814 mm^2 = 8.14 cm^2 at TSMC 4N maturity (D~0.08)
yield_pct(8.14, 0.08) # ~ 50% raw die yield

# With ~10% spare SMs giving recovery on isolated SM defects:
yield_pct(8.14, 0.08 * 0.55) # ~ 70% effective yield after binning

# Compare AD102 608 mm^2 at same D — consumer Ada flagship
yield_pct(6.08, 0.08) # ~ 60% raw — smaller is forgiving
The economic punchline

A GH100 wafer has roughly 65 candidate dies. After yield, ~30–45 are sellable. NVIDIA's wafer cost from TSMC for 4N has been quoted at $16–20k. That works out to roughly $400–600 of silicon per H100 — far less than retail. The rest of the BoM (HBM stacks, packaging, test) and gross margin make the rest. This is why a defective die that bins into H800 or H20 still earns money.

05

Blackwell's Dual-Die Solution — NV-HBI

When the reticle ceiling stopped rising, NVIDIA stopped trying. The Blackwell B200 is two reticle-sized dies, side by side on a single CoWoS-L interposer, connected edge-to-edge by a high-bandwidth silicon bridge that NVIDIA calls NV-HBI — NVIDIA High-Bandwidth Interface.

What NV-HBI delivers

Bandwidth

10 TB/s bidirectional across the die-to-die interface — about 2× what HBM3e delivers per stack. That is enough for cache-coherent traffic, L2 sharing, and CUDA kernel migration without observable cost.

Coherence

Software sees one GPU: one CUDA context, one device ID, one unified L2 cache (the two dies' L2 partitions are kept coherent). The driver does not split workloads across dies; the hardware presents a single SM array.

Physical implementation

Point-to-point silicon bridges (LSI — Local Silicon Interconnect) under the die-to-die boundary. Not a full silicon interposer between the dies — just a strip of bridges where the high-density traces need to land.

Why bridges, not full interposer

A full silicon interposer big enough for two reticle-sized dies plus 8 HBM stacks would be enormous (>1500 mm²), costly, and capacity-bound at TSMC. CoWoS-L with strategic bridges keeps the cost trajectory rational.

B200 package cross-section

B200 package — cross-section (schematic, not to scale) BGA balls organic substrate (BT/ABF) — routes to motherboard CoWoS-L mixed substrate — organic + silicon bridge tiles LSI bridge (NV-HBI: 10 TB/s) HBM3e 24 GB HBM3e HBM3e HBM3e 8× HBM3e stacks — 192 GB total, ~8 TB/s aggregate Die 0 ~800 mm² ~104 B tx Die 1 ~800 mm² ~104 B tx NV-HBI seam From software: one GPU, one CUDA context, unified L2 across both dies From hardware: two reticle-limit dies + LSI bridge + 8 HBM3e stacks on CoWoS-L
Why this matters above the silicon layer

The dual-die architecture is invisible to CUDA but visible in NCCL all-reduce timing and in memory access latency variance. A kernel that touches HBM stacks on the "far" die from the SM scheduling it pays an extra few hundred picoseconds and crosses NV-HBI. Aggregate effect is small but measurable on memory-bound workloads — expect roughly 1–3% delta on workloads that don't fit in L2 vs idealised single-die.

06

CoWoS — TSMC's Advanced Packaging Family

CoWoS (Chip-on-Wafer-on-Substrate) is TSMC's umbrella term for advanced packaging that puts logic dies and HBM stacks on a shared interposer over an organic substrate. There are three production variants in 2026, with very different cost / capacity / bandwidth trade-offs.

CoWoS-S (Silicon)

The original. A single large piece of silicon interposer sits below all the dies, routed with TSV grids and fine-pitch metal. Highest bandwidth per mm, highest cost.

  • Limit: interposer area capped at ~3× reticle (the interposer itself is patterned with photolithography).
  • Used by: H100, H200, AMD MI250X, MI300X.
  • Cost: highest of the three; capacity-constrained at TSMC.

CoWoS-L (LSI bridges)

The 2024+ workhorse. Replace most of the silicon interposer with cheap organic substrate; place small silicon bridges (LSI) only where you need ultra-fast die-to-die or die-to-HBM links.

  • Why it scales: organic is cheap and unbounded in area. Silicon is used only where the BW demand actually requires it.
  • Used by: Blackwell B200, Apple Vision Pro M2, Intel Falcon Shores.
  • Cost: lower than CoWoS-S at equivalent BW; better capacity scaling for 2026+.

CoWoS-R (RDL)

The cheapest variant. Uses an organic-only substrate with a high-density redistribution layer (RDL); no silicon bridges or interposer at all. Lower BW than -S or -L.

  • BW ceiling: a fraction of -S; suits chiplets that can tolerate narrower die-to-die links.
  • Used by: AMD client/server CPUs, mid-tier accelerators.
  • Cost: the cheap end — commodity-ish organic packaging.

Comparing the three at a glance

VariantInterposer typeDie-to-die BWMax areaCost classWhere it fits
CoWoS-SSilicon, fullHighest (TBs/s)~3× reticleHighestH100, H200, MI300X — max BW priority
CoWoS-LOrganic + silicon bridgesHigh (TBs/s on bridge)Effectively unboundedMid-highB200, Rubin — multi-die scaling
CoWoS-ROrganic + RDLModest (100s GB/s)Effectively unboundedLowestClient CPU chiplets, narrow-BW accelerators
The trade-off in one line

CoWoS-S delivers the highest BW per mm but is capacity-limited at TSMC because the interposer itself burns leading-node photolithography. CoWoS-L scales much better because most of the interposer is organic; you only burn silicon where you actually need the bandwidth. That is why every multi-die GPU launching from 2024 onward uses CoWoS-L, not CoWoS-S.

07

What Limits Capacity Right Now

As of 2026 the bottleneck on AI GPU supply is not silicon wafer capacity. TSMC has plenty of N4/N5-class wafer slots; HBM is also well-sourced from SK hynix, Micron, and (slowly) Samsung. The actual bottleneck is CoWoS line capacity — the advanced packaging step that ties dies and HBM together onto interposers.

The 2026 CoWoS picture

TSMC sites

  • Phoenix (AZ): packaging line under construction; HVM 2026/27.
  • Tainan (TW): the original CoWoS-S line; capacity expanded ~3× since 2023.
  • Taoyuan (TW): CoWoS-L line; the priority expansion 2024–26.
  • Zhunan (TW): additional advanced packaging fab brought online.

Other packagers

  • SK hynix: in-house TSV stacking + some advanced packaging for HBM-attached dies.
  • Samsung: I-Cube (silicon interposer) and X-Cube (3D vertical stack); used in some hyperscaler ASICs.
  • Amkor / ASE: OSAT capacity for less demanding RDL packaging.

Capacity figures (approximate, 2026)

2023 Q1
~10k WPM
Pre-boom
2024 Q1
~15k WPM
Doubling
2025 Q1
~22k WPM
Easing
2026 Q1
~30k WPM
Now
2027 (proj.)
~55–60k WPM (projected)
Near-double

Wafer-equivalents/month is an aggregated figure across CoWoS-S and CoWoS-L; the split is roughly 50/50 in 2026 and shifts toward L thereafter. One CoWoS wafer carries multiple GPU packages, but the carrier-wafer is the rate-limiting accounting unit.

Who's competing for the slot

The downstream consequence

Even if NVIDIA wanted to ship 5× as many H200s, they couldn't — CoWoS slots are pre-allocated, multi-quarter, and rationed across customers. That is why H100 was scarce through 2023–24 despite ample wafer supply. The packaging line, not the fab, sets the global ceiling on AI GPU shipments.

08

Coming Next — N3, A16, Backside Power

Three transitions queue up over the next 18–36 months: the move to N3-family for compute dies, the introduction of backside power delivery on TSMC A16 (and Intel 18A), and a fresh round of die-stacking on top of all of it.

N3 family

What N3 / N3E / N3P delivers

  • ~70% transistor density vs N5 (full library; SRAM scales much less).
  • ~30% lower power at iso-perf, or ~15% faster at iso-power.
  • N3B (Apple A17 Pro, M3) used aggressive EUV multi-patterning — expensive masks.
  • N3E relaxed several layers; cheaper masks, better yield. The volume node.
  • N3P (2025+): refresh of N3E with improved Fmax and slight density gain.

Where NVIDIA fits in

  • Rubin (2026): reportedly N3 chiplets in dual-die config, similar to Blackwell topology but with the smaller area / lower power that N3 unlocks.
  • RTX consumer (2026/27): consumer Blackwell-successor cards likely on N4P initially; N3 transition follows datacenter by ~12 months as usual.
  • Cost: N3 wafer is ~30% pricier than N4P at HVM. NVIDIA absorbs it because the perf-per-watt win is large.

A16 and backside power delivery

TSMC's A16 (announced for HVM ~2026) is the "1.6 nm" class node. Beyond the usual density and power gains it introduces something architecturally new: Super Power Rail, TSMC's name for backside power delivery network (BPDN, sometimes BSPDN).

Front-side stack (traditional) vs front+back split (A16 / Intel 18A) Traditional (N5 / N3): all metal on front side M10–M14 wide signal & power M5–M9 mid-pitch M1–M4 fine pitch + power tap-down Silicon (transistors) power & signal share metal stack A16 / 18A: power moves to back side front: signal-only metal stack tighter pitch, cleaner routing no power-rail compromise Silicon (transistors) backside power network (BSPDN) power delivered from below

What backside power delivery buys you

Roadmap snapshot

NodeYearFoundryKey featureLikely first NVIDIA part
N3 / N3E / N3P2024–26TSMC~70% density vs N5; no BSPDNRubin (datacenter, 2026)
A16 ("1.6 nm")2026–27TSMCSuper Power Rail (BSPDN)Rubin Ultra / next-gen flagship
Intel 18A2025–26IntelPowerVia BSPDN + RibbonFET (GAAFET)NVIDIA may diversify; not committed
N2 / N2P2025–27TSMCGAAFET (nanosheet); BSPDN on N2PRubin Ultra successor
09

Chiplet Trends — Industry Convergence

Multi-die GPUs are no longer the exception. Across the AI accelerator market the dominant 2025+ design pattern is "multiple compute dies on an advanced package, connected by a proprietary or open die-to-die fabric." Here is the lay of the land.

ProductVendorChiplet topologyDie-to-die fabric
B200NVIDIA2× compute diesNV-HBI (proprietary, 10 TB/s)
GB200 NVL72NVIDIA2× B200 GPUs + 1× Grace CPU per "superchip"; 36 superchips per rackNV-HBI within GPU; NVLink-C2C to Grace; NVLink 5 inter-GPU; NVSwitch fabric
MI300XAMD8× GPU compute dies (XCDs) + 4× IO diesInfinityFabric over silicon interposer (CoWoS-S)
MI325X / MI350XAMDSame XCD topology; updated HBM3eInfinityFabric / future UCIe
Falcon ShoresIntelTile-based; CPU + GPU tilesEMIB bridges (Intel's CoWoS-L equivalent)
Gaudi 3Intel2× compute tilesEMIB
M3 UltraApple2× M3 MaxUltraFusion (>2.5 TB/s)
TPU v5e/v5pGoogleSingle die, but pod-level torus is the unitOCS optical inter-chip switching

Why the convergence

NV-HBI vs UCIe — the standard war

NV-HBI (NVIDIA proprietary)

  • Targeted, optimised for NVIDIA's architecture.
  • 10 TB/s on B200; tight L2-coherent traffic patterns.
  • Locks in a single vendor — no third-party chiplets.
  • Comparable in spirit to AMD's InfinityFabric-on-package.

UCIe (open standard)

  • Universal Chiplet Interconnect Express; backed by Intel, AMD, Arm, Qualcomm, Samsung, TSMC.
  • Defines PHY + protocol layers for chiplet interop.
  • 2.0 spec (2024): up to 64 GT/s, 4 TB/s/mm shoreline density.
  • Goal: a chiplet marketplace where you can mix vendors. Reality: still mostly intra-vendor.
The underlying truth

The industry isn't really converging on chiplets — it has already converged. Every flagship 2025 AI accelerator is multi-die. The remaining argument is over which die-to-die fabric. NVIDIA, AMD, and Apple use proprietary fabrics; the rest of the market rallies around UCIe. Expect that pattern to hold through this decade: incumbents keep their fabrics, challengers standardise to compete.

10

What This Means for AI Engineers

You don't have to know which lithography stepper carved your H100 to use it. But the packaging-and-process layer leaks upward into observable behaviours that any platform engineer or LLM-serving operator will hit.

(a) Launches are gated by packaging, not litho

From 2023 onwards every major NVIDIA launch (H100, H200, B200) has been constrained by CoWoS allocation, not by silicon supply. Estimating ramp rates means tracking TSMC packaging investments, not just wafer starts.

(b) SKU segmentation is a yield artefact

H100 SXM5 / H100 PCIe / H800 / H20 are the same GH100 silicon at different bin tiers. When you specify hardware for a workload, choose by SM count + memory + clock, not by "which generation" — the underlying die may be identical.

(c) Multi-die GPUs introduce NUMA-like effects

Blackwell B200 presents as one GPU but accesses to "far" HBM across NV-HBI cost slightly more. CUDA hides this; profilers don't always. For latency-sensitive kernels you may see jitter that maps to die boundaries.

(d) CoWoS allocation drives product mix

Why was H100 a year scarce? Why did H200 ship roughly when it did? Both decisions came down to packaging slot allocation between H-class and B-class lines. The answer to "when can I get 1000 H200s" lives at TSMC's CoWoS coordinator's desk, not in NVIDIA's marketing schedule.

(e) Datacenter leads consumer by ~12 months

Foundry-tier process is expensive: TSMC 4N for H100 (2022) shipped on consumer cards (RTX 4090, also 4N) the same year, but datacenter capacity is consumed first. B200 (4NP, 2024) ships consumer (RTX 5090, 4NP) in 2025. Plan accordingly.

(f) FP8 / FP4 ride process gains too

Tensor-core support for FP8 (Hopper) and FP4 (Blackwell) isn't free silicon — the hardware runs because the process node delivered the density to spend on dedicated tensor pipes. Future MX-FP4 / MX-FP6 viability is largely a process-node story.

The takeaway in one sentence

The cadence of the AI accelerator market is the cadence of TSMC's advanced packaging line; everything else — FP4 support, dual-die GPUs, multi-quarter lead times — is downstream of that one fact.

11

Interactive: Process & Packaging Picker

Pick a workload class, target year, and unit volume. The picker estimates the foundry node, packaging type, die area, transistor budget, and approximate cost band — plus where the supply-chain bottleneck lands.

10k
Foundry node
—
Packaging
—
Die area (mm²)
—
Transistors (B)
—
Wafer cost (USD)
—
Retail band
—