Behind every NVIDIA GPU is a TSMC fab, an EUV scanner, a CoWoS line, and an absolute reticle limit of about 830 mm² per die. Walk through the process node history, why Blackwell broke through reticle limits with two dies, NV-HBI die-to-die links, CoWoS-S vs CoWoS-L, and what's coming next.
The fab-and-package layer of the NVIDIA stack — not because everyone needs to know the litho stepper sequence, but because a 2026 AI engineer who doesn't know what CoWoS-L is can't reason about why H200 was a year-long shortage.
Eight years, four real process node generations, two reticle-limit walls, and one chiplet pivot. The progression of NVIDIA's flagship datacenter dies tracks the limits of TSMC's bleeding-edge logic process more closely than any other product family in the industry.
| Family | Year | Foundry | Node | Top-die transistors | Top-die area |
|---|---|---|---|---|---|
| Pascal GP100 | 2016 | TSMC | 16FF | 15.3 B | 610 mm² |
| Volta GV100 | 2017 | TSMC | 12FFN | 21.1 B | 815 mm² |
| Turing TU102 | 2018 | TSMC | 12FFN | 18.6 B | 754 mm² |
| Ampere GA100 | 2020 | TSMC | 7N | 54.2 B | 826 mm² |
| Ada AD102 | 2022 | TSMC | 4N | 76.3 B | 608 mm² |
| Hopper GH100 | 2022 | TSMC | 4N | 80 B | 814 mm² |
| Blackwell B200 | 2024 | TSMC | 4NP | ~104 B × 2 = 208 B | ~800 mm² × 2 |
Notice GV100, GA100, GH100 all sit at 815, 826, 814 mm² — the reticle ceiling. NVIDIA's flagship strategy for a decade was simply "build the largest die TSMC can print." That approach hit the wall: Blackwell B200 is two ~800 mm² dies stitched together because one die could not be made bigger. The transistor count doubling between Hopper and Blackwell came almost entirely from going dual-die, not from process shrink.
"TSMC 4N" (used by Hopper, Ada, RTX 40-series) is a custom NVIDIA-tuned variant of TSMC's N4P. It is not a different node from N4P — the same transistor density, the same back-end-of-line metal pitches. The differences live in the cell library and physical-implementation flow. Slide 02 unpacks why this label is misleading.
NVIDIA's process labels have become deliberately misleading marketing artefacts. Knowing the real foundry-node alignment matters because it tells you the actual transistor density, power envelope, and (more practically) which AI accelerators are competing for the same wafer slots.
| Marketing label | Real TSMC alignment | Who uses it |
|---|---|---|
| NVIDIA 4N | N5 family, NVIDIA-customised; closer to N4 than to stock N5 | Hopper GH100, Ada AD102, RTX 40-series |
| NVIDIA 4NP | N4P with NVIDIA-specific physical tweaks | Blackwell B100/B200, RTX 50-series |
| AMD "5 nm" | Stock N5 | MI300A/X compute dies (with N6 IO chiplets) |
| Apple "3 nm" | N3B (M3, A17 Pro), N3E (M3 refresh, M4) | Apple silicon |
| Intel 18A | Intel internal node, ~TSMC N2 class with backside power (PowerVia) | Lunar Lake successors, Clearwater Forest |
The label was chosen carefully. 4N sounds adjacent to N4, which sounds newer than N5. In reality NVIDIA Hopper shipped on a customised N5-class process while Apple's contemporaneous M2 Max also ran on N5 — same generation, different marketing. AMD MI300X uses N5 + N6 chiplets; Apple M3/M4 uses N3B/N3E. Naming aside, what matters is who's competing for the same wafer slot at TSMC: in 2024, it was NVIDIA, AMD, Apple, and several hyperscaler ASICs all on N5/N4-class capacity.
EUV scanners print one reticle field at a time. The reticle is a stencil mask; the scanner steps it across the wafer, exposing one rectangular field per shot. The maximum field size is set by the optics — specifically the lens diameter and the scanner's slit dimensions — and it is roughly 26 mm × 33 mm = 858 mm². Practical usable die area is closer to ~830 mm² once scribe lanes and alignment marks are subtracted.
Volta, Ampere, Hopper sat permanently against the ceiling. Blackwell broke the wall not by building a bigger die but by building two and connecting them — the subject of slide 05.
Wafers carry random defects: lithography particles, mask flaws, etch micro-cracks, contamination. Each defect that lands in a critical area kills the die. Yield versus die area follows a negative-binomial model:
Y ≈ (1 + A·D / k)^−k — where A = die area (cm²), D = defect density (defects/cm²), k = clustering parameter (typically 2–5). At TSMC 4N maturity in 2024, D ≈ 0.08/cm². For an 800 mm² (8 cm²) die, base yield without redundancy is ~50%.
GH100 has 144 SMs physical on the silicon. NVIDIA never ships all 144 enabled — they bin parts based on which SMs failed test, then disable failures and clusters of failures to meet target SM counts. The result is a SKU ladder out of identical silicon:
| SKU | SMs enabled | HBM stacks | Notes |
|---|---|---|---|
| GH100 die (full) | 144 (physical) | 6 (physical) | Never shipped at full count |
| H100 SXM5 | 132 | 5 active (80 GB HBM3) | Top-bin shipped product; ~92% SM yield |
| H100 PCIe | 114 | 5 active (80 GB) | Lower-bin parts; lower clocks; PCIe form factor |
| H800 | 132 | 5 active | Export-restricted China SKU; capped NVLink BW |
| H20 | 78 | 5 active (96 GB HBM3) | China-market further-cut SKU; FP64 throttled |
# Murphy model with k=3 (typical TSMC mature node)
def yield_pct(area_cm2, D, k=3):
return (1 + area_cm2 * D / k) ** (-k) * 100
# GH100 814 mm^2 = 8.14 cm^2 at TSMC 4N maturity (D~0.08)
yield_pct(8.14, 0.08) # ~ 50% raw die yield
# With ~10% spare SMs giving recovery on isolated SM defects:
yield_pct(8.14, 0.08 * 0.55) # ~ 70% effective yield after binning
# Compare AD102 608 mm^2 at same D — consumer Ada flagship
yield_pct(6.08, 0.08) # ~ 60% raw — smaller is forgiving
A GH100 wafer has roughly 65 candidate dies. After yield, ~30–45 are sellable. NVIDIA's wafer cost from TSMC for 4N has been quoted at $16–20k. That works out to roughly $400–600 of silicon per H100 — far less than retail. The rest of the BoM (HBM stacks, packaging, test) and gross margin make the rest. This is why a defective die that bins into H800 or H20 still earns money.
When the reticle ceiling stopped rising, NVIDIA stopped trying. The Blackwell B200 is two reticle-sized dies, side by side on a single CoWoS-L interposer, connected edge-to-edge by a high-bandwidth silicon bridge that NVIDIA calls NV-HBI — NVIDIA High-Bandwidth Interface.
10 TB/s bidirectional across the die-to-die interface — about 2× what HBM3e delivers per stack. That is enough for cache-coherent traffic, L2 sharing, and CUDA kernel migration without observable cost.
Software sees one GPU: one CUDA context, one device ID, one unified L2 cache (the two dies' L2 partitions are kept coherent). The driver does not split workloads across dies; the hardware presents a single SM array.
Point-to-point silicon bridges (LSI — Local Silicon Interconnect) under the die-to-die boundary. Not a full silicon interposer between the dies — just a strip of bridges where the high-density traces need to land.
A full silicon interposer big enough for two reticle-sized dies plus 8 HBM stacks would be enormous (>1500 mm²), costly, and capacity-bound at TSMC. CoWoS-L with strategic bridges keeps the cost trajectory rational.
The dual-die architecture is invisible to CUDA but visible in NCCL all-reduce timing and in memory access latency variance. A kernel that touches HBM stacks on the "far" die from the SM scheduling it pays an extra few hundred picoseconds and crosses NV-HBI. Aggregate effect is small but measurable on memory-bound workloads — expect roughly 1–3% delta on workloads that don't fit in L2 vs idealised single-die.
CoWoS (Chip-on-Wafer-on-Substrate) is TSMC's umbrella term for advanced packaging that puts logic dies and HBM stacks on a shared interposer over an organic substrate. There are three production variants in 2026, with very different cost / capacity / bandwidth trade-offs.
The original. A single large piece of silicon interposer sits below all the dies, routed with TSV grids and fine-pitch metal. Highest bandwidth per mm, highest cost.
The 2024+ workhorse. Replace most of the silicon interposer with cheap organic substrate; place small silicon bridges (LSI) only where you need ultra-fast die-to-die or die-to-HBM links.
The cheapest variant. Uses an organic-only substrate with a high-density redistribution layer (RDL); no silicon bridges or interposer at all. Lower BW than -S or -L.
| Variant | Interposer type | Die-to-die BW | Max area | Cost class | Where it fits |
|---|---|---|---|---|---|
| CoWoS-S | Silicon, full | Highest (TBs/s) | ~3× reticle | Highest | H100, H200, MI300X — max BW priority |
| CoWoS-L | Organic + silicon bridges | High (TBs/s on bridge) | Effectively unbounded | Mid-high | B200, Rubin — multi-die scaling |
| CoWoS-R | Organic + RDL | Modest (100s GB/s) | Effectively unbounded | Lowest | Client CPU chiplets, narrow-BW accelerators |
CoWoS-S delivers the highest BW per mm but is capacity-limited at TSMC because the interposer itself burns leading-node photolithography. CoWoS-L scales much better because most of the interposer is organic; you only burn silicon where you actually need the bandwidth. That is why every multi-die GPU launching from 2024 onward uses CoWoS-L, not CoWoS-S.
As of 2026 the bottleneck on AI GPU supply is not silicon wafer capacity. TSMC has plenty of N4/N5-class wafer slots; HBM is also well-sourced from SK hynix, Micron, and (slowly) Samsung. The actual bottleneck is CoWoS line capacity — the advanced packaging step that ties dies and HBM together onto interposers.
Wafer-equivalents/month is an aggregated figure across CoWoS-S and CoWoS-L; the split is roughly 50/50 in 2026 and shifts toward L thereafter. One CoWoS wafer carries multiple GPU packages, but the carrier-wafer is the rate-limiting accounting unit.
Even if NVIDIA wanted to ship 5× as many H200s, they couldn't — CoWoS slots are pre-allocated, multi-quarter, and rationed across customers. That is why H100 was scarce through 2023–24 despite ample wafer supply. The packaging line, not the fab, sets the global ceiling on AI GPU shipments.
Three transitions queue up over the next 18–36 months: the move to N3-family for compute dies, the introduction of backside power delivery on TSMC A16 (and Intel 18A), and a fresh round of die-stacking on top of all of it.
TSMC's A16 (announced for HVM ~2026) is the "1.6 nm" class node. Beyond the usual density and power gains it introduces something architecturally new: Super Power Rail, TSMC's name for backside power delivery network (BPDN, sometimes BSPDN).
| Node | Year | Foundry | Key feature | Likely first NVIDIA part |
|---|---|---|---|---|
| N3 / N3E / N3P | 2024–26 | TSMC | ~70% density vs N5; no BSPDN | Rubin (datacenter, 2026) |
| A16 ("1.6 nm") | 2026–27 | TSMC | Super Power Rail (BSPDN) | Rubin Ultra / next-gen flagship |
| Intel 18A | 2025–26 | Intel | PowerVia BSPDN + RibbonFET (GAAFET) | NVIDIA may diversify; not committed |
| N2 / N2P | 2025–27 | TSMC | GAAFET (nanosheet); BSPDN on N2P | Rubin Ultra successor |
Multi-die GPUs are no longer the exception. Across the AI accelerator market the dominant 2025+ design pattern is "multiple compute dies on an advanced package, connected by a proprietary or open die-to-die fabric." Here is the lay of the land.
| Product | Vendor | Chiplet topology | Die-to-die fabric |
|---|---|---|---|
| B200 | NVIDIA | 2× compute dies | NV-HBI (proprietary, 10 TB/s) |
| GB200 NVL72 | NVIDIA | 2× B200 GPUs + 1× Grace CPU per "superchip"; 36 superchips per rack | NV-HBI within GPU; NVLink-C2C to Grace; NVLink 5 inter-GPU; NVSwitch fabric |
| MI300X | AMD | 8× GPU compute dies (XCDs) + 4× IO dies | InfinityFabric over silicon interposer (CoWoS-S) |
| MI325X / MI350X | AMD | Same XCD topology; updated HBM3e | InfinityFabric / future UCIe |
| Falcon Shores | Intel | Tile-based; CPU + GPU tiles | EMIB bridges (Intel's CoWoS-L equivalent) |
| Gaudi 3 | Intel | 2× compute tiles | EMIB |
| M3 Ultra | Apple | 2× M3 Max | UltraFusion (>2.5 TB/s) |
| TPU v5e/v5p | Single die, but pod-level torus is the unit | OCS optical inter-chip switching |
The industry isn't really converging on chiplets — it has already converged. Every flagship 2025 AI accelerator is multi-die. The remaining argument is over which die-to-die fabric. NVIDIA, AMD, and Apple use proprietary fabrics; the rest of the market rallies around UCIe. Expect that pattern to hold through this decade: incumbents keep their fabrics, challengers standardise to compete.
You don't have to know which lithography stepper carved your H100 to use it. But the packaging-and-process layer leaks upward into observable behaviours that any platform engineer or LLM-serving operator will hit.
From 2023 onwards every major NVIDIA launch (H100, H200, B200) has been constrained by CoWoS allocation, not by silicon supply. Estimating ramp rates means tracking TSMC packaging investments, not just wafer starts.
H100 SXM5 / H100 PCIe / H800 / H20 are the same GH100 silicon at different bin tiers. When you specify hardware for a workload, choose by SM count + memory + clock, not by "which generation" — the underlying die may be identical.
Blackwell B200 presents as one GPU but accesses to "far" HBM across NV-HBI cost slightly more. CUDA hides this; profilers don't always. For latency-sensitive kernels you may see jitter that maps to die boundaries.
Why was H100 a year scarce? Why did H200 ship roughly when it did? Both decisions came down to packaging slot allocation between H-class and B-class lines. The answer to "when can I get 1000 H200s" lives at TSMC's CoWoS coordinator's desk, not in NVIDIA's marketing schedule.
Foundry-tier process is expensive: TSMC 4N for H100 (2022) shipped on consumer cards (RTX 4090, also 4N) the same year, but datacenter capacity is consumed first. B200 (4NP, 2024) ships consumer (RTX 5090, 4NP) in 2025. Plan accordingly.
Tensor-core support for FP8 (Hopper) and FP4 (Blackwell) isn't free silicon — the hardware runs because the process node delivered the density to spend on dedicated tensor pipes. Future MX-FP4 / MX-FP6 viability is largely a process-node story.
The cadence of the AI accelerator market is the cadence of TSMC's advanced packaging line; everything else — FP4 support, dual-die GPUs, multi-quarter lead times — is downstream of that one fact.
Pick a workload class, target year, and unit volume. The picker estimates the foundry node, packaging type, die area, transistor budget, and approximate cost band — plus where the supply-chain bottleneck lands.