NVIDIA GPU Architectures Series — Presentation 01

The NVIDIA GPU Family Tree — Pascal to Blackwell

A grounded tour of every NVIDIA GPU family that matters for AI today: when each landed, what changed, and why each generation reshaped what we could train and serve.

PascalVoltaTuring AmpereAdaHopper BlackwellRubin DatacenterConsumer
Pascal → Volta → Turing → Ampere → Ada → Hopper → Blackwell → Rubin
00

Topics We'll Cover

Ten years of NVIDIA architectures, told as a single story: how each generation enabled the next class of model, and why the software-visible jumps (FP16 tensor cores, FP8, FP4, NVL72) matter more than the raw transistor count.

01

Why GPU Architecture Matters Now

Every public LLM milestone of the last seven years sat on top of a specific NVIDIA generation. The headline model didn't appear because someone had a clever idea in isolation — it appeared because a tensor-core or memory advance made training and serving that class of model affordable. The software follows the silicon by 12 to 24 months.

Model eraYearEnabling GPUWhat the silicon unlocked
BERT-era models, ResNet-50 at scale2018–19V100 (Volta; Google's own BERT ran on TPUs)First-gen tensor cores: FP16 mixed-precision training became standard.
GPT-2 / T5 / early multimodal2019–20V100, then A10032 GB HBM2 on V100; A100's 40/80 GB and 2:4 sparsity opened the 10B-parameter regime.
OPT-175B, BLOOM, LLaMA-12020–22A100 (Ampere)BF16 + TF32 + 3rd-gen tensor cores + NVLink 3 made dense 100B+ training routine.
GPT-4-class, Llama-2 70B2023–24H100 (Hopper)FP8 (Transformer Engine) doubled effective throughput; TMA cut prefill latency.
Frontier 2025 (GPT-4o, Claude 3.7, Llama-3.1 405B)2024–25H100, H200, early B200HBM3e at 4.8 TB/s on H200 turned long-context decode from bandwidth-starved to compute-bound.
Frontier 2026 (next-gen MoE, very-long-context agents)2025–26B200, GB200 NVL72FP4 doubles inference again; NVL72 makes a 72-GPU domain look like one machine for KV-heavy MoE.
The pattern

For each generation the cycle is the same: NVIDIA ships a new precision (FP16 → BF16 → TF32 → FP8 → FP4) along with a memory bandwidth uplift, the framework teams catch up over six to twelve months, and within a year the new class of model the silicon enabled is the public state of the art. You can't read the LLM roadmap without reading the GPU roadmap.

02

The Family Lineage — Pascal to Rubin

Eight generations in ten years. Each row below is one architecture family, with the top "datacenter" die that defined it for AI workloads. Consumer SKUs share the architecture name but use different (usually smaller, cheaper) dies and memory.

FamilyCodenameYearProcessTop dieTransistorsTop SKU
PascalGP1002016TSMC 16FFGP10015.3BTesla P100
VoltaGV1002017TSMC 12FFGV10021.1BTesla V100
TuringTU1022018TSMC 12FFTU10218.6BRTX 2080 Ti / T4 (TU104)
AmpereGA100 / GA10x2020TSMC 7N (GA100) / Samsung 8N (GA10x)GA10054.2BA100 80GB
Ada LovelaceAD1022022TSMC 4NAD10276.3BRTX 4090 / L40S
HopperGH1002022TSMC 4NGH10080BH100 / H200
BlackwellGB100 / B2002024TSMC 4NP (dual-die)2× reticle-limited~208B (2 dies)B200 / GB200
Rubin (planned)R1002026+TSMC 3nmTBATBARubin GPU + Vera CPU

Naming intuition

Why Turing matters in the lineage

Turing is often skipped in AI write-ups because GV100 had tensor cores first and GA100 had the first really useful ones. But Turing's TU104-based T4 was the inference workhorse of 2019–2022 — the first card cloud providers deployed at scale specifically for serving. INT8 / INT4 tensor cores and a 70 W TDP made it the default inference SKU long before anyone had heard of vLLM.

03

Process Node & Transistor Budget

Transistor counts climbed from 15 billion (P100, 2016) to 208 billion (B200, 2024) — about 14× in eight years. That's faster than Moore's law on paper, but only because Blackwell broke the reticle limit by bonding two dies together. The single-die curve flattens hard around Hopper.

Transistor count of top die per family (billions) 50B 100B 150B 200B Pascal GP100 (2016) 15.3B Volta GV100 (2017) 21.1B Turing TU102 (2018) 18.6B Ampere GA100 (2020) 54.2B Ada AD102 (2022) 76.3B Hopper GH100 (2022) 80B Blackwell B200 (2024) ~208B (2 dies) Rubin R100 (2026, planned) TBA

What the chart actually shows

The architectural takeaway

Moore's-law slowdown is the reason every modern NVIDIA generation is also an arithmetic innovation. You can't double FLOPS by doubling transistors any more, so you double them by halving precision: FP32 → FP16 (Volta) → BF16/TF32 (Ampere) → FP8 (Hopper) → FP4 (Blackwell). The die can't grow; the numbers shrink.

04

Die → Package → Board → System

"H100" can mean four different things depending on context: the silicon die (GH100), the package it's mounted on (SXM5 vs PCIe vs MGX), the baseboard that ties eight of them together (HGX), or the integrated system the OEM ships (DGX, OEM HGX, GB200 rack). Confusing them is the most common source of pricing and capability errors.

1. Chip (the die)

One piece of silicon, e.g. GH100 for Hopper or the dual-die B200 for Blackwell. This is what fabs ship. Defines compute capability, tensor-core formats, max HBM stacks. Yields and binning here drive the entire SKU stack.

2. Package (form factor)

The die plus HBM stacks on a substrate. Options: SXM5 (mezzanine, NVLink-native, 700 W), PCIe (drop-in card, no NVSwitch, 350 W), MGX (modular reference platforms), OAM (industry-standard module). Same die, very different system options.

3. Board (HGX baseboard)

The HGX H100/B200 8-GPU baseboard mounts eight SXM modules plus four NVSwitch chips. It is the building block OEMs (Supermicro, Dell, HPE, Lenovo) buy from NVIDIA and integrate into chassis. Same baseboard, dozens of OEM SKUs.

4. System (the chassis)

DGX H100/H200/B200: NVIDIA's reference 8-GPU server (HGX baseboard + 2× CPU + 8× ConnectX, ~10 kW). GB200 NVL72: a full rack with 72 B200 GPUs, 36 Grace CPUs, and an NVLink-Switch tray, all in one NVLink domain. Same chips inside; very different scale-up.

die
GH100
package
SXM5PCIeMGX
board
HGX H100 8-GPU baseboard+ 4× NVSwitch
system
DGX H100Supermicro 8U HGXDell XE9680
scale-up
SuperPOD (32× DGX)GB200 NVL72 (72 GPUs)
The buying-decision implication

"I have an H100" is ambiguous until you also say SXM or PCIe. SXM gets full NVLink 4 (900 GB/s/GPU) and 80 GB at 3.35 TB/s; PCIe is bridge-pair NVLink (600 GB/s) and 350 W instead of 700. For 70B+ training the SXM/HGX path is the only sensible one; for single-GPU inference PCIe is fine.

05

Pascal & Volta — Where AI GPUs Began

The two generations that turned NVIDIA from "graphics company that also does HPC" into "the AI hardware company". Pascal introduced the modern datacenter form factor; Volta introduced the tensor core.

Pascal — P100 (2016)

  • GP100, TSMC 16FF, 15.3B transistors, 56 SMs, 3584 CUDA cores.
  • First card with HBM2 on package: 16 GB at ~720 GB/s. A ~2.5× bandwidth jump over the GDDR5 of Maxwell (Tesla M40: 288 GB/s).
  • First NVLink (gen 1, 160 GB/s aggregate) and the first SXM2 mezzanine module.
  • No tensor cores. Mixed-precision FP16 was supported as a regular SIMD op, not a matrix-multiply-accumulate.
  • P100 was the GPU AlexNet's successors trained on. The original Transformer paper (2017) and early seq2seq translation are P100-era work (ResNet-152, Dec 2015, predates it).

Volta — V100 (2017)

  • GV100, TSMC 12FF, 21.1B transistors, 80 SMs, 5120 CUDA cores, 32 GB HBM2 at ~900 GB/s.
  • First tensor cores. 1st-gen, FP16 inputs into FP32 accumulator, 4×4×4 matrix per cycle per tensor core. ~125 TFLOPS FP16 vs ~15 TFLOPS FP32.
  • NVLink 2.0: 300 GB/s aggregate; six links per GPU instead of four.
  • Independent thread scheduling: every CUDA thread now had its own program counter, fixing a long-standing warp-divergence pain point.
  • V100 trained GPT-3 and much of the 2018–2020 wave of transformer pretraining (Google's BERT and T5 ran on TPUs). Still in production at AWS (p3 instances), Lambda, and academic clusters — the "AK-47 of AI training".
A naming aside

Up to and including Volta, NVIDIA called the datacenter line "Tesla" (Tesla P100, Tesla V100). With Ampere they dropped the brand to avoid confusion with the car company; A100 / H100 / B200 are just the SKU names now. So if you read a 2018 paper that says "trained on Tesla V100s", that is exactly the V100 you'd order today — the rebrand was cosmetic.

06

Turing — The Mixed-Precision Inference GPU

Turing is the generation that almost no one writes about in AI papers, yet ran more inference requests than any other family for several years. It also introduced two features that defined the next decade: RT cores for ray tracing and integer tensor cores for low-precision inference.

SKUDieMemoryRole
RTX 2080 TiTU10211 GB GDDR6Consumer flagship; first RT cores reach gamers.
RTX 2080 / 2070TU104 / TU1068 GB GDDR6Mainstream gaming.
T4TU10416 GB GDDR6The 70 W half-height inference card. Default cloud inference SKU 2019–2022.
Quadro RTX 6000 / 8000TU10224 / 48 GB GDDR6Workstation visualisation; some early ML use.

What changed under the hood

Why T4 mattered

T4 launched at $2,000-ish, drew 70 W with no external power connector, fit in any 1U server, and could do INT8 inference at ~130 TOPS. Compared with V100 (300 W, $10k+, dedicated server) it was the first NVIDIA datacenter card you could rack densely without rebuilding your power and cooling. It ran most of the BERT-era inference traffic at the major clouds. Even today (2026) you'll find T4s in production for embedding generation and small-model serving.

07

Ampere — The LLM Training Workhorse

Ampere is the generation that took LLMs from "interesting research" to "industrial product". GA100 (the datacenter die) and GA102/4/6/7 (the consumer dies) cover totally different process nodes and memory but share the architecture name.

GA100 — A100 (datacenter, TSMC 7N)

  • 54.2B transistors, 826 mm², 108 SMs, 6912 CUDA cores, 432 tensor cores.
  • 40 MB L2 cache, 6× A100's predecessor.
  • 40 GB HBM2 at 1.5 TB/s or 80 GB HBM2e at 2.0 TB/s.
  • 3rd-gen tensor cores: TF32 (FP32's 8-bit exponent with FP16's 10-bit mantissa; a drop-in for FP32 matmul), BF16 (8-bit exponent, training-friendly), and 2:4 structured sparsity (2× throughput on sparse matrices).
  • NVLink 3: 600 GB/s per GPU; SXM4 form factor.
  • MIG (Multi-Instance GPU): one A100 splits into up to seven hardware-isolated slices, each with its own SMs, L2, and HBM partition.

GA10x — RTX 30-series (consumer, Samsung 8N)

  • GA102 (3090, 3090 Ti): 28.3B transistors, 24 GB GDDR6X, ~936 GB/s.
  • Different fab (Samsung 8N, less efficient than TSMC 7N) means consumer Ampere actually runs hotter and thirstier than its datacenter cousin.
  • Same tensor-core formats as GA100 (TF32, BF16, FP16, INT8, INT4) but no MIG, no NVLink on most SKUs (3090/3090 Ti had NVLink, gone on 30-series Ti and below).
  • 3090 Ti at 24 GB became the canonical hobbyist LLM card 2022–2024 — first time normal humans could run 13B–30B locally.

The training story

GPT-3 (175B parameters, May 2020) was in fact trained on V100s, and PaLM on TPU v4. The canonical Ampere pretrains — OPT-175B, BLOOM, the LLaMA-1 family — all happened on A100 clusters with NVLink + InfiniBand fabric. The combination of 80 GB HBM, BF16 with stable accumulators, and 600 GB/s NVLink is what made dense 100B-parameter training routine. You still see A100 80 GB advertised on every cloud provider in 2026; the depreciation curve is unusually long.

Capability v. compute capability

Ampere is compute capability 8.0 (A100) or 8.6 (consumer GA10x) — not the same number. Some CUDA features (notably some MIG and async-copy paths) only exist on 8.0. This is the first generation where "Ampere" as a marketing label hides two genuinely different feature sets.

08

Ada Lovelace + Hopper — Parallel Evolution

2022 is the only year NVIDIA shipped two distinct architectures simultaneously: Ada Lovelace for graphics and prosumer, Hopper for the datacenter. They share a process node (TSMC 4N) but otherwise diverge sharply — same fab tooling, very different design priorities.

Ada Lovelace — AD102 / RTX 4090, L40S

  • 76.3B transistors on AD102, 608 mm², TSMC 4N.
  • 3rd-generation RT cores: opacity micromaps, displaced micro-meshes; ~2× ray-triangle throughput over Ampere.
  • 4th-gen tensor cores — same FP16/BF16/INT8 lineup, plus FP8 support on every Ada part (sm_89), RTX 4090 included; datacenter Ada (L40S) is where NVIDIA markets it.
  • Massive 96 MB L2 — about 16× the cache of GA102. Designed for ray-tracing locality, but also a huge win for LLM decode (KV cache hits in L2 more often).
  • GDDR6X (4090, 4080) at ~1 TB/s; no HBM, no NVLink on 4090.
  • RTX 4090: $1,600 consumer card with 24 GB. L40S: 48 GB, datacenter-licensed, ECC, FP8.

Hopper — GH100 / H100, H200

  • 80B transistors on GH100, 814 mm² (reticle-limited), TSMC 4N.
  • 4th-gen tensor cores with native FP8 (E4M3 and E5M2): 1979 TFLOPS FP8 dense on H100 SXM, 2× the BF16 throughput.
  • Transformer Engine: per-tensor scaling factors, layer-wise auto-promotion between FP8 and BF16; the software side of FP8 in one library.
  • TMA (Tensor Memory Accelerator): fire-and-forget async bulk copies between global and shared memory; massively reduces register pressure on attention kernels.
  • Thread-block clusters: groups of CTAs that share an SM cluster's L1 and a new Distributed Shared Memory address space — a true new tier in the memory hierarchy.
  • DPX instructions: dynamic-programming acceleration (alignment, RL value iteration) — niche but novel.
  • 50 MB L2, 80 GB HBM3 (H100, 3.35 TB/s) or 141 GB HBM3e (H200, 4.8 TB/s).
  • NVLink 4: 900 GB/s per GPU, 3rd-gen NVSwitch trays for 8-GPU and 256-GPU domains.
Why two architectures at once

Graphics and AI have different design constraints: ray tracing wants huge L2 and lots of small, latency-sensitive kernels; transformer training wants HBM bandwidth, FP8 matmul, and async TMA copies. NVIDIA could no longer afford to compromise either. Ada is the first "graphics-first" architecture since Maxwell; Hopper is the first "transformers-first" architecture, full stop.

09

Blackwell — Dual-Die, FP4, NVL72

Blackwell is the first generation where the answer to "how do we double again?" had to be three things at once: stop trying to make a bigger die, halve the precision again, and turn an entire rack into a cache-coherent NVLink domain. B200 is the headline SKU; GB200 NVL72 is what the architecture really enables.

The chip — B200

  • Two reticle-limited dies on TSMC 4NP, each ~104B transistors, total ~208B per package.
  • Linked by NV-HBI (NVIDIA High Bandwidth Interface) at 10 TB/s — transparent to software; the OS sees one GPU.
  • 5th-gen tensor cores with FP4 microscaling — both MX-FP4 (open OCP standard: 32-element blocks, E8M0 scale) and NVFP4 (NVIDIA's tighter variant: 16-element blocks with an E4M3 scale plus a per-tensor FP32 scale). 2× the FP8 throughput, 4× FP16; NVFP4 typically wins on accuracy, MX-FP4 on tooling portability.
  • 192 GB HBM3e (8 stacks × 24 GB), 8 TB/s aggregate bandwidth.
  • Up to 1000 W TDP on SXM6.

The fabric — NVLink 5 + NVL72

  • NVLink 5: 1.8 TB/s per GPU (2× NVLink 4).
  • GB200 NVL72: a full rack containing 72 B200 GPUs and 36 Grace CPUs, with an external NVLink Switch tray. All 72 GPUs in one cache-coherent NVLink domain.
  • Aggregate intra-rack NVLink bandwidth: 130 TB/s.
  • Allows 72-way tensor parallelism (or any combination of TP / EP / DP) without ever touching InfiniBand — latency profile of one machine.
  • Designed specifically for KV-heavy MoE inference (e.g. 400B+ models with sparse activation).

Why FP4 is the headline

FP4 has 16 representable values. Naively that destroys model quality. Microscaling saves it by giving each small block of FP4 values its own scale, so the dynamic range is per-block rather than per-tensor. Blackwell's Tensor Cores accelerate two flavours: MX-FP4 — the open OCP standard, with 32-element blocks and an E8M0 (8-bit, exponent-only) per-block scale — and NVFP4, NVIDIA's variant, with tighter 16-element blocks and an E4M3 (FP8) per-block scale on top of a per-tensor FP32 scale. NVFP4 is typically the more accurate of the two and is what TensorRT-LLM and the Transformer Engine default to; MX-FP4 wins where you need a vendor-neutral, OCP-compliant pipeline. Empirically FP4 inference on Blackwell lands within ~0.5 percentage points of FP8 on most benchmarks while running 2× faster. The serving cost of frontier LLMs roughly halves overnight when the kernels and quantisers catch up.

NVL72 in one sentence

NVL72 turns the question "do I have enough VRAM?" into "do I have enough rack?" — 72 × 186 GB = 13.4 TB of HBM3e in a single coherent address space, with 130 TB/s of intra-domain bandwidth. For a 400B-parameter model with a 1M-token context window, that is the difference between possible and not.

10

The Roadmap — Rubin and Beyond

NVIDIA telegraphs roadmaps about two years ahead. As of 2026, the public-facing pieces are Rubin (the GPU) and Vera (the Grace successor CPU). Past that the names exist but the silicon is speculative.

WhenCodenameWhat's announced
Late 2026 (planned)Rubin (R100)Successor to Blackwell. TSMC 3nm. New HBM (HBM4) and a wider NVLink 6 are expected. Continues the dual-die package model. NVIDIA has shown VR200 NVL144 as the rack-scale system.
2026VeraArm-based CPU successor to Grace, paired with Rubin in Vera Rubin superchips (analogous to Grace Hopper / Grace Blackwell).
2027 (announced, no detail)Rubin UltraRefresh of Rubin, larger HBM stacks. Equivalent to the H100→H200 step.
2028+ (named only)FeynmanGeneration after Rubin. No public technical detail. Don't plan around this yet.

Things to watch (not yet specified)

Scope of the rest of this series

Decks 02 through 10 in this series cover Pascal through Blackwell in detail: SM microarchitecture, tensor-core math, memory hierarchy, NVLink fabric, the Transformer Engine, FP4 microscaling, NVL72 topology, and the LLM-serving software stack on each. All of it is current as of 2026 and reflects shipping silicon, not roadmap. Where Rubin matters we'll flag it as forward-looking.

11

Interactive: Family Explorer

Pick a family and see its top-die spec at a glance, alongside the precision support badges that determine which software stacks run on it. The summary explains why that family mattered and what came next.

Reading the badges

Green = native tensor-core support. Amber = supported on a subset of SKUs. Red = not supported in tensor cores; software emulation is possible but defeats the point.

12

Cheat Sheet — Want X, Pick Y

The shortest, opinionated map from "what I want to do" to "what GPU family I should be looking at" in 2026.

I want to…PickWhy
Run a 7B model locally, no fussUsed RTX 3060 12 GB or 3090 24 GB (Ampere)FP16 / BF16 in tensor cores, plenty for q4_K_M, cheapest entry.
Cheap experimental fine-tune (LoRA on 7B–13B)A100 80 GB (Ampere)BF16 + 80 GB HBM2e; cloud-spot prices are low, training scripts work everywhere.
Consumer inference, 13–30B at homeRTX 4090 or RTX 5090 (Ada / Blackwell consumer)96 MB L2 (Ada) and GDDR7 (Blackwell) destroy decode latency. No NVLink, but you don't need it.
Frontier training (100B+ dense, multi-node)H100 / H200 cluster (Hopper) on HGX + InfiniBandFP8 + Transformer Engine + NVLink 4 (900 GB/s). The known-good frontier stack.
Production inference, 70B classH200 141 GB or B200 192 GBBandwidth (4.8–8 TB/s) is what bounds tok/s; capacity holds full FP8 weights + KV.
Production inference, MoE 400B+GB200 NVL7213.4 TB HBM in one NVLink domain at 130 TB/s. The only sane home for trillion-parameter MoE.
Slice a big GPU into many tenantsA100 / H100 + MIGHardware-isolated partitions, separate L2 and HBM slice each. Blackwell B200 also supports MIG.
Cheap inference fleet (L4-class)L40S 48 GB (Ada datacenter)FP8 on Ada, ECC, datacenter-licensed, 350 W. DP across many cards beats one big card per dollar.
The two software-visible cliffs

Most generation-to-generation improvements are gradual: a bit more bandwidth, a bit more L2, a bit faster NVLink. Two transitions are not gradual. The FP8 cliff at Hopper (2022) doubled effective inference throughput and made the Transformer Engine a real thing. The FP4 cliff at Blackwell (2024) doubled it again with FP4 microscaling (MX-FP4 and NVFP4). If you're choosing a card, ask: do my workloads need FP8? do they need FP4? Each "yes" maps cleanly to one architectural threshold.

One more thing

Buying GPUs in 2026 also means buying a fabric. A single H100 is fine; eight of them only behave as one machine because of NVLink + NVSwitch. Once you cross the boundary into multi-GPU territory, the interconnect generation matters as much as the compute generation. NVL72 is not "more H100s in a rack"; it's a categorically different machine.