A grounded tour of every NVIDIA GPU family that matters for AI today: when each landed, what changed, and why each generation reshaped what we could train and serve.
Ten years of NVIDIA architectures, told as a single story: how each generation enabled the next class of model, and why the software-visible jumps (FP16 tensor cores, FP8, FP4, NVL72) matter more than the raw transistor count.
Every public LLM milestone of the last seven years sat on top of a specific NVIDIA generation. The headline model didn't appear because someone had a clever idea in isolation — it appeared because a tensor-core or memory advance made training and serving that class of model affordable. The software follows the silicon by 12 to 24 months.
| Model era | Year | Enabling GPU | What the silicon unlocked |
|---|---|---|---|
| BERT-era models, ResNet-50 at scale | 2018–19 | V100 (Volta; Google's own BERT ran on TPUs) | First-gen tensor cores: FP16 mixed-precision training became standard. |
| GPT-2 / T5 / early multimodal | 2019–20 | V100, then A100 | 32 GB HBM2 on V100; A100's 40/80 GB and 2:4 sparsity opened the 10B-parameter regime. |
| OPT-175B, BLOOM, LLaMA-1 | 2020–22 | A100 (Ampere) | BF16 + TF32 + 3rd-gen tensor cores + NVLink 3 made dense 100B+ training routine. |
| GPT-4-class, Llama-2 70B | 2023–24 | H100 (Hopper) | FP8 (Transformer Engine) doubled effective throughput; TMA cut prefill latency. |
| Frontier 2025 (GPT-4o, Claude 3.7, Llama-3.1 405B) | 2024–25 | H100, H200, early B200 | HBM3e at 4.8 TB/s on H200 turned long-context decode from bandwidth-starved to compute-bound. |
| Frontier 2026 (next-gen MoE, very-long-context agents) | 2025–26 | B200, GB200 NVL72 | FP4 doubles inference again; NVL72 makes a 72-GPU domain look like one machine for KV-heavy MoE. |
For each generation the cycle is the same: NVIDIA ships a new precision (FP16 → BF16 → TF32 → FP8 → FP4) along with a memory bandwidth uplift, the framework teams catch up over six to twelve months, and within a year the new class of model the silicon enabled is the public state of the art. You can't read the LLM roadmap without reading the GPU roadmap.
Eight generations in ten years. Each row below is one architecture family, with the top "datacenter" die that defined it for AI workloads. Consumer SKUs share the architecture name but use different (usually smaller, cheaper) dies and memory.
| Family | Codename | Year | Process | Top die | Transistors | Top SKU |
|---|---|---|---|---|---|---|
| Pascal | GP100 | 2016 | TSMC 16FF | GP100 | 15.3B | Tesla P100 |
| Volta | GV100 | 2017 | TSMC 12FF | GV100 | 21.1B | Tesla V100 |
| Turing | TU102 | 2018 | TSMC 12FF | TU102 | 18.6B | RTX 2080 Ti / T4 (TU104) |
| Ampere | GA100 / GA10x | 2020 | TSMC 7N (GA100) / Samsung 8N (GA10x) | GA100 | 54.2B | A100 80GB |
| Ada Lovelace | AD102 | 2022 | TSMC 4N | AD102 | 76.3B | RTX 4090 / L40S |
| Hopper | GH100 | 2022 | TSMC 4N | GH100 | 80B | H100 / H200 |
| Blackwell | GB100 / B200 | 2024 | TSMC 4NP (dual-die) | 2× reticle-limited | ~208B (2 dies) | B200 / GB200 |
| Rubin (planned) | R100 | 2026+ | TSMC 3nm | TBA | TBA | Rubin GPU + Vera CPU |
Turing is often skipped in AI write-ups because GV100 had tensor cores first and GA100 had the first really useful ones. But Turing's TU104-based T4 was the inference workhorse of 2019–2022 — the first card cloud providers deployed at scale specifically for serving. INT8 / INT4 tensor cores and a 70 W TDP made it the default inference SKU long before anyone had heard of vLLM.
Transistor counts climbed from 15 billion (P100, 2016) to 208 billion (B200, 2024) — about 14× in eight years. That's faster than Moore's law on paper, but only because Blackwell broke the reticle limit by bonding two dies together. The single-die curve flattens hard around Hopper.
Moore's-law slowdown is the reason every modern NVIDIA generation is also an arithmetic innovation. You can't double FLOPS by doubling transistors any more, so you double them by halving precision: FP32 → FP16 (Volta) → BF16/TF32 (Ampere) → FP8 (Hopper) → FP4 (Blackwell). The die can't grow; the numbers shrink.
"H100" can mean four different things depending on context: the silicon die (GH100), the package it's mounted on (SXM5 vs PCIe vs MGX), the baseboard that ties eight of them together (HGX), or the integrated system the OEM ships (DGX, OEM HGX, GB200 rack). Confusing them is the most common source of pricing and capability errors.
One piece of silicon, e.g. GH100 for Hopper or the dual-die B200 for Blackwell. This is what fabs ship. Defines compute capability, tensor-core formats, max HBM stacks. Yields and binning here drive the entire SKU stack.
The die plus HBM stacks on a substrate. Options: SXM5 (mezzanine, NVLink-native, 700 W), PCIe (drop-in card, no NVSwitch, 350 W), MGX (modular reference platforms), OAM (industry-standard module). Same die, very different system options.
The HGX H100/B200 8-GPU baseboard mounts eight SXM modules plus four NVSwitch chips. It is the building block OEMs (Supermicro, Dell, HPE, Lenovo) buy from NVIDIA and integrate into chassis. Same baseboard, dozens of OEM SKUs.
DGX H100/H200/B200: NVIDIA's reference 8-GPU server (HGX baseboard + 2× CPU + 8× ConnectX, ~10 kW). GB200 NVL72: a full rack with 72 B200 GPUs, 36 Grace CPUs, and an NVLink-Switch tray, all in one NVLink domain. Same chips inside; very different scale-up.
"I have an H100" is ambiguous until you also say SXM or PCIe. SXM gets full NVLink 4 (900 GB/s/GPU) and 80 GB at 3.35 TB/s; PCIe is bridge-pair NVLink (600 GB/s) and 350 W instead of 700. For 70B+ training the SXM/HGX path is the only sensible one; for single-GPU inference PCIe is fine.
The two generations that turned NVIDIA from "graphics company that also does HPC" into "the AI hardware company". Pascal introduced the modern datacenter form factor; Volta introduced the tensor core.
Up to and including Volta, NVIDIA called the datacenter line "Tesla" (Tesla P100, Tesla V100). With Ampere they dropped the brand to avoid confusion with the car company; A100 / H100 / B200 are just the SKU names now. So if you read a 2018 paper that says "trained on Tesla V100s", that is exactly the V100 you'd order today — the rebrand was cosmetic.
Turing is the generation that almost no one writes about in AI papers, yet ran more inference requests than any other family for several years. It also introduced two features that defined the next decade: RT cores for ray tracing and integer tensor cores for low-precision inference.
| SKU | Die | Memory | Role |
|---|---|---|---|
| RTX 2080 Ti | TU102 | 11 GB GDDR6 | Consumer flagship; first RT cores reach gamers. |
| RTX 2080 / 2070 | TU104 / TU106 | 8 GB GDDR6 | Mainstream gaming. |
| T4 | TU104 | 16 GB GDDR6 | The 70 W half-height inference card. Default cloud inference SKU 2019–2022. |
| Quadro RTX 6000 / 8000 | TU102 | 24 / 48 GB GDDR6 | Workstation visualisation; some early ML use. |
T4 launched at $2,000-ish, drew 70 W with no external power connector, fit in any 1U server, and could do INT8 inference at ~130 TOPS. Compared with V100 (300 W, $10k+, dedicated server) it was the first NVIDIA datacenter card you could rack densely without rebuilding your power and cooling. It ran most of the BERT-era inference traffic at the major clouds. Even today (2026) you'll find T4s in production for embedding generation and small-model serving.
Ampere is the generation that took LLMs from "interesting research" to "industrial product". GA100 (the datacenter die) and GA102/4/6/7 (the consumer dies) cover totally different process nodes and memory but share the architecture name.
GPT-3 (175B parameters, May 2020) was in fact trained on V100s, and PaLM on TPU v4. The canonical Ampere pretrains — OPT-175B, BLOOM, the LLaMA-1 family — all happened on A100 clusters with NVLink + InfiniBand fabric. The combination of 80 GB HBM, BF16 with stable accumulators, and 600 GB/s NVLink is what made dense 100B-parameter training routine. You still see A100 80 GB advertised on every cloud provider in 2026; the depreciation curve is unusually long.
Ampere is compute capability 8.0 (A100) or 8.6 (consumer GA10x) — not the same number. Some CUDA features (notably some MIG and async-copy paths) only exist on 8.0. This is the first generation where "Ampere" as a marketing label hides two genuinely different feature sets.
2022 is the only year NVIDIA shipped two distinct architectures simultaneously: Ada Lovelace for graphics and prosumer, Hopper for the datacenter. They share a process node (TSMC 4N) but otherwise diverge sharply — same fab tooling, very different design priorities.
Graphics and AI have different design constraints: ray tracing wants huge L2 and lots of small, latency-sensitive kernels; transformer training wants HBM bandwidth, FP8 matmul, and async TMA copies. NVIDIA could no longer afford to compromise either. Ada is the first "graphics-first" architecture since Maxwell; Hopper is the first "transformers-first" architecture, full stop.
Blackwell is the first generation where the answer to "how do we double again?" had to be three things at once: stop trying to make a bigger die, halve the precision again, and turn an entire rack into a cache-coherent NVLink domain. B200 is the headline SKU; GB200 NVL72 is what the architecture really enables.
FP4 has 16 representable values. Naively that destroys model quality. Microscaling saves it by giving each small block of FP4 values its own scale, so the dynamic range is per-block rather than per-tensor. Blackwell's Tensor Cores accelerate two flavours: MX-FP4 — the open OCP standard, with 32-element blocks and an E8M0 (8-bit, exponent-only) per-block scale — and NVFP4, NVIDIA's variant, with tighter 16-element blocks and an E4M3 (FP8) per-block scale on top of a per-tensor FP32 scale. NVFP4 is typically the more accurate of the two and is what TensorRT-LLM and the Transformer Engine default to; MX-FP4 wins where you need a vendor-neutral, OCP-compliant pipeline. Empirically FP4 inference on Blackwell lands within ~0.5 percentage points of FP8 on most benchmarks while running 2× faster. The serving cost of frontier LLMs roughly halves overnight when the kernels and quantisers catch up.
NVL72 turns the question "do I have enough VRAM?" into "do I have enough rack?" — 72 × 186 GB = 13.4 TB of HBM3e in a single coherent address space, with 130 TB/s of intra-domain bandwidth. For a 400B-parameter model with a 1M-token context window, that is the difference between possible and not.
NVIDIA telegraphs roadmaps about two years ahead. As of 2026, the public-facing pieces are Rubin (the GPU) and Vera (the Grace successor CPU). Past that the names exist but the silicon is speculative.
| When | Codename | What's announced |
|---|---|---|
| Late 2026 (planned) | Rubin (R100) | Successor to Blackwell. TSMC 3nm. New HBM (HBM4) and a wider NVLink 6 are expected. Continues the dual-die package model. NVIDIA has shown VR200 NVL144 as the rack-scale system. |
| 2026 | Vera | Arm-based CPU successor to Grace, paired with Rubin in Vera Rubin superchips (analogous to Grace Hopper / Grace Blackwell). |
| 2027 (announced, no detail) | Rubin Ultra | Refresh of Rubin, larger HBM stacks. Equivalent to the H100→H200 step. |
| 2028+ (named only) | Feynman | Generation after Rubin. No public technical detail. Don't plan around this yet. |
Decks 02 through 10 in this series cover Pascal through Blackwell in detail: SM microarchitecture, tensor-core math, memory hierarchy, NVLink fabric, the Transformer Engine, FP4 microscaling, NVL72 topology, and the LLM-serving software stack on each. All of it is current as of 2026 and reflects shipping silicon, not roadmap. Where Rubin matters we'll flag it as forward-looking.
Pick a family and see its top-die spec at a glance, alongside the precision support badges that determine which software stacks run on it. The summary explains why that family mattered and what came next.
Green = native tensor-core support. Amber = supported on a subset of SKUs. Red = not supported in tensor cores; software emulation is possible but defeats the point.
The shortest, opinionated map from "what I want to do" to "what GPU family I should be looking at" in 2026.
| I want to… | Pick | Why |
|---|---|---|
| Run a 7B model locally, no fuss | Used RTX 3060 12 GB or 3090 24 GB (Ampere) | FP16 / BF16 in tensor cores, plenty for q4_K_M, cheapest entry. |
| Cheap experimental fine-tune (LoRA on 7B–13B) | A100 80 GB (Ampere) | BF16 + 80 GB HBM2e; cloud-spot prices are low, training scripts work everywhere. |
| Consumer inference, 13–30B at home | RTX 4090 or RTX 5090 (Ada / Blackwell consumer) | 96 MB L2 (Ada) and GDDR7 (Blackwell) destroy decode latency. No NVLink, but you don't need it. |
| Frontier training (100B+ dense, multi-node) | H100 / H200 cluster (Hopper) on HGX + InfiniBand | FP8 + Transformer Engine + NVLink 4 (900 GB/s). The known-good frontier stack. |
| Production inference, 70B class | H200 141 GB or B200 192 GB | Bandwidth (4.8–8 TB/s) is what bounds tok/s; capacity holds full FP8 weights + KV. |
| Production inference, MoE 400B+ | GB200 NVL72 | 13.4 TB HBM in one NVLink domain at 130 TB/s. The only sane home for trillion-parameter MoE. |
| Slice a big GPU into many tenants | A100 / H100 + MIG | Hardware-isolated partitions, separate L2 and HBM slice each. Blackwell B200 also supports MIG. |
| Cheap inference fleet (L4-class) | L40S 48 GB (Ada datacenter) | FP8 on Ada, ECC, datacenter-licensed, 350 W. DP across many cards beats one big card per dollar. |
Most generation-to-generation improvements are gradual: a bit more bandwidth, a bit more L2, a bit faster NVLink. Two transitions are not gradual. The FP8 cliff at Hopper (2022) doubled effective inference throughput and made the Transformer Engine a real thing. The FP4 cliff at Blackwell (2024) doubled it again with FP4 microscaling (MX-FP4 and NVFP4). If you're choosing a card, ask: do my workloads need FP8? do they need FP4? Each "yes" maps cleanly to one architectural threshold.
Buying GPUs in 2026 also means buying a fabric. A single H100 is fine; eight of them only behave as one machine because of NVLink + NVSwitch. Once you cross the boundary into multi-GPU territory, the interconnect generation matters as much as the compute generation. NVL72 is not "more H100s in a rack"; it's a categorically different machine.