Why is the L40S a datacenter card and the RTX 4090 a 'gaming' card when both are AD102? Why does the EULA matter? What replaced Quadro and Tesla? Walk through every NVIDIA product family — past and present — and the rules that govern where each card can legally run.
An end-to-end map of NVIDIA's product lineup — the names, the silicon shared between them, the licence boundaries, and how to pick the right SKU for any role.
NVIDIA today sells GPU silicon under three commercial pillars. The same physical die can appear in all three with different firmware, driver, BIOS, board layout, and licence — the differences are real and they matter.
Gaming-first. RTX 30/40/50 series. Air-cooled axial fans, display outputs (HDMI + DP), GeForce Experience, ShadowPlay, DLSS, GeForce Now cloud. No ECC, no SR-IOV, no MIG, no certified ISV drivers. The GeForce Driver EULA explicitly forbids datacenter deployment.
Formerly Quadro (1999–2020), then "NVIDIA RTX" (2020–2023, e.g. RTX A6000, RTX 6000 Ada), now "RTX Pro" (2025+). ECC GDDR, certified ISV drivers (Autodesk, SolidWorks, Adobe, ANSYS, DaVinci), full FP64 path on professional models, datacenter-licensed. Often offered with NVLink bridges.
Formerly Tesla branding (dropped 2020 to avoid confusion with the car company). Today: A100/A40/A30/A10/A2, H100/H200, L40/L40S/L4, B100/B200/B300. Bare PCIe cards or HGX/SXM baseboards, passive cooling, ECC mandatory, no display outputs, top-tier NVLink, MIG on the 100-tier, full datacenter EULA.
| Feature | GeForce | RTX Pro | Datacenter |
|---|---|---|---|
| Display outputs | 3–4 | 4 DP (mini) | None |
| Cooling | Axial fans | Blower / axial | Passive (chassis airflow) |
| Form factor | 2–3.5 slot | 2-slot blower | 2-slot FHFL or SXM module |
| ECC | No | Yes (GDDR ECC) | Yes (mandatory) |
| NVLink | 3090 only (last consumer) | Bridge on some Pro | NVSwitch fabric (SXM) |
| MIG | No | No | 100-tier (A100, H100, B200) |
| SR-IOV / vGPU | No | Yes | Yes |
| EULA | Consumer (no DC) | Workstation/DC OK | Datacenter OK |
The most important difference between the three pillars isn't silicon — it's the BIOS strap and the driver branch. The same AD102 die in a 4090, RTX 6000 Ada, and L40S has the same SMs and the same Tensor Cores. What changes is which features the firmware unlocks (ECC, vGPU, MIG, NVLink), how the driver enforces the licence, and whether the board ships with display outputs and consumer cooling.
Datacenter SKUs follow a strict letter-and-number scheme. Once you know the rules, the part number tells you the architecture, the tier, and whether it's a China-export variant.
| Arch | Die prefix | Era | Datacenter SKUs | Notes |
|---|---|---|---|---|
| Pascal | GP | 2016–17 | P100, P40, P4 | First "Tesla" cards aimed squarely at deep learning. P100 was first HBM2. |
| Volta | GV | 2017–18 | V100 | First Tensor Cores. SXM2 form factor, NVLink 2. |
| Turing | TU | 2018–19 | T4 (only major DC SKU) | 70 W, single-slot — mass-deployed inference card. |
| Ampere | GA | 2020–21 | A100, A40, A30, A10, A2 | Full inference-to-training span; A100 introduced MIG. |
| Ada Lovelace | AD | 2022–23 | L40, L40S, L4 | "L" replaces "A" mid-arch; reflects Lovelace branding. |
| Hopper | GH | 2023–24 | H100, H200, H800, H20 | FP8 + Transformer Engine. H800/H20 are China-only. |
| Blackwell | GB | 2025–26 | B100, B200, B300, GB200 | Dual-die package. FP4 (MX-FP4 and NVFP4). NVL72 rack as a unit. |
| Suffix | Role | Examples | Typical TDP |
|---|---|---|---|
| 100 | Top compute — HBM, NVLink, MIG, training-class | P100, V100, A100, H100, B100, B200 | 300–1000 W |
| 40 | Cost-optimised inference / visualisation — GDDR, no MIG | P40, A40, L40, L40S | 250–350 W |
| 30 | Mid-range inference — PCIe-only, GDDR ECC | A30 | 165 W |
| 10 | Single-slot inference | A10 | 150 W |
| 4 | Low-power inference, single-slot, 70 W | T4, L4 | 70–72 W |
| 2 | Edge / low-profile, < 60 W | A2 | 40–60 W |
Sales to China face US export controls (October 2022 and October 2023 BIS rules) capping FLOPS-per-package and interconnect bandwidth. NVIDIA's response was to add an '8' suffix — A800 (an A100 with NVLink throttled), H800 (H100 with NVLink throttled), and the further-cut H20. The naming itself is a regulatory artefact: '8' = China-tuned variant.
GeForce has been NVIDIA's consumer line since 1999. It's marketed at gamers, sold through Best Buy and Amazon, and runs the GeForce driver branch. The fact that the same silicon is wildly useful for LLM inference is a side effect — one that NVIDIA accepts but does not encourage.
| Generation | Top SKU | Arch | VRAM | Why it mattered for AI |
|---|---|---|---|---|
| 2017 | GTX 1080 Ti | Pascal (GP102) | 11 GB GDDR5X | First "everyone has one" CUDA card; FP32-only. |
| 2018 | RTX 2080 Ti | Turing (TU102) | 11 GB GDDR6 | First RT cores and first Tensor Cores in consumer hardware. |
| 2020 | RTX 3090 / 3090 Ti | Ampere (GA102) | 24 GB GDDR6X | 24 GB became the de-facto LLM hobbyist card; last consumer NVLink bridge. |
| 2022 | RTX 4090 | Ada (AD102) | 24 GB GDDR6X | ~2× perf/W of 3090. NVLink removed; multi-GPU is PCIe-only. |
| 2025 | RTX 5090 | Blackwell (GB202) | 32 GB GDDR7 | ~1.8× bandwidth of 4090; native FP4 tensor cores. |
Every GeForce card ships with the consumer driver branch (R5xx/R6xx New Feature Branch), GeForce Experience for game-ready optimisations, ShadowPlay capture, and the GeForce EULA. None ship with ECC, certified ISV drivers, or datacenter licence. The card itself doesn't know — the driver enforces.
The professional line has had three names in five years. Same line; new branding each generation.
| Card | Die | VRAM | NVLink | FP64 | Year |
|---|---|---|---|---|---|
| Quadro RTX 8000 | TU102 (Turing) | 48 GB GDDR6 ECC | Bridge | 1/32 rate | 2018 |
| RTX A6000 | GA102 (Ampere) | 48 GB GDDR6 ECC | Bridge | 1/64 rate | 2020 |
| RTX 6000 Ada | AD102 (Ada) | 48 GB GDDR6 ECC | None | 1/64 rate | 2022 |
| RTX Pro 6000 Blackwell | GB202 (Blackwell) | 96 GB GDDR7 ECC | None | 1/64 rate | 2025 |
The same GA102/AD102/GB202 silicon supports clamshell GDDR placement — chips on both sides of the PCB — doubling capacity with no die change. NVIDIA reserves clamshell layouts for Pro and datacenter SKUs to keep the consumer flagship's price segmentation intact.
The datacenter line is split into three rough lanes: top compute (HBM, training-class), inference value (GDDR), and edge/specialist. All datacenter cards share passive cooling, no display outputs, full ECC, and a datacenter-permitting EULA.
SXM module → NVSwitch fabric, 8× per HGX baseboard. PCIe versions exist but with reduced TDP and no NVLink fabric.
A2 is the smallest "real" datacenter card — designed for telco edge boxes and small inference appliances.
| Form factor | Used by | NVLink | TDP envelope | Cooling |
|---|---|---|---|---|
| SXM2/3/4/5 | V100, A100, H100, H200, B100, B200 | Full NVSwitch fabric | 400–1000 W | Mezzanine on HGX baseboard |
| PCIe FHFL 2-slot | A40, A100 PCIe, H100 PCIe, L40S, B200 PCIe | Bridge or none | 250–400 W | Passive (server airflow) |
| PCIe single-slot | T4, A10, L4, A2 | None | 40–150 W | Passive low-profile |
| SuperPOD / NVL72 | GB200 (rack as a unit) | NVLink 5 spine | 120 kW per rack | Liquid (DLC) |
Datacenter cards omit display engine PCB area to make room for HBM stacks (or extra GDDR clamshell), to remove a class of failure and a class of EMC certification, and to physically distinguish them from prosumer SKUs. The display engine still exists on-die, it just isn't bonded out. This is partly why the BIOS strap matters: the same AD102 in an L40S has the display path disabled in firmware.
US Bureau of Industry and Security (BIS) export controls have driven a parallel SKU family aimed exclusively at the Chinese market. The rules don't ban GPU sales — they cap the capability of any single package.
| SKU | Based on | What's throttled | What's preserved | Niche |
|---|---|---|---|---|
| A800 | A100 80 GB | NVLink 600 → 400 GB/s | Compute, HBM2e bandwidth | Training in China, 2022–2023 |
| H800 | H100 SXM | NVLink 900 → 400 GB/s | FP8 throughput, HBM3 bandwidth | Training in China, 2023 |
| H20 | H100 die, fewer SMs | FP8 ~6.7x cut to 296 TFLOPS (dense) | 96 GB HBM3, full bandwidth | LLM inference in China, 2024+ |
| L20 / L2 | L40 / L4 | SM count + clocks | VRAM, video engines | Visualisation / inference, China |
Compute-cap rules drive product segmentation by region, not just by tier. The H20 in particular demonstrates that bandwidth and capacity matter more than peak FLOPS for inference — NVIDIA could keep the memory subsystem intact while butchering compute, and the result is a perfectly serviceable LLM serving card that satisfies BIS thresholds.
NVIDIA's mobile/embedded line has two names. Tegra is the SoC silicon family (originally for phones and tablets, today for automotive and embedded); Jetson is the developer kit / module family that ships those SoCs as edge AI hardware. Same chip, different packaging and distribution channel.
An ARM SoC with an integrated NVIDIA GPU, video engines, and ISP. Generations: Tegra 2 / 3 / 4 (early Android), Tegra K1 (first Kepler GPU), X1 (Maxwell, used in Nintendo Switch), Xavier (Volta + Carmel ARM cores), Orin (Ampere SM + Cortex-A78AE), Thor (Blackwell SM, 2025+).
Core market: automotive (NVIDIA Drive AGX), industrial, robotics. Switch is the rare mass-consumer slot.
SoMs built around Tegra. Officially supported through JetPack (Linux for Tegra + CUDA + cuDNN + TensorRT for ARM). Sold as developer kits and as production modules.
Current line: Orin (Ampere). Next: Thor (Blackwell), 2025+, up to 2070 TFLOPS FP4 (sparse) for robotics + autonomous vehicles.
| Module | GPU | RAM | AI perf | TDP | Price |
|---|---|---|---|---|---|
| Jetson Orin Nano (8 GB) | Ampere, 1024 CUDA / 32 Tensor | 8 GB LPDDR5 | 40 TOPS (67 TOPS Super) | 15 W | ~$200 |
| Jetson Orin NX 8 GB | Ampere, 1024 CUDA | 8 GB LPDDR5 | 70 TOPS | 10–20 W | ~$400 |
| Jetson Orin NX 16 GB | Ampere, 1024 CUDA | 16 GB LPDDR5 | 100 TOPS | 10–25 W | ~$600 |
| Jetson AGX Orin 32 GB | Ampere, 1792 CUDA | 32 GB LPDDR5 | 200 TOPS | 15–40 W | ~$1,500 |
| Jetson AGX Orin 64 GB | Ampere, 2048 CUDA | 64 GB LPDDR5 | 275 TOPS | 15–60 W | ~$2,000 |
Jetson AGX Orin 64 GB has more usable memory than an RTX 4090 (64 GB unified vs 24 GB VRAM), albeit at much lower bandwidth (~204 GB/s LPDDR5 vs 1 TB/s GDDR6X). For a 30B INT4 model where decode is bandwidth-bound, AGX Orin will serve a few tok/s entirely on a 60 W power budget — a useful proposition for offline / edge LLM products. Jetson Thor will dramatically widen this niche.
NVIDIA enforces product segmentation through software licences, not silicon. Two distinct EULAs govern almost every GPU sold today.
Forbids "datacenter deployment" of consumer GeForce cards. The exact clause has evolved, but the operative term is "Datacenter Deployment".
Allowed:
Forbidden:
Applies to RTX Pro, A/H/L/B-series, and Tesla-branded cards. No datacenter restriction.
Allowed:
Comes with strings: SLA-grade driver branch (Production Branch), formal support contract via NVIDIA AI Enterprise (separate licence), validated server platforms.
If you're building a commercial multi-tenant service, the cost difference between a 4090 and an L40S is small relative to your time and risk. Use the workstation/datacenter SKU. If you're a single team running internal inference on company hardware, 4090s and 5090s are fine in practice, and a great deal cheaper. The EULA boundary is "do you sell GPU time to third parties" — not "do you have it in a rack".
The single most useful insight in this entire deck: marketing tier ≠ silicon tier. NVIDIA sells the same physical GPU die in three pillars at vastly different prices, with the differences driven by VRAM count, clocks, BIOS strap, and licence — not silicon.
| Card | Pillar | VRAM | ECC | NVLink | EULA | TDP |
|---|---|---|---|---|---|---|
| RTX 3080 Ti | Consumer | 12 GB GDDR6X | No | No | GeForce | 350 W |
| RTX 3090 | Consumer | 24 GB GDDR6X | No | Bridge | GeForce | 350 W |
| RTX 3090 Ti | Consumer | 24 GB GDDR6X | No | Bridge | GeForce | 450 W |
| RTX A6000 | Workstation | 48 GB GDDR6 ECC | Yes | Bridge | Pro/DC | 300 W |
| A40 | Datacenter | 48 GB GDDR6 ECC | Yes | Bridge | DC | 300 W |
| A10 | Datacenter | 24 GB GDDR6 | Yes | No | DC | 150 W |
| Card | Pillar | VRAM | ECC | FP8 | EULA | TDP |
|---|---|---|---|---|---|---|
| RTX 4080 Super (actually AD103) | Consumer | 16 GB GDDR6X | No | Yes (no TE) | GeForce | 320 W |
| RTX 4090 | Consumer | 24 GB GDDR6X | No | Yes (no TE) | GeForce | 450 W |
| RTX 6000 Ada | Workstation | 48 GB GDDR6 ECC | Yes | Yes | Pro/DC | 300 W |
| L40 | Datacenter | 48 GB GDDR6 ECC | Yes | Yes | DC | 300 W |
| L40S | Datacenter | 48 GB GDDR6 ECC | Yes | Yes (higher throughput) | DC | 350 W |
| Card | Die | Pillar | VRAM | Notes |
|---|---|---|---|---|
| RTX 5090 | GB202 | Consumer | 32 GB GDDR7 | GeForce flagship |
| RTX Pro 6000 Blackwell | GB202 | Workstation | 96 GB GDDR7 ECC | Same die, 3× VRAM via clamshell |
| B100 | GB100 (dual-die) | Datacenter | 192 GB HBM3e | 700 W |
| B200 | GB100 (dual-die) | Datacenter | 192 GB HBM3e | 1000 W — same silicon, higher power bin |
Don't buy the marketing tier — buy the firmware features you need. If your job needs ECC, NVLink, MIG, or datacenter EULA, those features are gated by SKU choice within the same silicon. If your job needs only single-stream raw throughput, the cheapest card on the right die wins. The L40S vs RTX 6000 Ada example is a good illustration: same AD102 die, same 48 GB ECC, both have FP8 tensor cores, but the L40S is the high-clock "Server" bin tuned for inference throughput — and for LLM serving in 2025+, that throughput uplift is often the deciding feature.
NVIDIA's driver isn't a single artefact — it's at least two parallel streams, and matching the driver to the silicon and the workload is essential operational hygiene.
~6-month cadence, longer support tail (~1 year). Validated against the NVIDIA AI Enterprise stack and certified server platforms. Used by datacenter operators and enterprise workstation users.
Examples (2024–2026): R535, R550, R570 PB.
Trade-off: conservative. Gets bug fixes; doesn't get the latest CUDA features as fast.
Faster cadence (monthly to quarterly). Adds new CUDA toolkit features as they ship. Used by GeForce/Studio consumers, gamers, and bleeding-edge developers.
Examples (2024–2026): R545, R555, R560, R565, R570 NFB, R575.
Trade-off: aggressive. Newest features, occasional regressions.
| Driver | Required for | Notes |
|---|---|---|
| R525+ | Hopper (H100) basic support | CUDA 12.0; Hopper FP8 via Transformer Engine library |
| R535+ | Hopper FP8 production | CUDA 12.2; cuBLASLt FP8 GEMMs |
| R550+ | Hopper TMA, thread-block clusters | CUDA 12.4; full Hopper feature set |
| R555+ | Blackwell preview | CUDA 12.5; B100/B200 early access |
| R560+ | Pre-Blackwell | CUDA 12.6; no sm_100/sm_120 target yet |
| R570+ | Blackwell + flash-attn 3 polished | CUDA 12.8+; first toolkit with sm_100/sm_120 (Blackwell), FP4 paths (MX-FP4 and NVFP4) |
Every CUDA kernel is compiled for a specific Compute Capability (CC). A wheel built only for sm_80 won't use Hopper's FP8 tensor cores even on an H100; you'd need sm_90. The same wheel running on Blackwell needs sm_100 (B100/B200) or sm_120 (consumer Blackwell).
| CC | Arch | What unlocks at this CC |
|---|---|---|
| sm_75 | Turing | 2nd-gen Tensor Cores (first were Volta, sm_70); adds INT8/INT4 |
| sm_80 | A100 | BF16, TF32, async copy, MIG |
| sm_86 | Consumer Ampere (3090) | BF16, TF32, no MIG |
| sm_89 | Ada (4090, L40S, RTX 6000 Ada) | FP8 tensor cores (all Ada); no Hopper Transformer Engine |
| sm_90 | Hopper (H100, H200) | FP8 native, TMA, thread-block clusters, DSMEM |
| sm_100 | Datacenter Blackwell (B100, B200) | FP4 (MX-FP4 and NVFP4), 5th-gen Tensor Cores, NVLink 5 |
| sm_120 | Consumer Blackwell (RTX 50, RTX Pro 6000 Blackwell) | FP4, GDDR7 |
# nvcc explicit target list:
nvcc -gencode arch=compute_80,code=sm_80 \
-gencode arch=compute_86,code=sm_86 \
-gencode arch=compute_89,code=sm_89 \
-gencode arch=compute_90,code=sm_90 \
-gencode arch=compute_100,code=sm_100 \
-gencode arch=compute_120,code=sm_120 \
-O3 kernel.cu -o kernel
# PyTorch / torch.compile environment hint:
export TORCH_CUDA_ARCH_LIST="8.0;8.6;8.9;9.0;10.0;12.0"
# Verify the GPU's CC at runtime:
nvidia-smi --query-gpu=compute_cap --format=csv
If a wheel was built only for sm_80 and you run it on a Blackwell card, the driver falls back to JIT-compiling PTX (slow first launch) or refuses if no PTX is embedded. Always verify the CC list embedded in your wheels matches your fleet — especially when migrating from H100 to B200 or from 4090 to 5090.
Pick a SKU. Get the family classification, the firmware features, the licence boundary, and a recommended role.
The most informative comparisons are within the same die: pick RTX 4090, then RTX 6000 Ada, then L40, then L40S. All AD102. The differences — ECC, EULA, FP8 enabled, VRAM, NVLink — are entirely firmware and PCB choices, not silicon. That's the whole game.