NVIDIA GPU Architectures Series — Presentation 12

GeForce, RTX Pro, Tesla & A/H/L/B — Decoding NVIDIA's Product Lineup

Why is the L40S a datacenter card and the RTX 4090 a 'gaming' card when both are AD102? Why does the EULA matter? What replaced Quadro and Tesla? Walk through every NVIDIA product family — past and present — and the rules that govern where each card can legally run.

GeForceRTX ProTesla T4A100H100 L40SJetsonEULA OEM
GeForce → Quadro → Tesla → RTX Pro → A/H/B/L → Jetson → Tegra
00

Topics We'll Cover

An end-to-end map of NVIDIA's product lineup — the names, the silicon shared between them, the licence boundaries, and how to pick the right SKU for any role.

01

The Three Pillars Today

NVIDIA today sells GPU silicon under three commercial pillars. The same physical die can appear in all three with different firmware, driver, BIOS, board layout, and licence — the differences are real and they matter.

Consumer (GeForce RTX)

Gaming-first. RTX 30/40/50 series. Air-cooled axial fans, display outputs (HDMI + DP), GeForce Experience, ShadowPlay, DLSS, GeForce Now cloud. No ECC, no SR-IOV, no MIG, no certified ISV drivers. The GeForce Driver EULA explicitly forbids datacenter deployment.

Workstation (RTX Pro)

Formerly Quadro (1999–2020), then "NVIDIA RTX" (2020–2023, e.g. RTX A6000, RTX 6000 Ada), now "RTX Pro" (2025+). ECC GDDR, certified ISV drivers (Autodesk, SolidWorks, Adobe, ANSYS, DaVinci), full FP64 path on professional models, datacenter-licensed. Often offered with NVLink bridges.

Datacenter (A/H/L/B)

Formerly Tesla branding (dropped 2020 to avoid confusion with the car company). Today: A100/A40/A30/A10/A2, H100/H200, L40/L40S/L4, B100/B200/B300. Bare PCIe cards or HGX/SXM baseboards, passive cooling, ECC mandatory, no display outputs, top-tier NVLink, MIG on the 100-tier, full datacenter EULA.

What separates them physically

FeatureGeForceRTX ProDatacenter
Display outputs3–44 DP (mini)None
CoolingAxial fansBlower / axialPassive (chassis airflow)
Form factor2–3.5 slot2-slot blower2-slot FHFL or SXM module
ECCNoYes (GDDR ECC)Yes (mandatory)
NVLink3090 only (last consumer)Bridge on some ProNVSwitch fabric (SXM)
MIGNoNo100-tier (A100, H100, B200)
SR-IOV / vGPUNoYesYes
EULAConsumer (no DC)Workstation/DC OKDatacenter OK
The hidden axis

The most important difference between the three pillars isn't silicon — it's the BIOS strap and the driver branch. The same AD102 die in a 4090, RTX 6000 Ada, and L40S has the same SMs and the same Tensor Cores. What changes is which features the firmware unlocks (ECC, vGPU, MIG, NVLink), how the driver enforces the licence, and whether the board ships with display outputs and consumer cooling.

02

The Naming System Decoded

Datacenter SKUs follow a strict letter-and-number scheme. Once you know the rules, the part number tells you the architecture, the tier, and whether it's a China-export variant.

ArchDie prefixEraDatacenter SKUsNotes
PascalGP2016–17P100, P40, P4First "Tesla" cards aimed squarely at deep learning. P100 was first HBM2.
VoltaGV2017–18V100First Tensor Cores. SXM2 form factor, NVLink 2.
TuringTU2018–19T4 (only major DC SKU)70 W, single-slot — mass-deployed inference card.
AmpereGA2020–21A100, A40, A30, A10, A2Full inference-to-training span; A100 introduced MIG.
Ada LovelaceAD2022–23L40, L40S, L4"L" replaces "A" mid-arch; reflects Lovelace branding.
HopperGH2023–24H100, H200, H800, H20FP8 + Transformer Engine. H800/H20 are China-only.
BlackwellGB2025–26B100, B200, B300, GB200Dual-die package. FP4 (MX-FP4 and NVFP4). NVL72 rack as a unit.

The numeric tier code

SuffixRoleExamplesTypical TDP
100Top compute — HBM, NVLink, MIG, training-classP100, V100, A100, H100, B100, B200300–1000 W
40Cost-optimised inference / visualisation — GDDR, no MIGP40, A40, L40, L40S250–350 W
30Mid-range inference — PCIe-only, GDDR ECCA30165 W
10Single-slot inferenceA10150 W
4Low-power inference, single-slot, 70 WT4, L470–72 W
2Edge / low-profile, < 60 WA240–60 W
China-export digit '8'

Sales to China face US export controls (October 2022 and October 2023 BIS rules) capping FLOPS-per-package and interconnect bandwidth. NVIDIA's response was to add an '8' suffix — A800 (an A100 with NVLink throttled), H800 (H100 with NVLink throttled), and the further-cut H20. The naming itself is a regulatory artefact: '8' = China-tuned variant.

Reading a model number

H + 100 = Hopper top-tier (HBM3, NVLink 4, MIG)
L + 40 + S = Lovelace (Ada) cost-tier with FP8 'Server' bin (L40S)
A + 800 = Ampere top-tier, China export-throttled NVLink (A800)
03

GeForce — Gaming First, AI by Accident

GeForce has been NVIDIA's consumer line since 1999. It's marketed at gamers, sold through Best Buy and Amazon, and runs the GeForce driver branch. The fact that the same silicon is wildly useful for LLM inference is a side effect — one that NVIDIA accepts but does not encourage.

Lineage of the flagship

GenerationTop SKUArchVRAMWhy it mattered for AI
2017GTX 1080 TiPascal (GP102)11 GB GDDR5XFirst "everyone has one" CUDA card; FP32-only.
2018RTX 2080 TiTuring (TU102)11 GB GDDR6First RT cores and first Tensor Cores in consumer hardware.
2020RTX 3090 / 3090 TiAmpere (GA102)24 GB GDDR6X24 GB became the de-facto LLM hobbyist card; last consumer NVLink bridge.
2022RTX 4090Ada (AD102)24 GB GDDR6X~2× perf/W of 3090. NVLink removed; multi-GPU is PCIe-only.
2025RTX 5090Blackwell (GB202)32 GB GDDR7~1.8× bandwidth of 4090; native FP4 tensor cores.

The tier grid — every Ada/Blackwell SKU

90 (top)
RTX 4090 24 GB RTX 5090 32 GB
80 (enthusiast)
RTX 4080 16 GB RTX 4080 Super 16 GB RTX 5080 16 GB
70 (mid)
RTX 4070 12 GB RTX 4070 Ti 12 GB RTX 4070 Ti Super 16 GB RTX 5070 / 5070 Ti
60 (entry)
RTX 4060 8 GB RTX 4060 Ti 8/16 GB RTX 5060 / 5060 Ti
50 (budget)
RTX 4050 (mobile) RTX 5050 (mobile)
What ships in the box

Every GeForce card ships with the consumer driver branch (R5xx/R6xx New Feature Branch), GeForce Experience for game-ready optimisations, ShadowPlay capture, and the GeForce EULA. None ship with ECC, certified ISV drivers, or datacenter licence. The card itself doesn't know — the driver enforces.

04

RTX Pro — The Workstation Line

The professional line has had three names in five years. Same line; new branding each generation.

1999–2020 — Quadro (NVS, FX, K, M, P, RTX-Quadro)
↓
2020–2023 — "NVIDIA RTX" (RTX A6000 Ampere, RTX 6000 Ada)
↓
2025+ — RTX Pro (RTX Pro 6000 Blackwell, RTX Pro 5000/4500/4000)

Flagship by generation

CardDieVRAMNVLinkFP64Year
Quadro RTX 8000TU102 (Turing)48 GB GDDR6 ECCBridge1/32 rate2018
RTX A6000GA102 (Ampere)48 GB GDDR6 ECCBridge1/64 rate2020
RTX 6000 AdaAD102 (Ada)48 GB GDDR6 ECCNone1/64 rate2022
RTX Pro 6000 BlackwellGB202 (Blackwell)96 GB GDDR7 ECCNone1/64 rate2025

What makes a card "Pro"

Why VRAM doubles on Pro

The same GA102/AD102/GB202 silicon supports clamshell GDDR placement — chips on both sides of the PCB — doubling capacity with no die change. NVIDIA reserves clamshell layouts for Pro and datacenter SKUs to keep the consumer flagship's price segmentation intact.

05

Datacenter — A/H/L/B Series

The datacenter line is split into three rough lanes: top compute (HBM, training-class), inference value (GDDR), and edge/specialist. All datacenter cards share passive cooling, no display outputs, full ECC, and a datacenter-permitting EULA.

Top compute (HBM)

  • A100 40 GB HBM2 / 80 GB HBM2e
  • H100 80 GB HBM3
  • H200 141 GB HBM3e
  • B100 192 GB HBM3e
  • B200 192 GB HBM3e (8 TB/s)
  • B300 288 GB HBM3e

SXM module → NVSwitch fabric, 8× per HGX baseboard. PCIe versions exist but with reduced TDP and no NVLink fabric.

Inference value (GDDR)

  • A30 24 GB HBM2 (PCIe, MIG)
  • A40 48 GB GDDR6 ECC
  • A10 24 GB GDDR6 (single slot)
  • L4 24 GB GDDR6, 72 W
  • L40 48 GB GDDR6 ECC
  • L40S 48 GB GDDR6 ECC, FP8 enabled
  • T4 16 GB GDDR6 (Turing legacy)

Specialist / edge

  • A2 16 GB GDDR6, 40–60 W, low-profile single-slot
  • T4 16 GB, 70 W — the workhorse of cloud inference 2019–2023

A2 is the smallest "real" datacenter card — designed for telco edge boxes and small inference appliances.

Form factors

Form factorUsed byNVLinkTDP envelopeCooling
SXM2/3/4/5V100, A100, H100, H200, B100, B200Full NVSwitch fabric400–1000 WMezzanine on HGX baseboard
PCIe FHFL 2-slotA40, A100 PCIe, H100 PCIe, L40S, B200 PCIeBridge or none250–400 WPassive (server airflow)
PCIe single-slotT4, A10, L4, A2None40–150 WPassive low-profile
SuperPOD / NVL72GB200 (rack as a unit)NVLink 5 spine120 kW per rackLiquid (DLC)
Why no display outputs

Datacenter cards omit display engine PCB area to make room for HBM stacks (or extra GDDR clamshell), to remove a class of failure and a class of EMC certification, and to physically distinguish them from prosumer SKUs. The display engine still exists on-die, it just isn't bonded out. This is partly why the BIOS strap matters: the same AD102 in an L40S has the display path disabled in firmware.

06

The China-Export Variants

US Bureau of Industry and Security (BIS) export controls have driven a parallel SKU family aimed exclusively at the Chinese market. The rules don't ban GPU sales — they cap the capability of any single package.

Timeline of the rules

SKUBased onWhat's throttledWhat's preservedNiche
A800A100 80 GBNVLink 600 → 400 GB/sCompute, HBM2e bandwidthTraining in China, 2022–2023
H800H100 SXMNVLink 900 → 400 GB/sFP8 throughput, HBM3 bandwidthTraining in China, 2023
H20H100 die, fewer SMsFP8 ~6.7x cut to 296 TFLOPS (dense)96 GB HBM3, full bandwidthLLM inference in China, 2024+
L20 / L2L40 / L4SM count + clocksVRAM, video enginesVisualisation / inference, China
The lesson

Compute-cap rules drive product segmentation by region, not just by tier. The H20 in particular demonstrates that bandwidth and capacity matter more than peak FLOPS for inference — NVIDIA could keep the memory subsystem intact while butchering compute, and the result is a perfectly serviceable LLM serving card that satisfies BIS thresholds.

07

Jetson & Tegra — Edge AI Family

NVIDIA's mobile/embedded line has two names. Tegra is the SoC silicon family (originally for phones and tablets, today for automotive and embedded); Jetson is the developer kit / module family that ships those SoCs as edge AI hardware. Same chip, different packaging and distribution channel.

Tegra (the SoC)

An ARM SoC with an integrated NVIDIA GPU, video engines, and ISP. Generations: Tegra 2 / 3 / 4 (early Android), Tegra K1 (first Kepler GPU), X1 (Maxwell, used in Nintendo Switch), Xavier (Volta + Carmel ARM cores), Orin (Ampere SM + Cortex-A78AE), Thor (Blackwell SM, 2025+).

Core market: automotive (NVIDIA Drive AGX), industrial, robotics. Switch is the rare mass-consumer slot.

Jetson (the dev module)

SoMs built around Tegra. Officially supported through JetPack (Linux for Tegra + CUDA + cuDNN + TensorRT for ARM). Sold as developer kits and as production modules.

Current line: Orin (Ampere). Next: Thor (Blackwell), 2025+, up to 2070 TFLOPS FP4 (sparse) for robotics + autonomous vehicles.

The Orin family today

ModuleGPURAMAI perfTDPPrice
Jetson Orin Nano (8 GB)Ampere, 1024 CUDA / 32 Tensor8 GB LPDDR540 TOPS (67 TOPS Super)15 W~$200
Jetson Orin NX 8 GBAmpere, 1024 CUDA8 GB LPDDR570 TOPS10–20 W~$400
Jetson Orin NX 16 GBAmpere, 1024 CUDA16 GB LPDDR5100 TOPS10–25 W~$600
Jetson AGX Orin 32 GBAmpere, 1792 CUDA32 GB LPDDR5200 TOPS15–40 W~$1,500
Jetson AGX Orin 64 GBAmpere, 2048 CUDA64 GB LPDDR5275 TOPS15–60 W~$2,000

The software stack

Why Jetson matters for LLMs

Jetson AGX Orin 64 GB has more usable memory than an RTX 4090 (64 GB unified vs 24 GB VRAM), albeit at much lower bandwidth (~204 GB/s LPDDR5 vs 1 TB/s GDDR6X). For a 30B INT4 model where decode is bandwidth-bound, AGX Orin will serve a few tok/s entirely on a 60 W power budget — a useful proposition for offline / edge LLM products. Jetson Thor will dramatically widen this niche.

08

The EULA — What You Can and Can't Do

NVIDIA enforces product segmentation through software licences, not silicon. Two distinct EULAs govern almost every GPU sold today.

GeForce Driver EULA

Forbids "datacenter deployment" of consumer GeForce cards. The exact clause has evolved, but the operative term is "Datacenter Deployment".

Allowed:

  • Home, office, single-tenant internal use
  • Personal AI/ML research, hobby projects, learning
  • Cryptocurrency mining (post-LHR removal in 2022)
  • Game development (any scale)
  • Blockchain processing (single-tenant)

Forbidden:

  • Multi-tenant cloud hosting
  • Public-facing AI services at scale on consumer hardware
  • Commercial datacenter colocation
  • SaaS GPU rental

Workstation / Datacenter Driver EULA

Applies to RTX Pro, A/H/L/B-series, and Tesla-branded cards. No datacenter restriction.

Allowed:

  • All of the above
  • Multi-tenant cloud / SaaS
  • Public-facing commercial AI services
  • Datacenter colocation
  • Bare-metal rental, GPU-as-a-service
  • Carrier-grade telco edge

Comes with strings: SLA-grade driver branch (Production Branch), formal support contract via NVIDIA AI Enterprise (separate licence), validated server platforms.

Enforcement reality

The pragmatic rule

If you're building a commercial multi-tenant service, the cost difference between a 4090 and an L40S is small relative to your time and risk. Use the workstation/datacenter SKU. If you're a single team running internal inference on company hardware, 4090s and 5090s are fine in practice, and a great deal cheaper. The EULA boundary is "do you sell GPU time to third parties" — not "do you have it in a rack".

09

The "Same Die, Different Card" Phenomenon

The single most useful insight in this entire deck: marketing tier ≠ silicon tier. NVIDIA sells the same physical GPU die in three pillars at vastly different prices, with the differences driven by VRAM count, clocks, BIOS strap, and licence — not silicon.

GA102 (Ampere) — one die, six cards

CardPillarVRAMECCNVLinkEULATDP
RTX 3080 TiConsumer12 GB GDDR6XNoNoGeForce350 W
RTX 3090Consumer24 GB GDDR6XNoBridgeGeForce350 W
RTX 3090 TiConsumer24 GB GDDR6XNoBridgeGeForce450 W
RTX A6000Workstation48 GB GDDR6 ECCYesBridgePro/DC300 W
A40Datacenter48 GB GDDR6 ECCYesBridgeDC300 W
A10Datacenter24 GB GDDR6YesNoDC150 W

AD102 (Ada Lovelace) — one die, five cards

CardPillarVRAMECCFP8EULATDP
RTX 4080 Super (actually AD103)Consumer16 GB GDDR6XNoYes (no TE)GeForce320 W
RTX 4090Consumer24 GB GDDR6XNoYes (no TE)GeForce450 W
RTX 6000 AdaWorkstation48 GB GDDR6 ECCYesYesPro/DC300 W
L40Datacenter48 GB GDDR6 ECCYesYesDC300 W
L40SDatacenter48 GB GDDR6 ECCYesYes (higher throughput)DC350 W

GB100 / GB202 (Blackwell)

CardDiePillarVRAMNotes
RTX 5090GB202Consumer32 GB GDDR7GeForce flagship
RTX Pro 6000 BlackwellGB202Workstation96 GB GDDR7 ECCSame die, 3× VRAM via clamshell
B100GB100 (dual-die)Datacenter192 GB HBM3e700 W
B200GB100 (dual-die)Datacenter192 GB HBM3e1000 W — same silicon, higher power bin
Buying logic

Don't buy the marketing tier — buy the firmware features you need. If your job needs ECC, NVLink, MIG, or datacenter EULA, those features are gated by SKU choice within the same silicon. If your job needs only single-stream raw throughput, the cheapest card on the right die wins. The L40S vs RTX 6000 Ada example is a good illustration: same AD102 die, same 48 GB ECC, both have FP8 tensor cores, but the L40S is the high-clock "Server" bin tuned for inference throughput — and for LLM serving in 2025+, that throughput uplift is often the deciding feature.

10

Driver Branches & Compute Capability

NVIDIA's driver isn't a single artefact — it's at least two parallel streams, and matching the driver to the silicon and the workload is essential operational hygiene.

Production Branch (PB)

~6-month cadence, longer support tail (~1 year). Validated against the NVIDIA AI Enterprise stack and certified server platforms. Used by datacenter operators and enterprise workstation users.

Examples (2024–2026): R535, R550, R570 PB.

Trade-off: conservative. Gets bug fixes; doesn't get the latest CUDA features as fast.

New Feature Branch (NFB)

Faster cadence (monthly to quarterly). Adds new CUDA toolkit features as they ship. Used by GeForce/Studio consumers, gamers, and bleeding-edge developers.

Examples (2024–2026): R545, R555, R560, R565, R570 NFB, R575.

Trade-off: aggressive. Newest features, occasional regressions.

Driver version ↔ CUDA toolkit support

DriverRequired forNotes
R525+Hopper (H100) basic supportCUDA 12.0; Hopper FP8 via Transformer Engine library
R535+Hopper FP8 productionCUDA 12.2; cuBLASLt FP8 GEMMs
R550+Hopper TMA, thread-block clustersCUDA 12.4; full Hopper feature set
R555+Blackwell previewCUDA 12.5; B100/B200 early access
R560+Pre-BlackwellCUDA 12.6; no sm_100/sm_120 target yet
R570+Blackwell + flash-attn 3 polishedCUDA 12.8+; first toolkit with sm_100/sm_120 (Blackwell), FP4 paths (MX-FP4 and NVFP4)

Compute Capability and kernel binaries

Every CUDA kernel is compiled for a specific Compute Capability (CC). A wheel built only for sm_80 won't use Hopper's FP8 tensor cores even on an H100; you'd need sm_90. The same wheel running on Blackwell needs sm_100 (B100/B200) or sm_120 (consumer Blackwell).

CCArchWhat unlocks at this CC
sm_75Turing2nd-gen Tensor Cores (first were Volta, sm_70); adds INT8/INT4
sm_80A100BF16, TF32, async copy, MIG
sm_86Consumer Ampere (3090)BF16, TF32, no MIG
sm_89Ada (4090, L40S, RTX 6000 Ada)FP8 tensor cores (all Ada); no Hopper Transformer Engine
sm_90Hopper (H100, H200)FP8 native, TMA, thread-block clusters, DSMEM
sm_100Datacenter Blackwell (B100, B200)FP4 (MX-FP4 and NVFP4), 5th-gen Tensor Cores, NVLink 5
sm_120Consumer Blackwell (RTX 50, RTX Pro 6000 Blackwell)FP4, GDDR7
Build wheels for the right CC
# nvcc explicit target list:
nvcc -gencode arch=compute_80,code=sm_80 \
     -gencode arch=compute_86,code=sm_86 \
     -gencode arch=compute_89,code=sm_89 \
     -gencode arch=compute_90,code=sm_90 \
     -gencode arch=compute_100,code=sm_100 \
     -gencode arch=compute_120,code=sm_120 \
     -O3 kernel.cu -o kernel

# PyTorch / torch.compile environment hint:
export TORCH_CUDA_ARCH_LIST="8.0;8.6;8.9;9.0;10.0;12.0"

# Verify the GPU's CC at runtime:
nvidia-smi --query-gpu=compute_cap --format=csv
The cliff

If a wheel was built only for sm_80 and you run it on a Blackwell card, the driver falls back to JIT-compiling PTX (slow first launch) or refuses if no PTX is embedded. Always verify the CC list embedded in your wheels matches your fleet — especially when migrating from H100 to B200 or from 4090 to 5090.

11

Interactive: Card-to-Use-Case Mapper

Pick a SKU. Get the family classification, the firmware features, the licence boundary, and a recommended role.

Family
—
VRAM
—
ECC
—
NVLink
MIG
—
DC EULA
—
Suggested role
—
Price band
—
Reading the picker

The most informative comparisons are within the same die: pick RTX 4090, then RTX 6000 Ada, then L40, then L40S. All AD102. The differences — ECC, EULA, FP8 enabled, VRAM, NVLink — are entirely firmware and PCB choices, not silicon. That's the whole game.