NVIDIA GPU Architectures Series — Presentation 29

Inside Ampere — 3rd-Gen Tensor Cores, 40 MB L2, MIG

Low-level deep dive into NVIDIA's 2020 architecture. Two physical dies on two different processes (GA100 on TSMC 7N, GA10x on Samsung 8N), the SM with 3rd-gen tensor cores adding TF32 / BF16 / 2:4 sparsity, the giant 40 MB L2 jump, NVLink 3.0, MIG hardware partitioning, and HBM2e at 2.0 TB/s.

GA100GA102GA104 A100RTX 3090RTX A6000 TSMC 7NSamsung 8N HBM2eGDDR6X NVLink 3.0MIG
GA100/GA10x → GPC → SM → TC3 → 40 MB L2 → HBM2e/GDDR6X → NVLink 3.0
00

Topics We'll Cover

01

Ampere at a Glance — Two Dies, Two Foundries

Ampere launched in May 2020 as two physically distinct die families on two foundries. GA100 is the HPC/AI die at TSMC 7N. GA102 / GA104 / GA106 / GA107 are graphics dies at Samsung 8N (a custom 8 nm process). Same architecture name, very different silicon.

DieFoundryTransAreaSMsTop SKU
GA100TSMC 7N54.2 B826 mm²128 (108 on A100)A100 80 GB SXM4
GA102Samsung 8N28.3 B628 mm²84 (all 84 on RTX 3090 Ti; 82 on 3090)RTX 3090 Ti / A6000 / A40
GA104Samsung 8N17.4 B392 mm²48 (all 48 on RTX 3070 Ti; 46 on 3070)RTX 3070 Ti / RTX A4000
GA106Samsung 8N12.0 B276 mm²30 (28 on RTX 3060)RTX 3060 12 GB
GA107Samsung 8N6.7 B200 mm²20RTX 3050 / A2

Datacenter (GA100) and consumer (GA10x) Ampere SMs are not the same SM. GA10x doubles FP32 cores per partition for graphics throughput; GA100 doubles FP64 cores instead. Read the next slides carefully.

02

GA100 Die — Block Diagram

Chip
GA100 die — 826 mm², 54.2 B transistors, TSMC 7N
GPCs
GPC0GPC1GPC2GPC3GPC4GPC5GPC6GPC7
TPCs / GPC
8 TPCs × 8 GPCs = 64 TPCs (128 SMs total; 108 enabled on A100)
L2
40 MB total (split across 8 partitions; ~5 MB per partition)
Memory
6-stack HBM2e die — 5 stacks active on A100 (1 disabled for yield), 5120-bit bus
IO
PCIe 4.0 x1612 NVLink 3.0
03

GA102 Die — The Consumer Cousin

GA102 powers the RTX 3080 / 3090 / 3090 Ti, the RTX A6000 / A40 workstation cards, and the GA102-based Tesla A40. Process Samsung 8N, 28.3 B transistors on 628 mm².

Chip
GA102 — 628 mm², 28.3 B transistors, Samsung 8N
GPCs
7 GPCs × 6 TPCs × 2 SMs = 84 SMs (all enabled on 3090 Ti; 82 on 3090)
L2
6 MB total (much smaller than GA100; graphics relies on lower-latency L1)
Memory
12 GDDR6X channels, 384-bit bus (RTX 3090/Ti, A6000)

Notable cuts vs GA100: FP64 lanes drop to 2/SM (1:64 of FP32); tensor cores stay 3rd-gen but fewer per SM at half the throughput; NVLink 3 reduced to a 4-link bridge on RTX 3090 / A6000, no NVSwitch.

04

The Ampere SM — Two Variants

GA100 SM (datacenter)

  • 4 partitions, 1 warp scheduler each
  • 64 FP32 cores (16/partition)
  • 64 INT32 cores (16/partition)
  • 32 FP64 cores (8/partition) — 1:2 FP64:FP32
  • 4 tensor cores (1/partition) — 3rd gen, much wider than Volta's
  • 32 LD/ST, 16 SFU
  • 192 KB L1+shared (configurable: 164/28, 132/60, 100/92, 64/128 KB)
  • 256 KB register file
  • No RT cores

GA10x SM (consumer/workstation)

  • 4 partitions, 1 warp scheduler each
  • 128 FP32 cores (32/partition) — the doubled-FP32 trick
  • 64 INT32 cores (16/partition; one of the two FP32 sets shares lanes with INT)
  • 2 FP64 cores per SM — 1:64 of FP32
  • 4 tensor cores (1/partition) — same 3rd gen but lower per-cycle throughput
  • 1 RT core per SM (2nd gen; adds hardware motion-blur interpolation)
  • 128 KB L1+shared
  • 256 KB register file
The "doubled FP32" trick — consumer only

GA10x partitions can dual-issue FP32 each cycle but only one of the two ports is also INT32-capable. Marketing counted "10496 CUDA cores" on the 3090 (10752 on the 3090 Ti); the realistic compute throughput depends on the FP/INT mix. On pure FP32 graphics shaders it matches; on mixed-INT/FP code it halves. GA100 keeps the more balanced one-FP32 + one-INT32 design that compute likes.

05

3rd-Generation Tensor Cores — TF32, BF16, 2:4 Sparsity

Each Ampere SM has only 4 tensor cores (Volta had 8) but each is 4× the throughput — net 2× FP16 per SM. New formats are the headline:

OperationA100 SXM4-80 peak (1410 MHz)RTX 3090 (1700 MHz)
FP32 FMA19.5 TFLOPS35.6 TFLOPS (FP32-doubled)
FP64 FMA9.7 TFLOPS0.55 TFLOPS
Tensor TF32156 TFLOPS35.6 TFLOPS
Tensor BF16 / FP16312 TFLOPS71 TFLOPS (FP32 accumulate; FP16-accumulate FP16 is 142)
Tensor BF16 / FP16 + 2:4 sparse624 TFLOPS142 TFLOPS
Tensor INT8624 TOPS284 TOPS
Tensor INT8 sparse1248 TOPS568 TOPS
06

L2 Cache Jump — 6 MB → 40 MB

The L2 grew nearly 7× from Volta to GA100 — from 6 MB to 40 MB. Practical effect: working sets that previously thrashed HBM now fit on-die. For ResNet-50 inference at batch 1, weights fit in L2 entirely on A100. For LLM decoding the KV cache spills past it, but L2 buffering of recently-touched cache lines doubles effective bandwidth.

Ampere also added residency control: cudaStreamAttrAccessPolicyWindow lets you mark an address range as "persistent", giving its lines higher L2 retention priority. Used by cuDNN to keep convolution weights pinned.

GA10x consumer dies have only 6 MB L2 — not the headline jump. Consumer Ampere relies on bigger per-SM L1 (128 KB vs GA100's 192 KB but with consumer-tuned L1 hit rates) and high-BW GDDR6X.

07

Process — TSMC 7N vs Samsung 8N

NVIDIA split its Ampere supply between two foundries to manage capacity and cost. TSMC 7N for GA100 (premium-priced compute die, high-margin) and Samsung 8N for GA10x graphics dies (cheaper, larger volume). 8N is technically Samsung's 10 nm-class 8 nm LPP — not equivalent to TSMC N7. Density numbers:

DieProcessTransAreaDensity (M/mm²)
GA100TSMC 7N54.2 B826 mm²65.6
GA102Samsung 8N28.3 B628 mm²45.1
GA104Samsung 8N17.4 B392 mm²44.4

The density gap (65.6 vs 45.1 M/mm²) explains why a 3090 needs more silicon than an A100 to look superficially similar — Samsung 8N is roughly an N12-class node by transistor density.

08

Voltage, Power, Form Factors

SKUForm factorTDPConnector
A100 SXM4 40/80 GBSXM4 mezzanine (HGX A100)400 WSXM4 socket (no aux)
A100 PCIe 40/80 GBPCIe 4.0 x16 dual-slot250–300 W1× 8-pin EPS
RTX 3090 TiPCIe (consumer triple-slot)450 W1× 16-pin (12VHPWR)
RTX 3090PCIe (consumer triple-slot)350 W2× 8-pin or 1× 12-pin (FE)
RTX A6000PCIe (workstation dual-slot)300 W1× 8-pin EPS
A40PCIe (datacenter dual-slot)300 W1× 8-pin EPS, ECC GDDR6
A2PCIe (single-slot, low-profile)40–60 Wslot-only, 10 SMs from GA107

RTX 3090 Ti debuted the 16-pin 12VHPWR connector (later renamed 12V-2×6 after melting incidents on Ada). Voltage rails: GA100 has separate VDD, VDD-HBM, VDD-NVLink-IO, VDD-PLL; GA10x has VDD, VDDQ-GDDR6X (1.35 V), VDD-PLL.

09

Memory — HBM2e (Datacenter), GDDR6X (Consumer)

HBM2e on GA100

  • 6-stack HBM2e die; 5 stacks active on A100 (1 disabled for yield)
  • Stack: 8-Hi, 8 GB (A100 40 GB) or 16 GB (A100 80 GB)
  • Pin rate 2.43 Gbps (40GB) or 3.2 Gbps (80GB)
  • Per-stack BW 311 GB/s (40GB) or 410 GB/s (80GB)
  • 5120-bit total bus (5 active stacks × 1024-bit)
  • Aggregate 1555 GB/s (40GB) / 2039 GB/s (80GB)
  • ECC SECDED + on-die ECC (HBM2e first added on-die ECC)

GDDR6X on GA102

  • 12 channels × 32-bit = 384-bit bus on RTX 3090/Ti and A6000
  • PAM4 signalling — 4 levels per symbol, doubles data per clock vs GDDR6's NRZ
  • 19 Gbps (3080), 19.5 Gbps (3090), 21 Gbps (3090 Ti)
  • Per-card 760–1008 GB/s
  • No ECC on consumer; ECC enabled on A40/A6000 GDDR6 (no GDDR6X for ECC parts)
  • PAM4 signalling cost: tighter eye, more power per bit, hotter VRAM
10

NVLink 3.0 & NVSwitch 2 (HGX A100)

NVLink 3.0 keeps per-link bandwidth the same as NVLink 2 (25 GB/s/dir, 50 GB/s bidirectional) but doubles the link count. A100 has 12 links → 300 GB/s/dir aggregate, 600 GB/s bidirectional. Signalling: 50 Gbps NRZ per lane, 4 lanes per link — reused the same 25 Gbps SerDes physical hardware as NVLink 2 but halved the lane count per link, with FEC (forward error correction) at the link layer to reach 50 Gbps reliably.

PropertyNVLink 2.0NVLink 3.0
Per-lane rate25 Gbps NRZ50 Gbps NRZ + FEC
Lanes per link84
Per-link BW (1-dir)25 GB/s25 GB/s
Wait, what?NVLink 3.0 keeps the per-link headline number but doubles link count: A100 has 12 links instead of V100's 6.
Total per-GPU (1-dir)150 GB/s (V100)300 GB/s (A100; 600 GB/s bidirectional)

HGX A100 baseboard: 8× A100 SXM4 + 6× NVSwitch 2.0 chips (each NVSwitch 2.0 has 36 NVLink 3 ports, 1.8 TB/s aggregate). Every A100 talks to every other at full 600 GB/s. Cross-node uses 8× ConnectX-6 HCAs, 200 Gb/s HDR InfiniBand each.

Consumer Ampere (RTX 3090 / A6000 / A40) carries a 4-link NVLink bridge for two-card pairing: 112.5 GB/s bidirectional aggregate. RTX 3080 and below have no NVLink.

11

Ampere's Innovations

12

Interactive: Ampere SKU Picker

Die / Foundry
—
SMs
—
FP32 TF
—
BF16 TC TF
—
VRAM
—
BW (GB/s)
—
TDP
—
NVLink / MIG
—