Low-level deep dive into NVIDIA's 2020 architecture. Two physical dies on two different processes (GA100 on TSMC 7N, GA10x on Samsung 8N), the SM with 3rd-gen tensor cores adding TF32 / BF16 / 2:4 sparsity, the giant 40 MB L2 jump, NVLink 3.0, MIG hardware partitioning, and HBM2e at 2.0 TB/s.
Ampere launched in May 2020 as two physically distinct die families on two foundries. GA100 is the HPC/AI die at TSMC 7N. GA102 / GA104 / GA106 / GA107 are graphics dies at Samsung 8N (a custom 8 nm process). Same architecture name, very different silicon.
| Die | Foundry | Trans | Area | SMs | Top SKU |
|---|---|---|---|---|---|
| GA100 | TSMC 7N | 54.2 B | 826 mm² | 128 (108 on A100) | A100 80 GB SXM4 |
| GA102 | Samsung 8N | 28.3 B | 628 mm² | 84 (all 84 on RTX 3090 Ti; 82 on 3090) | RTX 3090 Ti / A6000 / A40 |
| GA104 | Samsung 8N | 17.4 B | 392 mm² | 48 (all 48 on RTX 3070 Ti; 46 on 3070) | RTX 3070 Ti / RTX A4000 |
| GA106 | Samsung 8N | 12.0 B | 276 mm² | 30 (28 on RTX 3060) | RTX 3060 12 GB |
| GA107 | Samsung 8N | 6.7 B | 200 mm² | 20 | RTX 3050 / A2 |
Datacenter (GA100) and consumer (GA10x) Ampere SMs are not the same SM. GA10x doubles FP32 cores per partition for graphics throughput; GA100 doubles FP64 cores instead. Read the next slides carefully.
GA102 powers the RTX 3080 / 3090 / 3090 Ti, the RTX A6000 / A40 workstation cards, and the GA102-based Tesla A40. Process Samsung 8N, 28.3 B transistors on 628 mm².
Notable cuts vs GA100: FP64 lanes drop to 2/SM (1:64 of FP32); tensor cores stay 3rd-gen but fewer per SM at half the throughput; NVLink 3 reduced to a 4-link bridge on RTX 3090 / A6000, no NVSwitch.
GA10x partitions can dual-issue FP32 each cycle but only one of the two ports is also INT32-capable. Marketing counted "10496 CUDA cores" on the 3090 (10752 on the 3090 Ti); the realistic compute throughput depends on the FP/INT mix. On pure FP32 graphics shaders it matches; on mixed-INT/FP code it halves. GA100 keeps the more balanced one-FP32 + one-INT32 design that compute likes.
Each Ampere SM has only 4 tensor cores (Volta had 8) but each is 4× the throughput — net 2× FP16 per SM. New formats are the headline:
| Operation | A100 SXM4-80 peak (1410 MHz) | RTX 3090 (1700 MHz) |
|---|---|---|
| FP32 FMA | 19.5 TFLOPS | 35.6 TFLOPS (FP32-doubled) |
| FP64 FMA | 9.7 TFLOPS | 0.55 TFLOPS |
| Tensor TF32 | 156 TFLOPS | 35.6 TFLOPS |
| Tensor BF16 / FP16 | 312 TFLOPS | 71 TFLOPS (FP32 accumulate; FP16-accumulate FP16 is 142) |
| Tensor BF16 / FP16 + 2:4 sparse | 624 TFLOPS | 142 TFLOPS |
| Tensor INT8 | 624 TOPS | 284 TOPS |
| Tensor INT8 sparse | 1248 TOPS | 568 TOPS |
The L2 grew nearly 7× from Volta to GA100 — from 6 MB to 40 MB. Practical effect: working sets that previously thrashed HBM now fit on-die. For ResNet-50 inference at batch 1, weights fit in L2 entirely on A100. For LLM decoding the KV cache spills past it, but L2 buffering of recently-touched cache lines doubles effective bandwidth.
Ampere also added residency control: cudaStreamAttrAccessPolicyWindow lets you mark an address range as "persistent", giving its lines higher L2 retention priority. Used by cuDNN to keep convolution weights pinned.
GA10x consumer dies have only 6 MB L2 — not the headline jump. Consumer Ampere relies on bigger per-SM L1 (128 KB vs GA100's 192 KB but with consumer-tuned L1 hit rates) and high-BW GDDR6X.
NVIDIA split its Ampere supply between two foundries to manage capacity and cost. TSMC 7N for GA100 (premium-priced compute die, high-margin) and Samsung 8N for GA10x graphics dies (cheaper, larger volume). 8N is technically Samsung's 10 nm-class 8 nm LPP — not equivalent to TSMC N7. Density numbers:
| Die | Process | Trans | Area | Density (M/mm²) |
|---|---|---|---|---|
| GA100 | TSMC 7N | 54.2 B | 826 mm² | 65.6 |
| GA102 | Samsung 8N | 28.3 B | 628 mm² | 45.1 |
| GA104 | Samsung 8N | 17.4 B | 392 mm² | 44.4 |
The density gap (65.6 vs 45.1 M/mm²) explains why a 3090 needs more silicon than an A100 to look superficially similar — Samsung 8N is roughly an N12-class node by transistor density.
| SKU | Form factor | TDP | Connector |
|---|---|---|---|
| A100 SXM4 40/80 GB | SXM4 mezzanine (HGX A100) | 400 W | SXM4 socket (no aux) |
| A100 PCIe 40/80 GB | PCIe 4.0 x16 dual-slot | 250–300 W | 1× 8-pin EPS |
| RTX 3090 Ti | PCIe (consumer triple-slot) | 450 W | 1× 16-pin (12VHPWR) |
| RTX 3090 | PCIe (consumer triple-slot) | 350 W | 2× 8-pin or 1× 12-pin (FE) |
| RTX A6000 | PCIe (workstation dual-slot) | 300 W | 1× 8-pin EPS |
| A40 | PCIe (datacenter dual-slot) | 300 W | 1× 8-pin EPS, ECC GDDR6 |
| A2 | PCIe (single-slot, low-profile) | 40–60 W | slot-only, 10 SMs from GA107 |
RTX 3090 Ti debuted the 16-pin 12VHPWR connector (later renamed 12V-2×6 after melting incidents on Ada). Voltage rails: GA100 has separate VDD, VDD-HBM, VDD-NVLink-IO, VDD-PLL; GA10x has VDD, VDDQ-GDDR6X (1.35 V), VDD-PLL.
NVLink 3.0 keeps per-link bandwidth the same as NVLink 2 (25 GB/s/dir, 50 GB/s bidirectional) but doubles the link count. A100 has 12 links → 300 GB/s/dir aggregate, 600 GB/s bidirectional. Signalling: 50 Gbps NRZ per lane, 4 lanes per link — reused the same 25 Gbps SerDes physical hardware as NVLink 2 but halved the lane count per link, with FEC (forward error correction) at the link layer to reach 50 Gbps reliably.
| Property | NVLink 2.0 | NVLink 3.0 |
|---|---|---|
| Per-lane rate | 25 Gbps NRZ | 50 Gbps NRZ + FEC |
| Lanes per link | 8 | 4 |
| Per-link BW (1-dir) | 25 GB/s | 25 GB/s |
| Wait, what? | NVLink 3.0 keeps the per-link headline number but doubles link count: A100 has 12 links instead of V100's 6. | |
| Total per-GPU (1-dir) | 150 GB/s (V100) | 300 GB/s (A100; 600 GB/s bidirectional) |
HGX A100 baseboard: 8× A100 SXM4 + 6× NVSwitch 2.0 chips (each NVSwitch 2.0 has 36 NVLink 3 ports, 1.8 TB/s aggregate). Every A100 talks to every other at full 600 GB/s. Cross-node uses 8× ConnectX-6 HCAs, 200 Gb/s HDR InfiniBand each.
Consumer Ampere (RTX 3090 / A6000 / A40) carries a 4-link NVLink bridge for two-card pairing: 112.5 GB/s bidirectional aggregate. RTX 3080 and below have no NVLink.