The architecture that trained GPT-3 and BERT-Large, gave us BF16 and TF32 as drop-in FP32 replacements, introduced 2:4 structured sparsity and MIG hardware partitioning. Years on, the A100 is still the workhorse of mid-tier training and the RTX 3090 still serves quants at home.
A focused tour of the NVIDIA Ampere generation: the chips, the SM, the tensor-core format revolution, the partitioning story, and the practical numbers that matter for fine-tuning and inference today.
Ampere was NVIDIA's first architecture designed with transformer training as a primary workload. Volta (2017) had introduced tensor cores; Turing (2018) brought them to consumer cards; Ampere (2020) is where the format choices, capacity targets, and partitioning story finally fit the shape of a transformer pretraining run.
The headline trick: 3rd-gen tensor cores with TF32. TF32 is a 19-bit format with FP32-style 8-bit exponent and a 10-bit mantissa. It is a drop-in replacement for FP32 matmul — existing pretraining code “just works” at roughly 8× the throughput, with no autocast and no loss-scaling dance. The other formats — BF16 and INT8 with 2:4 structured sparsity — are what unlocked the GPT-3 era.
Two physical chips, very different worlds:
GA100 (A100 40 GB) launched May 2020. GA102 (RTX 3090) launched September 2020. The A100 80 GB upgrade was announced in November 2020 (PCIe 80 GB in mid-2021). Ampere is the architecture that defined the “LLM era” — OPT, BLOOM, Megatron-Turing NLG and the first wave of Llama all trained on A100 silicon (GPT-3 itself was trained on V100s).
GA100 on TSMC 7N. GA10x on Samsung 8N (a relabelled 10nm-class node). The split is why A100 hits 1.41 GHz at 400 W and GA102 needs 350 W to hit clock parity.
HBM2 / HBM2e on GA100 (1.5–2.0 TB/s). GDDR6X on GA10x (936 GB/s on 3090). The bandwidth gap is the single biggest performance differentiator across the family.
CC 8.0 on GA100; CC 8.6 on GA10x. They share TF32, BF16, FP16 tensor ops, but only GA100 has full FP64 tensor cores and MIG.
The A100 is the binned-down product. The full GA100 is bigger than what ships — NVIDIA leaves headroom for yield. Knowing the full layout helps explain odd numbers like “108 SMs”.
128 SMs total on the die, organised as 8 GPCs × 16 SMs. A100 ships with 108 enabled (binning leaves 20 disabled for yield). Each GPC contains 8 TPCs, each TPC two SMs.
826 mm2, 54.2 billion transistors on TSMC 7N. By comparison V100 (Volta) was 815 mm2 at 21.1 B transistors — Ampere ships 2.5× the transistors in essentially the same area.
6 HBM stack sites (5 active) on a 5120-bit memory bus. Two SKUs:
40 MB unified L2 — ~6.7× larger than V100's 6 MB. Split into two halves with a crossbar; the SM scheduler keeps load local.
12× NVLink 3.0 links, 25 GB/s/dir (50 GB/s bidirectional) each → 600 GB/s aggregate off-package.
HBM2e is fast but it isn't free. A 40 MB L2 lets transformer kernels — especially attention with KV-cache reuse — keep working sets on-die. cuDNN and cuBLAS were rewritten around this number. On Hopper this becomes 50 MB; on Blackwell B200, ~63 MB per die (126 MB total). The Ampere L2 set the template.
An Ampere SM is structured into 4 partitions. Each partition has its own warp scheduler, register file slice, and tensor core.
Pre-Ampere, every shared-memory load relayed through registers: global → reg → shared. On a transformer GEMM that means the register file becomes a strict gatekeeper for occupancy.
Ampere introduces cp.async in PTX: the SM issues an asynchronous copy that goes L2 → shared memory directly, bypassing the register file. Combined with a multi-stage pipeline (typically 3 stages), it lets a CUDA programmer prefetch the next tile while the current one is being multiplied — without burning registers on the in-flight data.
CUDA 11 + Ampere promote cooperative thread arrays from a curiosity to a routine. cooperative_groups::thread_block_tile<32>, multi-block sync via the grid group, and async memcpy with cuda::memcpy_async in libcu++ all assume Ampere as the floor.
Flash-Attention v1 (Dao et al, 2022) is essentially a cp.async-shaped algorithm. It tiles attention through SRAM and overlaps the next K/V load with the current softmax. The kernel that made attention IO-bound instead of compute-bound was, at the hardware level, an Ampere kernel before it became anything else.
Volta tensor cores did FP16 with FP32 accumulation. Turing added INT8 / INT4. Ampere kept all that and added three things that genuinely changed how people train transformers.
19-bit format: 1 sign + 8 exponent + 10 mantissa. Same dynamic range as FP32 (because of the 8-bit exponent), with FP16's mantissa precision.
Drop-in for FP32 matmul. Existing PyTorch / TensorFlow code with no autocast runs at roughly 8× FP32 throughput — the runtime simply rounds FP32 inputs to TF32 before the tensor-core MMA.
Default-on in cuDNN/cuBLAS from CUDA 11. Toggle with torch.backends.cuda.matmul.allow_tf32.
16-bit: 1 sign + 8 exponent + 7 mantissa. Same dynamic range as FP32 (the 8-bit exponent), but FP16 storage cost.
This is the format that made stable mixed-precision training routine. Unlike FP16, BF16 doesn't underflow gradients, so loss scaling becomes optional. Every modern transformer pretrain uses BF16.
Same throughput as FP16 on Ampere; no accuracy/speed tradeoff — pure win.
Hardware-supported structured sparsity. Every 4-element block of a weight matrix carries a 2-bit mask selecting 2 non-zero values. The tensor core skips the zeros.
Doubles throughput when the model has been pruned + retrained to the 2:4 pattern. Real-world adoption is patchy — aggressive quantisation often beats it — but the silicon is there.
A100 SXM4 80 GB, dense unless noted. Numbers from the NVIDIA A100 datasheet.
| Format | Dense TFLOPS | 2:4 Sparse TFLOPS | Notes |
|---|---|---|---|
| FP64 tensor | 19.5 | — | Genuine FP64 in tensor cores (HPC win) |
| FP32 (CUDA core) | 19.5 | — | No tensor-core FP32; use TF32 |
| TF32 | 156 | 312 | Drop-in FP32 replacement |
| BF16 / FP16 | 312 | 624 | Standard transformer training format |
| INT8 | 624 TOPS | 1248 TOPS | Inference |
| INT4 | 1248 TOPS | 2496 TOPS | Inference (rare in practice) |
FP8 is a Hopper feature (CC 9.0). On A100, FP8 inference is software-emulated — you save weight-storage bandwidth but the math still happens in BF16/FP16 tensor cores. Don't expect compute uplift from FP8 on Ampere.
Multi-Instance GPU is Ampere's most underrated feature. An A100 can be carved at the hardware level into up to seven isolated GPU instances, each with its own dedicated SMs, L2 slice, and HBM region. Not virtualisation — isolation is enforced by the memory crossbar, not by software.
The partition unit is one GPC (which on A100 corresponds to ~14 SMs, ~5 GB HBM, ~5 MB L2). Available shapes:
| Profile | GPCs | VRAM | SMs | L2 | Use case |
|---|---|---|---|---|---|
| 1g.5gb | 1 | 5 GB | ~14 | ~5 MB | Single small inference replica, dev sandbox |
| 2g.10gb | 2 | 10 GB | ~28 | ~10 MB | 7B inference, embedding service |
| 3g.20gb | 3 | 20 GB | ~42 | ~20 MB | 13B inference, mid-batch serving |
| 4g.20gb | 4 | 20 GB | ~56 | ~20 MB | (rare) compute-heavy with same RAM |
| 7g.40gb | 7 | 40 GB (or 80) | 98 (7 × 14) | 40 MB | The whole GPU, MIG mode “on” |
# enable MIG mode (drains running work; needs reset)
nvidia-smi -mig 1
# list available profiles
nvidia-smi mig -lgip
# create the slices on GPU 0
nvidia-smi mig -cgi 9,14,19,19 -C
# 9 = 3g.20gb, 14 = 2g.10gb, 19 = 1g.5gb (profile IDs vary by SKU)
# each slice gets its own UUID; pin a container to one
docker run --gpus '"device=MIG-GPU-..."' vllm/vllm-openai ...
No NVLink between MIG instances on the same physical GPU. No P2P. CUDA graphs work per slice but not across. MIG is for serving / dev-pool isolation, not for splitting a training job. Hopper extends MIG with TEE (confidential compute) and per-slice MPS; Blackwell B200 inherits both.
The other half of Ampere lives on Samsung 8N. Same 3rd-gen tensor cores, same TF32/BF16 story, but no HBM, no NVSwitch, and a quietly different SM that has been the source of years of confusion.
| Card | Die | CUDA cores | VRAM | Bandwidth | TDP | NVLink |
|---|---|---|---|---|---|---|
| RTX 3090 Ti | GA102 | 10752 | 24 GB GDDR6X | 1008 GB/s | 450 W | bridge |
| RTX 3090 | GA102 | 10496 | 24 GB GDDR6X | 936 GB/s | 350 W | bridge |
| RTX 3080 Ti | GA102 | 10240 | 12 GB GDDR6X | 912 GB/s | 350 W | none |
| RTX 3080 | GA102 | 8704 / 8960 | 10 / 12 GB | 760 / 912 GB/s | 320 / 350 W | none |
| RTX 3070 Ti | GA104 | 6144 | 8 GB GDDR6X | 608 GB/s | 290 W | none |
| RTX 3070 | GA104 | 5888 | 8 GB GDDR6 | 448 GB/s | 220 W | none |
| RTX 3060 Ti | GA104 | 4864 | 8 GB GDDR6 | 448 GB/s | 200 W | none |
| RTX 3060 12 GB | GA106 | 3584 | 12 GB GDDR6 | 360 GB/s | 170 W | none |
| RTX A6000 | GA102 | 10752 | 48 GB ECC GDDR6 | 768 GB/s | 300 W | bridge |
| A40 | GA102 | 10752 | 48 GB GDDR6 ECC | 696 GB/s | 300 W | bridge |
On GA100, a partition has 16 FP32 cores and 16 INT32 cores; an SM totals 64 FP32. On GA10x, NVIDIA doubled the FP32 count per partition: each partition has 16 FP32 + 16 dual-issue (FP32 or INT32) cores. An SM totals 128 FP32 when running pure FP32, but only 64 FP32 if INT32 is also live (because the dual-issue cores get co-opted).
This is why an RTX 3090 advertises 10496 CUDA cores when GA100 with the same silicon area would have ~5000. The marketing number is the FP32-only peak; real shader workloads alternate FP32 and INT32 (address arithmetic, indexing) and rarely sustain it. The tensor-core count is unchanged — 4 per SM, just like GA100.
For LLM inference (which is tensor-core bound when prefilling and bandwidth-bound when decoding), the doubled FP32 cores are roughly irrelevant. A 3090's 936 GB/s of GDDR6X matters far more than its 35 TFLOPS of FP32. For graphics workloads, the dual-issue path is genuinely useful. Don't compare cards by CUDA-core count alone.
The RTX A6000 and A40 are Ampere's 48 GB GDDR6 ECC options. Same GA102 die as a 3090, but with the full memory complement, ECC enabled, and a datacenter-blessed driver. The A6000 supports an NVLink bridge between two cards (112 GB/s) — not switch-class, but fine for 2-way TP on 70B models. Used 48 GB Ampere workstation cards remain a 2026 sweet spot for serious local work.
NVLink 3.0 doubles the per-lane signalling rate from NVLink 2.0 (50 vs 25 Gb/s) but halves the lanes per link, so each link stays at 25 GB/s per direction. A100 doubles the link count to 12 → 600 GB/s aggregate bidirectional. NVSwitch 2.0 is the fabric chip that ties them together.
Meta's OPT-175B (2022), BLOOM (2022), Megatron-Turing NLG (2022), and most large public LLMs from this period were trained on DGX A100 or HGX A100-derived systems. The 3D-parallelism recipe (TP × PP × DP) was tuned around exactly this baseboard.
Per layer, a tensor-parallel all-reduce needs to ship one full layer's hidden state. For Llama-2-70B with hidden dim 8192 and BF16: 8192 × seq × 2 bytes per token. At seq=4096 that's 67 MB per layer.
That's a 10–20× latency gap. With 80 transformer layers, each step accumulates ~80 × 75 µs = 6 ms on NVSwitch vs 80 ms on PCIe. NVLink is why TP=8 works on a single A100 node.
The PCIe form-factor A100 has only 3× NVLink 3 bridges (not 12), giving 600 GB/s only between paired cards via bridge, not arbitrary-pair. For real multi-GPU work always pick SXM4 if you can.
In November 2020, NVIDIA announced the A100 80 GB. Same die. Same SMs. Same NVLink. The change was on the memory side: doubled HBM capacity (40 GB HBM2 → 80 GB HBM2e) and a 31% bandwidth bump (1.5 → 2.0 TB/s) thanks to faster HBM2e clocks.
| Spec | A100 40 GB | A100 80 GB | Delta |
|---|---|---|---|
| HBM capacity | 40 GB | 80 GB | +100% |
| Memory bandwidth | 1.555 TB/s | 2.039 TB/s | +31% |
| SMs (enabled) | 108 | 108 | 0% |
| L2 cache | 40 MB | 40 MB | 0% |
| BF16 dense TFLOPS | 312 | 312 | 0% |
| BF16 sparse TFLOPS | 624 | 624 | 0% |
| NVLink 3 aggregate | 600 GB/s | 600 GB/s | 0% |
| TDP (SXM4) | 400 W | 400 W | 0% |
| MIG max instances | 7 | 7 | 0% |
| MIG max slice VRAM | 1g.5gb | 1g.10gb | +100% |
40 GB held 13B in BF16 fine, struggled at 30B, and was helpless at 70B+ without sharding. 80 GB removed sharding from the workflow for an entire class of model sizes:
A100 80 GB SXM4 modules are still going strong on the second-hand market. They are the price-performance sweet spot for fine-tuning Llama-3-70B, Mixtral, and similar mid-tier models. Hopper and Blackwell are faster per-dollar for new builds, but a used A100 80 GB cluster remains a credible 2026 platform — especially for INT4 inference, where the bottleneck is bandwidth and not FP8 compute.
Capacity and bandwidth, not flops, are the LLM constraint. The A100 80 GB upgrade ships the same compute and adds nothing-but-memory — and yet it changed which workloads were practical. This pattern recurs: H100 80 GB → H200 141 GB → B200 192 GB. Each generation, the headline win is HBM.
Where does Ampere sit in 2026? Pretraining at frontier scale has moved decisively to H100 / H200 / B200 — FP8 doubles compute and the Transformer Engine's per-tensor scaling is genuinely difficult to give up. But Ampere remains unbeatable in three places:
BF16 LoRA and full fine-tuning of 7B–70B models. The compute headroom (312 BF16 TFLOPS) is more than enough; the 80 GB HBM keeps optimiser state on-card. PyTorch FSDP works beautifully on DGX A100.
Pretraining models in the 1B–15B range. A 4× A100 80 GB node trains a 7B from scratch in days, not months, and the second-hand cost is a fraction of an H100 rig.
AWQ and GPTQ quantised serving. INT4 weights fit gigantic models in modest VRAM, and the bottleneck moves to bandwidth — where A100 80 GB's 2 TB/s holds up well.
You can “run” FP8 on A100 via libraries like vLLM's FP8-W8A8 path, but be clear about what's happening: the weights are stored as FP8 (saving HBM bandwidth on weight load), then dequantised into BF16 for tensor-core math. There is no compute uplift — throughput is the same as plain BF16. The win is purely a smaller working set.
For decode (memory-bound), this is still useful: you cut weight bandwidth in half and tok/s improves by 30–50%. For prefill (compute-bound), it's a wash. Hopper's native FP8 doubles compute as well; that's the meaningful gap.
Two A100 80 GB SXM4 with TP=2 over NVLink, vLLM 0.5+, batch sizes from solo user to mid-batch:
| Workload | Configuration | Tokens/s |
|---|---|---|
| Single user, 1 stream | seq=2048, INT4 weights, BF16 KV | 30–45 tok/s |
| Batch=8 concurrent | same | ~250 tok/s aggregate |
| Batch=32 concurrent | same | ~600 tok/s aggregate |
| Prefill burst (compute) | 4096 prompt, single seq | ~12000 tok/s prefill |
A single A100 80 GB at INT4 holds Llama-3-70B with ~40 GB headroom for KV-cache. Two cards via NVLink approximately double throughput on prefill and add ~70–80% on decode (TP overhead).
If your workload is FP8 pretraining, use Hopper. If it's FP4 inference, use Blackwell. If it's multi-trillion-token data curation with TE-fp8 numerics, use Hopper. For everything else — especially fine-tuning and INT4 serving — Ampere remains a remarkable price/perf sweet spot.
Ampere has the most mature software stack of any NVIDIA generation simply because it has been in production the longest. Five years of bug fixes, kernel tuning, and library work all assume CC 8.0 / 8.6 as the baseline.
torch.cuda.amp.autocast(dtype=torch.bfloat16) — mostly painlesstorch.backends.cuda.matmul.allow_tf32 = True — default on, gives free FP32-→TF32 speeduptorch.compile works, though TE / FP8 paths require HopperIf you're writing new transformer code, the right default is BF16 explicitly, not TF32. Reason: cuDNN/cuBLAS schedule a different (and faster) tensor-core kernel for BF16 matmuls than for TF32. TF32 is the “don't break my old code” format; BF16 is the “I'm writing this from scratch” format. On A100, BF16 matmul is faster than TF32 matmul (312 TFLOPS vs 156 TFLOPS) and accuracy is essentially identical for training transformers.
import torch
# keep TF32 enabled for any leftover FP32 codepaths
torch.backends.cuda.matmul.allow_tf32 = True
torch.backends.cudnn.allow_tf32 = True
# but actually train in BF16 explicitly
with torch.autocast(device_type='cuda', dtype=torch.bfloat16):
out = model(x)
loss = criterion(out, y)
loss.backward()
Pin to a stable production driver (R535 or R550 LTS branches are good for Ampere). Ampere works with newer drivers but you don't need them — only Blackwell and Hopper FP8 features need bleeding-edge driver minor versions. Long-running A100 fleets benefit from the LTS branches.
Pick an Ampere SKU; see the headline numbers, what it means for LLM capacity at FP16 and INT4, and what it's actually good for.
“Largest LLM” is a back-of-envelope estimate: VRAM ÷ 2 (B params) for FP16 weight-only fit, VRAM ÷ 0.6 for INT4. Real workloads need additional headroom for KV-cache, activations, and CUDA overhead — subtract 10–20%.