NVIDIA GPU Architectures Series — Presentation 06

Ampere — A100, RTX 30, and the Birth of the LLM Era

The architecture that trained GPT-3 and BERT-Large, gave us BF16 and TF32 as drop-in FP32 replacements, introduced 2:4 structured sparsity and MIG hardware partitioning. Years on, the A100 is still the workhorse of mid-tier training and the RTX 3090 still serves quants at home.

A100GA100GA102 RTX 3090A40A6000 BF16TF32 MIGNVLink 3
GA100 → SM → TC3 → MIG → NVLink 3 → BF16/TF32 → 2:4 → A100 80GB
00

Topics We'll Cover

A focused tour of the NVIDIA Ampere generation: the chips, the SM, the tensor-core format revolution, the partitioning story, and the practical numbers that matter for fine-tuning and inference today.

01

Ampere in One Page

Ampere was NVIDIA's first architecture designed with transformer training as a primary workload. Volta (2017) had introduced tensor cores; Turing (2018) brought them to consumer cards; Ampere (2020) is where the format choices, capacity targets, and partitioning story finally fit the shape of a transformer pretraining run.

The headline trick: 3rd-gen tensor cores with TF32. TF32 is a 19-bit format with FP32-style 8-bit exponent and a 10-bit mantissa. It is a drop-in replacement for FP32 matmul — existing pretraining code “just works” at roughly 8× the throughput, with no autocast and no loss-scaling dance. The other formats — BF16 and INT8 with 2:4 structured sparsity — are what unlocked the GPT-3 era.

Two physical chips, very different worlds:

Released

GA100 (A100 40 GB) launched May 2020. GA102 (RTX 3090) launched September 2020. The A100 80 GB upgrade was announced in November 2020 (PCIe 80 GB in mid-2021). Ampere is the architecture that defined the “LLM era” — OPT, BLOOM, Megatron-Turing NLG and the first wave of Llama all trained on A100 silicon (GPT-3 itself was trained on V100s).

Process

GA100 on TSMC 7N. GA10x on Samsung 8N (a relabelled 10nm-class node). The split is why A100 hits 1.41 GHz at 400 W and GA102 needs 350 W to hit clock parity.

Memory

HBM2 / HBM2e on GA100 (1.5–2.0 TB/s). GDDR6X on GA10x (936 GB/s on 3090). The bandwidth gap is the single biggest performance differentiator across the family.

Capability

CC 8.0 on GA100; CC 8.6 on GA10x. They share TF32, BF16, FP16 tensor ops, but only GA100 has full FP64 tensor cores and MIG.

02

GA100 — The Datacenter Die

The A100 is the binned-down product. The full GA100 is bigger than what ships — NVIDIA leaves headroom for yield. Knowing the full layout helps explain odd numbers like “108 SMs”.

Compute layout

128 SMs total on the die, organised as 8 GPCs × 16 SMs. A100 ships with 108 enabled (binning leaves 20 disabled for yield). Each GPC contains 8 TPCs, each TPC two SMs.

826 mm2, 54.2 billion transistors on TSMC 7N. By comparison V100 (Volta) was 815 mm2 at 21.1 B transistors — Ampere ships 2.5× the transistors in essentially the same area.

Memory and fabric

6 HBM stack sites (5 active) on a 5120-bit memory bus. Two SKUs:

  • A100 40 GB — 1.5 TB/s
  • A100 80 GB — 2.0 TB/s (HBM2e clocked higher)

40 MB unified L2 — ~6.7× larger than V100's 6 MB. Split into two halves with a crossbar; the SM scheduler keeps load local.

12× NVLink 3.0 links, 25 GB/s/dir (50 GB/s bidirectional) each → 600 GB/s aggregate off-package.

die
GA100 (826 mm², 54.2B txn)
8× GPC
GPC0GPC1 GPC2GPC3 GPC4GPC5 GPC6GPC7
16 SM/GPC
128 SM total 108 enabled on A100 20 disabled (yield)
L2
40 MB unified 2 partitions + xbar
HBM
6 stacks 5120-bit bus 1.5 / 2.0 TB/s
NVLink
12 lanes × 50 GB/s 600 GB/s total
Why a giant L2

HBM2e is fast but it isn't free. A 40 MB L2 lets transformer kernels — especially attention with KV-cache reuse — keep working sets on-die. cuDNN and cuBLAS were rewritten around this number. On Hopper this becomes 50 MB; on Blackwell B200, ~63 MB per die (126 MB total). The Ampere L2 set the template.

03

The Ampere SM

An Ampere SM is structured into 4 partitions. Each partition has its own warp scheduler, register file slice, and tensor core.

Per partition

  • 16 FP32 CUDA cores
  • 16 INT32 CUDA cores
  • 4 LD/ST units
  • 1 3rd-gen tensor core
  • 1 SFU (special function unit)
  • 16K register file slots

Per SM (×4)

  • 64 FP32 cores
  • 64 INT32 cores
  • 4 tensor cores (TC3)
  • 192 KB unified L1 / shared mem
  • 256 KB register file (64K × 32-bit)

Async copy: registers stop being a bottleneck

Pre-Ampere, every shared-memory load relayed through registers: global → reg → shared. On a transformer GEMM that means the register file becomes a strict gatekeeper for occupancy.

Ampere introduces cp.async in PTX: the SM issues an asynchronous copy that goes L2 → shared memory directly, bypassing the register file. Combined with a multi-stage pipeline (typically 3 stages), it lets a CUDA programmer prefetch the next tile while the current one is being multiplied — without burning registers on the in-flight data.

global mem
→
L2 (40 MB)
→
cp.async
→
shared mem
→
tensor core

Cooperative groups become first-class

CUDA 11 + Ampere promote cooperative thread arrays from a curiosity to a routine. cooperative_groups::thread_block_tile<32>, multi-block sync via the grid group, and async memcpy with cuda::memcpy_async in libcu++ all assume Ampere as the floor.

Why this matters for transformers

Flash-Attention v1 (Dao et al, 2022) is essentially a cp.async-shaped algorithm. It tiles attention through SRAM and overlaps the next K/V load with the current softmax. The kernel that made attention IO-bound instead of compute-bound was, at the hardware level, an Ampere kernel before it became anything else.

04

3rd-Gen Tensor Cores — The Format Revolution

Volta tensor cores did FP16 with FP32 accumulation. Turing added INT8 / INT4. Ampere kept all that and added three things that genuinely changed how people train transformers.

(a) TF32

19-bit format: 1 sign + 8 exponent + 10 mantissa. Same dynamic range as FP32 (because of the 8-bit exponent), with FP16's mantissa precision.

Drop-in for FP32 matmul. Existing PyTorch / TensorFlow code with no autocast runs at roughly 8× FP32 throughput — the runtime simply rounds FP32 inputs to TF32 before the tensor-core MMA.

Default-on in cuDNN/cuBLAS from CUDA 11. Toggle with torch.backends.cuda.matmul.allow_tf32.

(b) BF16

16-bit: 1 sign + 8 exponent + 7 mantissa. Same dynamic range as FP32 (the 8-bit exponent), but FP16 storage cost.

This is the format that made stable mixed-precision training routine. Unlike FP16, BF16 doesn't underflow gradients, so loss scaling becomes optional. Every modern transformer pretrain uses BF16.

Same throughput as FP16 on Ampere; no accuracy/speed tradeoff — pure win.

(c) 2:4 sparsity

Hardware-supported structured sparsity. Every 4-element block of a weight matrix carries a 2-bit mask selecting 2 non-zero values. The tensor core skips the zeros.

Doubles throughput when the model has been pruned + retrained to the 2:4 pattern. Real-world adoption is patchy — aggressive quantisation often beats it — but the silicon is there.

The Ampere throughput table

A100 SXM4 80 GB, dense unless noted. Numbers from the NVIDIA A100 datasheet.

FormatDense TFLOPS2:4 Sparse TFLOPSNotes
FP64 tensor19.5—Genuine FP64 in tensor cores (HPC win)
FP32 (CUDA core)19.5—No tensor-core FP32; use TF32
TF32156312Drop-in FP32 replacement
BF16 / FP16312624Standard transformer training format
INT8624 TOPS1248 TOPSInference
INT41248 TOPS2496 TOPSInference (rare in practice)
No FP8 here

FP8 is a Hopper feature (CC 9.0). On A100, FP8 inference is software-emulated — you save weight-storage bandwidth but the math still happens in BF16/FP16 tensor cores. Don't expect compute uplift from FP8 on Ampere.

05

MIG — Hardware Partitioning

Multi-Instance GPU is Ampere's most underrated feature. An A100 can be carved at the hardware level into up to seven isolated GPU instances, each with its own dedicated SMs, L2 slice, and HBM region. Not virtualisation — isolation is enforced by the memory crossbar, not by software.

Granularities

The partition unit is one GPC (which on A100 corresponds to ~14 SMs, ~5 GB HBM, ~5 MB L2). Available shapes:

ProfileGPCsVRAMSMsL2Use case
1g.5gb15 GB~14~5 MBSingle small inference replica, dev sandbox
2g.10gb210 GB~28~10 MB7B inference, embedding service
3g.20gb320 GB~42~20 MB13B inference, mid-batch serving
4g.20gb420 GB~56~20 MB(rare) compute-heavy with same RAM
7g.40gb740 GB (or 80)98 (7 × 14)40 MBThe whole GPU, MIG mode “on”

Why this matters in shared clusters

One A100 carved up — a typical mixed layout

A100 40 GB — 7 GPCs total, MIG enabled 3g.20gb 42 SMs 20 GB HBM 20 MB L2 vLLM 13B serving 2g.10gb 28 SMs 10 GB HBM 10 MB L2 7B inference 1g.5gb 14 SMs 5 GB embed 1g.5gb 14 SMs 5 GB dev
enable MIG and create a typical 3+2+1+1 layout
# enable MIG mode (drains running work; needs reset)
nvidia-smi -mig 1

# list available profiles
nvidia-smi mig -lgip

# create the slices on GPU 0
nvidia-smi mig -cgi 9,14,19,19 -C
# 9 = 3g.20gb, 14 = 2g.10gb, 19 = 1g.5gb (profile IDs vary by SKU)

# each slice gets its own UUID; pin a container to one
docker run --gpus '"device=MIG-GPU-..."' vllm/vllm-openai ...
MIG limitations

No NVLink between MIG instances on the same physical GPU. No P2P. CUDA graphs work per slice but not across. MIG is for serving / dev-pool isolation, not for splitting a training job. Hopper extends MIG with TEE (confidential compute) and per-slice MPS; Blackwell B200 inherits both.

06

GA10x — The Consumer Family

The other half of Ampere lives on Samsung 8N. Same 3rd-gen tensor cores, same TF32/BF16 story, but no HBM, no NVSwitch, and a quietly different SM that has been the source of years of confusion.

CardDieCUDA coresVRAMBandwidthTDPNVLink
RTX 3090 TiGA1021075224 GB GDDR6X1008 GB/s450 Wbridge
RTX 3090GA1021049624 GB GDDR6X936 GB/s350 Wbridge
RTX 3080 TiGA1021024012 GB GDDR6X912 GB/s350 Wnone
RTX 3080GA1028704 / 896010 / 12 GB760 / 912 GB/s320 / 350 Wnone
RTX 3070 TiGA10461448 GB GDDR6X608 GB/s290 Wnone
RTX 3070GA10458888 GB GDDR6448 GB/s220 Wnone
RTX 3060 TiGA10448648 GB GDDR6448 GB/s200 Wnone
RTX 3060 12 GBGA106358412 GB GDDR6360 GB/s170 Wnone
RTX A6000GA1021075248 GB ECC GDDR6768 GB/s300 Wbridge
A40GA1021075248 GB GDDR6 ECC696 GB/s300 Wbridge

Why the “CUDA core” count is misleading

On GA100, a partition has 16 FP32 cores and 16 INT32 cores; an SM totals 64 FP32. On GA10x, NVIDIA doubled the FP32 count per partition: each partition has 16 FP32 + 16 dual-issue (FP32 or INT32) cores. An SM totals 128 FP32 when running pure FP32, but only 64 FP32 if INT32 is also live (because the dual-issue cores get co-opted).

This is why an RTX 3090 advertises 10496 CUDA cores when GA100 with the same silicon area would have ~5000. The marketing number is the FP32-only peak; real shader workloads alternate FP32 and INT32 (address arithmetic, indexing) and rarely sustain it. The tensor-core count is unchanged — 4 per SM, just like GA100.

Practical consequence

For LLM inference (which is tensor-core bound when prefilling and bandwidth-bound when decoding), the doubled FP32 cores are roughly irrelevant. A 3090's 936 GB/s of GDDR6X matters far more than its 35 TFLOPS of FP32. For graphics workloads, the dual-issue path is genuinely useful. Don't compare cards by CUDA-core count alone.

Workstation cards are the sleeper picks

The RTX A6000 and A40 are Ampere's 48 GB GDDR6 ECC options. Same GA102 die as a 3090, but with the full memory complement, ECC enabled, and a datacenter-blessed driver. The A6000 supports an NVLink bridge between two cards (112 GB/s) — not switch-class, but fine for 2-way TP on 70B models. Used 48 GB Ampere workstation cards remain a 2026 sweet spot for serious local work.

07

NVLink 3 + DGX A100

NVLink 3.0 doubles the per-lane signalling rate from NVLink 2.0 (50 vs 25 Gb/s) but halves the lanes per link, so each link stays at 25 GB/s per direction. A100 doubles the link count to 12 → 600 GB/s aggregate bidirectional. NVSwitch 2.0 is the fabric chip that ties them together.

DGX A100 baseboard (HGX A100)

  • 8× A100 SXM4 (40 or 80 GB)
  • 6× NVSwitch 2.0 chips
  • Any-to-any 600 GB/s NVLink between any two GPUs
  • Total NVLink fabric: 4.8 TB/s aggregate

Cross-node networking

  • 9× Mellanox ConnectX-6 200 Gb/s NICs
  • 8 NICs for compute (one per GPU, RDMA), 1 for storage
  • 200 Gb/s = 25 GB/s/dir per NIC
  • InfiniBand HDR or 200GbE RoCE

The reference platform for the GPT-3 era

Meta's OPT-175B (2022), BLOOM (2022), Megatron-Turing NLG (2022), and most large public LLMs from this period were trained on DGX A100 or HGX A100-derived systems. The 3D-parallelism recipe (TP × PP × DP) was tuned around exactly this baseboard.

Worked example: TP=8 all-reduce on a 70B model

Per layer, a tensor-parallel all-reduce needs to ship one full layer's hidden state. For Llama-2-70B with hidden dim 8192 and BF16: 8192 × seq × 2 bytes per token. At seq=4096 that's 67 MB per layer.

That's a 10–20× latency gap. With 80 transformer layers, each step accumulates ~80 × 75 µs = 6 ms on NVSwitch vs 80 ms on PCIe. NVLink is why TP=8 works on a single A100 node.

A100·0
↔
NVSwitch ×6
↔
A100·1..7
PCIe A100 has half the NVLink

The PCIe form-factor A100 has only 3× NVLink 3 bridges (not 12), giving 600 GB/s only between paired cards via bridge, not arbitrary-pair. For real multi-GPU work always pick SXM4 if you can.

08

A100 80 GB Upgrade — Why HBM Matters

In November 2020, NVIDIA announced the A100 80 GB. Same die. Same SMs. Same NVLink. The change was on the memory side: doubled HBM capacity (40 GB HBM2 → 80 GB HBM2e) and a 31% bandwidth bump (1.5 → 2.0 TB/s) thanks to faster HBM2e clocks.

SpecA100 40 GBA100 80 GBDelta
HBM capacity40 GB80 GB+100%
Memory bandwidth1.555 TB/s2.039 TB/s+31%
SMs (enabled)1081080%
L2 cache40 MB40 MB0%
BF16 dense TFLOPS3123120%
BF16 sparse TFLOPS6246240%
NVLink 3 aggregate600 GB/s600 GB/s0%
TDP (SXM4)400 W400 W0%
MIG max instances770%
MIG max slice VRAM1g.5gb1g.10gb+100%

Why this was a big deal for LLMs

40 GB held 13B in BF16 fine, struggled at 30B, and was helpless at 70B+ without sharding. 80 GB removed sharding from the workflow for an entire class of model sizes:

2026 used market

A100 80 GB SXM4 modules are still going strong on the second-hand market. They are the price-performance sweet spot for fine-tuning Llama-3-70B, Mixtral, and similar mid-tier models. Hopper and Blackwell are faster per-dollar for new builds, but a used A100 80 GB cluster remains a credible 2026 platform — especially for INT4 inference, where the bottleneck is bandwidth and not FP8 compute.

The lesson

Capacity and bandwidth, not flops, are the LLM constraint. The A100 80 GB upgrade ships the same compute and adds nothing-but-memory — and yet it changed which workloads were practical. This pattern recurs: H100 80 GB → H200 141 GB → B200 192 GB. Each generation, the headline win is HBM.

09

Performance for Modern LLMs

Where does Ampere sit in 2026? Pretraining at frontier scale has moved decisively to H100 / H200 / B200 — FP8 doubles compute and the Transformer Engine's per-tensor scaling is genuinely difficult to give up. But Ampere remains unbeatable in three places:

1. Fine-tuning

BF16 LoRA and full fine-tuning of 7B–70B models. The compute headroom (312 BF16 TFLOPS) is more than enough; the 80 GB HBM keeps optimiser state on-card. PyTorch FSDP works beautifully on DGX A100.

2. Mid-scale training

Pretraining models in the 1B–15B range. A 4× A100 80 GB node trains a 7B from scratch in days, not months, and the second-hand cost is a fraction of an H100 rig.

3. INT4 inference

AWQ and GPTQ quantised serving. INT4 weights fit gigantic models in modest VRAM, and the bottleneck moves to bandwidth — where A100 80 GB's 2 TB/s holds up well.

FP8 emulation on Ampere

You can “run” FP8 on A100 via libraries like vLLM's FP8-W8A8 path, but be clear about what's happening: the weights are stored as FP8 (saving HBM bandwidth on weight load), then dequantised into BF16 for tensor-core math. There is no compute uplift — throughput is the same as plain BF16. The win is purely a smaller working set.

For decode (memory-bound), this is still useful: you cut weight bandwidth in half and tok/s improves by 30–50%. For prefill (compute-bound), it's a wash. Hopper's native FP8 doubles compute as well; that's the meaningful gap.

Worked numbers: Llama-3-70B AWQ-INT4 inference

Two A100 80 GB SXM4 with TP=2 over NVLink, vLLM 0.5+, batch sizes from solo user to mid-batch:

WorkloadConfigurationTokens/s
Single user, 1 streamseq=2048, INT4 weights, BF16 KV30–45 tok/s
Batch=8 concurrentsame~250 tok/s aggregate
Batch=32 concurrentsame~600 tok/s aggregate
Prefill burst (compute)4096 prompt, single seq~12000 tok/s prefill

A single A100 80 GB at INT4 holds Llama-3-70B with ~40 GB headroom for KV-cache. Two cards via NVLink approximately double throughput on prefill and add ~70–80% on decode (TP overhead).

When to skip Ampere

If your workload is FP8 pretraining, use Hopper. If it's FP4 inference, use Blackwell. If it's multi-trillion-token data curation with TE-fp8 numerics, use Hopper. For everything else — especially fine-tuning and INT4 serving — Ampere remains a remarkable price/perf sweet spot.

10

Software Stack on Ampere

Ampere has the most mature software stack of any NVIDIA generation simply because it has been in production the longest. Five years of bug fixes, kernel tuning, and library work all assume CC 8.0 / 8.6 as the baseline.

CUDA + libraries

  • CUDA 11.0 introduced Ampere support (May 2020)
  • cuDNN 8.x rewrote convolution + attention kernels for Ampere
  • cuBLAS 11+ uses TF32 by default for FP32 GEMM
  • NCCL 2.7+ exploits NVLink 3 + NVSwitch 2 bandwidth
  • CUDA 12.x remains fully supported on Ampere

PyTorch

  • torch.cuda.amp.autocast(dtype=torch.bfloat16) — mostly painless
  • torch.backends.cuda.matmul.allow_tf32 = True — default on, gives free FP32-→TF32 speedup
  • FSDP — the default sharded training API; Ampere is a first-class target
  • torch.compile works, though TE / FP8 paths require Hopper

The TF32 vs BF16 trap

If you're writing new transformer code, the right default is BF16 explicitly, not TF32. Reason: cuDNN/cuBLAS schedule a different (and faster) tensor-core kernel for BF16 matmuls than for TF32. TF32 is the “don't break my old code” format; BF16 is the “I'm writing this from scratch” format. On A100, BF16 matmul is faster than TF32 matmul (312 TFLOPS vs 156 TFLOPS) and accuracy is essentially identical for training transformers.

Inference / serving

recommended PyTorch defaults on Ampere (training)
import torch
# keep TF32 enabled for any leftover FP32 codepaths
torch.backends.cuda.matmul.allow_tf32 = True
torch.backends.cudnn.allow_tf32 = True

# but actually train in BF16 explicitly
with torch.autocast(device_type='cuda', dtype=torch.bfloat16):
    out = model(x)
    loss = criterion(out, y)
loss.backward()
Driver hygiene

Pin to a stable production driver (R535 or R550 LTS branches are good for Ampere). Ampere works with newer drivers but you don't need them — only Blackwell and Hopper FP8 features need bleeding-edge driver minor versions. Long-running A100 fleets benefit from the LTS branches.

11

Interactive: Ampere SKU Picker

Pick an Ampere SKU; see the headline numbers, what it means for LLM capacity at FP16 and INT4, and what it's actually good for.

“Largest LLM” is a back-of-envelope estimate: VRAM ÷ 2 (B params) for FP16 weight-only fit, VRAM ÷ 0.6 for INT4. Real workloads need additional headroom for KV-cache, activations, and CUDA overhead — subtract 10–20%.