NVIDIA GPU Architectures Series — Presentation 02

Inside the SM — How NVIDIA's Streaming Multiprocessor Evolved

The SM is the unit-cell of every NVIDIA GPU. Trace its evolution from Pascal's symmetric 64-FP32 design to Blackwell's tensor-core-dominated, cluster-aware monster — and see why each change unlocked new model sizes.

SM Warp scheduler Register file Shared memory L1 Tensor core TMA Cluster
Cores → Schedulers → Regs → Shmem → L1 → Tensor → TMA → Cluster
00

Topics We'll Cover

Twelve slides tracing the inner architecture of NVIDIA's Streaming Multiprocessor across nine years and six generations — from Pascal's symmetric design to Blackwell's microscaling tensor cores.

01

What is an SM?

The Streaming Multiprocessor is the unit-cell of every NVIDIA GPU. Replicate it 100–150× on a die, wire the copies together with an L2 crossbar and a memory controller, and you have a GPU. Everything in CUDA — threads, warps, blocks, shared memory, kernels — is defined relative to this one block.

What lives inside one SM

Scale at the die level

GPUSM countNotes
A100 (GA100)108 enabled128 physical, harvested for yield
H100 (GH100, SXM)132 enabled144 physical
B200 (Blackwell)148 enabled (80 physical per die × 2 dies)two reticle-limit dies on one package
RTX 5090 (GB202)170 enabled192 physical
Why this matters

Almost every architectural improvement NVIDIA has shipped since 2016 is a change inside the SM — a new datapath, a wider tensor core, a new memory tier, an extra scheduling level. The cross-GPU machinery (NVLink, NVSwitch, SHARP) matters at fleet scale, but the per-token compute economics live and die inside this one block.

02

Pascal SM (GP100, 2016)

Pascal — specifically the GP100 die used in the P100 datacenter card — was the last NVIDIA SM before tensor cores. It set the template that every successor would refine.

GP100 SM at a glance

SIMT, but lock-step

Pascal still inherited the original SIMT contract from Tesla/Fermi/Kepler: every thread in a warp shares one program counter and one call stack. Divergent branches were handled by predication and re-convergence stack, not by genuinely independent threads. Warp-synchronous code — __shfl_sync, ballot patterns — was correct because the hardware physically kept lanes stepping together.

What Pascal did well

Big FP64 numbers for traditional HPC. NVLink 1.0 between 8 P100s. First HBM2 in a NVIDIA part. Programmer model unchanged from Kepler/Maxwell — existing CUDA code just ran faster.

Pascal's blind spot

No matrix engine. ResNet-50 training on a P100 cluster was state of the art in 2016, but already the FP32 cores were the bottleneck for emerging deep-learning workloads. Volta would change that completely.

Reference

NVIDIA Tesla P100 Whitepaper (2016) is the canonical Pascal-SM reference. It documents the 64+32 layout, the two dual-issue schedulers, and the 64 KB shared memory.

03

Volta SM (GV100, 2017)

Volta is the largest single architectural step NVIDIA has taken since the original CUDA architecture (G80, 2006). It rewrote two things: the thread-scheduling contract and the arithmetic engine.

Independent thread scheduling

Each thread in a warp now has its own program counter and call stack. Divergent paths can interleave; lanes are no longer guaranteed to step together. The lock-step assumption that previously made warp-synchronous shuffles safe is now unsafe — CUDA 9 introduced explicit *_sync intrinsics with mask arguments to make synchronisation requirements visible. This was a foundational change: it lets fine-grained algorithms (lock-free queues, producer-consumer patterns) work correctly inside a warp.

Separate INT32 and FP32 datapaths

Tensor cores arrive

8 first-generation tensor cores per SM, each performing a 4×4×4 half-precision matrix-multiply-accumulate per cycle. Inputs FP16, accumulator FP32. NVIDIA reported >100 TFLOPS of tensor throughput per V100, an order of magnitude over the FP32 cores. PyTorch's autocast and Apex AMP were direct responses to Volta.

L1 unification

The 128 KB SRAM array is configurable as a split between shared memory (programmer-managed) and L1 cache (hardware-managed). Common splits were 96/32 or 64/64. The 256 KB register file size set on Volta has remained constant on every successor SM.

Programmer impact

Volta is the first SM where naive un-tuned matmul code is dramatically slower than tensor-core code. The pressure to use cuBLAS / cuDNN / cutlass kernels — rather than hand-written FP32 kernels — dates from this generation.

04

Turing SM (TU102, 2018)

Turing is a refinement of Volta for the consumer/workstation market. It kept Volta's separated INT/FP datapaths and tensor-core layout, then added two new things and refined a third.

2nd-generation tensor cores

RT cores arrive

One ray-tracing core per SM. Performs ray-AABB intersection and ray-triangle intersection in fixed-function hardware, plus BVH traversal acceleration. From the LLM perspective these are dead weight; from the graphics/visualisation/path-tracing perspective they are transformative. The point is that Turing was where NVIDIA started using the SM as a host for multiple kinds of fixed-function accelerator, not just the FP/INT/Tensor triad.

INT/FP concurrency becomes routine

Turing tightened the issue logic so that mixing integer addressing arithmetic with FP math is essentially free. Driver and compiler scheduling now assume the dual datapath.

L1/shared

Slightly trimmed to 96 KB unified L1+shared per SM. Still configurable, but with fewer split options than Volta.

Why Turing matters for LLMs

INT8 inference was born here. TensorRT's INT8 path, GPTQ-style 4-bit experiments, and the long arc toward cheaper inference all start on Turing tensor cores. Compute-capability 7.5 is still a supported deployment target in 2026.

05

Ampere SM (GA100, 2020)

Ampere — specifically the GA100 die used in the A100 — was the SM where deep learning became the primary workload assumption. Tensor-core throughput per SM doubled, new datatypes made FP32-trained networks usable without code changes, and asynchronous data movement entered the SM proper.

3rd-generation tensor cores

New datatypes

  • TF32 — 10-bit mantissa, 8-bit exponent, FP32-style range. Drop-in replacement for FP32 matmul, no code change.
  • BF16 — 8-bit exponent (FP32 range), 7-bit mantissa. The training datatype the rest of the industry has now standardised on.
  • INT8 / INT4 carried over from Turing.

Structured 2:4 sparsity

Within every group of 4 weights, 2 must be zero. The tensor core skips the zeros, doubling effective throughput at the cost of an off-line pruning pass. Free 2× for inference if the model retrains acceptably.

Balanced datacenter partitions

GA100 keeps Volta's balanced layout: each of the SM's 4 partitions has 16 FP32 + 16 INT32 + 8 FP64 cores, with 1 third-gen tensor core per partition (4 per SM total). Peak FP32 throughput per SM is therefore the same as V100 — the per-SM math gain is in the wider, denser tensor cores, not in the FP32 lane count. The "FP32 doubling" trick is GA10x-only (next slide).

Async copy — cp.async

The new cp.async instruction issues a register-pipelined DMA from L2 directly into shared memory. The thread does not stall on the data path. This is the prerequisite for the software-pipelined matmul kernels that became standard in cutlass 2.x. It also previews Hopper's TMA — which formalises the same idea at the SM level rather than the warp level.

L1+shared and registers

06

GA10x — Consumer Ampere's Twist

Important nuance: "Ampere" is two different SMs, depending on which die you're talking about. The datacenter GA100 die above and the consumer GA102/GA104 dies (RTX 30 series, A6000 workstation) ship a different SM partition.

Same generation, different SM

FeatureGA100 (datacenter)GA10x (consumer)
FP32 cores per partition1632 (doubled)
INT32 cores per partition1616 (one of the FP32 datapaths is dual-purpose)
FP64 cores per partition82 (token amount)
Tensor cores per SM4 (3rd gen)4 (3rd gen)
Tensor-core throughput per SMbaselinehalf of GA100's FP16 rate per SM (1024 vs 2048 FLOP/clk); 2:4 sparsity supported
L1+shared per SM192 KB128 KB
RT coresnone2nd gen, 1 per SM

What's actually going on

On GA10x, one of the two integer datapaths in each partition is now dual-FP32. The result: for FP32-heavy graphics shaders (which never used much INT) the per-SM peak FP32 number doubles. But the tensor-core count is unchanged, so for matmul-bound LLM inference the scaling is closer to that of GA100. The "FP32 doubling" is therefore a graphics-tuned trick — great for raster shaders, less impactful for tensor-bound kernels.

Reading the spec sheet

When NVIDIA quotes "10,496 CUDA cores" on a 3090 vs "6,912 CUDA cores" on an A100, those numbers are not directly comparable — they count the dual-FP32 partitions separately. For tensor-bound workloads, count SMs × tensor-core throughput, not CUDA cores.

07

Hopper SM (GH100, 2022)

Hopper is the SM that was designed for transformer training and serving. Four big additions, all of which are still defining features of the platform in 2026.

1. 4th-generation tensor cores — FP8

2. Tensor Memory Accelerator (TMA)

Hardware-accelerated, asynchronous, multi-dimensional memcpy. The SM submits a tile descriptor — base pointer, multi-dim extents, strides — and the TMA performs the global-to-shared (or shared-to-global) bulk copy on its own. The SM is freed; warps don't busy-wait on address arithmetic. Generalises Ampere's cp.async from a per-warp instruction to an SM-level engine.

3. Distributed Shared Memory (DSMEM) and Thread-Block Clusters

A new level introduced between the existing block and grid:

Old
thread warp block grid
Hopper
thread warp block cluster grid

Up to 16 cooperating blocks form a cluster, scheduled together on the same GPC. Blocks within a cluster can directly read each other's shared memory (DSMEM) over a fast on-die path — effectively turning a cluster into one shared-memory region 16× bigger than a single block's. Used by attention and large-tile matmul kernels to keep tiles fully on-chip.

4. DPX instructions

New per-SM instructions for dynamic-programming inner loops — min-of-min-plus-cost patterns of the kind found in Smith–Waterman alignment and Floyd–Warshall shortest paths. Niche but pulled the bioinformatics / route-finding workloads into the SM's first-class instruction set.

L1+shared and registers

08

Ada SM (AD102, 2022)

Ada is the consumer/workstation cousin of Hopper, sharing the timeline (2022) but a different SM emphasis: graphics first, then ML.

3rd-generation RT cores

4th-gen tensor cores — FP8

Ada's tensor cores are 4th-generation and support FP8 (E4M3, E5M2) across the whole family — consumer (RTX 4090, 4080, etc.) and workstation/datacenter (RTX 6000 Ada, L40, L40S) alike (compute capability 8.9). What Ada lacks vs Hopper is not FP8 itself but Hopper's surrounding plumbing: no TMA, no thread-block clusters / DSMEM, no second-gen Transformer Engine of the Hopper kind, and a smaller datacenter-scale memory subsystem. A 4090 happily runs FP8 kernels; what it doesn't run are Hopper-specific cluster/TMA kernels.

Big L2 is shared, not per-SM

AD102 has a 96 MB L2 — nearly 16× larger than GA102's. This is a die-level cache, not an SM resource: every SM accesses it through the on-die crossbar. It dramatically reduces DRAM traffic for graphics workloads (and for small LLMs with KV caches that fit). It's not the same as Hopper's per-SM 256 KB L1+shared.

Practical buyer's note

Ada and Hopper are siblings of the same year and Compute Capability lineage but solve different problems. Ada (RTX 4090, RTX 6000 Ada, L40S) sits in workstations; Hopper (H100, H200) sits in racks. The SM looks similar from a CUDA programmer's seat but the FP8 story differs — and Hopper has TMA, DSMEM and clusters that Ada doesn't.

09

Blackwell SM (B200, 2024)

Blackwell extends Hopper's SM template rather than rewriting it. The big changes are at the tensor-core level and around the inter-die / inter-GPU plumbing.

5th-generation tensor cores — microscaling FP4 and FP6

2nd-generation Transformer Engine

The Transformer Engine library now manages per-microblock scaling, not just per-tensor scaling. A model can be cast to MX-FP4 weights, MX-FP4 (or FP8) activations, and FP32 master weights in framework-managed wrappers. The library decides which datatype is safe per layer.

Beyond the SM — NVLink at SM granularity

Blackwell's 5th-gen NVLink raises bisection bandwidth significantly, but a more interesting change is NVLink-Sharp: SHARP-style in-network all-reduce reductions can now be triggered at SM granularity rather than at full-collective granularity. For mixture-of-experts routing (where dispatch is unstructured), this is meaningful.

Reliability — RAS engine

A new RAS engine sits at the SM/L2 boundary, providing fault isolation (so a single SM ECC failure doesn't take down a job) and richer telemetry to the management plane. It's an obvious nod to the realities of running multi-thousand-GPU jobs for weeks at a time.

L1+shared and registers

10

Resource Growth Visualised

Per-SM resources across six families. Register file size has been frozen at 256 KB since Volta — the action is in shared memory, tensor-core throughput, and per-tensor-core feature width. Tensor-core throughput is dense, per SM per clock, normalised to V100 = 1, in the lowest floating-point precision shipped on each generation (FP16 on Volta/Turing, BF16 on Ampere, FP8 on Hopper, FP4 on Blackwell); the tensor-core bars use a log2 height scale.

Per-SM resources by family (registers, shmem+L1, tensor-core throughput vs V100) Register file (KB) Shmem+L1 (KB) TC ratio 0 128 256 384 Pascal GP100 256 64 0.0 Volta GV100 256 128 1.0 Turing TU102 256 96 1.0 Ampere GA100 256 192 2 Hopper GH100 256 256 8 1st FP8 Blackwell B200 256 256 ~32 1st FP4
Reading the chart

Two stories. Storage grew steadily (shmem+L1 from 64 KB on Pascal to 256 KB on Hopper/Blackwell), then plateaued; the register file has been a constant 256 KB for seven years. Tensor-core throughput, in contrast, is a hockey-stick: each generation introduces a narrower datatype (FP16 → BF16 → FP8 → MX-FP4) that doubles or triples per-SM math without growing the SM itself. (Blackwell supports both MX-FP4 and NVFP4.) The "free lunch" of LLM inference cost reductions has come almost entirely from this column.

11

Interactive: SM Explorer

Pick a generation; the SVG below redraws an idealised SM floorplan and the metric grid updates with that generation's headline numbers. The diagram is schematic — it shows the four partitions, the warp scheduler in each, the register file slice, the FP/INT/Tensor/SFU functional units, and the unified L1+shared block at the bottom.

What the diagram does and doesn't show

This is a logical floorplan, not a physical one. The widths of FP32 vs INT32 vs tensor-core boxes are scaled to give each its share of the partition; on real silicon the tensor core is far larger than its outline suggests, and the register file is interleaved with the math units rather than sitting in a single slab. The point of the picture is to show which units are present in each generation and how they are grouped into the four partitions that have remained the SM's organising structure since Volta.