The SM is the unit-cell of every NVIDIA GPU. Trace its evolution from Pascal's symmetric 64-FP32 design to Blackwell's tensor-core-dominated, cluster-aware monster — and see why each change unlocked new model sizes.
Twelve slides tracing the inner architecture of NVIDIA's Streaming Multiprocessor across nine years and six generations — from Pascal's symmetric design to Blackwell's microscaling tensor cores.
The Streaming Multiprocessor is the unit-cell of every NVIDIA GPU. Replicate it 100–150× on a die, wire the copies together with an L2 crossbar and a memory controller, and you have a GPU. Everything in CUDA — threads, warps, blocks, shared memory, kernels — is defined relative to this one block.
| GPU | SM count | Notes |
|---|---|---|
| A100 (GA100) | 108 enabled | 128 physical, harvested for yield |
| H100 (GH100, SXM) | 132 enabled | 144 physical |
| B200 (Blackwell) | 148 enabled (80 physical per die × 2 dies) | two reticle-limit dies on one package |
| RTX 5090 (GB202) | 170 enabled | 192 physical |
Almost every architectural improvement NVIDIA has shipped since 2016 is a change inside the SM — a new datapath, a wider tensor core, a new memory tier, an extra scheduling level. The cross-GPU machinery (NVLink, NVSwitch, SHARP) matters at fleet scale, but the per-token compute economics live and die inside this one block.
Pascal — specifically the GP100 die used in the P100 datacenter card — was the last NVIDIA SM before tensor cores. It set the template that every successor would refine.
Pascal still inherited the original SIMT contract from Tesla/Fermi/Kepler: every thread in a warp shares one program counter and one call stack. Divergent branches were handled by predication and re-convergence stack, not by genuinely independent threads. Warp-synchronous code — __shfl_sync, ballot patterns — was correct because the hardware physically kept lanes stepping together.
Big FP64 numbers for traditional HPC. NVLink 1.0 between 8 P100s. First HBM2 in a NVIDIA part. Programmer model unchanged from Kepler/Maxwell — existing CUDA code just ran faster.
No matrix engine. ResNet-50 training on a P100 cluster was state of the art in 2016, but already the FP32 cores were the bottleneck for emerging deep-learning workloads. Volta would change that completely.
NVIDIA Tesla P100 Whitepaper (2016) is the canonical Pascal-SM reference. It documents the 64+32 layout, the two dual-issue schedulers, and the 64 KB shared memory.
Volta is the largest single architectural step NVIDIA has taken since the original CUDA architecture (G80, 2006). It rewrote two things: the thread-scheduling contract and the arithmetic engine.
Each thread in a warp now has its own program counter and call stack. Divergent paths can interleave; lanes are no longer guaranteed to step together. The lock-step assumption that previously made warp-synchronous shuffles safe is now unsafe — CUDA 9 introduced explicit *_sync intrinsics with mask arguments to make synchronisation requirements visible. This was a foundational change: it lets fine-grained algorithms (lock-free queues, producer-consumer patterns) work correctly inside a warp.
8 first-generation tensor cores per SM, each performing a 4×4×4 half-precision matrix-multiply-accumulate per cycle. Inputs FP16, accumulator FP32. NVIDIA reported >100 TFLOPS of tensor throughput per V100, an order of magnitude over the FP32 cores. PyTorch's autocast and Apex AMP were direct responses to Volta.
The 128 KB SRAM array is configurable as a split between shared memory (programmer-managed) and L1 cache (hardware-managed). Common splits were 96/32 or 64/64. The 256 KB register file size set on Volta has remained constant on every successor SM.
Volta is the first SM where naive un-tuned matmul code is dramatically slower than tensor-core code. The pressure to use cuBLAS / cuDNN / cutlass kernels — rather than hand-written FP32 kernels — dates from this generation.
Turing is a refinement of Volta for the consumer/workstation market. It kept Volta's separated INT/FP datapaths and tensor-core layout, then added two new things and refined a third.
One ray-tracing core per SM. Performs ray-AABB intersection and ray-triangle intersection in fixed-function hardware, plus BVH traversal acceleration. From the LLM perspective these are dead weight; from the graphics/visualisation/path-tracing perspective they are transformative. The point is that Turing was where NVIDIA started using the SM as a host for multiple kinds of fixed-function accelerator, not just the FP/INT/Tensor triad.
Turing tightened the issue logic so that mixing integer addressing arithmetic with FP math is essentially free. Driver and compiler scheduling now assume the dual datapath.
Slightly trimmed to 96 KB unified L1+shared per SM. Still configurable, but with fewer split options than Volta.
INT8 inference was born here. TensorRT's INT8 path, GPTQ-style 4-bit experiments, and the long arc toward cheaper inference all start on Turing tensor cores. Compute-capability 7.5 is still a supported deployment target in 2026.
Ampere — specifically the GA100 die used in the A100 — was the SM where deep learning became the primary workload assumption. Tensor-core throughput per SM doubled, new datatypes made FP32-trained networks usable without code changes, and asynchronous data movement entered the SM proper.
Within every group of 4 weights, 2 must be zero. The tensor core skips the zeros, doubling effective throughput at the cost of an off-line pruning pass. Free 2× for inference if the model retrains acceptably.
GA100 keeps Volta's balanced layout: each of the SM's 4 partitions has 16 FP32 + 16 INT32 + 8 FP64 cores, with 1 third-gen tensor core per partition (4 per SM total). Peak FP32 throughput per SM is therefore the same as V100 — the per-SM math gain is in the wider, denser tensor cores, not in the FP32 lane count. The "FP32 doubling" trick is GA10x-only (next slide).
The new cp.async instruction issues a register-pipelined DMA from L2 directly into shared memory. The thread does not stall on the data path. This is the prerequisite for the software-pipelined matmul kernels that became standard in cutlass 2.x. It also previews Hopper's TMA — which formalises the same idea at the SM level rather than the warp level.
Important nuance: "Ampere" is two different SMs, depending on which die you're talking about. The datacenter GA100 die above and the consumer GA102/GA104 dies (RTX 30 series, A6000 workstation) ship a different SM partition.
| Feature | GA100 (datacenter) | GA10x (consumer) |
|---|---|---|
| FP32 cores per partition | 16 | 32 (doubled) |
| INT32 cores per partition | 16 | 16 (one of the FP32 datapaths is dual-purpose) |
| FP64 cores per partition | 8 | 2 (token amount) |
| Tensor cores per SM | 4 (3rd gen) | 4 (3rd gen) |
| Tensor-core throughput per SM | baseline | half of GA100's FP16 rate per SM (1024 vs 2048 FLOP/clk); 2:4 sparsity supported |
| L1+shared per SM | 192 KB | 128 KB |
| RT cores | none | 2nd gen, 1 per SM |
On GA10x, one of the two integer datapaths in each partition is now dual-FP32. The result: for FP32-heavy graphics shaders (which never used much INT) the per-SM peak FP32 number doubles. But the tensor-core count is unchanged, so for matmul-bound LLM inference the scaling is closer to that of GA100. The "FP32 doubling" is therefore a graphics-tuned trick — great for raster shaders, less impactful for tensor-bound kernels.
When NVIDIA quotes "10,496 CUDA cores" on a 3090 vs "6,912 CUDA cores" on an A100, those numbers are not directly comparable — they count the dual-FP32 partitions separately. For tensor-bound workloads, count SMs × tensor-core throughput, not CUDA cores.
Hopper is the SM that was designed for transformer training and serving. Four big additions, all of which are still defining features of the platform in 2026.
Hardware-accelerated, asynchronous, multi-dimensional memcpy. The SM submits a tile descriptor — base pointer, multi-dim extents, strides — and the TMA performs the global-to-shared (or shared-to-global) bulk copy on its own. The SM is freed; warps don't busy-wait on address arithmetic. Generalises Ampere's cp.async from a per-warp instruction to an SM-level engine.
A new level introduced between the existing block and grid:
Up to 16 cooperating blocks form a cluster, scheduled together on the same GPC. Blocks within a cluster can directly read each other's shared memory (DSMEM) over a fast on-die path — effectively turning a cluster into one shared-memory region 16× bigger than a single block's. Used by attention and large-tile matmul kernels to keep tiles fully on-chip.
New per-SM instructions for dynamic-programming inner loops — min-of-min-plus-cost patterns of the kind found in Smith–Waterman alignment and Floyd–Warshall shortest paths. Niche but pulled the bioinformatics / route-finding workloads into the SM's first-class instruction set.
Ada is the consumer/workstation cousin of Hopper, sharing the timeline (2022) but a different SM emphasis: graphics first, then ML.
Ada's tensor cores are 4th-generation and support FP8 (E4M3, E5M2) across the whole family — consumer (RTX 4090, 4080, etc.) and workstation/datacenter (RTX 6000 Ada, L40, L40S) alike (compute capability 8.9). What Ada lacks vs Hopper is not FP8 itself but Hopper's surrounding plumbing: no TMA, no thread-block clusters / DSMEM, no second-gen Transformer Engine of the Hopper kind, and a smaller datacenter-scale memory subsystem. A 4090 happily runs FP8 kernels; what it doesn't run are Hopper-specific cluster/TMA kernels.
AD102 has a 96 MB L2 — nearly 16× larger than GA102's. This is a die-level cache, not an SM resource: every SM accesses it through the on-die crossbar. It dramatically reduces DRAM traffic for graphics workloads (and for small LLMs with KV caches that fit). It's not the same as Hopper's per-SM 256 KB L1+shared.
Ada and Hopper are siblings of the same year and Compute Capability lineage but solve different problems. Ada (RTX 4090, RTX 6000 Ada, L40S) sits in workstations; Hopper (H100, H200) sits in racks. The SM looks similar from a CUDA programmer's seat but the FP8 story differs — and Hopper has TMA, DSMEM and clusters that Ada doesn't.
Blackwell extends Hopper's SM template rather than rewriting it. The big changes are at the tensor-core level and around the inter-die / inter-GPU plumbing.
The Transformer Engine library now manages per-microblock scaling, not just per-tensor scaling. A model can be cast to MX-FP4 weights, MX-FP4 (or FP8) activations, and FP32 master weights in framework-managed wrappers. The library decides which datatype is safe per layer.
Blackwell's 5th-gen NVLink raises bisection bandwidth significantly, but a more interesting change is NVLink-Sharp: SHARP-style in-network all-reduce reductions can now be triggered at SM granularity rather than at full-collective granularity. For mixture-of-experts routing (where dispatch is unstructured), this is meaningful.
A new RAS engine sits at the SM/L2 boundary, providing fault isolation (so a single SM ECC failure doesn't take down a job) and richer telemetry to the management plane. It's an obvious nod to the realities of running multi-thousand-GPU jobs for weeks at a time.
Per-SM resources across six families. Register file size has been frozen at 256 KB since Volta — the action is in shared memory, tensor-core throughput, and per-tensor-core feature width. Tensor-core throughput is dense, per SM per clock, normalised to V100 = 1, in the lowest floating-point precision shipped on each generation (FP16 on Volta/Turing, BF16 on Ampere, FP8 on Hopper, FP4 on Blackwell); the tensor-core bars use a log2 height scale.
Two stories. Storage grew steadily (shmem+L1 from 64 KB on Pascal to 256 KB on Hopper/Blackwell), then plateaued; the register file has been a constant 256 KB for seven years. Tensor-core throughput, in contrast, is a hockey-stick: each generation introduces a narrower datatype (FP16 → BF16 → FP8 → MX-FP4) that doubles or triples per-SM math without growing the SM itself. (Blackwell supports both MX-FP4 and NVFP4.) The "free lunch" of LLM inference cost reductions has come almost entirely from this column.
Pick a generation; the SVG below redraws an idealised SM floorplan and the metric grid updates with that generation's headline numbers. The diagram is schematic — it shows the four partitions, the warp scheduler in each, the register file slice, the FP/INT/Tensor/SFU functional units, and the unified L1+shared block at the bottom.
This is a logical floorplan, not a physical one. The widths of FP32 vs INT32 vs tensor-core boxes are scaled to give each its share of the partition; on real silicon the tensor core is far larger than its outline suggests, and the register file is interleaved with the math units rather than sitting in a single slab. The point of the picture is to show which units are present in each generation and how they are grouped into the four partitions that have remained the SM's organising structure since Volta.