NVIDIA GPU Architectures Series — Presentation 27

Inside Volta — First Tensor Cores, Independent Thread Scheduling

A low-level look at GV100 — the 2017 die that introduced 1st-generation tensor cores, separated INT and FP datapaths, and gave every thread its own program counter. Block diagrams, arithmetic, process, voltage, clocks, capacities, and the NVLink 2.0 / NVSwitch 1 fabric.

GV100V100 TSMC 12FFHBM2 Tensor Core 1st genFP16x2 NVLink 2.0NVSwitch 1 ITS
GV100 → GPC → SM → TC1 → L2 → HBM2 → NVLink 2.0
00

Topics We'll Cover

01

Volta at a Glance — Datacenter-Only

Volta launched in 2017 (V100 announced May 2017; Titan V in December) as a single-die, datacenter-only architecture. There is no consumer Volta. The graphics path skipped Volta entirely — Pascal's GP102 carried the high end until Turing arrived a year later. Volta was a deep-learning bet by NVIDIA.

SKUForm factorSMsTensor coresMemoryTDP
Tesla V100 SXM2 16 GBSXM2 mezzanine8064016 GB HBM2 @ 900 GB/s300 W
Tesla V100 SXM3 32 GBSXM3 (DGX-2)8064032 GB HBM2 @ 900 GB/s350 W
Tesla V100 PCIe 16/32 GBPCIe8064016/32 GB HBM2250 W
Quadro GV100workstation8064032 GB HBM2250 W
Titan Vconsumer (HEDT)8064012 GB HBM2250 W
Tesla V100SPCIe (refresh 2019)8064032 GB HBM2 @ 1134 GB/s250 W
02

GV100 Die — Block Diagram

Chip
GV100 die — 815 mm², 21.1 B transistors
GPCs
GPC0GPC1GPC2GPC3GPC4GPC5
TPCs / GPC
7 TPCs × 6 GPCs = 42 TPCs (84 SMs total; 80 enabled)
L2
6 MB total (8 partitions of 768 KB)
Memory
4 stacks HBM2 — 4096-bit bus
IO
PCIe 3.0 x166 NVLink 2.0
03

The Volta SM — Four Partitions, Split FP/INT

The Volta SM is reorganised relative to Pascal's two-partition design. Four partitions, each with its own warp scheduler and dispatch unit. Critically, FP32 and INT32 datapaths are separate and concurrent — they can issue together in the same cycle, doubling effective issue rate when an integer address calculation overlaps a floating-point operation.

Per-SM compute

  • 64 FP32 cores (16 per partition)
  • 64 INT32 cores (16 per partition) — separate datapath, concurrent issue
  • 32 FP64 cores (8 per partition) — 1:2 of FP32
  • 8 tensor cores (2 per partition) — first generation
  • 16 SFUs, 32 LD/ST, 4 texture units
  • 4 warp schedulers, 1 dispatch each

Per-SM memory

  • 256 KB register file (65536 32-bit, 16K per partition)
  • 128 KB unified L1 + shared (configurable: 96 KB shared / 32 KB L1, or other splits)
  • L1 latency dropped from ~28 cycles (Pascal) to ~20 cycles
  • 4× bandwidth from L1 vs Pascal
04

First-Generation Tensor Cores

The headline feature. Each Volta SM has 8 tensor cores; each performs a fused 4×4×4 matrix multiply-accumulate per cycle: D = A·B + C, where A and B are 4×4 FP16 matrices and C/D are 4×4 FP32 (or FP16) matrices — 64 FMAs per tensor core per cycle.

OperationPer SM per cycleV100 SXM2 peak (1530 MHz)
FP32 FMA (concurrent with INT)64 (+64 INT32)15.7 TFLOPS
FP64 FMA327.8 TFLOPS
FP16x2 packed12831.4 TFLOPS
Tensor core HMMA (FP16->FP32)512125 TFLOPS

Programming interface: PTX wmma.mma.sync.aligned.m16n16k16.f32.f16.f16.f32; C++ nvcuda::wmma. Tensor cores were a ~12× FLOP increase (125 vs 10.6 TFLOPS) over Pascal's FP32 path on transformer-style workloads — the moment "deep learning hardware" stopped being marketing.

05

Independent Thread Scheduling

Pre-Volta SIMT: all 32 lanes of a warp share one program counter. Divergent branches mask lanes, taken-and-not-taken paths run sequentially. Critical sections were unsafe (a thread waiting on a lock could starve threads in the same warp holding it → deadlock).

Volta gives each thread its own PC and call stack. The warp can re-converge opportunistically; threads on different branches can interleave at instruction granularity. Spinlocks, fine-grained synchronisation, and producer-consumer within a warp now work correctly.

Cost: explicit warp sync

Code that relied on lock-step now needs __shfl_sync(mask, x, lane) with explicit masks (and __ballot_sync, __activemask). The non-sync intrinsics are deprecated. The CUDA 9 release that landed alongside Volta documented every breakage; legacy kernels that ignored the change silently produced wrong results.

06

Process — TSMC 12FF

GV100 is TSMC 12FF — a half-node refinement of 16FF+ with denser standard cells and improved transistor performance, but same metal pitch — so density gains over GP100 came mostly from cell rework, not lithography. NVIDIA called it 12 nm; TSMC's official name was 12FFN (FinFET with NVIDIA-specific tweaks).

MetricGP100 (Pascal)GV100 (Volta)
ProcessTSMC 16FF+TSMC 12FF
Transistors15.3 B21.1 B
Die area610 mm²815 mm²
Density25.1 M/mm²25.9 M/mm²
SMs6084 (40% more)

815 mm² was very close to TSMC's reticle limit (~858 mm², 26 × 33 mm). GV100 was at the ceiling of what one die could be at the 12FF yield curve.

07

Voltage, Power, Package (SXM2/SXM3)

V100 SXM2 introduced the SXM2 form factor proper — the same socket Pascal P100 used, at the same 300 W. The SXM3 variant in DGX-2 pushed to 350 W to support 32 GB HBM2 stacks at higher pin rate.

08

Clocks & Speeds

SKUBaseBoostMemoryBW
V100 SXM2 16/321290 MHz1530 MHzHBM2 1.75 Gbps/pin900 GB/s
V100 SXM3 32 (DGX-2)1312 MHz1530 MHzHBM2 1.75 Gbps/pin900 GB/s
V100 PCIe 16/321230 MHz1380 MHzHBM2 1.75 Gbps/pin900 GB/s
V100S PCIe 32 (2019)1245 MHz1597 MHzHBM2 2.21 Gbps/pin1134 GB/s
Titan V1200 MHz1455 MHzHBM2 1.7 Gbps/pin653 GB/s (3 stacks)

Titan V is the odd one: same GV100 die but with one HBM2 stack disabled (3-of-4 stacks → 12 GB, 3072-bit bus, 653 GB/s) for yield reasons.

09

Memory — HBM2 16/32 GB

V100 was the first NVIDIA part to ship 32 GB HBM2, using 8-Hi stacks (8 dies per stack at 8 Gb/die) introduced mid-cycle in March 2018. Bus width unchanged: 4 stacks × 1024 bit = 4096 bit. ECC SECDED is mandatory on Tesla parts.

HBM2 stack details

  • 4 or 8 DRAM dies per stack (4-Hi or 8-Hi)
  • 8 channels × 128-bit per channel = 1024-bit per stack
  • Pin rate 1.75 Gbps (V100); 2.21 Gbps (V100S)
  • Per-stack BW 224 GB/s (1.75 Gbps) or 283 GB/s (V100S)
  • Capacity 4 GB (4-Hi) or 8 GB (8-Hi)

L2 cache

  • 6 MB total (was 4 MB on GP100)
  • 8 partitions of 768 KB
  • Each L2 partition associated with a 64-bit memory channel
  • Crossbar between L2 and SMs
  • ~2× bandwidth/clock vs Pascal L2
10

NVLink 2.0 & NVSwitch 1 (DGX-2)

NVLink 2.0 doubles signalling rate: 25 Gbps NRZ per lane, 8 lanes per link → 25 GB/s/dir per link. Volta GPUs have 6 NVLinks — 150 GB/s/dir aggregate, ~9× PCIe 3.0 x16. New protocol features: address translation services (ATS) so GPU can directly address pageable host memory; cache-coherent atomics across NVLink.

PropertyNVLink 1.0 (Pascal)NVLink 2.0 (Volta)
Lane rate20 Gbps NRZ25 Gbps NRZ
Lanes per link88
Per-link BW (1-dir)20 GB/s25 GB/s
Links per GPU46
Aggregate per GPU80 GB/s150 GB/s
Topologycube-mesh (DGX-1)NVSwitch fabric (DGX-2)
DGX-2 — NVSwitch debuts

DGX-2 (March 2018) paired 16 V100s with 12 NVSwitch-1 chips. Every GPU could talk to every other at full 300 GB/s bidirectional NVLink bandwidth simultaneously — the first time NVIDIA built a true GPU fabric. This made TP=16 useful for the first time and set the architectural precedent for HGX A100/H100/B200.

11

Volta's Innovations

12

Interactive: Volta SKU Picker

FP32 TFLOPS
—
FP64 TFLOPS
—
Tensor TFLOPS
—
VRAM
—
BW (GB/s)
—
TDP
—
NVLink
—
Form factor
—