A low-level look at GV100 — the 2017 die that introduced 1st-generation tensor cores, separated INT and FP datapaths, and gave every thread its own program counter. Block diagrams, arithmetic, process, voltage, clocks, capacities, and the NVLink 2.0 / NVSwitch 1 fabric.
Volta launched in 2017 (V100 announced May 2017; Titan V in December) as a single-die, datacenter-only architecture. There is no consumer Volta. The graphics path skipped Volta entirely — Pascal's GP102 carried the high end until Turing arrived a year later. Volta was a deep-learning bet by NVIDIA.
| SKU | Form factor | SMs | Tensor cores | Memory | TDP |
|---|---|---|---|---|---|
| Tesla V100 SXM2 16 GB | SXM2 mezzanine | 80 | 640 | 16 GB HBM2 @ 900 GB/s | 300 W |
| Tesla V100 SXM3 32 GB | SXM3 (DGX-2) | 80 | 640 | 32 GB HBM2 @ 900 GB/s | 350 W |
| Tesla V100 PCIe 16/32 GB | PCIe | 80 | 640 | 16/32 GB HBM2 | 250 W |
| Quadro GV100 | workstation | 80 | 640 | 32 GB HBM2 | 250 W |
| Titan V | consumer (HEDT) | 80 | 640 | 12 GB HBM2 | 250 W |
| Tesla V100S | PCIe (refresh 2019) | 80 | 640 | 32 GB HBM2 @ 1134 GB/s | 250 W |
The Volta SM is reorganised relative to Pascal's two-partition design. Four partitions, each with its own warp scheduler and dispatch unit. Critically, FP32 and INT32 datapaths are separate and concurrent — they can issue together in the same cycle, doubling effective issue rate when an integer address calculation overlaps a floating-point operation.
The headline feature. Each Volta SM has 8 tensor cores; each performs a fused 4×4×4 matrix multiply-accumulate per cycle: D = A·B + C, where A and B are 4×4 FP16 matrices and C/D are 4×4 FP32 (or FP16) matrices — 64 FMAs per tensor core per cycle.
| Operation | Per SM per cycle | V100 SXM2 peak (1530 MHz) |
|---|---|---|
| FP32 FMA (concurrent with INT) | 64 (+64 INT32) | 15.7 TFLOPS |
| FP64 FMA | 32 | 7.8 TFLOPS |
| FP16x2 packed | 128 | 31.4 TFLOPS |
| Tensor core HMMA (FP16->FP32) | 512 | 125 TFLOPS |
Programming interface: PTX wmma.mma.sync.aligned.m16n16k16.f32.f16.f16.f32; C++ nvcuda::wmma. Tensor cores were a ~12× FLOP increase (125 vs 10.6 TFLOPS) over Pascal's FP32 path on transformer-style workloads — the moment "deep learning hardware" stopped being marketing.
Pre-Volta SIMT: all 32 lanes of a warp share one program counter. Divergent branches mask lanes, taken-and-not-taken paths run sequentially. Critical sections were unsafe (a thread waiting on a lock could starve threads in the same warp holding it → deadlock).
Volta gives each thread its own PC and call stack. The warp can re-converge opportunistically; threads on different branches can interleave at instruction granularity. Spinlocks, fine-grained synchronisation, and producer-consumer within a warp now work correctly.
Code that relied on lock-step now needs __shfl_sync(mask, x, lane) with explicit masks (and __ballot_sync, __activemask). The non-sync intrinsics are deprecated. The CUDA 9 release that landed alongside Volta documented every breakage; legacy kernels that ignored the change silently produced wrong results.
GV100 is TSMC 12FF — a half-node refinement of 16FF+ with denser standard cells and improved transistor performance, but same metal pitch — so density gains over GP100 came mostly from cell rework, not lithography. NVIDIA called it 12 nm; TSMC's official name was 12FFN (FinFET with NVIDIA-specific tweaks).
| Metric | GP100 (Pascal) | GV100 (Volta) |
|---|---|---|
| Process | TSMC 16FF+ | TSMC 12FF |
| Transistors | 15.3 B | 21.1 B |
| Die area | 610 mm² | 815 mm² |
| Density | 25.1 M/mm² | 25.9 M/mm² |
| SMs | 60 | 84 (40% more) |
815 mm² was very close to TSMC's reticle limit (~858 mm², 26 × 33 mm). GV100 was at the ceiling of what one die could be at the 12FF yield curve.
V100 SXM2 introduced the SXM2 form factor proper — the same socket Pascal P100 used, at the same 300 W. The SXM3 variant in DGX-2 pushed to 350 W to support 32 GB HBM2 stacks at higher pin rate.
| SKU | Base | Boost | Memory | BW |
|---|---|---|---|---|
| V100 SXM2 16/32 | 1290 MHz | 1530 MHz | HBM2 1.75 Gbps/pin | 900 GB/s |
| V100 SXM3 32 (DGX-2) | 1312 MHz | 1530 MHz | HBM2 1.75 Gbps/pin | 900 GB/s |
| V100 PCIe 16/32 | 1230 MHz | 1380 MHz | HBM2 1.75 Gbps/pin | 900 GB/s |
| V100S PCIe 32 (2019) | 1245 MHz | 1597 MHz | HBM2 2.21 Gbps/pin | 1134 GB/s |
| Titan V | 1200 MHz | 1455 MHz | HBM2 1.7 Gbps/pin | 653 GB/s (3 stacks) |
Titan V is the odd one: same GV100 die but with one HBM2 stack disabled (3-of-4 stacks → 12 GB, 3072-bit bus, 653 GB/s) for yield reasons.
V100 was the first NVIDIA part to ship 32 GB HBM2, using 8-Hi stacks (8 dies per stack at 8 Gb/die) introduced mid-cycle in March 2018. Bus width unchanged: 4 stacks × 1024 bit = 4096 bit. ECC SECDED is mandatory on Tesla parts.
NVLink 2.0 doubles signalling rate: 25 Gbps NRZ per lane, 8 lanes per link → 25 GB/s/dir per link. Volta GPUs have 6 NVLinks — 150 GB/s/dir aggregate, ~9× PCIe 3.0 x16. New protocol features: address translation services (ATS) so GPU can directly address pageable host memory; cache-coherent atomics across NVLink.
| Property | NVLink 1.0 (Pascal) | NVLink 2.0 (Volta) |
|---|---|---|
| Lane rate | 20 Gbps NRZ | 25 Gbps NRZ |
| Lanes per link | 8 | 8 |
| Per-link BW (1-dir) | 20 GB/s | 25 GB/s |
| Links per GPU | 4 | 6 |
| Aggregate per GPU | 80 GB/s | 150 GB/s |
| Topology | cube-mesh (DGX-1) | NVSwitch fabric (DGX-2) |
DGX-2 (March 2018) paired 16 V100s with 12 NVSwitch-1 chips. Every GPU could talk to every other at full 300 GB/s bidirectional NVLink bandwidth simultaneously — the first time NVIDIA built a true GPU fabric. This made TP=16 useful for the first time and set the architectural precedent for HGX A100/H100/B200.
mma.sync PTX. Every later tensor-core generation extends this same surface.