A low-level look at NVIDIA's 2016 architecture: the GP100 die that introduced HBM2 and NVLink 1.0 to the world, and the GP102/104/106 consumer dies that powered the GTX 10 series. Block diagrams, arithmetic, voltages, clocks, capacities, and the standards Pascal first put in the field.
A low-level tour of NVIDIA's Pascal generation — the die plan, the SM, arithmetic units, process, voltage and clocks, memory, NVLink, and the things Pascal introduced to the rest of the GPU industry.
Pascal launched in 2016 as the first 16 nm NVIDIA architecture, and the first to ship HBM2 and NVLink. Two distinct die families: GP100 for HPC/AI on TSMC 16FF+ with HBM2, and GP102/104/106/107/108 for graphics on the same node with GDDR5/5X.
| Die | SMs | FP32 cores | Memory | SKU | TDP |
|---|---|---|---|---|---|
| GP100 | 60 (56 enabled) | 3584 | 16 GB HBM2 @ 720 GB/s | Tesla P100 | 300 W (SXM) |
| GP102 | 30 (28 enabled) | 3584 | 11–12 GB GDDR5X | GTX 1080 Ti / Titan Xp / Quadro P6000 | 250 W |
| GP104 | 20 | 2560 | 8 GB GDDR5/5X | GTX 1080 / 1070 / Tesla P4 | 180 W |
| GP106 | 10 | 1280 | 3–6 GB GDDR5 | GTX 1060 | 120 W |
| GP107 / GP108 | 6 / 3 | 768 / 384 | 2–4 GB GDDR5 | GTX 1050 / 1030 | 30–75 W |
GP100's SM is FP-heavy and includes FP64 (1:2 of FP32) and FP16x2 vector ops — designed for scientific computing and the still-young deep-learning workload. GP10x graphics dies drop FP64 to a 1:32 token rate and skip FP16x2 in favour of bigger ROPs and texture units. One architecture, two physical implementations.
GP100 is organised as a hierarchy of clusters. The top-level chip-wide units are the L2 cache slices, the memory controllers, and the host interface; everything else fans out from a small number of GPCs (Graphics Processing Clusters).
Each GPC contains a raster engine, 5 TPCs, and per-GPC fixed-function graphics blocks. Each TPC is two SMs sharing texture units. Total: 60 SMs on a 610 mm² die.
The GP100 SM is split into two partitions, each with its own warp scheduler. Total per SM:
GP10x graphics SMs differ: 128 FP32 cores per SM (4 partitions of 32) but only 4 FP64 cores per SM (1:32 rate) and no FP16x2 dual-issue. Larger 96 KB shared memory is statically partitioned with a 48 KB L1.
Pascal predates tensor cores. Every matmul runs on the FP32 SIMT lanes — no fused multiply-accumulate at higher density. The architecture's headline arithmetic feature was FP16x2: pack two FP16 values into one 32-bit register and execute one packed instruction per cycle, doubling the FLOP rate over scalar FP16. This was the first hardware support for half-precision deep learning at full throughput.
| Operation | Per SM per cycle | P100 peak (1480 MHz boost) |
|---|---|---|
| FP32 FMA | 64 | 10.6 TFLOPS |
| FP64 FMA | 32 | 5.3 TFLOPS |
| FP16x2 packed | 128 | 21.2 TFLOPS |
| INT32 / INT16 | 64 (shared with FP32 datapath) | 10.6 TIPS (on FP32 path) |
FP16 was emerging as the storage format for neural-net training weights. Without packed math, FP16 lanes ran at the same rate as FP32 — you got memory savings but no speed-up. FP16x2 gave a real 2× FLOP gain. This was Pascal's bridge from scientific computing into deep learning, and it set the stage for proper tensor cores in Volta.
Pascal was NVIDIA's first FinFET node, an enormous density jump over Maxwell's planar 28 nm. GP100, GP104, GP106 on TSMC 16FF+; GP107, GP108 on Samsung 14LPP (lower-cost, lower-leakage variant). 16FF+ delivered roughly 2× transistor density and 30% lower switching energy versus 28HPM Maxwell.
| Die | Process | Transistors | Area | Density (M/mm²) |
|---|---|---|---|---|
| GP100 | TSMC 16FF+ | 15.3 B | 610 mm² | 25.1 |
| GP102 | TSMC 16FF+ | 11.8 B | 471 mm² | 25.1 |
| GP104 | TSMC 16FF+ | 7.2 B | 314 mm² | 22.9 |
| GP106 | TSMC 16FF+ | 4.4 B | 200 mm² | 22.0 |
| GP107 | Samsung 14LPP | 3.3 B | 132 mm² | 25.0 |
Pascal's core voltage runs 0.80–1.06 V depending on boost state, with a per-die voltage curve baked into the GPU's bootstrap. Power delivery on P100 SXM uses a 16-phase digital VRM at 700 A peak; reference GTX 1080 Ti boards use 7+2 phase designs at ~250 A.
| SKU | TDP | Connector | Package |
|---|---|---|---|
| Tesla P100 (SXM2) | 300 W | SXM2 mezzanine | 2.5D HBM2 + GP100 on Si interposer |
| Tesla P100 (PCIe) | 250 W | 2× 8-pin PCIe | same package, different board |
| GTX 1080 Ti | 250 W | 1× 8-pin + 1× 6-pin | flip-chip BGA, GDDR5X around die |
| GTX 1080 / 1070 | 180 W / 150 W | 1× 8-pin | flip-chip BGA |
| Tesla P4 | 75 W (no aux) | PCIe slot only | passively cooled, datacenter inference |
P100 was the first NVIDIA part to use a 2.5D silicon-interposer package — GPU die plus four HBM2 stacks bonded onto a passive silicon interposer mounted on an organic substrate. CoWoS-S was new at TSMC; yields were challenging through 2016.
| SKU | Base | Boost | Memory clock | Effective bandwidth |
|---|---|---|---|---|
| P100 SXM2 | 1328 MHz | 1480 MHz | HBM2 1.4 Gbps/pin | 720 GB/s |
| GTX 1080 Ti | 1480 MHz | 1582 MHz | GDDR5X 11 Gbps/pin | 484 GB/s |
| GTX 1080 | 1607 MHz | 1733 MHz | GDDR5X 10 Gbps/pin | 320 GB/s |
| Titan Xp | 1405 MHz | 1582 MHz | GDDR5X 11.4 Gbps/pin | 547 GB/s |
| Tesla P4 | 810 MHz | 1063 MHz | GDDR5 6 Gbps/pin | 192 GB/s |
Pascal introduced GPU Boost 3.0 — each voltage point on the V/F curve is independently programmable, allowing per-state overclocking. Modern nvidia-smi -q -d CLOCK still reports the same V/F state machine that Pascal first exposed.
Pascal shipped the industry's first HBM2 product. P100 connects to 4 stacks of 4-Hi HBM2 over a 4096-bit bus; each stack is 4 GB at launch. GP10x graphics dies use GDDR5 or the new GDDR5X (Micron's quad-data-rate variant, still NRZ signalling, doubling the per-pin rate at the same clock as GDDR5).
NVLink 1.0 is Pascal's other industry-first. Each link is 8 differential pairs × 20 Gbps NRZ = 20 GB/s/dir = 40 GB/s/link bidirectional. P100 has 4 NVLinks for an aggregate 80 GB/s/dir, 160 GB/s bidirectional — ~5× PCIe 3.0 x16. Topology: in DGX-1, eight P100s wired in a hybrid cube-mesh; no NVSwitch yet, so with 4 links per GPU only 16 of the 28 GPU pairs are directly linked; the other 12 take 2 hops.
| Property | NVLink 1.0 | PCIe 3.0 x16 |
|---|---|---|
| Per-direction BW | 20 GB/s/link | 16 GB/s |
| Links per P100 | 4 | 1 host link |
| Aggregate per GPU | 80 GB/s | 16 GB/s |
| Latency (D2D) | ~1 µs | ~1 µs |
| Coherent | No (raw P2P) | No |
| Cable / signalling | NRZ, 20 Gbps/lane, twin-ax | NRZ, 8 GT/s |
IBM's POWER9 was the only host CPU that connected to NVLink directly — OpenPOWER systems like the Summit supercomputer at Oak Ridge ran all CPU-GPU traffic over NVLink, not PCIe. Volta inherited this; x86 hosts were stuck with PCIe until Grace arrived in 2023.
Single host link at 8 GT/s, 16 GB/s/dir. Resizable BAR not yet a thing in 2016 mainstream BIOS. ATS (Address Translation Services) supported.
Dedicated H.264 encode + HEVC encode (8b only); HEVC + VP9 decode. ~2× 1080p60 streams encode per chip. Tesla P4 became the cloud transcoding workhorse.
Up to 4 displays, DisplayPort 1.4, HDMI 2.0b, HDR10. Pascal added simultaneous multi-projection and lens-matched shading for early VR headsets.
nvidia-smi today.