NVIDIA GPU Architectures Series — Presentation 26

Inside Pascal — The First HBM and NVLink GPU

A low-level look at NVIDIA's 2016 architecture: the GP100 die that introduced HBM2 and NVLink 1.0 to the world, and the GP102/104/106 consumer dies that powered the GTX 10 series. Block diagrams, arithmetic, voltages, clocks, capacities, and the standards Pascal first put in the field.

GP100GP102P100 GTX 1080 TiTSMC 16FF+ HBM2GDDR5XNVLink 1.0 FP16x2PCIe 3.0
Die → GPC → SM → FP/INT → L2 → HBM2 → NVLink 1.0
00

Topics We'll Cover

A low-level tour of NVIDIA's Pascal generation — the die plan, the SM, arithmetic units, process, voltage and clocks, memory, NVLink, and the things Pascal introduced to the rest of the GPU industry.

01

Pascal at a Glance

Pascal launched in 2016 as the first 16 nm NVIDIA architecture, and the first to ship HBM2 and NVLink. Two distinct die families: GP100 for HPC/AI on TSMC 16FF+ with HBM2, and GP102/104/106/107/108 for graphics on the same node with GDDR5/5X.

DieSMsFP32 coresMemorySKUTDP
GP10060 (56 enabled)358416 GB HBM2 @ 720 GB/sTesla P100300 W (SXM)
GP10230 (28 enabled)358411–12 GB GDDR5XGTX 1080 Ti / Titan Xp / Quadro P6000250 W
GP1042025608 GB GDDR5/5XGTX 1080 / 1070 / Tesla P4180 W
GP1061012803–6 GB GDDR5GTX 1060120 W
GP107 / GP1086 / 3768 / 3842–4 GB GDDR5GTX 1050 / 103030–75 W
Why two die families?

GP100's SM is FP-heavy and includes FP64 (1:2 of FP32) and FP16x2 vector ops — designed for scientific computing and the still-young deep-learning workload. GP10x graphics dies drop FP64 to a 1:32 token rate and skip FP16x2 in favour of bigger ROPs and texture units. One architecture, two physical implementations.

02

GP100 Die — Block Diagram

GP100 is organised as a hierarchy of clusters. The top-level chip-wide units are the L2 cache slices, the memory controllers, and the host interface; everything else fans out from a small number of GPCs (Graphics Processing Clusters).

Chip
GP100 die — 610 mm², 15.3 B transistors
GPCs
GPC0GPC1GPC2GPC3GPC4GPC5
TPCs / GPC
5 TPCs × 6 GPCs = 30 TPCs
SMs / TPC
2 SMs × 30 TPCs = 60 SMs (56 enabled on P100)
Memory
8 × 512 KB L2 = 4 MB total4 stacks HBM2 — 4096-bit bus
IO
PCIe 3.0 x164 NVLink 1.0 lanes

Each GPC contains a raster engine, 5 TPCs, and per-GPC fixed-function graphics blocks. Each TPC is two SMs sharing texture units. Total: 60 SMs on a 610 mm² die.

03

The Pascal SM (GP100 Variant)

The GP100 SM is split into two partitions, each with its own warp scheduler. Total per SM:

Compute Resources

  • 64 FP32 cores (32 per partition)
  • 32 FP64 cores (16 per partition) — 1:2 FP64:FP32 rate
  • FP16x2 dual-issue — treats two FP16 values as one packed register
  • 16 LD/ST units, 16 SFUs, 4 texture units
  • 2 warp schedulers, 2 dispatch units each

Memory Resources

  • 256 KB register file (65536 32-bit registers)
  • 64 KB shared memory — statically partitioned from L1
  • 24 KB L1 / texture cache
  • 4 KB instruction cache

GP10x graphics SMs differ: 128 FP32 cores per SM (4 partitions of 32) but only 4 FP64 cores per SM (1:32 rate) and no FP16x2 dual-issue. Larger 96 KB shared memory is statically partitioned with a 48 KB L1.

04

Arithmetic Units & FP16x2

Pascal predates tensor cores. Every matmul runs on the FP32 SIMT lanes — no fused multiply-accumulate at higher density. The architecture's headline arithmetic feature was FP16x2: pack two FP16 values into one 32-bit register and execute one packed instruction per cycle, doubling the FLOP rate over scalar FP16. This was the first hardware support for half-precision deep learning at full throughput.

OperationPer SM per cycleP100 peak (1480 MHz boost)
FP32 FMA6410.6 TFLOPS
FP64 FMA325.3 TFLOPS
FP16x2 packed12821.2 TFLOPS
INT32 / INT1664 (shared with FP32 datapath)10.6 TIPS (on FP32 path)
Why FP16x2 mattered

FP16 was emerging as the storage format for neural-net training weights. Without packed math, FP16 lanes ran at the same rate as FP32 — you got memory savings but no speed-up. FP16x2 gave a real 2× FLOP gain. This was Pascal's bridge from scientific computing into deep learning, and it set the stage for proper tensor cores in Volta.

05

Process — TSMC 16FF+ & Samsung 14LPP

Pascal was NVIDIA's first FinFET node, an enormous density jump over Maxwell's planar 28 nm. GP100, GP104, GP106 on TSMC 16FF+; GP107, GP108 on Samsung 14LPP (lower-cost, lower-leakage variant). 16FF+ delivered roughly 2× transistor density and 30% lower switching energy versus 28HPM Maxwell.

DieProcessTransistorsAreaDensity (M/mm²)
GP100TSMC 16FF+15.3 B610 mm²25.1
GP102TSMC 16FF+11.8 B471 mm²25.1
GP104TSMC 16FF+7.2 B314 mm²22.9
GP106TSMC 16FF+4.4 B200 mm²22.0
GP107Samsung 14LPP3.3 B132 mm²25.0
06

Voltage, Power, Package

Pascal's core voltage runs 0.80–1.06 V depending on boost state, with a per-die voltage curve baked into the GPU's bootstrap. Power delivery on P100 SXM uses a 16-phase digital VRM at 700 A peak; reference GTX 1080 Ti boards use 7+2 phase designs at ~250 A.

SKUTDPConnectorPackage
Tesla P100 (SXM2)300 WSXM2 mezzanine2.5D HBM2 + GP100 on Si interposer
Tesla P100 (PCIe)250 W2× 8-pin PCIesame package, different board
GTX 1080 Ti250 W1× 8-pin + 1× 6-pinflip-chip BGA, GDDR5X around die
GTX 1080 / 1070180 W / 150 W1× 8-pinflip-chip BGA
Tesla P475 W (no aux)PCIe slot onlypassively cooled, datacenter inference

P100 was the first NVIDIA part to use a 2.5D silicon-interposer package — GPU die plus four HBM2 stacks bonded onto a passive silicon interposer mounted on an organic substrate. CoWoS-S was new at TSMC; yields were challenging through 2016.

07

Clocks & Speeds

SKUBaseBoostMemory clockEffective bandwidth
P100 SXM21328 MHz1480 MHzHBM2 1.4 Gbps/pin720 GB/s
GTX 1080 Ti1480 MHz1582 MHzGDDR5X 11 Gbps/pin484 GB/s
GTX 10801607 MHz1733 MHzGDDR5X 10 Gbps/pin320 GB/s
Titan Xp1405 MHz1582 MHzGDDR5X 11.4 Gbps/pin547 GB/s
Tesla P4810 MHz1063 MHzGDDR5 6 Gbps/pin192 GB/s

Pascal introduced GPU Boost 3.0 — each voltage point on the V/F curve is independently programmable, allowing per-state overclocking. Modern nvidia-smi -q -d CLOCK still reports the same V/F state machine that Pascal first exposed.

08

Memory — HBM2 First Time, GDDR5X Second

Pascal shipped the industry's first HBM2 product. P100 connects to 4 stacks of 4-Hi HBM2 over a 4096-bit bus; each stack is 4 GB at launch. GP10x graphics dies use GDDR5 or the new GDDR5X (Micron's quad-data-rate variant, still NRZ signalling, doubling the per-pin rate at the same clock as GDDR5).

HBM2 (P100)

  • 4 stacks × 4 GB = 16 GB
  • 4 stacks × 8 channels × 128-bit = 4096-bit bus
  • Pin rate 1.4 Gbps → 720 GB/s
  • 2.5D Si interposer, micro-bumps, TSVs
  • ECC SECDED on by default

GDDR5X (1080 Ti / 1080 / Titan Xp)

  • 352-bit bus on GP102 (11 channels × 32-bit)
  • 11 Gbps/pin using QDR (NRZ; PAM4 arrived later with GDDR6X)
  • Up to 12 GB capacity (Titan Xp)
  • 484–547 GB/s effective
  • No ECC on consumer SKUs
09

NVLink 1.0 — The Original

NVLink 1.0 is Pascal's other industry-first. Each link is 8 differential pairs × 20 Gbps NRZ = 20 GB/s/dir = 40 GB/s/link bidirectional. P100 has 4 NVLinks for an aggregate 80 GB/s/dir, 160 GB/s bidirectional — ~5× PCIe 3.0 x16. Topology: in DGX-1, eight P100s wired in a hybrid cube-mesh; no NVSwitch yet, so with 4 links per GPU only 16 of the 28 GPU pairs are directly linked; the other 12 take 2 hops.

PropertyNVLink 1.0PCIe 3.0 x16
Per-direction BW20 GB/s/link16 GB/s
Links per P10041 host link
Aggregate per GPU80 GB/s16 GB/s
Latency (D2D)~1 µs~1 µs
CoherentNo (raw P2P)No
Cable / signallingNRZ, 20 Gbps/lane, twin-axNRZ, 8 GT/s
P9 + NVLink

IBM's POWER9 was the only host CPU that connected to NVLink directly — OpenPOWER systems like the Summit supercomputer at Oak Ridge ran all CPU-GPU traffic over NVLink, not PCIe. Volta inherited this; x86 hosts were stuck with PCIe until Grace arrived in 2023.

10

PCIe 3.0, Display, NVENC/NVDEC

PCIe 3.0 x16

Single host link at 8 GT/s, 16 GB/s/dir. Resizable BAR not yet a thing in 2016 mainstream BIOS. ATS (Address Translation Services) supported.

NVENC / NVDEC

Dedicated H.264 encode + HEVC encode (8b only); HEVC + VP9 decode. ~2× 1080p60 streams encode per chip. Tesla P4 became the cloud transcoding workhorse.

Display engine

Up to 4 displays, DisplayPort 1.4, HDMI 2.0b, HDR10. Pascal added simultaneous multi-projection and lens-matched shading for early VR headsets.

11

Pascal's Innovations — What It Put on the Map

12

Interactive: Pascal SKU Picker

Die
—
SMs
—
FP32 TFLOPS
—
FP16x2 TFLOPS
—
VRAM
—
BW (GB/s)
—
TDP
—
NVLink
—