NVIDIA GPU Architectures Series — Presentation 28

Inside Turing — RT Cores, INT8/INT4 Tensor Cores, GDDR6

A low-level look at the 2018 architecture that brought consumer-class hardware ray tracing, second-generation tensor cores with INT8/INT4, and the T4 inference card that defined cloud GPU economics for years. Block diagrams, arithmetic, process, voltages, clocks, capacities, and standards.

TU102TU104TU106TU116 T4RTX 2080 Ti TSMC 12FFNGDDR6 RT 1st genTC 2nd gen
TU10x → GPC → SM → TC2 → RT1 → GDDR6 → NVLink 2.0
00

Topics We'll Cover

01

Turing at a Glance — Five Dies, One Architecture

Turing launched in September 2018 on the same TSMC 12FFN node as Volta, but with a far more diverse die family. TU102 / 104 / 106 are full RTX dies with RT and tensor cores; TU116 / 117 are budget dies that drop both, becoming the GTX 16-series. The same generation also produced the Tesla T4, NVIDIA's most successful datacenter inference card.

DieSMsRT coresTensor coresTop SKUTDP
TU10272 (68 enabled)72576RTX 2080 Ti / Titan RTX / Quadro RTX 8000250–280 W
TU10448 (46 enabled)48384RTX 2080 / 2070 Super / T4175–215 W (T4 = 70 W)
TU1063636288RTX 2070 / 2060175–185 W
TU116240 (no RT)0 (no TC)GTX 1660 Ti / 1660 Super120 W
TU1171600GTX 165075 W
02

TU102 Die — Block Diagram

Chip
TU102 die — 754 mm², 18.6 B transistors
GPCs
GPC0GPC1GPC2GPC3GPC4GPC5
TPCs / GPC
6 TPCs × 6 GPCs = 36 TPCs (72 SMs)
L2
6 MB total (12 partitions of 512 KB)
Memory
12 GDDR6 controllers, 384-bit bus
IO
PCIe 3.0 x162 NVLink 2.0 (bridge only on 2080 Ti / Titan / Quadro)
03

The Turing SM — Adds RT Core, Splits FP/INT (again)

Turing inherits Volta's 4-partition layout and concurrent INT/FP datapaths. New: a dedicated RT core per SM (1 per SM, not per partition). Tensor cores upgraded to 2nd generation. FP64 cut to a 1:32 token rate — this is a graphics architecture that retains compute-friendly tensor cores, not a HPC chip.

Per-SM compute (TU10x with RT)

  • 64 FP32 cores (16 per partition)
  • 64 INT32 cores (16 per partition) — concurrent with FP32
  • 2 FP64 cores per SM — 1:32 of FP32
  • 8 tensor cores (2 per partition) — 2nd generation
  • 1 RT core per SM — 1st generation
  • 16 LD/ST, 16 SFU
  • 4 warp schedulers

Per-SM memory

  • 256 KB register file
  • 96 KB unified L1 + shared (configurable 32+64 or 64+32)
  • 2× L1 BW vs Volta
  • 4 KB instruction cache

TU116/117 SMs drop the RT and tensor cores entirely — just FP32 + INT32 + FP64 lanes. They re-add INT8 dot-product (DP4A) for inference at low cost.

04

2nd-Generation Tensor Cores — INT8 & INT4

Volta's 1st-gen tensor cores were FP16-only. Turing adds INT8 (2× FP16) and INT4 (4× FP16) at the same MMA shape, opening the door to quantised inference at multi-TOPS rates on small chips.

OperationPer SM per cycleRTX 2080 Ti peak (1545 MHz)Tesla T4 peak (1590 MHz)
FP32 FMA6413.4 TFLOPS8.1 TFLOPS
FP16x2 (no TC)12826.9 TFLOPS16.2 TFLOPS
Tensor FP16512108 TFLOPS (FP16 accumulate; 53.8 with FP32 accumulate)65 TFLOPS
Tensor INT81024215 TOPS130 TOPS
Tensor INT42048431 TOPS260 TOPS

PTX additions: mma.sync.aligned.m8n8k16.s32.s8.s8.s32 (INT8 with INT32 accumulate) and mma.sync.aligned.m8n8k32.s32.s4.s4.s32 (INT4). DLSS 1.0/2.0 rendering on consumer Turing rides the same hardware.

05

1st-Generation RT Cores — BVH in Hardware

Each Turing RT core does two things: BVH (bounding-volume hierarchy) traversal and ray-triangle intersection testing, in fixed-function hardware. Without RT cores, the same operations run as compute kernels on the SM and consume ~10× the time per ray. Top-die RTX 2080 Ti claims ~10 Gigarays/sec.

Why fixed-function

BVH traversal is divergent and pointer-chasing — pathological for SIMT lanes. A dedicated unit with its own memory pipeline retires one ray-box test per cycle without warp divergence cost.

Programming model

Exposed via DXR (DirectX Raytracing), Vulkan KHR_ray_tracing, OptiX 7. The GPU builds the BVH (TLAS / BLAS), the RT core traverses it, the SM runs hit/miss shaders.

For AI workflows, RT cores became relevant later via path-traced synthetic data generation (Omniverse) — a use case 2018 didn't anticipate.

06

Process & Voltages

All Turing dies on TSMC 12FFN — same node as Volta. Density per mm² lower than GV100 because graphics-die layouts tend to be I/O-heavier. Core voltage 0.71–1.07 V.

DieTransistorsAreaDensity (M/mm²)
TU10218.6 B754 mm²24.7
TU10413.6 B545 mm²25.0
TU10610.8 B445 mm²24.3
TU1166.6 B284 mm²23.2
TU1174.7 B200 mm²23.5
07

Clocks, Power, Form Factors

SKUBaseBoostMemoryBWTDP
RTX 2080 Ti1350 MHz1545 MHz11 GB GDDR6 14 Gbps616 GB/s250 W (FE 260 W)
Titan RTX1350 MHz1770 MHz24 GB GDDR6 14 Gbps672 GB/s280 W
Quadro RTX 80001395 MHz1770 MHz48 GB GDDR6 ECC672 GB/s295 W
RTX 2080 Super1650 MHz1815 MHz8 GB GDDR6 15.5 Gbps496 GB/s250 W
Tesla T4585 MHz1590 MHz16 GB GDDR6 10 Gbps320 GB/s70 W (slot only)
08

Memory — GDDR6, the New Mainstream

Turing introduced GDDR6 at 14 Gbps/pin (later 15.5 Gbps Super refresh). Compared to Pascal's GDDR5X, GDDR6 uses regular DDR signalling at 14 Gbps PAM2 with two channels per chip (16-bit each = 32-bit per chip). Same capacity per chip; cleaner signal eye than GDDR5X's quirky QDR.

SKUCapacityBusPin rateBW
RTX 2080 Ti11 GB352-bit14 Gbps616 GB/s
Titan RTX / Quadro RTX 600024 GB384-bit14 Gbps672 GB/s
Quadro RTX 800048 GB ECC384-bit14 Gbps672 GB/s
RTX 2080 Super8 GB256-bit15.5 Gbps496 GB/s
Tesla T416 GB ECC256-bit10 Gbps320 GB/s
09

NVLink 2.0 (Bridge), PCIe 3.0

Turing was the first consumer arch with NVLink. RTX 2080 Ti, Titan RTX, and Quadro RTX 6000/8000 carry an NVLink 2.0 bridge connector for two-card configurations: 2 links × 25 GB/s/dir = 50 GB/s/dir aggregate, 100 GB/s bidirectional. RTX 2080 and 2080 Super carry a single link (25 GB/s/dir); the RTX 2070 and below have none. From RTX 30-series onward, only the 3090 / Quadro / A6000 retain bridges; from RTX 40-series, no consumer card has NVLink at all.

Host link

PCIe 3.0 x16 host link — 16 GB/s/dir — same as Pascal/Volta. Turing did not introduce PCIe 4.0; that came with Ampere.

10

Tesla T4 — The 70 W Inference Card That Won the Cloud

The T4 deserves its own slide. A single-slot, full-height card on TU104, 70 W slot-only (no aux power), passively cooled, 16 GB GDDR6 ECC, 130 INT8 TOPS, 65 FP16 TFLOPS via tensor cores. Every major cloud provider (AWS, GCP, Azure, OCI) deployed T4s by the tens of thousands for inference workloads from 2019–2023.

Why it won

Best perf/W in the datacenter for INT8 inference. ~3× lower op-cost than V100 at similar throughput. Slot-only power = retrofit any 1U/2U server.

What replaced it

Ada Lovelace L4 (24 GB, 72 W, adds FP8, 4× tensor throughput, GDDR6) — same form factor, generational bump.

Legacy in 2026

T4 still cheap-and-cheerful for embedding inference, OCR, ASR. Doesn't run modern LLMs well (no FP8, weak tensor cores by 2026 standards).

11

Turing's Innovations

12

Interactive: Turing SKU Picker

Die
—
SMs
—
FP32 TF
—
Tensor FP16 TF
—
INT8 TOPS
—
VRAM
—
BW
—
TDP
—