A low-level look at the 2018 architecture that brought consumer-class hardware ray tracing, second-generation tensor cores with INT8/INT4, and the T4 inference card that defined cloud GPU economics for years. Block diagrams, arithmetic, process, voltages, clocks, capacities, and standards.
Turing launched in September 2018 on the same TSMC 12FFN node as Volta, but with a far more diverse die family. TU102 / 104 / 106 are full RTX dies with RT and tensor cores; TU116 / 117 are budget dies that drop both, becoming the GTX 16-series. The same generation also produced the Tesla T4, NVIDIA's most successful datacenter inference card.
| Die | SMs | RT cores | Tensor cores | Top SKU | TDP |
|---|---|---|---|---|---|
| TU102 | 72 (68 enabled) | 72 | 576 | RTX 2080 Ti / Titan RTX / Quadro RTX 8000 | 250–280 W |
| TU104 | 48 (46 enabled) | 48 | 384 | RTX 2080 / 2070 Super / T4 | 175–215 W (T4 = 70 W) |
| TU106 | 36 | 36 | 288 | RTX 2070 / 2060 | 175–185 W |
| TU116 | 24 | 0 (no RT) | 0 (no TC) | GTX 1660 Ti / 1660 Super | 120 W |
| TU117 | 16 | 0 | 0 | GTX 1650 | 75 W |
Turing inherits Volta's 4-partition layout and concurrent INT/FP datapaths. New: a dedicated RT core per SM (1 per SM, not per partition). Tensor cores upgraded to 2nd generation. FP64 cut to a 1:32 token rate — this is a graphics architecture that retains compute-friendly tensor cores, not a HPC chip.
TU116/117 SMs drop the RT and tensor cores entirely — just FP32 + INT32 + FP64 lanes. They re-add INT8 dot-product (DP4A) for inference at low cost.
Volta's 1st-gen tensor cores were FP16-only. Turing adds INT8 (2× FP16) and INT4 (4× FP16) at the same MMA shape, opening the door to quantised inference at multi-TOPS rates on small chips.
| Operation | Per SM per cycle | RTX 2080 Ti peak (1545 MHz) | Tesla T4 peak (1590 MHz) |
|---|---|---|---|
| FP32 FMA | 64 | 13.4 TFLOPS | 8.1 TFLOPS |
| FP16x2 (no TC) | 128 | 26.9 TFLOPS | 16.2 TFLOPS |
| Tensor FP16 | 512 | 108 TFLOPS (FP16 accumulate; 53.8 with FP32 accumulate) | 65 TFLOPS |
| Tensor INT8 | 1024 | 215 TOPS | 130 TOPS |
| Tensor INT4 | 2048 | 431 TOPS | 260 TOPS |
PTX additions: mma.sync.aligned.m8n8k16.s32.s8.s8.s32 (INT8 with INT32 accumulate) and mma.sync.aligned.m8n8k32.s32.s4.s4.s32 (INT4). DLSS 1.0/2.0 rendering on consumer Turing rides the same hardware.
Each Turing RT core does two things: BVH (bounding-volume hierarchy) traversal and ray-triangle intersection testing, in fixed-function hardware. Without RT cores, the same operations run as compute kernels on the SM and consume ~10× the time per ray. Top-die RTX 2080 Ti claims ~10 Gigarays/sec.
BVH traversal is divergent and pointer-chasing — pathological for SIMT lanes. A dedicated unit with its own memory pipeline retires one ray-box test per cycle without warp divergence cost.
Exposed via DXR (DirectX Raytracing), Vulkan KHR_ray_tracing, OptiX 7. The GPU builds the BVH (TLAS / BLAS), the RT core traverses it, the SM runs hit/miss shaders.
For AI workflows, RT cores became relevant later via path-traced synthetic data generation (Omniverse) — a use case 2018 didn't anticipate.
All Turing dies on TSMC 12FFN — same node as Volta. Density per mm² lower than GV100 because graphics-die layouts tend to be I/O-heavier. Core voltage 0.71–1.07 V.
| Die | Transistors | Area | Density (M/mm²) |
|---|---|---|---|
| TU102 | 18.6 B | 754 mm² | 24.7 |
| TU104 | 13.6 B | 545 mm² | 25.0 |
| TU106 | 10.8 B | 445 mm² | 24.3 |
| TU116 | 6.6 B | 284 mm² | 23.2 |
| TU117 | 4.7 B | 200 mm² | 23.5 |
| SKU | Base | Boost | Memory | BW | TDP |
|---|---|---|---|---|---|
| RTX 2080 Ti | 1350 MHz | 1545 MHz | 11 GB GDDR6 14 Gbps | 616 GB/s | 250 W (FE 260 W) |
| Titan RTX | 1350 MHz | 1770 MHz | 24 GB GDDR6 14 Gbps | 672 GB/s | 280 W |
| Quadro RTX 8000 | 1395 MHz | 1770 MHz | 48 GB GDDR6 ECC | 672 GB/s | 295 W |
| RTX 2080 Super | 1650 MHz | 1815 MHz | 8 GB GDDR6 15.5 Gbps | 496 GB/s | 250 W |
| Tesla T4 | 585 MHz | 1590 MHz | 16 GB GDDR6 10 Gbps | 320 GB/s | 70 W (slot only) |
Turing introduced GDDR6 at 14 Gbps/pin (later 15.5 Gbps Super refresh). Compared to Pascal's GDDR5X, GDDR6 uses regular DDR signalling at 14 Gbps PAM2 with two channels per chip (16-bit each = 32-bit per chip). Same capacity per chip; cleaner signal eye than GDDR5X's quirky QDR.
| SKU | Capacity | Bus | Pin rate | BW |
|---|---|---|---|---|
| RTX 2080 Ti | 11 GB | 352-bit | 14 Gbps | 616 GB/s |
| Titan RTX / Quadro RTX 6000 | 24 GB | 384-bit | 14 Gbps | 672 GB/s |
| Quadro RTX 8000 | 48 GB ECC | 384-bit | 14 Gbps | 672 GB/s |
| RTX 2080 Super | 8 GB | 256-bit | 15.5 Gbps | 496 GB/s |
| Tesla T4 | 16 GB ECC | 256-bit | 10 Gbps | 320 GB/s |
Turing was the first consumer arch with NVLink. RTX 2080 Ti, Titan RTX, and Quadro RTX 6000/8000 carry an NVLink 2.0 bridge connector for two-card configurations: 2 links × 25 GB/s/dir = 50 GB/s/dir aggregate, 100 GB/s bidirectional. RTX 2080 and 2080 Super carry a single link (25 GB/s/dir); the RTX 2070 and below have none. From RTX 30-series onward, only the 3090 / Quadro / A6000 retain bridges; from RTX 40-series, no consumer card has NVLink at all.
PCIe 3.0 x16 host link — 16 GB/s/dir — same as Pascal/Volta. Turing did not introduce PCIe 4.0; that came with Ampere.
The T4 deserves its own slide. A single-slot, full-height card on TU104, 70 W slot-only (no aux power), passively cooled, 16 GB GDDR6 ECC, 130 INT8 TOPS, 65 FP16 TFLOPS via tensor cores. Every major cloud provider (AWS, GCP, Azure, OCI) deployed T4s by the tens of thousands for inference workloads from 2019–2023.
Best perf/W in the datacenter for INT8 inference. ~3× lower op-cost than V100 at similar throughput. Slot-only power = retrofit any 1U/2U server.
Ada Lovelace L4 (24 GB, 72 W, adds FP8, 4× tensor throughput, GDDR6) — same form factor, generational bump.
T4 still cheap-and-cheerful for embedding inference, OCR, ASR. Doesn't run modern LLMs well (no FP8, weak tensor cores by 2026 standards).