NVIDIA GPU Architectures Series — Presentation 30

Inside Ada Lovelace — AD102, 96 MB L2, 3rd-Gen RT

A low-level look at the 2022 architecture that put a 96 MB L2 cache on a graphics die, gave us 3rd-gen RT cores with opacity and displaced micro-meshes, and added an FP8 datapath to 4th-gen tensor cores across the whole Ada line. Block diagrams, arithmetic, TSMC 4N process, voltages, clocks, GDDR6X memory, no-NVLink consequences.

AD102AD103AD104AD106 RTX 4090L40SL4 TSMC 4NGDDR6X RT 3rd genDLSS 3OFA
AD10x → GPC → SM → TC4 → RT3 → 96 MB L2 → GDDR6X
00

Topics We'll Cover

01

Ada at a Glance — Consumer + Workstation, No HPC Die

Ada Lovelace launched in October 2022 as a graphics/workstation/inference architecture only. There is no datacenter HPC die in this generation — that workload is served by Hopper (GH100), launched in parallel on the same TSMC 4N node. AD102 to AD107 cover RTX 40 series consumer, RTX 6000 Ada / RTX 5000 Ada workstation, and L40 / L40S / L4 datacenter inference.

DieSMsTensor coresRT coresTop SKUTDP
AD102144 (128 enabled on 4090)576 (512 on 4090)144 (128 on 4090)RTX 4090 / RTX 6000 Ada / L40S300–450 W
AD1038032080RTX 4080 Super / RTX 4080320 W
AD1046024060RTX 4070 Ti / 4070 / L472–285 W
AD1063614436RTX 4060 Ti160–165 W
AD107249624RTX 4060 / 4050 mobile115 W
02

AD102 Die — Block Diagram

Chip
AD102 — 608 mm², 76.3 B transistors, TSMC 4N
GPCs
GPC0GPC1GPC2GPC3GPC4GPC5GPC6GPC7GPC8GPC9GPC10GPC11
SMs / GPC
12 SMs × 12 GPCs = 144 SMs (128 enabled on RTX 4090)
L2
96 MB L2 — 16× the GA102 L2
Memory
12 GDDR6X channels, 384-bit bus
IO
PCIe 4.0 x16no NVLink (consumer or RTX 6000 Ada)

Density: 76.3 B / 608 mm² = 125.5 M/mm² — nearly 2× the GA100 density and 2.8× the Samsung 8N GA102. The TSMC 4N node was a huge generational lift, especially for AD102 which managed to fit ~1.7× the SMs of GA102 (144 vs 84) on 3% less silicon.

03

The Ada SM — Doubled FP32 Plus FP8-Capable Tensor Cores

The Ada SM inherits GA10x's "doubled FP32" partition layout with new RT and tensor cores. FP8 (E4M3 / E5M2) is exposed in tensor cores across the entire Ada line — including the RTX 4090. What consumer Ada lacks vs Hopper is the Transformer Engine's dynamic per-tensor scaling, not the FP8 silicon itself.

Per-SM compute

  • 4 partitions, 1 warp scheduler each
  • 128 FP32 cores (32/partition; doubled-FP32 like GA10x)
  • 64 INT32 cores (16/partition)
  • 2 FP64 cores per SM — 1:64 of FP32
  • 4 tensor cores (4th gen, FP8-capable)
  • 1 RT core (3rd gen) per SM
  • 16 LD/ST, 16 SFU

Per-SM memory

  • 256 KB register file
  • 128 KB L1 + shared (configurable 100/28, 64/64, 28/100)
  • ~2× L1 BW vs Ampere
  • 4 KB instruction cache
FP8 vs Hopper's Transformer Engine

Tensor cores on Ada are physically capable of FP8 E4M3 / E5M2 on every Ada SKU (sm_89), RTX 4090 included. What Ada does not have is Hopper's Transformer Engine — the dynamic per-tensor scaling that automatically picks E4M3 vs E5M2 and tracks amax statistics. On Ada you do FP8 with explicit static scaling (e.g. via TransformerEngine library, TensorRT-LLM). Datacenter Ada (L40 / L40S / RTX 6000 Ada) markets the FP8 throughput; consumer Ada exposes the same instructions, just without the Transformer Engine's scheduling layer.

04

4th-Generation Tensor Cores — FP8 Across the Line

OperationRTX 4090 (2520 MHz)L40S (2520 MHz)
FP32 FMA82.6 TFLOPS91.6 TFLOPS
Tensor TF3282.6 TFLOPS183 TFLOPS
Tensor BF16 / FP16165 TFLOPS (FP32 acc.)362 TFLOPS
Tensor BF16 / FP16 + 2:4 sparse330 TFLOPS733 TFLOPS
Tensor INT8661 TOPS733 TOPS
Tensor FP8330 TFLOPS (FP32 acc.; 661 with FP16 acc.)733 TFLOPS
Tensor FP8 + 2:4 sparse660 TFLOPS1466 TFLOPS

FP8 (E4M3 / E5M2) is exposed on every Ada SKU (sm_89). Ada does not have Hopper's WGMMA — matmul shapes stay at the older 16×8×16 sm_89 family, no warp-group async, and no Transformer Engine for dynamic per-tensor scaling. So an L40S running FP8 inference still under-utilises its tensor cores compared to an H100 doing the same FLOPs — but it does so at much lower cost.

05

3rd-Generation RT Cores — OMM & DMM

Generation count: Turing = 1st, Ampere = 2nd, Ada = 3rd, Blackwell = 4th. Two new fixed-function blocks added to the Ada RT core:

Opacity Micromaps (OMM)

Encode the alpha-mask of foliage / hair / mesh triangles inside the BVH. RT core skips fully-transparent sub-triangles without invoking any-hit shaders. ~2× ray-throughput on alpha-heavy scenes.

Displaced Micro-Meshes (DMM)

Add a triangle + a displacement map; the RT core synthesises millions of micro-triangles in hardware without the BVH bloating. Crucial for cinematic content; not relevant to most AI workloads.

RT throughput on RTX 4090 ≈ 191 RT-TFLOPS — 2.8× RTX 3090 Ti's. Also added: Shader Execution Reordering (SER), software-controlled coherence to reduce divergent shader execution after BVH hits.

06

DLSS 3 & the Optical Flow Accelerator

Ada introduces the OFA — Optical Flow Accelerator, a fixed-function unit that computes per-pixel motion vectors between two rendered frames at ≈ 305 TOPS on AD102. DLSS 3 uses the OFA + a small CNN running on tensor cores to synthesise an entirely new intermediate frame in < 4 ms, doubling effective frame rate at modest latency cost. DLSS 3.5 (2023) added Ray Reconstruction — an NN denoiser running on tensor cores that replaces hand-tuned ray-tracing denoisers.

The OFA exists on Ampere too but is roughly 2.5× weaker; only Ada is fast enough for real-time frame generation.

07

96 MB L2 — Why This Big?

96 MB on AD102 is the largest L2 ever shipped on a graphics-class die. Reason: ray tracing has terrible memory locality. BVH traversal jumps across the data structure pointer-chasing; every miss to GDDR6X is ~400 cycles. A huge L2 caches the BVH and most texture working sets.

For LLM workloads the 96 MB L2 helps less — weights are bigger than L2 by orders of magnitude. But KV-cache lines for the most-recent tokens do cache well, and embedding lookups (vector search, recommendation) benefit hugely.

DieL2Comment
AD10296 MBFull L2; cut to 72 MB on RTX 4090.
AD10364 MBRTX 4080.
AD10448 MBRTX 4070 Ti / L4.
AD10632 MBRTX 4060 Ti.
AD10732 MBRTX 4060.
08

Process & Voltages — TSMC 4N

All Ada dies on TSMC 4N — the same NVIDIA-customised N4 variant used by Hopper. Density numbers:

DieProcessTransAreaDensity (M/mm²)
AD102TSMC 4N76.3 B608 mm²125.5
AD103TSMC 4N45.9 B378 mm²121.4
AD104TSMC 4N35.8 B295 mm²121.4
AD106TSMC 4N22.9 B190 mm²120.5
AD107TSMC 4N18.9 B146 mm²129.5

Core voltage range 0.85–1.10 V depending on boost. Ada's V/F curve is much aggressive than Ampere's; sustained boost clocks of 2520–2850 MHz are routine on RTX 4090 with adequate cooling.

09

Clocks, Power, the 12VHPWR Saga

SKUBaseBoostMemoryBWTDP
RTX 40902235 MHz2520 MHz24 GB GDDR6X 21 Gbps1008 GB/s450 W
RTX 4080 Super2295 MHz2550 MHz16 GB GDDR6X 23 Gbps736 GB/s320 W
RTX 40802205 MHz2505 MHz16 GB GDDR6X 22.4 Gbps716 GB/s320 W
RTX 4070 Ti Super2340 MHz2610 MHz16 GB GDDR6X 21 Gbps672 GB/s285 W
RTX 40701920 MHz2475 MHz12 GB GDDR6X 21 Gbps504 GB/s200 W
RTX 6000 Ada915 MHz2505 MHz48 GB GDDR6 ECC 20 Gbps960 GB/s300 W
L40S1110 MHz2520 MHz48 GB GDDR6 ECC 18 Gbps864 GB/s350 W
L40735 MHz2490 MHz48 GB GDDR6 ECC 18 Gbps864 GB/s300 W
L4795 MHz2040 MHz24 GB GDDR6 ECC 12.5 Gbps300 GB/s72 W
12VHPWR → 12V-2×6

The original 16-pin 12VHPWR connector on the RTX 4090 carries 600 W. Several user reports of melted connectors in 2022-23 prompted the PCIe 5.x revision to 12V-2×6: shortened sense pins so a partially-seated connector simply doesn't power on. Drop-in compatible with the same cable.

10

Memory — GDDR6X Up to 22.4 Gbps

Ada uses GDDR6X (PAM4) on consumer/workstation flagship cards and standard GDDR6 (NRZ) on inference cards (L40 / L40S / L4). GDDR7 didn't arrive until Blackwell.

GDDR6X variants

  • RTX 4090: 21 Gbps PAM4 → 1008 GB/s on 384-bit bus
  • RTX 4080 Super: 23 Gbps PAM4 → 736 GB/s on 256-bit
  • RTX 4080: 22.4 Gbps PAM4 → 716 GB/s on 256-bit
  • No ECC on consumer GDDR6X SKUs

GDDR6 ECC variants

  • L40S: 18 Gbps NRZ → 864 GB/s on 384-bit (48 GB ECC)
  • RTX 6000 Ada: 20 Gbps → 960 GB/s on 384-bit (48 GB ECC)
  • L4: 12.5 Gbps → 300 GB/s on 192-bit (24 GB ECC)
  • Side-band ECC, ~6% capacity overhead
11

No NVLink & What That Costs You

Ada was the first generation since Pascal where the consumer flagship has no NVLink at all. RTX 3090 had a 4-link bridge; RTX 4090 has none. Workstation Ada (RTX 6000 Ada) has no NVLink either; L40 / L40S / L4 datacenter cards have no NVLink — multi-GPU on those goes over PCIe 4 P2P only.

NVIDIA's stated rationale for removing NVLink from consumer cards: most gamers never used it, the connector cost was non-trivial, and segmenting NVLink for workstation/datacenter is a deliberate product-tier strategy.

12

Interactive: Ada SKU Picker

Die
—
SMs
—
FP32 TF
—
BF16 TC TF
—
FP8 TC TF
—
VRAM
—
BW (GB/s)
—
TDP / NVLink
—