A low-level look at the 2022 architecture that put a 96 MB L2 cache on a graphics die, gave us 3rd-gen RT cores with opacity and displaced micro-meshes, and added an FP8 datapath to 4th-gen tensor cores across the whole Ada line. Block diagrams, arithmetic, TSMC 4N process, voltages, clocks, GDDR6X memory, no-NVLink consequences.
Ada Lovelace launched in October 2022 as a graphics/workstation/inference architecture only. There is no datacenter HPC die in this generation — that workload is served by Hopper (GH100), launched in parallel on the same TSMC 4N node. AD102 to AD107 cover RTX 40 series consumer, RTX 6000 Ada / RTX 5000 Ada workstation, and L40 / L40S / L4 datacenter inference.
| Die | SMs | Tensor cores | RT cores | Top SKU | TDP |
|---|---|---|---|---|---|
| AD102 | 144 (128 enabled on 4090) | 576 (512 on 4090) | 144 (128 on 4090) | RTX 4090 / RTX 6000 Ada / L40S | 300–450 W |
| AD103 | 80 | 320 | 80 | RTX 4080 Super / RTX 4080 | 320 W |
| AD104 | 60 | 240 | 60 | RTX 4070 Ti / 4070 / L4 | 72–285 W |
| AD106 | 36 | 144 | 36 | RTX 4060 Ti | 160–165 W |
| AD107 | 24 | 96 | 24 | RTX 4060 / 4050 mobile | 115 W |
Density: 76.3 B / 608 mm² = 125.5 M/mm² — nearly 2× the GA100 density and 2.8× the Samsung 8N GA102. The TSMC 4N node was a huge generational lift, especially for AD102 which managed to fit ~1.7× the SMs of GA102 (144 vs 84) on 3% less silicon.
The Ada SM inherits GA10x's "doubled FP32" partition layout with new RT and tensor cores. FP8 (E4M3 / E5M2) is exposed in tensor cores across the entire Ada line — including the RTX 4090. What consumer Ada lacks vs Hopper is the Transformer Engine's dynamic per-tensor scaling, not the FP8 silicon itself.
Tensor cores on Ada are physically capable of FP8 E4M3 / E5M2 on every Ada SKU (sm_89), RTX 4090 included. What Ada does not have is Hopper's Transformer Engine — the dynamic per-tensor scaling that automatically picks E4M3 vs E5M2 and tracks amax statistics. On Ada you do FP8 with explicit static scaling (e.g. via TransformerEngine library, TensorRT-LLM). Datacenter Ada (L40 / L40S / RTX 6000 Ada) markets the FP8 throughput; consumer Ada exposes the same instructions, just without the Transformer Engine's scheduling layer.
| Operation | RTX 4090 (2520 MHz) | L40S (2520 MHz) |
|---|---|---|
| FP32 FMA | 82.6 TFLOPS | 91.6 TFLOPS |
| Tensor TF32 | 82.6 TFLOPS | 183 TFLOPS |
| Tensor BF16 / FP16 | 165 TFLOPS (FP32 acc.) | 362 TFLOPS |
| Tensor BF16 / FP16 + 2:4 sparse | 330 TFLOPS | 733 TFLOPS |
| Tensor INT8 | 661 TOPS | 733 TOPS |
| Tensor FP8 | 330 TFLOPS (FP32 acc.; 661 with FP16 acc.) | 733 TFLOPS |
| Tensor FP8 + 2:4 sparse | 660 TFLOPS | 1466 TFLOPS |
FP8 (E4M3 / E5M2) is exposed on every Ada SKU (sm_89). Ada does not have Hopper's WGMMA — matmul shapes stay at the older 16×8×16 sm_89 family, no warp-group async, and no Transformer Engine for dynamic per-tensor scaling. So an L40S running FP8 inference still under-utilises its tensor cores compared to an H100 doing the same FLOPs — but it does so at much lower cost.
Generation count: Turing = 1st, Ampere = 2nd, Ada = 3rd, Blackwell = 4th. Two new fixed-function blocks added to the Ada RT core:
Encode the alpha-mask of foliage / hair / mesh triangles inside the BVH. RT core skips fully-transparent sub-triangles without invoking any-hit shaders. ~2× ray-throughput on alpha-heavy scenes.
Add a triangle + a displacement map; the RT core synthesises millions of micro-triangles in hardware without the BVH bloating. Crucial for cinematic content; not relevant to most AI workloads.
RT throughput on RTX 4090 ≈ 191 RT-TFLOPS — 2.8× RTX 3090 Ti's. Also added: Shader Execution Reordering (SER), software-controlled coherence to reduce divergent shader execution after BVH hits.
Ada introduces the OFA — Optical Flow Accelerator, a fixed-function unit that computes per-pixel motion vectors between two rendered frames at ≈ 305 TOPS on AD102. DLSS 3 uses the OFA + a small CNN running on tensor cores to synthesise an entirely new intermediate frame in < 4 ms, doubling effective frame rate at modest latency cost. DLSS 3.5 (2023) added Ray Reconstruction — an NN denoiser running on tensor cores that replaces hand-tuned ray-tracing denoisers.
The OFA exists on Ampere too but is roughly 2.5× weaker; only Ada is fast enough for real-time frame generation.
96 MB on AD102 is the largest L2 ever shipped on a graphics-class die. Reason: ray tracing has terrible memory locality. BVH traversal jumps across the data structure pointer-chasing; every miss to GDDR6X is ~400 cycles. A huge L2 caches the BVH and most texture working sets.
For LLM workloads the 96 MB L2 helps less — weights are bigger than L2 by orders of magnitude. But KV-cache lines for the most-recent tokens do cache well, and embedding lookups (vector search, recommendation) benefit hugely.
| Die | L2 | Comment |
|---|---|---|
| AD102 | 96 MB | Full L2; cut to 72 MB on RTX 4090. |
| AD103 | 64 MB | RTX 4080. |
| AD104 | 48 MB | RTX 4070 Ti / L4. |
| AD106 | 32 MB | RTX 4060 Ti. |
| AD107 | 32 MB | RTX 4060. |
All Ada dies on TSMC 4N — the same NVIDIA-customised N4 variant used by Hopper. Density numbers:
| Die | Process | Trans | Area | Density (M/mm²) |
|---|---|---|---|---|
| AD102 | TSMC 4N | 76.3 B | 608 mm² | 125.5 |
| AD103 | TSMC 4N | 45.9 B | 378 mm² | 121.4 |
| AD104 | TSMC 4N | 35.8 B | 295 mm² | 121.4 |
| AD106 | TSMC 4N | 22.9 B | 190 mm² | 120.5 |
| AD107 | TSMC 4N | 18.9 B | 146 mm² | 129.5 |
Core voltage range 0.85–1.10 V depending on boost. Ada's V/F curve is much aggressive than Ampere's; sustained boost clocks of 2520–2850 MHz are routine on RTX 4090 with adequate cooling.
| SKU | Base | Boost | Memory | BW | TDP |
|---|---|---|---|---|---|
| RTX 4090 | 2235 MHz | 2520 MHz | 24 GB GDDR6X 21 Gbps | 1008 GB/s | 450 W |
| RTX 4080 Super | 2295 MHz | 2550 MHz | 16 GB GDDR6X 23 Gbps | 736 GB/s | 320 W |
| RTX 4080 | 2205 MHz | 2505 MHz | 16 GB GDDR6X 22.4 Gbps | 716 GB/s | 320 W |
| RTX 4070 Ti Super | 2340 MHz | 2610 MHz | 16 GB GDDR6X 21 Gbps | 672 GB/s | 285 W |
| RTX 4070 | 1920 MHz | 2475 MHz | 12 GB GDDR6X 21 Gbps | 504 GB/s | 200 W |
| RTX 6000 Ada | 915 MHz | 2505 MHz | 48 GB GDDR6 ECC 20 Gbps | 960 GB/s | 300 W |
| L40S | 1110 MHz | 2520 MHz | 48 GB GDDR6 ECC 18 Gbps | 864 GB/s | 350 W |
| L40 | 735 MHz | 2490 MHz | 48 GB GDDR6 ECC 18 Gbps | 864 GB/s | 300 W |
| L4 | 795 MHz | 2040 MHz | 24 GB GDDR6 ECC 12.5 Gbps | 300 GB/s | 72 W |
The original 16-pin 12VHPWR connector on the RTX 4090 carries 600 W. Several user reports of melted connectors in 2022-23 prompted the PCIe 5.x revision to 12V-2×6: shortened sense pins so a partially-seated connector simply doesn't power on. Drop-in compatible with the same cable.
Ada uses GDDR6X (PAM4) on consumer/workstation flagship cards and standard GDDR6 (NRZ) on inference cards (L40 / L40S / L4). GDDR7 didn't arrive until Blackwell.
Ada was the first generation since Pascal where the consumer flagship has no NVLink at all. RTX 3090 had a 4-link bridge; RTX 4090 has none. Workstation Ada (RTX 6000 Ada) has no NVLink either; L40 / L40S / L4 datacenter cards have no NVLink — multi-GPU on those goes over PCIe 4 P2P only.
NVIDIA's stated rationale for removing NVLink from consumer cards: most gamers never used it, the connector cost was non-trivial, and segmenting NVLink for workstation/datacenter is a deliberate product-tier strategy.