The 2022 architecture that gave gamers DLSS 3, gave workstation users the L40S as a cheap FP8 server card, and gave hobbyist AI a 24 GB RTX 4090 that runs Llama-class models locally — all on the same TSMC 4N process as the data-centre H100.
Ada Lovelace is two architectures wearing one label: a gaming flagship with new ray-tracing hardware and DLSS 3 frame generation, and a quietly aggressive datacenter inference card called the L40S. This deck walks both narratives.
Announced September 2022, shipped October 2022. The architecture name honours Ada Lovelace, the 19th-century English mathematician widely credited as the first programmer. Built on TSMC 4N — the same custom node as Hopper — it lifted transistor density enough to put 76 billion transistors on the flagship die.
| Property | AD102 (top die) |
|---|---|
| Process | TSMC 4N (custom 5 nm class) |
| Transistors | 76.3 billion |
| Die size | 608 mm² |
| SMs (full die) | 144 |
| SMs (RTX 4090) | 128 enabled |
| SMs (RTX 6000 Ada / L40S) | 142 enabled |
| L2 cache | 96 MB (16× the RTX 3090) |
| Memory bus | 384-bit GDDR6X (consumer) / GDDR6 ECC (datacenter) |
| NVLink | None on any Ada SKU (consumer, RTX 6000 Ada or L-series) |
| Released | October 2022 |
RTX 40-series, headlined by the RTX 4090. The story is 3rd-generation RT cores, DLSS 3 with the new Optical Flow Accelerator, GDDR6X, and a giant 96 MB L2 designed for ray-tracing locality. FP8 tensor cores are present, but with no Transformer Engine integration on consumer.
L40S, L40, L4 and RTX 6000 Ada. Same silicon, ECC GDDR6, datacenter EULA-compliant cooling, and the same 4th-gen tensor cores with FP8. The L40S became NVIDIA's cheap FP8 inference card — less than half an H100's price for a meaningful slice of its inference throughput.
Ada is the first architecture where a hobbyist could plausibly run a Llama-class model locally at decent speed (24 GB RTX 4090, GDDR6X ≈ 1008 GB/s) and where a small company could stand up FP8 inference on commodity-priced datacenter cards (L40S). The same fabric, sold to two completely different markets — with VRAM capacity, ECC, datacenter licensing, and Transformer Engine support drawing the line between buckets.
AD102 is the largest Ada die. Two views — compute and memory — explain almost everything you need to know about the 4090 and the L40S.
Going from 6 MB on Ampere to 96 MB on Ada was the single biggest architectural change. With the previous L2, ray-tracing kernels and large fused-attention kernels were dominated by L2 misses out to GDDR. With 96 MB, an entire BVH stack and a working slice of the attention KV-cache can sit on-die, cutting effective bandwidth pressure dramatically.
The Ada SM is structurally similar to GA10x (consumer Ampere) — same four-partition layout, same doubled-FP32 trick — but with newer tensor cores, newer RT core, and a much bigger L2 behind it.
Like GA10x before it, Ada inherited the trick where one of the two 32-wide datapaths per partition can do FP32 or INT32 each cycle. Mixed integer-heavy workloads run at half rate; FP-heavy graphics workloads run at the headline 128 FP32 per SM. Pure ML matmul lives in the tensor cores, so this matters mainly for graphics shaders and CUDA scalar code.
Ada's RT cores are the third generation (after Turing's first and Ampere's second). The headline number is roughly 2× ray-triangle throughput vs Ampere per SM, but the more interesting changes are two new BVH primitives that change what the RT core processes, not just how fast.
Foliage, hair, fences, chain-link — geometry that is mostly transparent. Pre-Ada, every ray hitting the bounding triangle had to invoke the alpha-test shader to decide whether the hit was real.
OMMs encode the alpha mask of each triangle directly inside the BVH, at sub-triangle resolution. The RT core can resolve "ray missed the opaque part" without ever calling a shader — eliminating millions of redundant shader invocations per frame.
A single triangle plus a displacement map encodes millions of micro-triangles that the RT core can ray-trace directly — without the BVH containing every micro-triangle explicitly.
This is a 10×+ memory reduction for highly detailed geometry (terrain, bricks, foliage), and it lets the RT core trace much higher-poly scenes than the BVH could otherwise hold in VRAM.
A separate, important addition. After divergent ray hits, the warp partitions out into many different shader paths — murderous for SIMT efficiency. SER is a software-controlled hint that lets the application reorder threads inside a warp so that threads taking the same shader path are grouped together. Effectively a programmable coherence sort. Big wins in path-traced engines (Cyberpunk's Overdrive mode reported up to 25% from SER alone).
DLSS 3 is the marquee Ada feature on the consumer side, and it is fundamentally different from DLSS 2. DLSS 2 was super-sampling: render at low resolution, upscale with a tensor-core neural network. DLSS 3 adds a second mechanism on top — Frame Generation — that synthesises wholly new in-between frames.
A dedicated hardware unit on Ada that computes optical flow vectors between two real rendered frames — per-pixel motion estimates including for objects without motion vectors (transparent particles, shadows, UI).
OFA exists on Ampere too, but Ada's OFA is roughly 2.5× faster — enough to run inside a single frame budget at 60+ Hz.
Frame Generation buys frame-rate, not responsiveness. Because the synthesised frame is between two real frames, the engine must hold N+1 briefly before showing it — net latency is similar to not using FG at the lower base rate. DLSS 3 is paired with NVIDIA Reflex (low-latency mode) to claw back some of that. Practical impact: great for cinematic / single-player, less attractive for competitive shooters.
Released 2023. Replaces hand-tuned ray-tracing denoisers (which historically blurred reflections and crawled at object edges) with a tensor-core neural network trained to reconstruct the noisy ray-traced image directly. Requires a path-traced renderer to shine; Cyberpunk Overdrive and Alan Wake 2 were the canonical demonstrations.
NVIDIA gates Frame Generation to Ada because of OFA throughput. Ampere has the units but not the speed. Whether that is genuinely a hardware limit or a marketing line has been debated; benchmarks show Ampere's OFA is meaningfully slower, but a software-only fallback would not be impossible. As of 2026 the gate remains.
The full launched and "Super" stack as it stood through 2024–2025, before Blackwell consumer cards arrived. The 4090 dominates absolute performance; the 4070 Ti Super and 4080 Super occupy the mid-tier; the 4060 / 4060 Ti are the budget end with notably narrow memory buses.
| SKU | CUDA cores | VRAM | BW (GB/s) | TDP | NVLink |
|---|---|---|---|---|---|
| RTX 4090 | 16 384 | 24 GB GDDR6X | 1008 | 450 W | none |
| RTX 4080 Super | 10 240 | 16 GB GDDR6X | 736 | 320 W | none |
| RTX 4070 Ti Super | 8 448 | 16 GB GDDR6X | 672 | 285 W | none |
| RTX 4070 | 5 888 | 12 GB GDDR6X | 504 | 200 W | none |
| RTX 4060 | 3 072 | 8 GB GDDR6 | 272 | 115 W | none |
The bridge connector was physically removed from the PCB. Multi-GPU with two RTX 4090s is therefore PCIe-only — PCIe 4 x16 P2P at ~32 GB/s vs ~112 GB/s of NVLink 3 on the previous 3090 / A6000 bridge. This is the single biggest blocker for tensor-parallel inference on consumer Ada and the reason the L40S and RTX 6000 Ada exist as separate products.
For local AI hobbyists in 2023–2024 the 4090 was extraordinary not because of raw FLOPS (the 3090 was already enough for many tasks) but because 24 GB on a single consumer GPU at 1008 GB/s was something no other card at that price point offered. It put 7–8B models comfortably into BF16 and 30B-class models into INT4 reach (13B BF16 is 26 GB and 70B INT4 ~40 GB, so both need two cards).
The L40S is the most under-appreciated card NVIDIA shipped during the Ada generation, and it became one of the most popular datacenter GPUs in 2024–2026 for cost-optimised LLM serving where bandwidth is not the bottleneck.
| Property | L40S |
|---|---|
| Die | AD102 (142 of 144 SMs enabled) |
| CUDA cores | 18 176 |
| Tensor cores | 568 (4th-gen) |
| RT cores | 142 (3rd-gen) |
| FP8 tensor path | supported (as on all Ada); ECC + datacenter licence are the real differentiators |
| VRAM | 48 GB GDDR6 with ECC |
| Bandwidth | ~864 GB/s |
| TDP | 350 W, passive cooling, datacenter form factor |
| NVLink | none |
| Datacenter EULA | compliant (unlike RTX 4090) |
Compute-bound prefill of long prompts; small-model batched inference where you stream concurrent requests; vector embedding services; vision encoders. Anywhere bandwidth-per-token is not the limit, the L40S buys you most of an H100's behaviour for half the money. Where it loses: long-context single-stream decode (bandwidth-bound), training (FP32 / FP64 paths much thinner), multi-GPU TP (no NVLink).
Three more workstation/datacenter Ada SKUs round out the family. They share the same AD102 / AD104 silicon but target very different deployment envelopes.
| SKU | VRAM | BW (GB/s) | TDP | FP8 | NVLink | Niche |
|---|---|---|---|---|---|---|
| L40 | 48 GB GDDR6 ECC | ~864 | 300 W | yes | none | Pre-L40S graphics+inference card; superseded by L40S on perf. |
| L40S | 48 GB GDDR6 ECC | ~864 | 350 W | yes | none | Cheap FP8 inference (slide 07). |
| L4 | 24 GB GDDR6 | 300 | 72 W single-slot | yes | none | Video transcoding + edge inference. |
| RTX 6000 Ada | 48 GB GDDR6 ECC | 960 | 300 W | yes | none | Workstation flagship; ECC, pro drivers. |
The original "graphics + inference" Ada datacenter card. 48 GB ECC, 300 W. Once the L40S (same form factor + FP8) launched, the L40 was effectively superseded for inference workloads and is now found mostly in deployed VDI / virtual workstation farms.
72 W single-slot low-profile card. Designed for video transcoding farms (NVENC / NVDEC) and edge inference. 24 GB GDDR6, 300 GB/s — bandwidth is modest but the power and form-factor envelope are unmatched. Drops into 1U servers and dense edge appliances.
The workstation flagship. AD102 with 142 SMs (18,176 CUDA cores), 48 GB GDDR6 ECC, ~960 GB/s, 300 W, blower cooler. No NVLink (RTX 6000 Ada datasheet: “NVLink: No”), so two-card setups run over PCIe like every other Ada part.
Both use AD102 with 142 of 144 SMs enabled and 4th-gen tensor cores (FP8 supported on both). The L40S has higher clocks, a higher 350 W TDP, and is tuned for AI throughput; the RTX 6000 Ada targets workstation graphics with pro drivers and active cooling. Buyers in 2024–2025 routinely picked L40S for AI inference and RTX 6000 Ada for content/CAD/sim work.
Of all the design choices in Ada, removing NVLink from the consumer cards had the most far-reaching consequence for the local AI community. The 3090 had an NVLink bridge; every 40-series card lost it. The reason matters and the implications cascade.
The dual gold-fingers along the top edge of the 3090 PCB — physically gone. No Ada consumer card has the connector. No Ada SKU at all — consumer, workstation or datacenter — exposes NVLink.
Plausibly motivated to push pro/datacenter buyers towards the L40S and RTX 6000 Ada rather than two 4090s.
This was the question every serious local-AI builder asked in 2023–2024. The architectures are identical (AD102), so it comes down to memory and link:
| Option | VRAM | BW (GB/s) | TP scaling | Power | Verdict |
|---|---|---|---|---|---|
| 2 × RTX 4090 | 48 GB total (24+24) | 1008 each | ~1.3–1.5× over PCIe 4 | 900 W combined | Higher peak BW, painful TP scaling, no datacenter licence. |
| 1 × RTX 6000 Ada | 48 GB on one card | 960 | n/a (single GPU) | 300 W | Same VRAM in one address space — no link tax, lower power, ECC, datacenter-licensed. |
For any TP-aware workload (vLLM TP=2 of a 70B-class model), the RTX 6000 Ada wins comfortably despite identical underlying silicon. For two-replica DP-style serving (one model per card behind a load balancer), the 4090 pair wins on raw aggregate bandwidth and absolute throughput.
Blackwell partly relented: the RTX PRO 6000 Blackwell puts 96 GB GDDR7 on a single card (still without NVLink), which sidesteps the question for many buyers. Consumer Blackwell (RTX 50-series) reportedly continues to have no NVLink bridge. Deck 09 covers Blackwell.
A frank assessment, two years after launch. Ada is two architectures — the consumer line is a single-GPU local-AI machine, and the L40S is a cheap datacenter inference card. Each excels at different things.
| Model / quant | Hardware | Throughput |
|---|---|---|
| Llama-3-8B BF16 | 1 × RTX 4090 | ~55 tok/s/stream (ceiling 1008 ÷ 16 GB ≈ 63) · batched ~3000 tok/s aggregate |
| Llama-3-70B AWQ-INT4 | 2 × RTX 4090, PCIe TP=2 | ~22 tok/s/stream |
| Llama-3-70B FP8 | 2 × L40S, PCIe TP=2 | ~20 tok/s/stream (ceiling 864 ÷ 35 GB per card ≈ 25) · batched ~600 tok/s |
| Llama-3-8B FP8 | 1 × L40S | ~95 tok/s/stream · batched ~4500 tok/s |
| SDXL image, 30 steps | 1 × RTX 4090 | ~3.5 s per 1024×1024 image |
If your workload is single-stream decode of a model that fits in 24–48 GB, Ada is the best price/performance NVIDIA offers from 2022–2026. If you need to span 70B+ with tensor parallelism, the lack of NVLink hurts and you should look at H100/H200 (Hopper deck) or RTX PRO 6000 Blackwell (Blackwell deck) instead.
Pick an Ada-generation SKU. The panel summarises capability, deployment licence, multi-GPU prospects, and the largest LLM you can comfortably host at FP16 vs INT4.
"Largest LLM" is a rough envelope: FP16 assumes ~70% of VRAM goes to weights and the rest to KV/runtime; INT4 assumes ~80%. In practice context length, batch size, and KV-cache dtype shift these by ~20% either way. Use this to filter out impossible matches, not as a precise spec.