NVIDIA GPU Architectures Series — Presentation 08

Ada Lovelace — RTX 40, L40S, and Consumer-Class AI

The 2022 architecture that gave gamers DLSS 3, gave workstation users the L40S as a cheap FP8 server card, and gave hobbyist AI a 24 GB RTX 4090 that runs Llama-class models locally — all on the same TSMC 4N process as the data-centre H100.

AD102RTX 4090RTX 5090 (no) L40SL4RT cores 3 DLSS 3OFA FP8GDDR6X
AD102 → SM → 3rd-gen RT → 4th-gen TC → DLSS 3 → L40S → No NVLink
00

Topics We'll Cover

Ada Lovelace is two architectures wearing one label: a gaming flagship with new ray-tracing hardware and DLSS 3 frame generation, and a quietly aggressive datacenter inference card called the L40S. This deck walks both narratives.

01

Ada in One Page

Announced September 2022, shipped October 2022. The architecture name honours Ada Lovelace, the 19th-century English mathematician widely credited as the first programmer. Built on TSMC 4N — the same custom node as Hopper — it lifted transistor density enough to put 76 billion transistors on the flagship die.

PropertyAD102 (top die)
ProcessTSMC 4N (custom 5 nm class)
Transistors76.3 billion
Die size608 mm²
SMs (full die)144
SMs (RTX 4090)128 enabled
SMs (RTX 6000 Ada / L40S)142 enabled
L2 cache96 MB (16× the RTX 3090)
Memory bus384-bit GDDR6X (consumer) / GDDR6 ECC (datacenter)
NVLinkNone on any Ada SKU (consumer, RTX 6000 Ada or L-series)
ReleasedOctober 2022

Two parallel narratives in one architecture

Consumer / Gaming

RTX 40-series, headlined by the RTX 4090. The story is 3rd-generation RT cores, DLSS 3 with the new Optical Flow Accelerator, GDDR6X, and a giant 96 MB L2 designed for ray-tracing locality. FP8 tensor cores are present, but with no Transformer Engine integration on consumer.

Workstation / Datacenter

L40S, L40, L4 and RTX 6000 Ada. Same silicon, ECC GDDR6, datacenter EULA-compliant cooling, and the same 4th-gen tensor cores with FP8. The L40S became NVIDIA's cheap FP8 inference card — less than half an H100's price for a meaningful slice of its inference throughput.

Why it matters

Ada is the first architecture where a hobbyist could plausibly run a Llama-class model locally at decent speed (24 GB RTX 4090, GDDR6X ≈ 1008 GB/s) and where a small company could stand up FP8 inference on commodity-priced datacenter cards (L40S). The same fabric, sold to two completely different markets — with VRAM capacity, ECC, datacenter licensing, and Transformer Engine support drawing the line between buckets.

02

AD102 — The Die

AD102 is the largest Ada die. Two views — compute and memory — explain almost everything you need to know about the 4090 and the L40S.

Compute

  • 12 GPCs (Graphics Processing Clusters)
  • 12 SMs per GPC → 144 SMs full die
  • RTX 4090: 128 SMs enabled (16 disabled for yield)
  • RTX 6000 Ada / L40S: 142 SMs enabled (18,176 CUDA cores)
  • 128 CUDA cores per SM → 18 432 CUDA cores on full die; 16 384 on 4090
  • 4 tensor cores per SM → 576 4th-gen tensor cores full die
  • 1 RT core per SM → 144 3rd-gen RT cores full die

Memory

  • 384-bit GDDR6X bus on RTX 4090 (12 channels × 32-bit)
  • 21 Gbps per pin → ~1008 GB/s aggregate bandwidth
  • L40S uses GDDR6 ECC → ~864 GB/s at 48 GB
  • 96 MB L2 cache — 16× the 6 MB on the RTX 3090, designed for ray-tracing access locality and large attention windows
  • No HBM on any Ada SKU — pure GDDR family
  • No NVLink on any 40-series consumer card, nor on RTX 6000 Ada or the L-series
AD102 floorplan — conceptual GPC0 (12 SMs) GPC1 (12 SMs) GPC2 (12 SMs) GPC3 (12 SMs) L2 slice 24 MB L2 slice 24 MB L2 slice 24 MB L2 slice 24 MB GPC4..7 + GPC8..11 GDDR6X 384-bit, 21 Gbps → 1008 GB/s 96 MB L2 total
The L2 explosion

Going from 6 MB on Ampere to 96 MB on Ada was the single biggest architectural change. With the previous L2, ray-tracing kernels and large fused-attention kernels were dominated by L2 misses out to GDDR. With 96 MB, an entire BVH stack and a working slice of the attention KV-cache can sit on-die, cutting effective bandwidth pressure dramatically.

03

Ada SM — Inside a Streaming Multiprocessor

The Ada SM is structurally similar to GA10x (consumer Ampere) — same four-partition layout, same doubled-FP32 trick — but with newer tensor cores, newer RT core, and a much bigger L2 behind it.

Per-SM resources

  • 4 partitions per SM
  • Per partition: 32 FP32 cores (consumer-doubled, like GA10x)
  • Per partition: 16 INT32 cores (FP32-or-INT32 dual-issue)
  • Per partition: 1 4th-gen tensor core
  • Per SM: 1 3rd-gen RT core
  • Per partition: 4 LD/ST units, 1 SFU (special-function)
  • 128 KB L1 / shared memory per SM (configurable split)

Tensor core nuances

  • Per-SM FP16/BF16 tensor throughput is about a quarter of a Hopper SM's per clock (1024 vs 4096 FLOP/clk dense)
  • Ada has more SMs on the full die than H100 (144 vs 132), but H100 wins absolute on bandwidth, NVLink, and Transformer Engine
  • FP8 tensor cores are present on all Ada SKUs (RTX 4090 included) — same E4M3 / E5M2 formats as Hopper
  • What's missing on consumer Ada is the Transformer Engine (auto FP8 scaling) integration that Hopper has
  • Sparsity (2:4) supported on all Ada SKUs
Ada SM — 4 partitions + shared L1 + 1 RT core Partition 0 32 FP32 16 INT32 4th-gen Tensor 4 LD/ST + 1 SFU warp sched + reg file Partition 1 32 FP32 16 INT32 4th-gen Tensor Partition 2 32 FP32 16 INT32 4th-gen Tensor Partition 3 32 FP32 16 INT32 4th-gen Tensor 128 KB L1 / shared (configurable split) 1 3rd-gen RT core
Practical: doubled FP32, halved INT32

Like GA10x before it, Ada inherited the trick where one of the two 32-wide datapaths per partition can do FP32 or INT32 each cycle. Mixed integer-heavy workloads run at half rate; FP-heavy graphics workloads run at the headline 128 FP32 per SM. Pure ML matmul lives in the tensor cores, so this matters mainly for graphics shaders and CUDA scalar code.

04

3rd-Gen RT Cores — OMM, DMM, SER

Ada's RT cores are the third generation (after Turing's first and Ampere's second). The headline number is roughly 2× ray-triangle throughput vs Ampere per SM, but the more interesting changes are two new BVH primitives that change what the RT core processes, not just how fast.

(a) Opacity Micromaps (OMM)

Foliage, hair, fences, chain-link — geometry that is mostly transparent. Pre-Ada, every ray hitting the bounding triangle had to invoke the alpha-test shader to decide whether the hit was real.

OMMs encode the alpha mask of each triangle directly inside the BVH, at sub-triangle resolution. The RT core can resolve "ray missed the opaque part" without ever calling a shader — eliminating millions of redundant shader invocations per frame.

(b) Displaced Micro-Meshes (DMM)

A single triangle plus a displacement map encodes millions of micro-triangles that the RT core can ray-trace directly — without the BVH containing every micro-triangle explicitly.

This is a 10×+ memory reduction for highly detailed geometry (terrain, bricks, foliage), and it lets the RT core trace much higher-poly scenes than the BVH could otherwise hold in VRAM.

SER — Shader Execution Reordering

A separate, important addition. After divergent ray hits, the warp partitions out into many different shader paths — murderous for SIMT efficiency. SER is a software-controlled hint that lets the application reorder threads inside a warp so that threads taking the same shader path are grouped together. Effectively a programmable coherence sort. Big wins in path-traced engines (Cyberpunk's Overdrive mode reported up to 25% from SER alone).

RT throughput — relative, per SM (rays/s, normalised) Turing (1st) 1.0× Ampere (2nd) ~2× Ada (3rd) ~4× (incl. OMM/DMM) +SER (Cyberpunk OD) ~5× effective Note: real games depend heavily on BVH layout and shader divergence.
05

DLSS 3 + Frame Generation

DLSS 3 is the marquee Ada feature on the consumer side, and it is fundamentally different from DLSS 2. DLSS 2 was super-sampling: render at low resolution, upscale with a tensor-core neural network. DLSS 3 adds a second mechanism on top — Frame Generation — that synthesises wholly new in-between frames.

The Optical Flow Accelerator (OFA)

A dedicated hardware unit on Ada that computes optical flow vectors between two real rendered frames — per-pixel motion estimates including for objects without motion vectors (transparent particles, shadows, UI).

OFA exists on Ampere too, but Ada's OFA is roughly 2.5× faster — enough to run inside a single frame budget at 60+ Hz.

Frame Generation pipeline

  • Render frame N and N+1 normally (with DLSS 2 upscaling)
  • OFA produces a flow field N→N+1
  • A small NN on the tensor cores fuses flow + game motion vectors + depth + colour to synthesise an intermediate frame
  • Display sequence: N → synthesised → N+1 → ...
  • Effective frame rate doubles; rendered work nearly the same

The latency tradeoff

Frame Generation buys frame-rate, not responsiveness. Because the synthesised frame is between two real frames, the engine must hold N+1 briefly before showing it — net latency is similar to not using FG at the lower base rate. DLSS 3 is paired with NVIDIA Reflex (low-latency mode) to claw back some of that. Practical impact: great for cinematic / single-player, less attractive for competitive shooters.

DLSS 3.5 — Ray Reconstruction

Released 2023. Replaces hand-tuned ray-tracing denoisers (which historically blurred reflections and crawled at object edges) with a tensor-core neural network trained to reconstruct the noisy ray-traced image directly. Requires a path-traced renderer to shine; Cyberpunk Overdrive and Alan Wake 2 were the canonical demonstrations.

DLSS 3 is Ada-only by feature gate, not capability

NVIDIA gates Frame Generation to Ada because of OFA throughput. Ampere has the units but not the speed. Whether that is genuinely a hardware limit or a marketing line has been debated; benchmarks show Ampere's OFA is meaningfully slower, but a software-only fallback would not be impossible. As of 2026 the gate remains.

06

RTX 40 Consumer Family

The full launched and "Super" stack as it stood through 2024–2025, before Blackwell consumer cards arrived. The 4090 dominates absolute performance; the 4070 Ti Super and 4080 Super occupy the mid-tier; the 4060 / 4060 Ti are the budget end with notably narrow memory buses.

SKUCUDA coresVRAMBW (GB/s)TDPNVLink
RTX 409016 38424 GB GDDR6X1008450 Wnone
RTX 4080 Super10 24016 GB GDDR6X736320 Wnone
RTX 4070 Ti Super8 44816 GB GDDR6X672285 Wnone
RTX 40705 88812 GB GDDR6X504200 Wnone
RTX 40603 0728 GB GDDR6272115 Wnone
No NVLink on any RTX 40

The bridge connector was physically removed from the PCB. Multi-GPU with two RTX 4090s is therefore PCIe-only — PCIe 4 x16 P2P at ~32 GB/s vs ~112 GB/s of NVLink 3 on the previous 3090 / A6000 bridge. This is the single biggest blocker for tensor-parallel inference on consumer Ada and the reason the L40S and RTX 6000 Ada exist as separate products.

The 4090 specifically

VRAM was the killer feature

For local AI hobbyists in 2023–2024 the 4090 was extraordinary not because of raw FLOPS (the 3090 was already enough for many tasks) but because 24 GB on a single consumer GPU at 1008 GB/s was something no other card at that price point offered. It put 7–8B models comfortably into BF16 and 30B-class models into INT4 reach (13B BF16 is 26 GB and 70B INT4 ~40 GB, so both need two cards).

07

L40S — The Sleeper Inference Card

The L40S is the most under-appreciated card NVIDIA shipped during the Ada generation, and it became one of the most popular datacenter GPUs in 2024–2026 for cost-optimised LLM serving where bandwidth is not the bottleneck.

PropertyL40S
DieAD102 (142 of 144 SMs enabled)
CUDA cores18 176
Tensor cores568 (4th-gen)
RT cores142 (3rd-gen)
FP8 tensor pathsupported (as on all Ada); ECC + datacenter licence are the real differentiators
VRAM48 GB GDDR6 with ECC
Bandwidth~864 GB/s
TDP350 W, passive cooling, datacenter form factor
NVLinknone
Datacenter EULAcompliant (unlike RTX 4090)

Why it became the sleeper hit

Where the L40S beats the H100 on cost-per-token

Compute-bound prefill of long prompts; small-model batched inference where you stream concurrent requests; vector embedding services; vision encoders. Anywhere bandwidth-per-token is not the limit, the L40S buys you most of an H100's behaviour for half the money. Where it loses: long-context single-stream decode (bandwidth-bound), training (FP32 / FP64 paths much thinner), multi-GPU TP (no NVLink).

08

L40 / L4 / RTX 6000 Ada

Three more workstation/datacenter Ada SKUs round out the family. They share the same AD102 / AD104 silicon but target very different deployment envelopes.

SKUVRAMBW (GB/s)TDPFP8NVLinkNiche
L4048 GB GDDR6 ECC~864300 WyesnonePre-L40S graphics+inference card; superseded by L40S on perf.
L40S48 GB GDDR6 ECC~864350 WyesnoneCheap FP8 inference (slide 07).
L424 GB GDDR630072 W single-slotyesnoneVideo transcoding + edge inference.
RTX 6000 Ada48 GB GDDR6 ECC960300 WyesnoneWorkstation flagship; ECC, pro drivers.

Each card's actual job

L40

The original "graphics + inference" Ada datacenter card. 48 GB ECC, 300 W. Once the L40S (same form factor + FP8) launched, the L40 was effectively superseded for inference workloads and is now found mostly in deployed VDI / virtual workstation farms.

L4

72 W single-slot low-profile card. Designed for video transcoding farms (NVENC / NVDEC) and edge inference. 24 GB GDDR6, 300 GB/s — bandwidth is modest but the power and form-factor envelope are unmatched. Drops into 1U servers and dense edge appliances.

RTX 6000 Ada

The workstation flagship. AD102 with 142 SMs (18,176 CUDA cores), 48 GB GDDR6 ECC, ~960 GB/s, 300 W, blower cooler. No NVLink (RTX 6000 Ada datasheet: “NVLink: No”), so two-card setups run over PCIe like every other Ada part.

L40S vs RTX 6000 Ada — same silicon, different positioning

Both use AD102 with 142 of 144 SMs enabled and 4th-gen tensor cores (FP8 supported on both). The L40S has higher clocks, a higher 350 W TDP, and is tuned for AI throughput; the RTX 6000 Ada targets workstation graphics with pro drivers and active cooling. Buyers in 2024–2025 routinely picked L40S for AI inference and RTX 6000 Ada for content/CAD/sim work.

09

No NVLink — A Pivotal Decision

Of all the design choices in Ada, removing NVLink from the consumer cards had the most far-reaching consequence for the local AI community. The 3090 had an NVLink bridge; every 40-series card lost it. The reason matters and the implications cascade.

What NVIDIA removed

The dual gold-fingers along the top edge of the 3090 PCB — physically gone. No Ada consumer card has the connector. No Ada SKU at all — consumer, workstation or datacenter — exposes NVLink.

Plausibly motivated to push pro/datacenter buyers towards the L40S and RTX 6000 Ada rather than two 4090s.

The throughput cost on TP=2

  • NVLink 3 bridge on 3090/A6000: ~112 GB/s aggregate, two-card
  • PCIe 4 x16 P2P on 4090: ~32 GB/s peak, often less in practice (root-complex topology dependent)
  • TP=2 scaling on Llama-70B: ~1.9× with NVLink, ~1.3–1.5× with PCIe 4 P2P alone

Two RTX 4090s vs one RTX 6000 Ada

This was the question every serious local-AI builder asked in 2023–2024. The architectures are identical (AD102), so it comes down to memory and link:

OptionVRAMBW (GB/s)TP scalingPowerVerdict
2 × RTX 409048 GB total (24+24)1008 each~1.3–1.5× over PCIe 4900 W combinedHigher peak BW, painful TP scaling, no datacenter licence.
1 × RTX 6000 Ada48 GB on one card960n/a (single GPU)300 WSame VRAM in one address space — no link tax, lower power, ECC, datacenter-licensed.

For any TP-aware workload (vLLM TP=2 of a 70B-class model), the RTX 6000 Ada wins comfortably despite identical underlying silicon. For two-replica DP-style serving (one model per card behind a load balancer), the 4090 pair wins on raw aggregate bandwidth and absolute throughput.

Successor preview

Blackwell partly relented: the RTX PRO 6000 Blackwell puts 96 GB GDDR7 on a single card (still without NVLink), which sidesteps the question for many buyers. Consumer Blackwell (RTX 50-series) reportedly continues to have no NVLink bridge. Deck 09 covers Blackwell.

10

Ada for LLMs — Where It Shines, Where It Doesn't

A frank assessment, two years after launch. Ada is two architectures — the consumer line is a single-GPU local-AI machine, and the L40S is a cheap datacenter inference card. Each excels at different things.

Where Ada shines

  • Single-GPU consumer inference — 24 GB on a 4090 fits 8B BF16 / 32B INT4 cleanly
  • Cheap FP8 inference — L40S at half H100 price, datacenter-licensed
  • Vector search / RAG — bandwidth (1008 GB/s) is plenty for embedding retrieval
  • Video / vision pipelines — NVENC/NVDEC and FP8 on L40S are excellent
  • Hobbyist fine-tuning — LoRA on 7–13B models on a single 4090 is now routine

Where Ada falls short

  • Multi-GPU TP — no NVLink on consumer cards, painful PCIe scaling
  • Pretraining — FP8 lacks Transformer Engine integration; no NVSwitch fabric; FP4 absent (Blackwell only)
  • HPC FP64 — very thin (Ada is graphics-first)
  • Largest models — 70B-FP8 needs 2× L40S over PCIe; H100 SXM still better there
  • Long-context decode — bandwidth is the limit; HBM3 cards win on tok/s

Worked numbers — what you actually get

Model / quantHardwareThroughput
Llama-3-8B BF161 × RTX 4090~55 tok/s/stream (ceiling 1008 ÷ 16 GB ≈ 63) · batched ~3000 tok/s aggregate
Llama-3-70B AWQ-INT42 × RTX 4090, PCIe TP=2~22 tok/s/stream
Llama-3-70B FP82 × L40S, PCIe TP=2~20 tok/s/stream (ceiling 864 ÷ 35 GB per card ≈ 25) · batched ~600 tok/s
Llama-3-8B FP81 × L40S~95 tok/s/stream · batched ~4500 tok/s
SDXL image, 30 steps1 × RTX 4090~3.5 s per 1024×1024 image
Honest bottom line

If your workload is single-stream decode of a model that fits in 24–48 GB, Ada is the best price/performance NVIDIA offers from 2022–2026. If you need to span 70B+ with tensor parallelism, the lack of NVLink hurts and you should look at H100/H200 (Hopper deck) or RTX PRO 6000 Blackwell (Blackwell deck) instead.

11

Interactive: Ada SKU Picker

Pick an Ada-generation SKU. The panel summarises capability, deployment licence, multi-GPU prospects, and the largest LLM you can comfortably host at FP16 vs INT4.

VRAM
—
BW (GB/s)
—
TDP (W)
—
FP8
—
NVLink
—
ECC
—
Datacenter licence
—
Largest LLM (FP16 / INT4)
—
How to read the picker

"Largest LLM" is a rough envelope: FP16 assumes ~70% of VRAM goes to weights and the rest to KV/runtime; INT4 assumes ~80%. In practice context length, batch size, and KV-cache dtype shift these by ~20% either way. Use this to filter out impossible matches, not as a precise spec.