NVIDIA GPU Architectures Series — Presentation 24

Power & Thermal — Delivering 1000 W to a Card and Removing It Again

A B200 dissipates 1000 watts in roughly 1,600 mm² of silicon (two ~800 mm² dies). Walk through how NVIDIA gets that power into the package — VRMs, multi-rail design, decoupling, transient response — and how the heat gets out — vapour chambers, direct-to-chip liquid cooling, immersion. Plus GPU Boost, thermal throttling, power capping and TGP.

VRMTGPGPU Boost P-statesThermalVapour chamber D2C liquidImmersion Power cappingGSP
PSU → Cable → VRM → Rail → Die → Heat → Cold plate → Coolant → Cluster
00

Topics We'll Cover

The GPU is increasingly an electrical and thermal engineering problem before it is a software one. This deck follows a single watt from the wall socket, through the PSU and 12V-2x6 cable, into the on-card VRMs and rails, into the die — and then back out as heat through the heat-spreader, cold plate, coolant loop, and ultimately the facility's chilled-water plant.

01

Power in 2026 — Why It Matters

The headline numbers tell you the story before any of the engineering does. A consumer flagship now draws what an entire workstation drew a decade ago, and a single rack of accelerators draws what a small office building draws.

RTX 4090
450 W
consumer
RTX 5090
575 W
consumer
H100 SXM5
700 W
datacenter
B200
1000 W
datacenter
GB200 superchip
2700 W (2×B200 + Grace)
superchip

Scaling out to a rack

An NVL72 rack is 36 GB200 superchips packaged with NVLink switches, PSUs, and a closed water loop in a single 19-inch frame. It draws roughly 120 kW. A datacenter row of eight NVL72 racks draws nearly 1 MW (about 580 kW of it the 576 GPUs themselves) — before networking outside the racks, storage, or the cooling and power-conversion overhead.

The constraint has shifted

Power is now the primary constraint on AI compute scaling — not silicon area, not foundry capacity, not even HBM supply. Hyperscalers buy land near hydroelectric dams and nuclear reactors specifically to deploy. New US AI campuses are sized in gigawatts, which is the language of national-grid load.

What that means for the engineering

02

Power Delivery — From PSU to Die

The journey from the wall socket to a transistor switching at 2.5 GHz crosses six or seven distinct voltage domains. Most of them are invisible to the user but each one is engineered to a tight specification.

230 V AC wall PSU 12 V rail ~94% eff 12V-2x6 600 W max On-card VRM 8–30 phases ~500 kHz switching MOSFET pair + inductor + driver IC interleaved phases Decoupling MLCC + POSCAP tantalum bulk Die 0.7–1.1 V hundreds of A

The legacy ATX path versus 12V-2x6

For a long time consumer PSUs delivered 12 V on three or more separate 8-pin EPS/PCIe connectors. Each 8-pin is rated at 150 W — so a 450 W card needed three of them, plus an adapter. The 12V-2x6 connector (a quiet revision of the original 12VHPWR shipped with the RTX 4090) consolidates that into a single 16-pin header rated for 600 W on one cable, with four sense pins that let the card negotiate its allowance with the PSU.

Inside the on-card VRM

A single 12 V → 1 V buck converter at 700 A is impossible — the inductor would saturate, the MOSFETs would melt, and the ripple would be unusable. So the design splits the conversion across multiple parallel phases, each handling roughly 50–80 A. A B200 board carries on the order of 20 phases for the GPU core alone.

Numbers worth remembering

A B200 at 1000 W on a 0.85 V core rail draws roughly 1180 A. Spread across 20 VRM phases that is ~60 A per phase — well within modern DrMOS ratings. Running the same card at 12 V directly would require a fraction of that current but no transistor in 2026 can switch from 0–1 in a fast enough window to drive a digital logic gate; the buck conversion is mandatory.

03

Multi-Rail Design

A modern datacenter card is not a single voltage domain. It is six to twelve separate regulated supplies, each tuned for a specific consumer on the package, each with its own controller IC, sequencing requirement, and tolerance budget.

Why split the rails at all?

Representative rail map for a Hopper-class card

RailVoltageApprox currentConsumer
VDD core0.85 V (variable)~700 ASMs, tensor cores, scheduler, L2
VDDQ-HBM1.1 V~80 A per stackHBM core array (per-stack)
VDDIO-HBM0.4 V~30 A per stackHBM PHY/IO drivers
VDDIO-NVL0.85 V~40 ANVLink SerDes PHYs
VDDA-PLL0.9 V~5 APLLs, analog clock tree
VDD181.8 V~10 Athermal/I²C/aux analog
VDD333.3 V~3 AEEPROM, GSP boot, fan tach

The controllers

Each rail has its own digital PWM controller talking to its own array of power stages. Common silicon vendors: Renesas (formerly Intersil), Monolithic Power Systems (MPS), Infineon (formerly International Rectifier). All speak PMBus/I²C to the GSP firmware on the card, which orchestrates startup sequencing, telemetry, and per-rail margining.

The sequence on cold boot

VDD33 first → GSP boots → GSP brings up VDDA-PLL and VDD18 → PLLs lock → VDDIO-HBM and VDDIO-NVL come up → VDDQ-HBM → HBM training runs → finally VDD core ramps under closed-loop control. Each step has a millisecond-scale guard window before the next is permitted. Reverse the order at the wrong step and the card will not POST — or will POST and brick.

04

Transient Response — The Hard Part

The hardest electrical problem on a modern GPU board is not steady-state efficiency; it is the transient. A single tensor-core matmul issued by the scheduler can take the chip from a few amps of leakage to several hundred amps of switching activity in a clock cycle or two — well under 100 nanoseconds.

Why VRMs cannot keep up

A buck phase switching at 500 kHz has a control loop bandwidth of roughly 50–100 kHz. That means the VRM cannot react to changes faster than ~10 microseconds, and certainly not on a 100-nanosecond timescale. If the load steps before the VRM can respond, the rail voltage droops until the inductor current ramps up to meet the demand. Droop the rail too far and the chip's minimum-operating-voltage spec is violated; logic glitches; PLLs lose lock; the workload hangs.

The decoupling network bridges the gap

A tiered capacitor network lives between the VRM phases and the package, each tier handling a different time-constant of disturbance:

< 10 ns
on-die MIMcaps
package MLCCs
10–100 ns
0402 MLCC array
~100 nF each, ~200 caps
100 ns–1 μs
POSCAP / polymer
~330 μF each
1–10 μs
tantalum bulk
~1 mF total
> 10 μs
VRM phases
control loop catches up

The cost of getting it wrong

Insufficient decoupling shows up as voltage droop that pushes the rail under its minimum-operating-voltage line. The chip responds in one of two ways:

Reference vs OEM

NVIDIA's reference designs are aggressive on decoupling. They specify a particular MLCC count, a particular POSCAP family, a particular package-level placement, and validate against a worst-case di/dt step. OEM partner cards sometimes cut corners on capacitor count to reduce BOM cost — the 2020 RTX 3080 launch was a public case study, where some partner boards crashed under FurMark transients while NVIDIA's Founders Edition did not. The fix was always more POSCAPs near the package.

05

GPU Boost & Clock/Voltage Curves

Every modern NVIDIA GPU runs at variable frequency. Idle, an H100 might sit at 800 MHz drawing 50 W; under sustained load, it reaches 1.98 GHz drawing its full 700 W. A consumer 4090 will idle at ~210 MHz and boost to 2.5 GHz or more.

What GPU Boost actually does

GPU Boost is a closed-loop firmware algorithm running inside the on-board GSP (GPU System Processor). Each die is binned at the factory with a per-part frequency-versus-voltage curve measured at production test. At runtime the GSP picks the highest stable clock that satisfies all of these constraints simultaneously:

Because the curve is per-part, two physically identical 4090s in the same chassis will boost differently. The luckier silicon does more work for the same watts.

Inspecting it on a live system

$ nvidia-smi -q -d CLOCK,POWER,TEMPERATURE
# --- Selected output, annotated ---
Clocks
    Graphics                          : 1980 MHz   # current SM clock under load
    SM                                : 1980 MHz
    Memory                            : 1593 MHz   # HBM clock (fixed)
    Video                             : 1740 MHz

Max Clocks
    Graphics                          : 1980 MHz   # max permitted

Default Applications Clocks
    Graphics                          : 1755 MHz   # conservative default

Power Readings
    Power Management                  : Supported
    Power Draw                        : 689.42 W
    Current Power Limit               : 700.00 W   # TGP cap from firmware
    Min Power Limit                   : 200.00 W
    Max Power Limit                   : 700.00 W

Temperature
    GPU Current Temp                  : 71 C
    GPU Shutdown Temp                 : 95 C    # hard cutoff
    GPU Slowdown Temp                 : 90 C    # aggressive throttle
    GPU Max Operating Temp            : 88 C
    Memory Current Temp               : 78 C    # HBM hotspot
    Memory Max Operating Temp         : 95 C

Datacenter vs consumer tuning

The two tiers tune their boost behaviour very differently:

Datacenter (H100, B200, GB200)

Pinned conservatively closer to the design point. The card is expected to run at full clock continuously for weeks. Headroom is held back so that thermal margin survives a hot inlet day or a partial fan failure. The boost curve is flat and predictable.

Consumer (RTX 4090, 5090)

Pushed aggressively close to the limit. Workloads are bursty: a frame, a generation, a benchmark. The card races to peak clock, hits the thermal wall after 30–90 seconds, and rides the throttle envelope. Two cards of the same SKU may differ by 5–10% on sustained workloads.

06

Power States (P-states) and Idle Power

NVIDIA exposes a discrete ladder of performance states, P0 through P12, that the GSP firmware moves between based on observed workload. P0 is full boost; P12 is the deepest idle. The transitions are fast (microseconds for clock changes, milliseconds for the voltage rail to settle), but with one important caveat that bites operators.

P-stateApprox clockApprox draw (4090)Use
P02520 MHz450 Wfull boost, sustained compute
P21950 MHz~280 Wcompute (CUDA workload, modest)
P51200 MHz~70 Wlight graphics, video decode
P8210 MHz~25 Wdisplay only, no compute
P12~210 MHz~5 Wdeep idle, headless, ASPM-able

Persistence mode — the production-server lever

By default on consumer cards, when no CUDA process holds the device the driver unloads and the card drops to P8 or P12. Re-loading the driver on the next process start takes around 10 seconds. For a server that handles spiky inference traffic, that delay is wholly unacceptable: the first request after a quiet period stalls for ten seconds, the user sees a timeout, the autoscaler panics.

The fix is persistence mode:

enable persistence on every reboot
# one-shot
$ sudo nvidia-smi -pm 1
Enabled persistence mode for GPU 00000000:01:00.0.

# systemd unit (the proper way)
$ sudo systemctl enable --now nvidia-persistenced

# verify
$ nvidia-smi --query-gpu=persistence_mode --format=csv
persistence_mode
Enabled

With persistence on, the driver stays loaded permanently, the card holds at P2 or P0 instead of dropping to P12, and process startup is back to milliseconds. This is mandatory on production inference hosts. Datacenter cards typically have it on by default; consumer cards do not.

The trade-off

Persistence mode increases idle power, by roughly 30–50 W per card, because the driver pins clocks higher so that the next request finds the rails already up. On a workstation that's a non-issue. In a 1000-card datacenter, leaving persistence on for cards that are genuinely idle 80% of the time is 30–50 kW of pure waste — more than a small office uses for everything. The right policy is workload-aware: persistence on for serving fleets, off for batch-only training pools.

07

Power Capping

Operator-level power capping lets you trade a small amount of throughput for a large amount of power. It is one of the most useful operational knobs on the platform, and it is criminally under-used.

$ sudo nvidia-smi -pl 350
# Cap RTX 4090 from 450 W to 350 W
$ sudo nvidia-smi -pl 350
Power limit for GPU 00000000:01:00.0 was set to 350.00 W from 450.00 W.
All done.

# Verify the floor and ceiling
$ nvidia-smi --query-gpu=power.limit,power.min_limit,power.max_limit --format=csv
power.limit [W], power.min_limit [W], power.max_limit [W]
350.00 W, 100.00 W, 450.00 W

# Recent driver versions also accept a floor:
$ sudo nvidia-smi --power-limit-min=200 --power-limit=350

The economics

Power scales roughly with V²f, while throughput scales linearly with f. So pulling clock down a small amount drops power a large amount — the relationship is supra-linear. Empirical numbers for a 4090:

CapThroughput deltaPower deltaPerf-per-W gain
450 W (stock)baselinebaseline1.00×
400 W−2%−11%1.10×
350 W−5%−22%1.22×
300 W−10%−33%1.34×
250 W−20%−44%1.43×

The PL is enforced inside the GSP firmware. When the operator sets a 350 W cap, the GSP's closed-loop algorithm dynamically scales clocks down to keep total board draw within budget. The user space sees nothing — the workload just runs slightly slower with no other behavioural change.

When to use it

A free 20% on inference fleets

Inference latency is dominated by memory bandwidth, not compute. Capping a fleet of H100s from 700 W to 550 W typically loses ~3% on tok/s because HBM is already at its bandwidth wall. You bank ~21% of the energy bill for almost no SLA cost. The first thing to do on a new serving fleet is benchmark at three or four cap settings and pick the knee of the curve.

08

Cooling Spectrum — Air to Immersion

There are roughly five practical approaches to extracting the heat that a power-delivery network has just deposited into the silicon. They span four orders of magnitude in cost and two in capability. The right choice depends almost entirely on the watts-per-card you're trying to remove.

(a) Air, axial fan

Cards: consumer/workstation (RTX 4090, RTX PRO 6000 Workstation). Two or three downward-facing axial fans, vapour chamber over the die, finned aluminium heat sink.

Limit: ~250–450 W per card depending on case airflow.

Failure mode: recirculation inside the case — hot exhaust gets sucked back in. Fine for one card; fatal for two or more in the same chassis.

(b) Air, blower

Cards: 1U/2U datacenter (legacy A40, A100 PCIe blower variants). A radial impeller pulls air axially through the heat sink and exhausts it out the back of the chassis.

Limit: ~400 W per card. Loud (typically 60+ dBA at full load).

OK for: A100 PCIe in HGX nodes, professional workstations with multiple cards.

(c) Air, passive (SXM/HGX)

Cards: SXM4/SXM5 modules (H100 SXM5, H200) on a server baseboard. No on-card fan. The chassis itself contains a wall of high-static-pressure fans pushing 35–50 CFM front-to-back across the bare heat sinks.

Limit: ~700 W per card, the H100 SXM5 standard.

Cost: chassis fans alone draw 1–2 kW per node; cooling overhead is significant.

(d) Direct-to-chip liquid (D2C)

Cards: B200, GB200 (NVL72 mandates D2C). A copper cold plate with internal microchannels is clamped to the package; 30–50°C facility water flows through it and carries the heat to a CDU (Coolant Distribution Unit) and the chilled-water plant.

Limit: ~1500 W per card sustained, with thermal margin.

Cost: facility water plumbing, leak-detection, dripless quick-disconnects, redundant pumps. Substantial first-time capex; modest opex once running.

(e) Immersion

Cards: any. The whole server is submerged in a tank of dielectric fluid (typically a fluorocarbon or synthetic mineral oil). Heat is removed by either single-phase pumped flow or two-phase boiling.

Limit: ~3000 W per card and beyond.

Cost: very expensive and operationally awkward (every service call drains a tank). Used at hyperscale where facility PUE wins justify the friction.

The vapour-chamber detail

All air-cooled cards rely on a vapour chamber as the first thermal stage. A flat copper enclosure containing a small amount of working fluid sits between the die and the fin stack; the fluid evaporates at the hotspot, rises to the cool fin side, condenses, and wicks back. It moves heat from a 800 mm² die out to a 12000 mm² sink with an effective conductance ~50× better than solid copper.

Where the line falls

Axial fan
≤ 450 W
consumer
Blower
≤ 400 W
workstation
Passive HGX
≤ 700 W
SXM5
D2C liquid
≤ 1500 W
B200/GB200
Immersion
≤ 3000 W+
hyperscale
09

Thermal Throttling — The Last Defence

If power-delivery is the offensive line, throttling is the goal-keeper of last resort. It exists because cooling can fail in many ways — a fan dies, an inlet temp rises, a coolant-loop CDU loses pressure — and the silicon must protect itself before any of those become permanent.

Sensors and thresholds

The GSP firmware runs a control loop sampling on the order of kilohertz across multiple thermal sensors:

Two failure modes operators see

(a) Fast throttle — predictable

Sustained load + insufficient cooling. Within 30–90 seconds the chip hits its throttle threshold. Clocks drop ~30%. Throughput drops ~20–30%. The behaviour is stable: tok/s settles at a new lower number that you can plan capacity around. Most painful, easiest to diagnose.

(b) Slow degradation — nasty

Cooling is marginal. Clocks ride a few hundred MHz below peak, hovering near throttle. Some seconds the chip clears, others it doesn't. Variable performance — tok/s wanders. Latency p99 spikes. SLAs miss randomly. Far harder to debug because no single sample looks unhealthy.

The diagnostic incantation

live throttle inspection
$ nvidia-smi --query-gpu=clocks.current.graphics,temperature.gpu,clocks_throttle_reasons.active --format=csv -l 1

clocks.current.graphics [MHz], temperature.gpu [C], clocks_throttle_reasons.active
1980, 71, 0x0000000000000001     # 0x1 = idle, no throttle
1980, 82, 0x0000000000000001
1755, 88, 0x0000000000000004     # 0x4 = SW thermal slowdown
1410, 91, 0x000000000000000c     # 0xc = HW thermal + SW thermal
1245, 94, 0x000000000000000c     # dropping 30%, near shutdown

The throttle-reason bitmap

BitMaskReasonMeaning
00x1GpuIdlenot throttling, just idle
10x2ApplicationsClocksSettingoperator-set frequency, not an alarm
20x4SwPowerCapoperator -pl cap holding clocks down
30x8HwSlowdownhardware thermal slowdown, urgent
60x40SwThermalSlowdownfirmware-driven thermal cap
70x80HwThermalSlowdownHW thermal trip, very urgent
80x100HwPowerBrakeSlowdownexternal PowerBrake# pin asserted (PSU/UPS event)
90x200SyncBoostmulti-GPU clock-sync, operational
The non-obvious one

HwPowerBrakeSlowdown (0x100) is asserted by an external pin held by the chassis or UPS — not by the GPU itself. If you see it firing on a healthy-looking card, your PSU or facility power is brown-ing out under load. The GPU is doing the right thing; the failure is upstream.

10

Datacenter Engineering — From Card to Row

Once you put eight cards in a chassis and 18 chassis in a rack, the engineering changes character. The card-level problems all still exist; on top of them sit airflow, power distribution, plumbing, and PUE.

From card heat to chassis heat

A 700 W card removes ~2390 BTU/hr. (One watt is 3.412 BTU/hr; the conversion is mechanical.) A 4U HGX chassis with eight H100s at 700 W each is 5.6 kW of heat — ~19,100 BTU/hr — that has to leave the rack into the hot aisle.

For air-cooled removal at a typical 15 °C aisle delta-T, the rule of thumb is roughly 0.026 CFM per watt:

For closed-loop water at 25 °C inlet and 35 °C return, a similar rule of thumb is ~0.85 GPM per kW (using water's specific heat). So:

Power distribution at rack scale

A 30–120 kW rack does not run on a household plug.

PUE — Power Usage Effectiveness

PUE is the ratio of total facility power to useful IT power: a PUE of 1.0 means every watt drawn from the grid ends up as compute. Real datacenters live between these poles:

Facility classTypical PUEComment
Hyperscale, cold climate, free cooling1.05–1.15Outside air does most of the work
Modern hyperscale, mixed climate1.15–1.30Hot aisle containment + chilled water
Enterprise / colo, chilled water1.30–1.50CRAC units, room-level cooling
Legacy raised-floor1.5–2.0+Inefficient airflow, oversized chillers

PUE is dominated by cooling overhead, with power-conversion efficiency a distant second. Moving from air-cooled HGX to D2C-cooled NVL72 typically buys 0.10–0.15 of PUE. At hyperscale with a multi-MW deployment, that pays for the entire D2C plumbing capex within the first year.

11

Interactive: Power & Cooling Calculator

Pick a GPU, a chassis density, a cooling approach, and an inlet temperature. The calculator estimates total heat load, the airflow or coolant flow you'd need, the thermal margin you have, and whether the configuration is feasible. All maths is inlined in the page so you can read it.

Per-card TGP
—
Chassis power
—
Required flow
—
Thermal margin
—
Throttle risk
—
Facility power
—
verdict: select inputs to compute

How the maths works

inlined formulas (no library calls)
// Per-card TGP table (W)
const tgp = {
    rtx4090: 450, rtx5090: 575, rtxpro6000bw: 600,
    a100: 400, h100: 700, h200: 700,
    b200: 1000, gb200: 2700
};

// Cooling capability per card (W) at 25 C inlet
const cap = {
    axial: 450, blower: 400, passive: 700,
    d2c: 1500, immersion: 3000
};

// Inlet derating (warm/hot air loses cooling capability)
const inletScale = (T) => T <= 18 ? 1.10 :
                          T <= 25 ? 1.00 :
                                       0.78;  // 35 C: ~22% capability loss

// Required airflow (CFM) at 15 C delta-T: 0.026 CFM per watt
const cfm = (W) => W * 0.026;

// Required water flow (GPM) at 10 C delta-T: ~0.85 GPM per kW
const gpm = (W) => (W / 1000) * 0.85;

// Facility overhead: each cooling type carries a PUE-like factor
const pue = {
    axial: 1.45, blower: 1.40, passive: 1.30,
    d2c:   1.12, immersion: 1.06
};
The headline

Try 8× B200 in a passive HGX chassis at 35 °C. Per-card TGP is 1000 W, effective passive capability at 35 °C inlet is roughly 546 W in this model — which is why air-cooled DGX B200 needs a 10U chassis moving ~1,550 CFM, and why GB200 NVL72 ships with mandatory D2C liquid: the engineering envelope at 1000 W per card has run off the end of what air can do, regardless of how hard you spin the fans.