A B200 dissipates 1000 watts in roughly 1,600 mm² of silicon (two ~800 mm² dies). Walk through how NVIDIA gets that power into the package — VRMs, multi-rail design, decoupling, transient response — and how the heat gets out — vapour chambers, direct-to-chip liquid cooling, immersion. Plus GPU Boost, thermal throttling, power capping and TGP.
The GPU is increasingly an electrical and thermal engineering problem before it is a software one. This deck follows a single watt from the wall socket, through the PSU and 12V-2x6 cable, into the on-card VRMs and rails, into the die — and then back out as heat through the heat-spreader, cold plate, coolant loop, and ultimately the facility's chilled-water plant.
The headline numbers tell you the story before any of the engineering does. A consumer flagship now draws what an entire workstation drew a decade ago, and a single rack of accelerators draws what a small office building draws.
An NVL72 rack is 36 GB200 superchips packaged with NVLink switches, PSUs, and a closed water loop in a single 19-inch frame. It draws roughly 120 kW. A datacenter row of eight NVL72 racks draws nearly 1 MW (about 580 kW of it the 576 GPUs themselves) — before networking outside the racks, storage, or the cooling and power-conversion overhead.
Power is now the primary constraint on AI compute scaling — not silicon area, not foundry capacity, not even HBM supply. Hyperscalers buy land near hydroelectric dams and nuclear reactors specifically to deploy. New US AI campuses are sized in gigawatts, which is the language of national-grid load.
The journey from the wall socket to a transistor switching at 2.5 GHz crosses six or seven distinct voltage domains. Most of them are invisible to the user but each one is engineered to a tight specification.
For a long time consumer PSUs delivered 12 V on three or more separate 8-pin EPS/PCIe connectors. Each 8-pin is rated at 150 W — so a 450 W card needed three of them, plus an adapter. The 12V-2x6 connector (a quiet revision of the original 12VHPWR shipped with the RTX 4090) consolidates that into a single 16-pin header rated for 600 W on one cable, with four sense pins that let the card negotiate its allowance with the PSU.
A single 12 V → 1 V buck converter at 700 A is impossible — the inductor would saturate, the MOSFETs would melt, and the ripple would be unusable. So the design splits the conversion across multiple parallel phases, each handling roughly 50–80 A. A B200 board carries on the order of 20 phases for the GPU core alone.
A B200 at 1000 W on a 0.85 V core rail draws roughly 1180 A. Spread across 20 VRM phases that is ~60 A per phase — well within modern DrMOS ratings. Running the same card at 12 V directly would require a fraction of that current but no transistor in 2026 can switch from 0–1 in a fast enough window to drive a digital logic gate; the buck conversion is mandatory.
A modern datacenter card is not a single voltage domain. It is six to twelve separate regulated supplies, each tuned for a specific consumer on the package, each with its own controller IC, sequencing requirement, and tolerance budget.
| Rail | Voltage | Approx current | Consumer |
|---|---|---|---|
| VDD core | 0.85 V (variable) | ~700 A | SMs, tensor cores, scheduler, L2 |
| VDDQ-HBM | 1.1 V | ~80 A per stack | HBM core array (per-stack) |
| VDDIO-HBM | 0.4 V | ~30 A per stack | HBM PHY/IO drivers |
| VDDIO-NVL | 0.85 V | ~40 A | NVLink SerDes PHYs |
| VDDA-PLL | 0.9 V | ~5 A | PLLs, analog clock tree |
| VDD18 | 1.8 V | ~10 A | thermal/I²C/aux analog |
| VDD33 | 3.3 V | ~3 A | EEPROM, GSP boot, fan tach |
Each rail has its own digital PWM controller talking to its own array of power stages. Common silicon vendors: Renesas (formerly Intersil), Monolithic Power Systems (MPS), Infineon (formerly International Rectifier). All speak PMBus/I²C to the GSP firmware on the card, which orchestrates startup sequencing, telemetry, and per-rail margining.
VDD33 first → GSP boots → GSP brings up VDDA-PLL and VDD18 → PLLs lock → VDDIO-HBM and VDDIO-NVL come up → VDDQ-HBM → HBM training runs → finally VDD core ramps under closed-loop control. Each step has a millisecond-scale guard window before the next is permitted. Reverse the order at the wrong step and the card will not POST — or will POST and brick.
The hardest electrical problem on a modern GPU board is not steady-state efficiency; it is the transient. A single tensor-core matmul issued by the scheduler can take the chip from a few amps of leakage to several hundred amps of switching activity in a clock cycle or two — well under 100 nanoseconds.
A buck phase switching at 500 kHz has a control loop bandwidth of roughly 50–100 kHz. That means the VRM cannot react to changes faster than ~10 microseconds, and certainly not on a 100-nanosecond timescale. If the load steps before the VRM can respond, the rail voltage droops until the inductor current ramps up to meet the demand. Droop the rail too far and the chip's minimum-operating-voltage spec is violated; logic glitches; PLLs lose lock; the workload hangs.
A tiered capacitor network lives between the VRM phases and the package, each tier handling a different time-constant of disturbance:
Insufficient decoupling shows up as voltage droop that pushes the rail under its minimum-operating-voltage line. The chip responds in one of two ways:
NVIDIA's reference designs are aggressive on decoupling. They specify a particular MLCC count, a particular POSCAP family, a particular package-level placement, and validate against a worst-case di/dt step. OEM partner cards sometimes cut corners on capacitor count to reduce BOM cost — the 2020 RTX 3080 launch was a public case study, where some partner boards crashed under FurMark transients while NVIDIA's Founders Edition did not. The fix was always more POSCAPs near the package.
Every modern NVIDIA GPU runs at variable frequency. Idle, an H100 might sit at 800 MHz drawing 50 W; under sustained load, it reaches 1.98 GHz drawing its full 700 W. A consumer 4090 will idle at ~210 MHz and boost to 2.5 GHz or more.
GPU Boost is a closed-loop firmware algorithm running inside the on-board GSP (GPU System Processor). Each die is binned at the factory with a per-part frequency-versus-voltage curve measured at production test. At runtime the GSP picks the highest stable clock that satisfies all of these constraints simultaneously:
-pl override).Because the curve is per-part, two physically identical 4090s in the same chassis will boost differently. The luckier silicon does more work for the same watts.
# --- Selected output, annotated ---
Clocks
Graphics : 1980 MHz # current SM clock under load
SM : 1980 MHz
Memory : 1593 MHz # HBM clock (fixed)
Video : 1740 MHz
Max Clocks
Graphics : 1980 MHz # max permitted
Default Applications Clocks
Graphics : 1755 MHz # conservative default
Power Readings
Power Management : Supported
Power Draw : 689.42 W
Current Power Limit : 700.00 W # TGP cap from firmware
Min Power Limit : 200.00 W
Max Power Limit : 700.00 W
Temperature
GPU Current Temp : 71 C
GPU Shutdown Temp : 95 C # hard cutoff
GPU Slowdown Temp : 90 C # aggressive throttle
GPU Max Operating Temp : 88 C
Memory Current Temp : 78 C # HBM hotspot
Memory Max Operating Temp : 95 C
The two tiers tune their boost behaviour very differently:
Pinned conservatively closer to the design point. The card is expected to run at full clock continuously for weeks. Headroom is held back so that thermal margin survives a hot inlet day or a partial fan failure. The boost curve is flat and predictable.
Pushed aggressively close to the limit. Workloads are bursty: a frame, a generation, a benchmark. The card races to peak clock, hits the thermal wall after 30–90 seconds, and rides the throttle envelope. Two cards of the same SKU may differ by 5–10% on sustained workloads.
NVIDIA exposes a discrete ladder of performance states, P0 through P12, that the GSP firmware moves between based on observed workload. P0 is full boost; P12 is the deepest idle. The transitions are fast (microseconds for clock changes, milliseconds for the voltage rail to settle), but with one important caveat that bites operators.
| P-state | Approx clock | Approx draw (4090) | Use |
|---|---|---|---|
| P0 | 2520 MHz | 450 W | full boost, sustained compute |
| P2 | 1950 MHz | ~280 W | compute (CUDA workload, modest) |
| P5 | 1200 MHz | ~70 W | light graphics, video decode |
| P8 | 210 MHz | ~25 W | display only, no compute |
| P12 | ~210 MHz | ~5 W | deep idle, headless, ASPM-able |
By default on consumer cards, when no CUDA process holds the device the driver unloads and the card drops to P8 or P12. Re-loading the driver on the next process start takes around 10 seconds. For a server that handles spiky inference traffic, that delay is wholly unacceptable: the first request after a quiet period stalls for ten seconds, the user sees a timeout, the autoscaler panics.
The fix is persistence mode:
# one-shot
$ sudo nvidia-smi -pm 1
Enabled persistence mode for GPU 00000000:01:00.0.
# systemd unit (the proper way)
$ sudo systemctl enable --now nvidia-persistenced
# verify
$ nvidia-smi --query-gpu=persistence_mode --format=csv
persistence_mode
Enabled
With persistence on, the driver stays loaded permanently, the card holds at P2 or P0 instead of dropping to P12, and process startup is back to milliseconds. This is mandatory on production inference hosts. Datacenter cards typically have it on by default; consumer cards do not.
Persistence mode increases idle power, by roughly 30–50 W per card, because the driver pins clocks higher so that the next request finds the rails already up. On a workstation that's a non-issue. In a 1000-card datacenter, leaving persistence on for cards that are genuinely idle 80% of the time is 30–50 kW of pure waste — more than a small office uses for everything. The right policy is workload-aware: persistence on for serving fleets, off for batch-only training pools.
Operator-level power capping lets you trade a small amount of throughput for a large amount of power. It is one of the most useful operational knobs on the platform, and it is criminally under-used.
# Cap RTX 4090 from 450 W to 350 W
$ sudo nvidia-smi -pl 350
Power limit for GPU 00000000:01:00.0 was set to 350.00 W from 450.00 W.
All done.
# Verify the floor and ceiling
$ nvidia-smi --query-gpu=power.limit,power.min_limit,power.max_limit --format=csv
power.limit [W], power.min_limit [W], power.max_limit [W]
350.00 W, 100.00 W, 450.00 W
# Recent driver versions also accept a floor:
$ sudo nvidia-smi --power-limit-min=200 --power-limit=350
Power scales roughly with V²f, while throughput scales linearly with f. So pulling clock down a small amount drops power a large amount — the relationship is supra-linear. Empirical numbers for a 4090:
| Cap | Throughput delta | Power delta | Perf-per-W gain |
|---|---|---|---|
| 450 W (stock) | baseline | baseline | 1.00× |
| 400 W | −2% | −11% | 1.10× |
| 350 W | −5% | −22% | 1.22× |
| 300 W | −10% | −33% | 1.34× |
| 250 W | −20% | −44% | 1.43× |
The PL is enforced inside the GSP firmware. When the operator sets a 350 W cap, the GSP's closed-loop algorithm dynamically scales clocks down to keep total board draw within budget. The user space sees nothing — the workload just runs slightly slower with no other behavioural change.
nvidia-smi -pl into emergency response: cap the entire fleet by 30% rather than shed load.Inference latency is dominated by memory bandwidth, not compute. Capping a fleet of H100s from 700 W to 550 W typically loses ~3% on tok/s because HBM is already at its bandwidth wall. You bank ~21% of the energy bill for almost no SLA cost. The first thing to do on a new serving fleet is benchmark at three or four cap settings and pick the knee of the curve.
There are roughly five practical approaches to extracting the heat that a power-delivery network has just deposited into the silicon. They span four orders of magnitude in cost and two in capability. The right choice depends almost entirely on the watts-per-card you're trying to remove.
Cards: consumer/workstation (RTX 4090, RTX PRO 6000 Workstation). Two or three downward-facing axial fans, vapour chamber over the die, finned aluminium heat sink.
Limit: ~250–450 W per card depending on case airflow.
Failure mode: recirculation inside the case — hot exhaust gets sucked back in. Fine for one card; fatal for two or more in the same chassis.
Cards: 1U/2U datacenter (legacy A40, A100 PCIe blower variants). A radial impeller pulls air axially through the heat sink and exhausts it out the back of the chassis.
Limit: ~400 W per card. Loud (typically 60+ dBA at full load).
OK for: A100 PCIe in HGX nodes, professional workstations with multiple cards.
Cards: SXM4/SXM5 modules (H100 SXM5, H200) on a server baseboard. No on-card fan. The chassis itself contains a wall of high-static-pressure fans pushing 35–50 CFM front-to-back across the bare heat sinks.
Limit: ~700 W per card, the H100 SXM5 standard.
Cost: chassis fans alone draw 1–2 kW per node; cooling overhead is significant.
Cards: B200, GB200 (NVL72 mandates D2C). A copper cold plate with internal microchannels is clamped to the package; 30–50°C facility water flows through it and carries the heat to a CDU (Coolant Distribution Unit) and the chilled-water plant.
Limit: ~1500 W per card sustained, with thermal margin.
Cost: facility water plumbing, leak-detection, dripless quick-disconnects, redundant pumps. Substantial first-time capex; modest opex once running.
Cards: any. The whole server is submerged in a tank of dielectric fluid (typically a fluorocarbon or synthetic mineral oil). Heat is removed by either single-phase pumped flow or two-phase boiling.
Limit: ~3000 W per card and beyond.
Cost: very expensive and operationally awkward (every service call drains a tank). Used at hyperscale where facility PUE wins justify the friction.
All air-cooled cards rely on a vapour chamber as the first thermal stage. A flat copper enclosure containing a small amount of working fluid sits between the die and the fin stack; the fluid evaporates at the hotspot, rises to the cool fin side, condenses, and wicks back. It moves heat from a 800 mm² die out to a 12000 mm² sink with an effective conductance ~50× better than solid copper.
If power-delivery is the offensive line, throttling is the goal-keeper of last resort. It exists because cooling can fail in many ways — a fan dies, an inlet temp rises, a coolant-loop CDU loses pressure — and the silicon must protect itself before any of those become permanent.
The GSP firmware runs a control loop sampling on the order of kilohertz across multiple thermal sensors:
Sustained load + insufficient cooling. Within 30–90 seconds the chip hits its throttle threshold. Clocks drop ~30%. Throughput drops ~20–30%. The behaviour is stable: tok/s settles at a new lower number that you can plan capacity around. Most painful, easiest to diagnose.
Cooling is marginal. Clocks ride a few hundred MHz below peak, hovering near throttle. Some seconds the chip clears, others it doesn't. Variable performance — tok/s wanders. Latency p99 spikes. SLAs miss randomly. Far harder to debug because no single sample looks unhealthy.
$ nvidia-smi --query-gpu=clocks.current.graphics,temperature.gpu,clocks_throttle_reasons.active --format=csv -l 1
clocks.current.graphics [MHz], temperature.gpu [C], clocks_throttle_reasons.active
1980, 71, 0x0000000000000001 # 0x1 = idle, no throttle
1980, 82, 0x0000000000000001
1755, 88, 0x0000000000000004 # 0x4 = SW thermal slowdown
1410, 91, 0x000000000000000c # 0xc = HW thermal + SW thermal
1245, 94, 0x000000000000000c # dropping 30%, near shutdown
| Bit | Mask | Reason | Meaning |
|---|---|---|---|
| 0 | 0x1 | GpuIdle | not throttling, just idle |
| 1 | 0x2 | ApplicationsClocksSetting | operator-set frequency, not an alarm |
| 2 | 0x4 | SwPowerCap | operator -pl cap holding clocks down |
| 3 | 0x8 | HwSlowdown | hardware thermal slowdown, urgent |
| 6 | 0x40 | SwThermalSlowdown | firmware-driven thermal cap |
| 7 | 0x80 | HwThermalSlowdown | HW thermal trip, very urgent |
| 8 | 0x100 | HwPowerBrakeSlowdown | external PowerBrake# pin asserted (PSU/UPS event) |
| 9 | 0x200 | SyncBoost | multi-GPU clock-sync, operational |
HwPowerBrakeSlowdown (0x100) is asserted by an external pin held by the chassis or UPS — not by the GPU itself. If you see it firing on a healthy-looking card, your PSU or facility power is brown-ing out under load. The GPU is doing the right thing; the failure is upstream.
Once you put eight cards in a chassis and 18 chassis in a rack, the engineering changes character. The card-level problems all still exist; on top of them sit airflow, power distribution, plumbing, and PUE.
A 700 W card removes ~2390 BTU/hr. (One watt is 3.412 BTU/hr; the conversion is mechanical.) A 4U HGX chassis with eight H100s at 700 W each is 5.6 kW of heat — ~19,100 BTU/hr — that has to leave the rack into the hot aisle.
For air-cooled removal at a typical 15 °C aisle delta-T, the rule of thumb is roughly 0.026 CFM per watt:
For closed-loop water at 25 °C inlet and 35 °C return, a similar rule of thumb is ~0.85 GPM per kW (using water's specific heat). So:
A 30–120 kW rack does not run on a household plug.
PUE is the ratio of total facility power to useful IT power: a PUE of 1.0 means every watt drawn from the grid ends up as compute. Real datacenters live between these poles:
| Facility class | Typical PUE | Comment |
|---|---|---|
| Hyperscale, cold climate, free cooling | 1.05–1.15 | Outside air does most of the work |
| Modern hyperscale, mixed climate | 1.15–1.30 | Hot aisle containment + chilled water |
| Enterprise / colo, chilled water | 1.30–1.50 | CRAC units, room-level cooling |
| Legacy raised-floor | 1.5–2.0+ | Inefficient airflow, oversized chillers |
PUE is dominated by cooling overhead, with power-conversion efficiency a distant second. Moving from air-cooled HGX to D2C-cooled NVL72 typically buys 0.10–0.15 of PUE. At hyperscale with a multi-MW deployment, that pays for the entire D2C plumbing capex within the first year.
Pick a GPU, a chassis density, a cooling approach, and an inlet temperature. The calculator estimates total heat load, the airflow or coolant flow you'd need, the thermal margin you have, and whether the configuration is feasible. All maths is inlined in the page so you can read it.
// Per-card TGP table (W)
const tgp = {
rtx4090: 450, rtx5090: 575, rtxpro6000bw: 600,
a100: 400, h100: 700, h200: 700,
b200: 1000, gb200: 2700
};
// Cooling capability per card (W) at 25 C inlet
const cap = {
axial: 450, blower: 400, passive: 700,
d2c: 1500, immersion: 3000
};
// Inlet derating (warm/hot air loses cooling capability)
const inletScale = (T) => T <= 18 ? 1.10 :
T <= 25 ? 1.00 :
0.78; // 35 C: ~22% capability loss
// Required airflow (CFM) at 15 C delta-T: 0.026 CFM per watt
const cfm = (W) => W * 0.026;
// Required water flow (GPM) at 10 C delta-T: ~0.85 GPM per kW
const gpm = (W) => (W / 1000) * 0.85;
// Facility overhead: each cooling type carries a PUE-like factor
const pue = {
axial: 1.45, blower: 1.40, passive: 1.30,
d2c: 1.12, immersion: 1.06
};
Try 8× B200 in a passive HGX chassis at 35 °C. Per-card TGP is 1000 W, effective passive capability at 35 °C inlet is roughly 546 W in this model — which is why air-cooled DGX B200 needs a 10U chassis moving ~1,550 CFM, and why GB200 NVL72 ships with mandatory D2C liquid: the engineering envelope at 1000 W per card has run off the end of what air can do, regardless of how hard you spin the fans.