NVIDIA GPU Architectures Series — Presentation 11

DGX, HGX, MGX — NVIDIA's Datacenter Reference Platforms

Behind every frontier model is a reference platform. Walk through DGX (turnkey systems), HGX (8-GPU baseboards used by every OEM), MGX (modular reference for custom servers), OAM (industry standard), and the SuperPOD/BasePOD cluster blueprints — and learn what each actually is, who builds it, and what's inside.

DGXHGXMGX OAMSuperPODBasePOD DGX CloudOEMs NVL72Reference Architecture
GPU → HGX baseboard → DGX/OEM → BasePOD → SuperPOD → NVL72 → Cloud
00

Topics We'll Cover

NVIDIA does not just sell chips. It publishes a whole stack of reference platforms — from GPU module up to fully wired clusters — and OEMs and hyperscalers build everything around those references. This deck disentangles the names so you know exactly what you are buying.

01

The Reference-Platform Pyramid

The same six layers appear in every NVIDIA datacenter generation. NVIDIA designs everything from chip up to rack reference; OEMs (Supermicro, Dell, Lenovo, HPE, Foxconn, Wiwynn) integrate one or more of those layers into a product you can actually purchase.

Cluster — SuperPOD Rack — NVL72 / BasePOD rack System — DGX / OEM server Baseboard — HGX (8 GPUs + 4 NVSwitches) Module — SXM / PCIe / OAM / MGX socket Chip — H100 / H200 / B100 / B200 / GB200 / GB10 100s–1000s GPUs 8–72 GPUs 1–8 GPUs/box 8 GPUs + fabric 1 GPU package Silicon die

NVIDIA designs

Chip, module, HGX baseboard, NVSwitch tray, NVL72 rack, BasePOD/SuperPOD reference architecture, networking (Quantum-2 IB, Spectrum-X Ethernet), management software stack (Base Command, NVIDIA AI Enterprise).

OEMs integrate

Chassis, PSU, cooling loop, BMC firmware, cabling and serviceability. Some (Supermicro/Dell) build full SuperPOD-class systems. Others (Foxconn/Wiwynn) build hyperscale white-box for the cloud providers under their own SKUs.

Key idea

"DGX", "HGX" and "MGX" are not three product lines for the same thing — they live at different levels of the pyramid. HGX is a baseboard. DGX is a fully assembled NVIDIA system that contains an HGX. MGX is a modular reference at the system level that competes with DGX as a starting point for OEMs.

02

SXM, PCIe and OAM Modules

Before you ever see "HGX" or "DGX", you need to know how a GPU is physically packaged. Datacenter GPUs come in three module form factors: NVIDIA's proprietary SXM, the standard PCIe add-in card, and the Open Compute Project's vendor-neutral OAM.

SXM

NVIDIA-proprietary mezzanine. Bolts down to a baseboard via a custom socket (Mirror Mezzanine Connector). Carries full 18 NVLink lanes and high-current power rails. SXM5 (Hopper) lands an H100 at 700 W; SXM6 (Blackwell) goes to 1000 W per module on B200. Always the top SKU.

  • NVLink-rich: 900 GB/s/GPU on H100, 1.8 TB/s on B200
  • Liquid- or air-cooled cold plate
  • Only sold as 8-pack on HGX baseboards

PCIe

Standard add-in card. Same chip, but limited to PCIe Gen5 x16 (~64 GB/s) for host I/O, with optional 2-card NVLink bridge (only between exactly two adjacent cards). 350–400 W typical (H100 PCIe is 350 W). Easy retrofit into existing servers.

  • Drop-in for any PCIe x16 slot with auxiliary power
  • NVLink bridge gives 600 GB/s peer-only between 2 cards
  • Lower clocks/TDP than SXM equivalents

OAM

OCP Accelerator Module. Universal mezzanine spec from the Open Compute Project — same footprint accepted by AMD MI300X, Intel Gaudi 3 and (for some SKUs) NVIDIA. Hyperscalers love it because the chassis is interchangeable across vendors.

  • Vendor-neutral pinout
  • Used by AMD/Intel for primary parts; NVIDIA prefers SXM
  • UBB (Universal Baseboard) sister spec defines the 8-OAM board
ModuleOwnerPower envelopeInter-GPU linkWhere you see it
SXM5 (H100/H200)NVIDIA700 WNVLink 4 — 900 GB/sHGX H100/H200, DGX H100/H200
SXM6 (B200)NVIDIA1000 WNVLink 5 — 1.8 TB/sHGX B200, DGX B200, GB200 NVL72
PCIe (H100/H200/L40S)PCI-SIG300–400 WPCIe Gen5; optional NVLink bridgeOEM 1U/2U boxes, MGX builds
OAM 1.5/2.0OCPup to 1000 Wvendor-defined fabric on UBBAMD MI300X, Intel Gaudi 3 systems
Density vs. flexibility

Datacenter customers prefer SXM/OAM for density — fewer cables, shared baseboard power, integrated NVLink/UBB fabric. They prefer PCIe for retrofit — you can drop one or two cards into a server you already own without redesigning the chassis. Mixing them in the same workload is rarely worth the firmware pain.

03

HGX — The 8-GPU Baseboard

HGX is the single most-deployed building block in AI infrastructure today, and the most misunderstood. It is a baseboard, not a system. NVIDIA ships you a populated PCB the size of a pizza box; the OEM puts it in a chassis with PSUs, fans, NICs, and a host node, and sells you a server.

What's on an HGX H100 baseboard

HGX H100 baseboard (8× SXM5 + 4× NVSwitch) H100 #080 GB HBM3 H100 #180 GB HBM3 H100 #280 GB HBM3 H100 #380 GB HBM3 H100 #480 GB HBM3 H100 #580 GB HBM3 H100 #680 GB HBM3 H100 #780 GB HBM3 NVSw 0 NVSw 1 NVSw 2 NVSw 3

HGX variants — same socket, different generations

HGX boardGPUsHBM/GPUNVLink BW/GPUTotal HBMFP8 PFLOPS (dense)
HGX A1008× A100 SXM440 GB HBM2 / 80 GB HBM2e600 GB/s320–640 GBn/a (FP16: ~2.5)
HGX H1008× H100 SXM580 GB HBM3900 GB/s640 GB~16
HGX H2008× H200 SXM5141 GB HBM3e900 GB/s1.13 TB~16
HGX B2008× B200 SXM6192 GB HBM3e1.8 TB/s1.54 TB~36
HGX B3008× B300 SXM6288 GB HBM3e1.8 TB/s2.30 TB~36
Why every OEM uses HGX

Designing 8 GPUs of NVLink fabric on a custom PCB is hard, expensive, and takes a year. By accepting HGX, the OEM inherits a validated electrical, thermal, and firmware design and only has to design the chassis around it. That is why a Supermicro HGX H200 box and a Dell PowerEdge XE9680 H200 box have radically different chassis but identical GPU performance.

04

DGX — NVIDIA's Reference Server

DGX is the only complete, fully-assembled AI server sold by NVIDIA itself. Every other vendor's box uses HGX or MGX as a starting point. DGX is the "this is exactly what we intend, validated end-to-end" version, with NVIDIA-supplied firmware, NICs, DPUs, NVMe storage, and software stack.

The DGX lineage

YearSystemGPUsWhat was new
2016DGX-18× P100, then refreshed to 8× V100First "deep learning supercomputer in a box"; hybrid cube-mesh NVLink
2018DGX-216× V100First NVSwitch (gen 1); single 16-GPU NVLink domain
2020DGX A1008× A100Mellanox ConnectX-6, MIG, 5 petaFLOPS AI performance, 6× NVSwitch 2
2022DGX H1008× H100 SXM5FP8 Transformer Engine, 4× NVSwitch 3, 900 GB/s/GPU
2024DGX H2008× H200 SXM5HBM3e (1.13 TB total) for long-context inference
2024DGX B2008× B200 SXM6Blackwell, 1.54 TB HBM3e, FP4 Transformer Engine v2 (MX-FP4 + NVFP4)
2024DGX GB200 NVL7272× B200 in 1 rackRack-scale NVLink domain (see slide 06)

What you actually get in a DGX H200

Why pay the DGX premium

List price of a DGX H200 is ~$300–400 k vs. ~$250–300 k for an OEM HGX H200 box with the same 8 GPUs. You're paying for: NVIDIA-validated firmware (no surprises with new CUDA), full software stack (Base Command, AI Enterprise), NVIDIA-direct support (you call NVIDIA, not your OEM), and bundled NICs/DPUs/NVMe. Hyperscalers usually pass; enterprises with no GPU operations team usually take it.

05

MGX — Modular Reference Architecture

Announced at COMPUTEX 2023, MGX is NVIDIA's modular reference architecture for building datacenter, edge, AI training, AI inference, HPC and enterprise servers from a small set of common building blocks. Where HGX is one fixed 8-GPU baseboard, MGX is a spec — a socket-and-board standard with many valid configurations.

What MGX standardises

  • Chassis dimensions (1U, 2U, 4U variants)
  • Power and cooling envelopes
  • CPU socket area — Grace ARM or x86 (Intel/AMD)
  • GPU module bays — SXM, PCIe, or MGX-socket Blackwell
  • NIC/DPU slots — ConnectX-7/8 or BlueField-3
  • NVLink-C2C and NVLink-Switch cabling between trays

Why OEMs adopt it

  • Faster TTM — ~6 months to ship a new MGX system vs. ~18 months for a from-scratch design
  • Multiple SKUs from one chassis — same box, swap CPU+GPU combos
  • Validated thermal/power profiles — NVIDIA pre-tested the worst case
  • Ecosystem accessory market — NICs, storage trays, cable harnesses are interchangeable

MGX flavours — the same chassis ships as many SKUs

WorkloadCPUGPUNetworkingTypical OEM SKU
AI training2× Xeon / EPYCHGX H200 / B2008× CX-7 NDRSupermicro SYS-A21GE-NBRT
AI inference1× Grace 72-core4× L40S or 1× B2002× CX-7 / BF-3QCT MGX-AI
Data analytics2× EPYC2× H100 PCIe2× CX-7GIGABYTE G593-SD0
Edge / 5G1× Grace 72-core2× L4 / L40SBF-3 + 100 GbEASRock Rack 1U2N4G-Genoa
HPC sim2× Grace ARM (GH200)built-in HopperNDR + SlingshotHPE Cray Supercomputing EX
DGX vs. HGX vs. MGX in one line

DGX = NVIDIA-built turnkey server. HGX = NVIDIA's fixed 8-GPU baseboard for OEM training boxes. MGX = NVIDIA's flexible chassis spec from which OEMs build many SKUs (training, inference, edge, enterprise).

06

GB200 NVL72 — The Rack Is the Computer

NVL72 is the most aggressive integration NVIDIA has ever shipped: a single 19-inch rack containing 72 Blackwell GPUs in one NVLink domain. Designed for trillion-parameter MoE training and inference at densities no air-cooled box can reach.

GB200 NVL72 rack Compute tray (2× Grace + 4× B200) Compute tray NVSwitch tray Compute tray Compute tray NVSwitch tray Compute tray Compute tray NVSwitch tray Compute tray Compute tray NVSwitch tray Compute tray Compute tray NVSwitch tray Compute tray Compute tray NVSwitch tray 18× compute trays each: 2× Grace + 4× B200 total: 36 Grace + 72 B200 9× NVLink-Switch trays each: 2× NVSwitch 4 ASIC 100% non-blocking 72-way mesh Rack totals 13.4 TB unified HBM3e (72×186 GB) 130 TB/s aggregate NVLink ~720 PFLOPS FP8 / 1.4 EFLOPS FP4 (sparse) ~120 kW per rack, liquid cooled External InfiniBand spine to other NVL72s 8× CX-7 / NIC per tray for IB

Why one NVLink domain matters

72 GPUs that share one NVLink domain look to software like one big GPU with 13.4 TB of unified HBM. Tensor parallel and expert parallel collectives stay on NVLink (1.8 TB/s per GPU, 900 GB/s each direction) instead of stepping down to InfiniBand (50 GB/s). For a trillion-parameter MoE this is the difference between profitable and pointless.

Why liquid cooling

72× B200 at 1000 W is 72 kW just for GPUs; with Grace, NICs, switches, and PSU losses you reach ~120 kW per rack. Air cooling caps out near ~40 kW/rack in most datacenters, so direct-to-chip warm-water liquid (typically ~32°C inlet) is mandatory. NVIDIA publishes manifold and CDU references; OEMs build the loop.

07

BasePOD — The Smaller Cluster Reference

BasePOD is NVIDIA's reference architecture for the 2–32 node enterprise cluster: pre-validated topology, storage, and software, with a small enough footprint that a single corporate datacenter can host it. Think of it as "DGX SuperPOD lite" for non-hyperscalers.

The reference shipping list

Use cases

  • Enterprise AI labs (banks, pharma, energy) needing DGX-grade reliability without SuperPOD scale
  • University faculty clusters
  • Sovereign-AI proof-of-concepts before scaling to full SuperPOD
  • Pre-production environments mirroring SuperPOD topology at smaller scale

Sizing rule of thumb

  • 4 nodes (32 GPUs) — team of ~50 ML engineers, fine-tuning and serving
  • 8 nodes (64 GPUs) — small frontier-model post-training
  • 16 nodes (128 GPUs) — production-grade training of mid-scale models (~30B)
  • 32 nodes (256 GPUs) — the inflection point: above this, jump to SuperPOD
BasePOD vs. random clusters

You can absolutely buy 8 DGX nodes and wire them up yourself with arbitrary IB switches and storage. BasePOD just pre-answers all the design questions (cabling diagrams, switch firmware, NCCL tuning, MOFED versions, storage QoS) so the cluster works on day one and NVIDIA support won't blame your topology when a job hangs.

08

SuperPOD — Hyperscale Cluster Blueprint

SuperPOD is NVIDIA's reference for clusters of 32 nodes and up, with a published "scalable unit" (SU) design that can be tiled to thousands of GPUs. This is the architecture behind Meta, Microsoft, OpenAI, Anthropic, xAI, and the major sovereign-AI builds.

Generational layouts

GenerationNodeSU sizeMax SUs / SuperPODGPU count
SuperPOD A100DGX A100 (8 GPU)20 nodes (160 GPUs)7up to 1120 A100
SuperPOD H100DGX H100 (8 GPU)32 nodes (256 GPUs)4up to 1024 H100
SuperPOD H200DGX H200 (8 GPU)32 nodes (256 GPUs)4up to 1024 H200
SuperPOD GB200GB200 NVL72 rack (72 GPU)8 racks (576 GPUs)4up to 2304 B200
SuperPOD GB300GB300 NVL72 rack (72 GPU)8 racks (576 GPUs)up to 8up to 4608 B300

What makes a SuperPOD a SuperPOD (and not just "lots of GPUs")

Why customers still use the reference

You could redesign every aspect of a SuperPOD — pick your own switches, your own storage, your own scheduler. Hyperscalers do this all the time. But for everyone else, the reference saves 12–18 months of engineering, ships with NVIDIA support, and matches the configuration their published benchmarks were measured on. That last point matters: if your cluster doesn't look like the reference, "this MoE should train in N hours" stops being a useful prediction.

09

DGX Cloud, DGX Spark — Same Stack, New Form Factors

NVIDIA has wrapped the DGX brand around two adjacent products that are not boxes you buy and rack: a cloud-native DGX-as-a-service, and a desk-side DGX Spark. Both share the same CUDA software stack as a real DGX SuperPOD — the value proposition is "develop here, run there".

DGX Cloud

Bare-metal access to NVIDIA-validated DGX clusters, hosted in OCI, Azure, GCP and (separately) AWS. NVIDIA designs the cluster to SuperPOD spec, the cloud provider operates the datacenter, the customer pays NVIDIA for time on the cluster.

  • Multi-node training with NVLink-rich H100/H200/B200 nodes
  • Bring your own model and data; NVIDIA provides Base Command, AI Enterprise, NeMo
  • Marketed at frontier-model startups and Fortune-500 AI teams that need elastic training capacity
  • Differs from generic cloud GPU instances: dedicated nodes, RDMA fabric, validated software

DGX Spark on desks now

The personal AI workstation announced as Project DIGITS in early 2025 and shipping under the DGX Spark name. 1× Grace + 1× Blackwell on a single GB10 superchip, 128 GB unified LPDDR5x, ~1 PFLOPS at FP4, ~$3–5 k.

  • Bridges hobbyist and enterprise stack — identical CUDA/Triton/NeMo software as DGX H200
  • Holds 70B-class FP8 models or 200B-class FP4 models in unified memory
  • Not a training workstation: bandwidth is LPDDR5x not HBM, ~273 GB/s aggregate
  • Two units chained over ConnectX provide 256 GB unified for 200B-class models at home
The continuum

The point of DGX Spark is not raw FLOPS — an RTX 6000 Ada beats it on dense compute — but memory capacity and identical software stack. Develop a NeMo training script on a Spark, scale it 1:1 onto a DGX Cloud H200 cluster, ship the resulting model to a customer running DGX B200 on-prem. Same CUDA, same NCCL, same Base Command, same PyTorch — just bigger numbers.

10

OEM Ecosystem — Who Actually Builds These

NVIDIA does not ship enough DGX boxes to feed the market. The vast majority of HGX baseboards go to a handful of OEMs and ODMs who build chassis around them. Here's who ships what in 2026.

VendorWhat they shipPosition
SupermicroHGX H200 / B200 systems (SYS-A21GE), MGX builds across many SKUs, GB200 NVL72 racksHighest volume Tier-1; aggressive on time-to-market
DellPowerEdge XE9680 (HGX H100/H200), XE9712 (GB200 NVL72)Enterprise sales channel; bundled with PowerScale storage
LenovoThinkSystem SR685a V3 (HGX H200/B200), SR780a V3 (GB200 NVL72)Strong in EMEA enterprise and HPC
HPECray EX with H100/H200 for supercomputers, ProLiant XD685 (HGX), Compute XD (GB200)HPC heritage, exascale system builds
FoxconnHyperscale HGX/MGX white-box for cloud providers under their own SKUsODM; not retail-facing, but largest unit shipper
WiwynnHyperscale HGX/MGX white-box (Meta, Microsoft)ODM; deep relationships with hyperscale buyers
GIGABYTEG593 series (HGX H200/B200), MGX-based AI inference SKUsMid-tier, price-aggressive
ASUSESC N8/N16 (HGX), MGX-based AI training and inference serversMid-tier OEM, strong APAC presence
ASRock RackMGX 1U/2U inference and edge SKUs, mid-tier HGXNiche price-aggressive builder
QCT (Quanta)MGX-based AI inference, edge, and analytics SKUs; ODM hyperscaleCross-over: own brand + ODM contracts

The pricing flow

NVIDIA designs & sells GPUs / HGX / MGX spec OEM / ODM chassis, PSU, cooling, BMC, validation Distributor / SI deployment, finance, support contracts Customer enterprise, hyperscaler, sovereign AI → → → $ to NVIDIA + OEM margin + SI margin / svcs final price
Where the margin actually lives

NVIDIA captures the vast majority of system value via GPU and HGX board pricing — gross margins above 75% on the silicon. The OEM adds a single-digit-to-low-teens margin for chassis, integration, validation and support. The SI / VAR layer adds another 5–15% for deployment and ongoing services. A "$400 k DGX-class" system is approximately $300 k of NVIDIA content, $50–70 k of OEM content, and the rest in services and support contracts.

11

Interactive: Pick Your Platform

Pick a workload, GPU class, and node count. The picker recommends the most appropriate NVIDIA reference platform, sums aggregate VRAM and FP8 PFLOPS, sizes the NVLink domain, and gives a rough $/node figure to anchor expectations.

Platform
—
Aggregate VRAM
—
Aggregate FP8
—
NVLink domain
—
$ / node
—
Notes
—
Verdict: Choose options above and press Compute.
How to read the recommendation

The picker is opinionated: pretraining frontier models always points at GB200 NVL72; production 70B inference points at HGX OEM boxes; edge inference points at L40S- or Spark-class kit. The aggregate FP8 PFLOPS (sparse, as NVIDIA quotes them) and VRAM are theoretical peaks — expect 50–70% sustained on real workloads. The $/node is mid-2026 list-equivalent, not your discounted contract price.