Behind every frontier model is a reference platform. Walk through DGX (turnkey systems), HGX (8-GPU baseboards used by every OEM), MGX (modular reference for custom servers), OAM (industry standard), and the SuperPOD/BasePOD cluster blueprints — and learn what each actually is, who builds it, and what's inside.
NVIDIA does not just sell chips. It publishes a whole stack of reference platforms — from GPU module up to fully wired clusters — and OEMs and hyperscalers build everything around those references. This deck disentangles the names so you know exactly what you are buying.
The same six layers appear in every NVIDIA datacenter generation. NVIDIA designs everything from chip up to rack reference; OEMs (Supermicro, Dell, Lenovo, HPE, Foxconn, Wiwynn) integrate one or more of those layers into a product you can actually purchase.
Chip, module, HGX baseboard, NVSwitch tray, NVL72 rack, BasePOD/SuperPOD reference architecture, networking (Quantum-2 IB, Spectrum-X Ethernet), management software stack (Base Command, NVIDIA AI Enterprise).
Chassis, PSU, cooling loop, BMC firmware, cabling and serviceability. Some (Supermicro/Dell) build full SuperPOD-class systems. Others (Foxconn/Wiwynn) build hyperscale white-box for the cloud providers under their own SKUs.
"DGX", "HGX" and "MGX" are not three product lines for the same thing — they live at different levels of the pyramid. HGX is a baseboard. DGX is a fully assembled NVIDIA system that contains an HGX. MGX is a modular reference at the system level that competes with DGX as a starting point for OEMs.
Before you ever see "HGX" or "DGX", you need to know how a GPU is physically packaged. Datacenter GPUs come in three module form factors: NVIDIA's proprietary SXM, the standard PCIe add-in card, and the Open Compute Project's vendor-neutral OAM.
NVIDIA-proprietary mezzanine. Bolts down to a baseboard via a custom socket (Mirror Mezzanine Connector). Carries full 18 NVLink lanes and high-current power rails. SXM5 (Hopper) lands an H100 at 700 W; SXM6 (Blackwell) goes to 1000 W per module on B200. Always the top SKU.
Standard add-in card. Same chip, but limited to PCIe Gen5 x16 (~64 GB/s) for host I/O, with optional 2-card NVLink bridge (only between exactly two adjacent cards). 350–400 W typical (H100 PCIe is 350 W). Easy retrofit into existing servers.
OCP Accelerator Module. Universal mezzanine spec from the Open Compute Project — same footprint accepted by AMD MI300X, Intel Gaudi 3 and (for some SKUs) NVIDIA. Hyperscalers love it because the chassis is interchangeable across vendors.
| Module | Owner | Power envelope | Inter-GPU link | Where you see it |
|---|---|---|---|---|
| SXM5 (H100/H200) | NVIDIA | 700 W | NVLink 4 — 900 GB/s | HGX H100/H200, DGX H100/H200 |
| SXM6 (B200) | NVIDIA | 1000 W | NVLink 5 — 1.8 TB/s | HGX B200, DGX B200, GB200 NVL72 |
| PCIe (H100/H200/L40S) | PCI-SIG | 300–400 W | PCIe Gen5; optional NVLink bridge | OEM 1U/2U boxes, MGX builds |
| OAM 1.5/2.0 | OCP | up to 1000 W | vendor-defined fabric on UBB | AMD MI300X, Intel Gaudi 3 systems |
Datacenter customers prefer SXM/OAM for density — fewer cables, shared baseboard power, integrated NVLink/UBB fabric. They prefer PCIe for retrofit — you can drop one or two cards into a server you already own without redesigning the chassis. Mixing them in the same workload is rarely worth the firmware pain.
HGX is the single most-deployed building block in AI infrastructure today, and the most misunderstood. It is a baseboard, not a system. NVIDIA ships you a populated PCB the size of a pizza box; the OEM puts it in a chassis with PSUs, fans, NICs, and a host node, and sells you a server.
| HGX board | GPUs | HBM/GPU | NVLink BW/GPU | Total HBM | FP8 PFLOPS (dense) |
|---|---|---|---|---|---|
| HGX A100 | 8× A100 SXM4 | 40 GB HBM2 / 80 GB HBM2e | 600 GB/s | 320–640 GB | n/a (FP16: ~2.5) |
| HGX H100 | 8× H100 SXM5 | 80 GB HBM3 | 900 GB/s | 640 GB | ~16 |
| HGX H200 | 8× H200 SXM5 | 141 GB HBM3e | 900 GB/s | 1.13 TB | ~16 |
| HGX B200 | 8× B200 SXM6 | 192 GB HBM3e | 1.8 TB/s | 1.54 TB | ~36 |
| HGX B300 | 8× B300 SXM6 | 288 GB HBM3e | 1.8 TB/s | 2.30 TB | ~36 |
Designing 8 GPUs of NVLink fabric on a custom PCB is hard, expensive, and takes a year. By accepting HGX, the OEM inherits a validated electrical, thermal, and firmware design and only has to design the chassis around it. That is why a Supermicro HGX H200 box and a Dell PowerEdge XE9680 H200 box have radically different chassis but identical GPU performance.
DGX is the only complete, fully-assembled AI server sold by NVIDIA itself. Every other vendor's box uses HGX or MGX as a starting point. DGX is the "this is exactly what we intend, validated end-to-end" version, with NVIDIA-supplied firmware, NICs, DPUs, NVMe storage, and software stack.
| Year | System | GPUs | What was new |
|---|---|---|---|
| 2016 | DGX-1 | 8× P100, then refreshed to 8× V100 | First "deep learning supercomputer in a box"; hybrid cube-mesh NVLink |
| 2018 | DGX-2 | 16× V100 | First NVSwitch (gen 1); single 16-GPU NVLink domain |
| 2020 | DGX A100 | 8× A100 | Mellanox ConnectX-6, MIG, 5 petaFLOPS AI performance, 6× NVSwitch 2 |
| 2022 | DGX H100 | 8× H100 SXM5 | FP8 Transformer Engine, 4× NVSwitch 3, 900 GB/s/GPU |
| 2024 | DGX H200 | 8× H200 SXM5 | HBM3e (1.13 TB total) for long-context inference |
| 2024 | DGX B200 | 8× B200 SXM6 | Blackwell, 1.54 TB HBM3e, FP4 Transformer Engine v2 (MX-FP4 + NVFP4) |
| 2024 | DGX GB200 NVL72 | 72× B200 in 1 rack | Rack-scale NVLink domain (see slide 06) |
List price of a DGX H200 is ~$300–400 k vs. ~$250–300 k for an OEM HGX H200 box with the same 8 GPUs. You're paying for: NVIDIA-validated firmware (no surprises with new CUDA), full software stack (Base Command, AI Enterprise), NVIDIA-direct support (you call NVIDIA, not your OEM), and bundled NICs/DPUs/NVMe. Hyperscalers usually pass; enterprises with no GPU operations team usually take it.
Announced at COMPUTEX 2023, MGX is NVIDIA's modular reference architecture for building datacenter, edge, AI training, AI inference, HPC and enterprise servers from a small set of common building blocks. Where HGX is one fixed 8-GPU baseboard, MGX is a spec — a socket-and-board standard with many valid configurations.
| Workload | CPU | GPU | Networking | Typical OEM SKU |
|---|---|---|---|---|
| AI training | 2× Xeon / EPYC | HGX H200 / B200 | 8× CX-7 NDR | Supermicro SYS-A21GE-NBRT |
| AI inference | 1× Grace 72-core | 4× L40S or 1× B200 | 2× CX-7 / BF-3 | QCT MGX-AI |
| Data analytics | 2× EPYC | 2× H100 PCIe | 2× CX-7 | GIGABYTE G593-SD0 |
| Edge / 5G | 1× Grace 72-core | 2× L4 / L40S | BF-3 + 100 GbE | ASRock Rack 1U2N4G-Genoa |
| HPC sim | 2× Grace ARM (GH200) | built-in Hopper | NDR + Slingshot | HPE Cray Supercomputing EX |
DGX = NVIDIA-built turnkey server. HGX = NVIDIA's fixed 8-GPU baseboard for OEM training boxes. MGX = NVIDIA's flexible chassis spec from which OEMs build many SKUs (training, inference, edge, enterprise).
NVL72 is the most aggressive integration NVIDIA has ever shipped: a single 19-inch rack containing 72 Blackwell GPUs in one NVLink domain. Designed for trillion-parameter MoE training and inference at densities no air-cooled box can reach.
72 GPUs that share one NVLink domain look to software like one big GPU with 13.4 TB of unified HBM. Tensor parallel and expert parallel collectives stay on NVLink (1.8 TB/s per GPU, 900 GB/s each direction) instead of stepping down to InfiniBand (50 GB/s). For a trillion-parameter MoE this is the difference between profitable and pointless.
72× B200 at 1000 W is 72 kW just for GPUs; with Grace, NICs, switches, and PSU losses you reach ~120 kW per rack. Air cooling caps out near ~40 kW/rack in most datacenters, so direct-to-chip warm-water liquid (typically ~32°C inlet) is mandatory. NVIDIA publishes manifold and CDU references; OEMs build the loop.
BasePOD is NVIDIA's reference architecture for the 2–32 node enterprise cluster: pre-validated topology, storage, and software, with a small enough footprint that a single corporate datacenter can host it. Think of it as "DGX SuperPOD lite" for non-hyperscalers.
You can absolutely buy 8 DGX nodes and wire them up yourself with arbitrary IB switches and storage. BasePOD just pre-answers all the design questions (cabling diagrams, switch firmware, NCCL tuning, MOFED versions, storage QoS) so the cluster works on day one and NVIDIA support won't blame your topology when a job hangs.
SuperPOD is NVIDIA's reference for clusters of 32 nodes and up, with a published "scalable unit" (SU) design that can be tiled to thousands of GPUs. This is the architecture behind Meta, Microsoft, OpenAI, Anthropic, xAI, and the major sovereign-AI builds.
| Generation | Node | SU size | Max SUs / SuperPOD | GPU count |
|---|---|---|---|---|
| SuperPOD A100 | DGX A100 (8 GPU) | 20 nodes (160 GPUs) | 7 | up to 1120 A100 |
| SuperPOD H100 | DGX H100 (8 GPU) | 32 nodes (256 GPUs) | 4 | up to 1024 H100 |
| SuperPOD H200 | DGX H200 (8 GPU) | 32 nodes (256 GPUs) | 4 | up to 1024 H200 |
| SuperPOD GB200 | GB200 NVL72 rack (72 GPU) | 8 racks (576 GPUs) | 4 | up to 2304 B200 |
| SuperPOD GB300 | GB300 NVL72 rack (72 GPU) | 8 racks (576 GPUs) | up to 8 | up to 4608 B300 |
You could redesign every aspect of a SuperPOD — pick your own switches, your own storage, your own scheduler. Hyperscalers do this all the time. But for everyone else, the reference saves 12–18 months of engineering, ships with NVIDIA support, and matches the configuration their published benchmarks were measured on. That last point matters: if your cluster doesn't look like the reference, "this MoE should train in N hours" stops being a useful prediction.
NVIDIA has wrapped the DGX brand around two adjacent products that are not boxes you buy and rack: a cloud-native DGX-as-a-service, and a desk-side DGX Spark. Both share the same CUDA software stack as a real DGX SuperPOD — the value proposition is "develop here, run there".
Bare-metal access to NVIDIA-validated DGX clusters, hosted in OCI, Azure, GCP and (separately) AWS. NVIDIA designs the cluster to SuperPOD spec, the cloud provider operates the datacenter, the customer pays NVIDIA for time on the cluster.
The personal AI workstation announced as Project DIGITS in early 2025 and shipping under the DGX Spark name. 1× Grace + 1× Blackwell on a single GB10 superchip, 128 GB unified LPDDR5x, ~1 PFLOPS at FP4, ~$3–5 k.
The point of DGX Spark is not raw FLOPS — an RTX 6000 Ada beats it on dense compute — but memory capacity and identical software stack. Develop a NeMo training script on a Spark, scale it 1:1 onto a DGX Cloud H200 cluster, ship the resulting model to a customer running DGX B200 on-prem. Same CUDA, same NCCL, same Base Command, same PyTorch — just bigger numbers.
NVIDIA does not ship enough DGX boxes to feed the market. The vast majority of HGX baseboards go to a handful of OEMs and ODMs who build chassis around them. Here's who ships what in 2026.
| Vendor | What they ship | Position |
|---|---|---|
| Supermicro | HGX H200 / B200 systems (SYS-A21GE), MGX builds across many SKUs, GB200 NVL72 racks | Highest volume Tier-1; aggressive on time-to-market |
| Dell | PowerEdge XE9680 (HGX H100/H200), XE9712 (GB200 NVL72) | Enterprise sales channel; bundled with PowerScale storage |
| Lenovo | ThinkSystem SR685a V3 (HGX H200/B200), SR780a V3 (GB200 NVL72) | Strong in EMEA enterprise and HPC |
| HPE | Cray EX with H100/H200 for supercomputers, ProLiant XD685 (HGX), Compute XD (GB200) | HPC heritage, exascale system builds |
| Foxconn | Hyperscale HGX/MGX white-box for cloud providers under their own SKUs | ODM; not retail-facing, but largest unit shipper |
| Wiwynn | Hyperscale HGX/MGX white-box (Meta, Microsoft) | ODM; deep relationships with hyperscale buyers |
| GIGABYTE | G593 series (HGX H200/B200), MGX-based AI inference SKUs | Mid-tier, price-aggressive |
| ASUS | ESC N8/N16 (HGX), MGX-based AI training and inference servers | Mid-tier OEM, strong APAC presence |
| ASRock Rack | MGX 1U/2U inference and edge SKUs, mid-tier HGX | Niche price-aggressive builder |
| QCT (Quanta) | MGX-based AI inference, edge, and analytics SKUs; ODM hyperscale | Cross-over: own brand + ODM contracts |
NVIDIA captures the vast majority of system value via GPU and HGX board pricing — gross margins above 75% on the silicon. The OEM adds a single-digit-to-low-teens margin for chassis, integration, validation and support. The SI / VAR layer adds another 5–15% for deployment and ongoing services. A "$400 k DGX-class" system is approximately $300 k of NVIDIA content, $50–70 k of OEM content, and the rest in services and support contracts.
Pick a workload, GPU class, and node count. The picker recommends the most appropriate NVIDIA reference platform, sums aggregate VRAM and FP8 PFLOPS, sizes the NVLink domain, and gives a rough $/node figure to anchor expectations.
The picker is opinionated: pretraining frontier models always points at GB200 NVL72; production 70B inference points at HGX OEM boxes; edge inference points at L40S- or Spark-class kit. The aggregate FP8 PFLOPS (sparse, as NVIDIA quotes them) and VRAM are theoretical peaks — expect 50–70% sustained on real workloads. The $/node is mid-2026 list-equivalent, not your discounted contract price.