NVIDIA GPU Architectures Series — Presentation 13

NVIDIA Networking — InfiniBand, ConnectX, BlueField, and the Cross-Node Fabric

Cross-rack tensor parallelism, multi-trillion-parameter MoE, distributed checkpointing, and inference at fleet scale all require NVIDIA's networking stack — InfiniBand or Ethernet, ConnectX HCAs, BlueField DPUs, Quantum and Spectrum switches, GPUDirect RDMA and SHARP. Here's what each is and how they stack up.

InfiniBandNDRXDR ConnectX-7ConnectX-8BlueField-3 Quantum-2Spectrum-X GPUDirectSHARPRoCE
HCA → Switch → Topology → GPUDirect → SHARP → DPU → Storage
00

Topics We'll Cover

The fabric is the second computer. Once a model leaves a single NVL72, every all-reduce becomes a network problem — this deck walks through the components that make it tractable.

01

Why Networking Matters for LLMs

Inside a single NVL72 or DGX node, NVLink and NVSwitch carry the load. The moment a model spills across the rack — pipeline stages, data-parallel replicas, expert shards in an MoE — every gradient step has to cross a slower fabric, and that fabric becomes the bottleneck.

Training

An 8-GPU tensor-parallel group fits inside one NVLink domain — cheap. Pipeline parallel and data parallel hop nodes — expensive. For a 405B Llama-class model, each layer's all-reduce is roughly 1–2 GB, and with ~100 layers and several reductions per step the fabric throughput dictates how often a GPU sits idle waiting for gradients.

Inference

KV-cache shipping for distributed serving (prefill on one node, decode on another), model weight loading from PFS at fleet startup, telemetry and tracing, request routing — all of it crosses the fabric. At 100k-GPU scale even "background" traffic is several Tbps.

How NVIDIA got here

NVIDIA didn't build this stack from scratch. In April 2020 they completed the acquisition of Mellanox for ~$7B, picking up the dominant InfiniBand vendor along with ConnectX HCAs, the Quantum/Spectrum switch lines, and the BlueField SmartNIC roadmap. Six years on, that one acquisition is arguably as strategic as any GPU generation — without it the AI factory pitch has no plumbing.

The framing for this deck

NVLink is for the rack. Networking is for everything beyond it. The rest of the deck walks the stack from NIC to switch to topology to software offloads, with the numbers that matter.

02

InfiniBand vs Ethernet — The Trade-off

NVIDIA sells both. Frontier AI clusters have historically chosen InfiniBand; hyperscalers increasingly push Ethernet because it's the fabric they already operate everything else on. Both can deliver an LLM training workload — the trade-offs are around lossless behaviour, latency, vendor lock-in, and price.

InfiniBand

  • Lossless by design via credit-based flow control — a sender never transmits without buffer credits at the receiver.
  • Microsecond latency end-to-end (single-digit µs through switches).
  • RDMA native — verbs, queue pairs, memory regions are all first-class.
  • Single-vendor — NVIDIA/Mellanox has the only credible high-end stack.
  • More expensive per port, especially the cables and switches.

Ethernet (with RoCE)

  • Ubiquitous, multi-vendor, runs the rest of the datacenter.
  • RDMA over Converged Ethernet (RoCE v2) bolts RDMA semantics onto UDP/IP.
  • Lossless requires tuning — PFC (Priority Flow Control) + ECN with careful queue sizing and watchdogs.
  • Slightly higher latency, more jitter.
  • Cheaper switches and cables, plus existing operations skill.

Where each lands in 2026

Lossless ≠ loss-free

Both fabrics are functionally lossless on a healthy day. The difference is what happens when something goes wrong. InfiniBand's credit scheme prevents drops by construction. Ethernet's PFC reacts to congestion and can deadlock or trigger pause storms if you've misconfigured a single switch. Spectrum-X's whole pitch is offloading the hard parts of that tuning to BlueField.

03

ConnectX HCA Lineage

The Host Channel Adapter (HCA) is the NIC: one PCIe card per server, or in AI clusters one per GPU. ConnectX is the Mellanox/NVIDIA brand. Generation tracks the IB speed standard.

GenerationYearPer-port speedStandardUse case / platform
ConnectX-52017100 Gb/sEDR / HDR100Pre-Ampere clusters, late P100 era.
ConnectX-62020200 Gb/sHDRDGX A100 baseline (8× CX-6 per node).
ConnectX-6 Dx2020200 Gb/sHDR / 200GbECloud + telco; security and TLS offload.
ConnectX-72022400 Gb/sNDRDGX H100 / H200 (8× CX-7 per node).
ConnectX-82024800 Gb/sXDRDGX B300 / GB300 NVL72 (DGX B200 still uses 8× CX-7).

The 1:1:1 rule

1 GPU ↔ 1 HCA ↔ 1 IB cable → leaf switch

For AI training clusters the rule is one dedicated HCA per GPU. A DGX H100 has 8 H100s and 8 ConnectX-7 cards; a DGX B200 has 8 B200s and 8 ConnectX-7 cards (ConnectX-8 arrives with B300). Plus there are typically additional ConnectX or BlueField NICs for the storage fabric and management. The compute fabric is sized so that a GPU never waits on its NIC.

Form factors

Why per-GPU NICs

Sharing one NIC across multiple GPUs adds head-of-line blocking and PCIe contention. Per-GPU NICs let NCCL's ring or tree algorithm send data along the shortest possible path and let GPUDirect RDMA work without crossing the host's main PCIe complex.

04

Quantum InfiniBand Switches

Quantum is NVIDIA's InfiniBand switch line. Each generation tracks the IB speed standard and roughly doubles per-port bandwidth and switch radix.

SwitchYearSpeedPortsAnchor cluster
Quantum (QM8700)2018HDR 200G40-portDGX A100 SuperPODs.
Quantum-2 (QM9700)2021NDR 400G64-portDGX H100 SuperPOD — the workhorse for Hopper-era clusters.
Quantum-X800 (Q3400)2024XDR 800G144-port effective (multi-host)DGX B200 / GB200 NVL72 SuperPODs.

Topologies you'll meet

Fat-tree (Clos)

The canonical AI topology. Two or three layers of switches: leaf, spine, and (for the largest pods) super-spine. Non-blocking 1:1 oversubscription means any GPU can talk to any other GPU at full line rate. Standard for DGX SuperPOD up to roughly 1k–4k GPUs.

  • Predictable bandwidth.
  • Many ECMP paths → adaptive routing thrives.
  • Cable count grows fast: ~1.5× the leaf-port count at full bisection.

Dragonfly+

Used in very large clusters (10k+ GPUs). Switches grouped into "groups", with all-to-all inside a group and selective links between groups. Lower cable count than fat-tree, but routing is harder and worst-case hop count is higher.

  • Cheaper at scale.
  • Needs adaptive routing to avoid hot links.
  • Used in some of the largest national-lab IB fabrics.

Adaptive routing

Quantum-2 and Quantum-X800 both ship with hardware adaptive routing: the switch picks the least-congested path per packet rather than a static hash. Combined with packet-level reordering at the receiver HCA, this avoids the "ECMP hash collision" pathology that plagues vanilla Ethernet at AI scale.

Quantum-X800 in one line

Doubles the per-port speed (NDR 400 → XDR 800), more than doubles the effective radix, and was designed alongside the GB200 NVL72 platform — one switch tier can carry the entire NVL72 to NVL72 traffic in a multi-rack pod.

05

Spectrum & Spectrum-X — Ethernet for AI

Spectrum is NVIDIA's Ethernet switch line (the Mellanox-origin SN-series). Spectrum-X is the AI-specific reference architecture: Spectrum switches plus ConnectX/BlueField NICs plus a software stack that delivers InfiniBand-class behaviour over Ethernet.

Spectrum (vanilla)

Conventional Ethernet switches: SN3000, SN4000, SN5000 series at 200/400/800 GbE. Used for general datacenter networking, storage, and management traffic. Same silicon family as the AI variant; what differs is how you operate it.

Spectrum-X (AI-tuned stack)

The combination that makes Ethernet credibly carry training traffic:

  • Spectrum-4 / Spectrum-5 switches with deep buffers and adaptive routing.
  • ConnectX-7 / ConnectX-8 or BlueField-3 NICs that participate in congestion control, packet reordering, and telemetry.
  • RoCE v2 with PFC + DCQCN tuned by NVIDIA.
  • End-to-end congestion control offloaded to BlueField, freeing the host CPU.

Where it has been deployed

Why operators want it

One fabric for AI plus general-purpose plus storage means one team, one toolchain, one set of cables, and inter-op with everything else in the datacenter. The cost is some single-digit-percent of throughput vs InfiniBand and a more complex tuning story — offset by sticking with the network they already run.

Ultra Ethernet on the horizon

The Ultra Ethernet Consortium (AMD, Broadcom, Cisco, Meta, Microsoft, plus dozens of others) is standardising AI Ethernet semantics — new transport, packet spraying, in-network telemetry. NVIDIA participates and is broadly compatible. Spectrum-X is NVIDIA's pre-standard answer; Ultra Ethernet is the multi-vendor convergence point. Expect them to merge over a few years.

06

GPUDirect — Bypassing the CPU

The fastest data path is the one that doesn't touch the host CPU. GPUDirect is the umbrella term for three closely related technologies that let the GPU and the NIC (or another GPU, or NVMe) talk to each other directly.

GPUDirect P2P

GPU-to-GPU within a node across PCIe or NVLink without staging in host RAM. cudaMemcpyPeer or NCCL P2P transport. On NVLink it's effectively free; on PCIe it depends on IOMMU and ACS configuration.

GPUDirect RDMA

GPU-to-GPU across nodes. The remote NIC reads or writes GPU VRAM directly via DMA. The host CPU is not on the fast path. Latency drops from ~50 µs (host bounce) to ~6 µs (direct). Required for any credible cross-node TP / PP.

GPUDirect Storage (GDS)

NVMe and PFS reads land directly in GPU VRAM via the cuFile API. Bypasses the host page cache entirely. Critical for fast checkpoint loading, model warm-start, and dataset streaming at scale.

The data path, with and without RDMA

Without RDMA: GPU → PCIe → host RAM → CPU memcpy → PCIe → NIC → wire ↓ With RDMA: GPU → PCIe → NIC → wire

Enabling it

Why this matters quantitatively

A 2 GB all-reduce on a 100k-GPU cluster touches the network billions of times per training run. Saving 40 µs per hop by removing the host bounce is the difference between a step that finishes in 80 ms and one that finishes in 130 ms — a 60% throughput delta on a multi-month run.

07

SHARP — In-Network Reductions

Scalable Hierarchical Aggregation and Reduction Protocol. The idea is simple and powerful: move the all-reduce into the switch instead of doing it at the endpoints.

How a normal all-reduce works

GPU 0 ↔ GPU 1 ↔ GPU 2 ↔ GPU 3 ↔ …

Ring or tree algorithm: each GPU sends a chunk, receives a chunk, sums, forwards. With N GPUs an all-reduce moves roughly 2×(N−1)/N of the tensor across the wire. As N grows the bandwidth wall dominates.

How SHARP works

GPUs 0..N each push tensor chunk ↓ Quantum switch sums across ports in hardware ↓ Switch broadcasts the result back to every GPU

The reduction itself happens in the switch silicon. Each GPU sends its data once and receives the result once — halving the bandwidth needed versus a ring all-reduce, and reducing the latency at scale because it's a single hierarchical step rather than O(N) sequential exchanges.

Versions and integration

When SHARP changes the answer

Below ~256 GPUs SHARP is a nice-to-have. Above that — especially in fat-tree topologies with many spines — it's the difference between an all-reduce that finishes in 100 µs and one that finishes in 250 µs. For LLM training where every layer triggers one, that compounds enormously.

08

BlueField DPU — Programmable SmartNIC

BlueField is the Data Processing Unit: a NIC plus ARM CPU cores plus dedicated accelerators on a single PCIe card. It runs its own Linux, owns its own networking stack, and offloads work the host CPU used to do.

Network
ConnectX-7 NIC subsystem (400 Gb/s) Hardware packet pipeline RoCE / IB transport
Compute
16× ARM Cortex-A78 DDR5 memory PCIe Gen5 root
Accel
Crypto (IPsec, TLS) Regex / DPI Storage (NVMe-oF, compression) Telemetry

Generations

GenYearSpeedARM coresNotes
BlueField-1201925 Gb/sModestFirst generation; mainly storage offload pilots.
BlueField-22020200 Gb/s8× A72Production hypervisor offload at hyperscalers.
BlueField-32023400 Gb/s16× A78Anchor of Spectrum-X. Networking + storage + security offload at line rate.
BlueField-42025+800 Gb/s (announced)more cores, integrated AI accelRoadmap; pairs with ConnectX-8 generation hosts.

What people actually use it for

DPU as architectural shift

The host CPU stops being the place where networking, storage, and security run. It becomes purely a compute engine for the application (or, in AI clusters, for orchestrating GPUs). The DPU owns the rest. This is why hyperscalers care about it more than NVIDIA's GPU customers do — for AI nodes the host CPU was already underused.

09

NCCL & the All-Reduce Path

NCCL (NVIDIA Collective Communications Library) is what every framework actually calls. PyTorch torch.distributed with the nccl backend, JAX with pjrt-gpu, vLLM's tensor-parallel transport — all roads end at NCCL.

A typical all-reduce in PyTorch
import torch
import torch.distributed as dist

dist.init_process_group(backend="nccl")
rank, world = dist.get_rank(), dist.get_world_size()
torch.cuda.set_device(rank % torch.cuda.device_count())

grad = torch.randn(1_000_000, device="cuda", dtype=torch.bfloat16)
dist.all_reduce(grad, op=dist.ReduceOp.SUM)   # the network does the work
grad /= world

What happens under that one line

source GPU VRAM (gradient buffer) ↓ PCIe Gen5 to ConnectX-7 / -8 HCA ↓ InfiniBand NDR/XDR (or RoCE on Spectrum-X) ↓ Quantum-2 / Quantum-X800 leaf switch ↓ SHARP reduction tree across the spine ↓ Result broadcast back to every leaf ↓ Destination GPU VRAM via GPUDirect RDMA

Topology detection

NCCL discovers the local topology automatically by reading nvidia-smi topo -m output, ibstat, and the kernel's PCIe tree. It picks ring for large messages (best bandwidth efficiency) and tree for small messages or wide world sizes (best latency). With SHARP enabled it picks collnet for matching ops.

Knobs that matter

The 80/20 of cluster perf

Most "our cluster is slow" tickets resolve to one of: GPU-to-NIC traffic crossing a CPU socket, GPUDirect RDMA not actually enabled, NCCL falling back from collnet because SHARP wasn't loaded, or the wrong HCA pinning. NCCL_DEBUG=INFO answers all four in the first 200 lines of log.

10

Numbers — Bandwidth, Latency, $$$ at Scale

Whenever fabric sizing comes up, the same handful of numbers anchor the discussion. Bandwidth per port, end-to-end latency, and price per port together set the ceiling and the cost.

Per-link bandwidth — GB/s, log-ish scale PCIe 5 x16 (intra-node) 64 GB/s IB NDR 400 (per port) 50 GB/s RoCE 400 (per port) ~50 GB/s IB XDR 800 (per port) 100 GB/s RoCE 800 (per port) ~100 GB/s NVLink 5 (intra-NVL72) 1800 GB/s log-ish — NVLink dwarfs the fabric for a reason: it carries TP, the fabric carries DP/PP/EP.
FabricBW per portLatencyIndicative priceBest for
PCIe 5 x1664 GB/s~1 µsincluded on hostIntra-node only.
NVLink 51.8 TB/s~150 nsincluded on B200/GB200Intra-NVL72 TP and EP.
InfiniBand NDR 40050 GB/s~2 µs end-to-end~$2k/HCA + ~$50k/switch portHopper-era SuperPOD.
InfiniBand XDR 800100 GB/s~2 µs~$3–5k/HCA + higher per portBlackwell-era SuperPOD.
RoCE 400 / 800 (Spectrum-X)50 / 100 GB/s~3–4 µsroughly 30% cheaper than IBHyperscaler AI Ethernet.

Cost intuition at scale

The fabric is the second computer

For the largest training runs, the network is a peer of the GPU rack, not a peripheral. Treating it that way — budgeting for it, profiling it, owning the topology choice end-to-end — is what separates working clusters from the ones that paper-launch and never reach claimed flops.

11

Interactive: Cluster Fabric Sizer

Pick a cluster size, NIC class, fabric, and workload. The sizer estimates per-GPU and total fabric bandwidth, switch radix needed, an indicative cost, and tags whether the fabric is likely to bottleneck.

Per-GPU BW
—
Total fabric BW
—
Switch radix needed
—
Est. fabric cost
—
Latency tier
—
Sizer caveats

Costs are indicative list-price ballparks — real deals vary 30–50% with volume and bundling. The radix calculation assumes a 1:1 non-blocking fat-tree; real deployments often run 2:1 oversubscription on the spine to save cost when the workload tolerates it. Always sanity-check against the NVIDIA reference architecture for your exact GPU count.