Cross-rack tensor parallelism, multi-trillion-parameter MoE, distributed checkpointing, and inference at fleet scale all require NVIDIA's networking stack — InfiniBand or Ethernet, ConnectX HCAs, BlueField DPUs, Quantum and Spectrum switches, GPUDirect RDMA and SHARP. Here's what each is and how they stack up.
The fabric is the second computer. Once a model leaves a single NVL72, every all-reduce becomes a network problem — this deck walks through the components that make it tractable.
Inside a single NVL72 or DGX node, NVLink and NVSwitch carry the load. The moment a model spills across the rack — pipeline stages, data-parallel replicas, expert shards in an MoE — every gradient step has to cross a slower fabric, and that fabric becomes the bottleneck.
An 8-GPU tensor-parallel group fits inside one NVLink domain — cheap. Pipeline parallel and data parallel hop nodes — expensive. For a 405B Llama-class model, each layer's all-reduce is roughly 1–2 GB, and with ~100 layers and several reductions per step the fabric throughput dictates how often a GPU sits idle waiting for gradients.
KV-cache shipping for distributed serving (prefill on one node, decode on another), model weight loading from PFS at fleet startup, telemetry and tracing, request routing — all of it crosses the fabric. At 100k-GPU scale even "background" traffic is several Tbps.
NVIDIA didn't build this stack from scratch. In April 2020 they completed the acquisition of Mellanox for ~$7B, picking up the dominant InfiniBand vendor along with ConnectX HCAs, the Quantum/Spectrum switch lines, and the BlueField SmartNIC roadmap. Six years on, that one acquisition is arguably as strategic as any GPU generation — without it the AI factory pitch has no plumbing.
NVLink is for the rack. Networking is for everything beyond it. The rest of the deck walks the stack from NIC to switch to topology to software offloads, with the numbers that matter.
NVIDIA sells both. Frontier AI clusters have historically chosen InfiniBand; hyperscalers increasingly push Ethernet because it's the fabric they already operate everything else on. Both can deliver an LLM training workload — the trade-offs are around lossless behaviour, latency, vendor lock-in, and price.
Both fabrics are functionally lossless on a healthy day. The difference is what happens when something goes wrong. InfiniBand's credit scheme prevents drops by construction. Ethernet's PFC reacts to congestion and can deadlock or trigger pause storms if you've misconfigured a single switch. Spectrum-X's whole pitch is offloading the hard parts of that tuning to BlueField.
The Host Channel Adapter (HCA) is the NIC: one PCIe card per server, or in AI clusters one per GPU. ConnectX is the Mellanox/NVIDIA brand. Generation tracks the IB speed standard.
| Generation | Year | Per-port speed | Standard | Use case / platform |
|---|---|---|---|---|
| ConnectX-5 | 2017 | 100 Gb/s | EDR / HDR100 | Pre-Ampere clusters, late P100 era. |
| ConnectX-6 | 2020 | 200 Gb/s | HDR | DGX A100 baseline (8× CX-6 per node). |
| ConnectX-6 Dx | 2020 | 200 Gb/s | HDR / 200GbE | Cloud + telco; security and TLS offload. |
| ConnectX-7 | 2022 | 400 Gb/s | NDR | DGX H100 / H200 (8× CX-7 per node). |
| ConnectX-8 | 2024 | 800 Gb/s | XDR | DGX B300 / GB300 NVL72 (DGX B200 still uses 8× CX-7). |
For AI training clusters the rule is one dedicated HCA per GPU. A DGX H100 has 8 H100s and 8 ConnectX-7 cards; a DGX B200 has 8 B200s and 8 ConnectX-7 cards (ConnectX-8 arrives with B300). Plus there are typically additional ConnectX or BlueField NICs for the storage fabric and management. The compute fabric is sized so that a GPU never waits on its NIC.
Sharing one NIC across multiple GPUs adds head-of-line blocking and PCIe contention. Per-GPU NICs let NCCL's ring or tree algorithm send data along the shortest possible path and let GPUDirect RDMA work without crossing the host's main PCIe complex.
Quantum is NVIDIA's InfiniBand switch line. Each generation tracks the IB speed standard and roughly doubles per-port bandwidth and switch radix.
| Switch | Year | Speed | Ports | Anchor cluster |
|---|---|---|---|---|
| Quantum (QM8700) | 2018 | HDR 200G | 40-port | DGX A100 SuperPODs. |
| Quantum-2 (QM9700) | 2021 | NDR 400G | 64-port | DGX H100 SuperPOD — the workhorse for Hopper-era clusters. |
| Quantum-X800 (Q3400) | 2024 | XDR 800G | 144-port effective (multi-host) | DGX B200 / GB200 NVL72 SuperPODs. |
The canonical AI topology. Two or three layers of switches: leaf, spine, and (for the largest pods) super-spine. Non-blocking 1:1 oversubscription means any GPU can talk to any other GPU at full line rate. Standard for DGX SuperPOD up to roughly 1k–4k GPUs.
Used in very large clusters (10k+ GPUs). Switches grouped into "groups", with all-to-all inside a group and selective links between groups. Lower cable count than fat-tree, but routing is harder and worst-case hop count is higher.
Quantum-2 and Quantum-X800 both ship with hardware adaptive routing: the switch picks the least-congested path per packet rather than a static hash. Combined with packet-level reordering at the receiver HCA, this avoids the "ECMP hash collision" pathology that plagues vanilla Ethernet at AI scale.
Doubles the per-port speed (NDR 400 → XDR 800), more than doubles the effective radix, and was designed alongside the GB200 NVL72 platform — one switch tier can carry the entire NVL72 to NVL72 traffic in a multi-rack pod.
Spectrum is NVIDIA's Ethernet switch line (the Mellanox-origin SN-series). Spectrum-X is the AI-specific reference architecture: Spectrum switches plus ConnectX/BlueField NICs plus a software stack that delivers InfiniBand-class behaviour over Ethernet.
Conventional Ethernet switches: SN3000, SN4000, SN5000 series at 200/400/800 GbE. Used for general datacenter networking, storage, and management traffic. Same silicon family as the AI variant; what differs is how you operate it.
The combination that makes Ethernet credibly carry training traffic:
One fabric for AI plus general-purpose plus storage means one team, one toolchain, one set of cables, and inter-op with everything else in the datacenter. The cost is some single-digit-percent of throughput vs InfiniBand and a more complex tuning story — offset by sticking with the network they already run.
The Ultra Ethernet Consortium (AMD, Broadcom, Cisco, Meta, Microsoft, plus dozens of others) is standardising AI Ethernet semantics — new transport, packet spraying, in-network telemetry. NVIDIA participates and is broadly compatible. Spectrum-X is NVIDIA's pre-standard answer; Ultra Ethernet is the multi-vendor convergence point. Expect them to merge over a few years.
The fastest data path is the one that doesn't touch the host CPU. GPUDirect is the umbrella term for three closely related technologies that let the GPU and the NIC (or another GPU, or NVMe) talk to each other directly.
GPU-to-GPU within a node across PCIe or NVLink without staging in host RAM. cudaMemcpyPeer or NCCL P2P transport. On NVLink it's effectively free; on PCIe it depends on IOMMU and ACS configuration.
GPU-to-GPU across nodes. The remote NIC reads or writes GPU VRAM directly via DMA. The host CPU is not on the fast path. Latency drops from ~50 µs (host bounce) to ~6 µs (direct). Required for any credible cross-node TP / PP.
NVMe and PFS reads land directly in GPU VRAM via the cuFile API. Bypasses the host page cache entirely. Critical for fast checkpoint loading, model warm-start, and dataset streaming at scale.
nvidia-peermem or nv_peer_mem kernel module loaded on every host.ibstat.nvidia-smi topo -m should show PIX or NV# between the GPU and HCA, not SYS. NCCL_DEBUG=INFO will print "NET/IB: Using direct GPU-to-NIC RDMA".A 2 GB all-reduce on a 100k-GPU cluster touches the network billions of times per training run. Saving 40 µs per hop by removing the host bounce is the difference between a step that finishes in 80 ms and one that finishes in 130 ms — a 60% throughput delta on a multi-month run.
Scalable Hierarchical Aggregation and Reduction Protocol. The idea is simple and powerful: move the all-reduce into the switch instead of doing it at the endpoints.
Ring or tree algorithm: each GPU sends a chunk, receives a chunk, sums, forwards. With N GPUs an all-reduce moves roughly 2×(N−1)/N of the tensor across the wire. As N grows the bandwidth wall dominates.
The reduction itself happens in the switch silicon. Each GPU sends its data once and receives the result once — halving the bandwidth needed versus a ring all-reduce, and reducing the latency at scale because it's a single hierarchical step rather than O(N) sequential exchanges.
NCCL_COLLNET_ENABLE=1 turns it on; NCCL falls back gracefully if topology doesn't support it.Below ~256 GPUs SHARP is a nice-to-have. Above that — especially in fat-tree topologies with many spines — it's the difference between an all-reduce that finishes in 100 µs and one that finishes in 250 µs. For LLM training where every layer triggers one, that compounds enormously.
BlueField is the Data Processing Unit: a NIC plus ARM CPU cores plus dedicated accelerators on a single PCIe card. It runs its own Linux, owns its own networking stack, and offloads work the host CPU used to do.
| Gen | Year | Speed | ARM cores | Notes |
|---|---|---|---|---|
| BlueField-1 | 2019 | 25 Gb/s | Modest | First generation; mainly storage offload pilots. |
| BlueField-2 | 2020 | 200 Gb/s | 8× A72 | Production hypervisor offload at hyperscalers. |
| BlueField-3 | 2023 | 400 Gb/s | 16× A78 | Anchor of Spectrum-X. Networking + storage + security offload at line rate. |
| BlueField-4 | 2025+ | 800 Gb/s (announced) | more cores, integrated AI accel | Roadmap; pairs with ConnectX-8 generation hosts. |
The host CPU stops being the place where networking, storage, and security run. It becomes purely a compute engine for the application (or, in AI clusters, for orchestrating GPUs). The DPU owns the rest. This is why hyperscalers care about it more than NVIDIA's GPU customers do — for AI nodes the host CPU was already underused.
NCCL (NVIDIA Collective Communications Library) is what every framework actually calls. PyTorch torch.distributed with the nccl backend, JAX with pjrt-gpu, vLLM's tensor-parallel transport — all roads end at NCCL.
import torch
import torch.distributed as dist
dist.init_process_group(backend="nccl")
rank, world = dist.get_rank(), dist.get_world_size()
torch.cuda.set_device(rank % torch.cuda.device_count())
grad = torch.randn(1_000_000, device="cuda", dtype=torch.bfloat16)
dist.all_reduce(grad, op=dist.ReduceOp.SUM) # the network does the work
grad /= world
NCCL discovers the local topology automatically by reading nvidia-smi topo -m output, ibstat, and the kernel's PCIe tree. It picks ring for large messages (best bandwidth efficiency) and tree for small messages or wide world sizes (best latency). With SHARP enabled it picks collnet for matching ops.
NCCL_IB_HCA — pin which HCA each GPU uses (e.g. mlx5_0:1,mlx5_1:1). Critical when the auto-detected mapping is wrong.NCCL_IB_TIMEOUT / NCCL_IB_RETRY_CNT — raise on flaky fabrics or large clusters; default values trip easily on 10k+ GPUs.NCCL_NET_GDR_LEVEL — how aggressively to use GPUDirect RDMA based on PCIe topology distance.NCCL_COLLNET_ENABLE=1 — turn on SHARP via the collnet plugin.NCCL_NVLS_ENABLE=1 — NVLink-Sharp on Hopper / Blackwell, the intra-NVL72 equivalent of SHARP using NVSwitch.NCCL_DEBUG=INFO, NCCL_DEBUG_SUBSYS=NET,COLL — first thing to set when debugging.Most "our cluster is slow" tickets resolve to one of: GPU-to-NIC traffic crossing a CPU socket, GPUDirect RDMA not actually enabled, NCCL falling back from collnet because SHARP wasn't loaded, or the wrong HCA pinning. NCCL_DEBUG=INFO answers all four in the first 200 lines of log.
Whenever fabric sizing comes up, the same handful of numbers anchor the discussion. Bandwidth per port, end-to-end latency, and price per port together set the ceiling and the cost.
| Fabric | BW per port | Latency | Indicative price | Best for |
|---|---|---|---|---|
| PCIe 5 x16 | 64 GB/s | ~1 µs | included on host | Intra-node only. |
| NVLink 5 | 1.8 TB/s | ~150 ns | included on B200/GB200 | Intra-NVL72 TP and EP. |
| InfiniBand NDR 400 | 50 GB/s | ~2 µs end-to-end | ~$2k/HCA + ~$50k/switch port | Hopper-era SuperPOD. |
| InfiniBand XDR 800 | 100 GB/s | ~2 µs | ~$3–5k/HCA + higher per port | Blackwell-era SuperPOD. |
| RoCE 400 / 800 (Spectrum-X) | 50 / 100 GB/s | ~3–4 µs | roughly 30% cheaper than IB | Hyperscaler AI Ethernet. |
For the largest training runs, the network is a peer of the GPU rack, not a peripheral. Treating it that way — budgeting for it, profiling it, owning the topology choice end-to-end — is what separates working clusters from the ones that paper-launch and never reach claimed flops.
Pick a cluster size, NIC class, fabric, and workload. The sizer estimates per-GPU and total fabric bandwidth, switch radix needed, an indicative cost, and tags whether the fabric is likely to bottleneck.
Costs are indicative list-price ballparks — real deals vary 30–50% with volume and bundling. The radix calculation assumes a 1:1 non-blocking fat-tree; real deployments often run 2:1 oversubscription on the spine to save cost when the workload tolerates it. Always sanity-check against the NVIDIA reference architecture for your exact GPU count.