NVIDIA GPU Architectures Series — Presentation 33

Inside the DGX Spark — GB10, 128 GB Unified, 1 PFLOP on Your Desk

A low-level look at NVIDIA's personal AI workstation. The GB10 superchip pairs a 20-core ARM CPU with a Blackwell GPU through NVLink-C2C, exposes 128 GB of LPDDR5x as one coherent memory pool, and delivers ~1 PFLOP of FP4 in a 170 W desktop box. Block diagram, board layout, ports, cooling, and the silicon decisions that define what the box can do.

DGX SparkGB10 GraceBlackwell Cortex-X925Cortex-A725 LPDDR5xNVLink-C2C FP4ConnectX-7
CPU ↔ C2C ↔ GPU → Unified LPDDR5x → 200 GbE → NVMe → Display
00

Topics We'll Cover

01

DGX Spark at a Glance

Announced as Project DIGITS at CES 2025, then renamed DGX Spark by mid-2025 and shipped in volume from Q3 2025. Spark is NVIDIA's first personal AI workstation — the same software stack as DGX H100/B200 in a desktop chassis the size of a Mac Mini, at consumer prices.

SpecValue
SoCNVIDIA GB10 (Grace + Blackwell on one package)
CPU20 ARM cores: 10× Cortex-X925 (perf) + 10× Cortex-A725 (efficiency)
GPUBlackwell-class with 5th-gen tensor cores supporting NVFP4 / MX-FP4 / MX-FP6 / FP8
Memory128 GB unified LPDDR5x at ~273 GB/s
Peak compute~1 PFLOP sparse FP4 / ~500 TFLOPS dense FP4 / ~250 TFLOPS FP8
Storage1 or 4 TB NVMe Gen5 SSD (slot, user-replaceable on some SKUs)
NetworkingConnectX-7 with 2× QSFP ports (200 Gb/s) + 1× 10 GbE RJ-45 + Wi-Fi 7
Display1× HDMI 2.1a (no DisplayPort)
USB4× USB Type-C (one is the power input)
System TDPGB10 SoC 140 W; 240 W power supply
Chassis~150×150×50 mm (sub-1 L), ~1.2 kg
Price~$3,000–$4,500 depending on SSD
OSDGX OS (Ubuntu-based ARM64)
02

The GB10 Superchip — Block Diagram

Unlike the rack-scale GB200 (1 Grace + 2 B200 GPUs as separate dies on one board), GB10 puts everything on one SoC: CPU, GPU, NVLink-C2C, NPU helper, and IO — in a single package on TSMC 4NP.

Package
GB10 SoC — ~250 mm², TSMC 4NP, MediaTek + NVIDIA collaboration
CPU complex
10× Cortex-X92510× Cortex-A725L3 mesh (~24 MB)
GPU
Blackwell SM array (~1 PFLOP sparse FP4)
C2C
NVLink-C2C in-package — coherent CPU↔GPU
Memory controller
Unified LPDDR5x at ~273 GB/s — serves both CPU and GPU
IO
PCIe 5 root complexUSB-CHDMI

Heritage: GB10 is to DGX Spark what Apple's M-series is to a Mac Mini — a CPU+GPU+memory SoC purpose-built for the form factor. NVIDIA partnered with MediaTek on the SoC platform; the CPU complex is licensed ARM IP, while the GPU is NVIDIA-designed Blackwell.

03

The CPU — 20 Cortex-X925 / A725 Cores

GB10's CPU complex is a 2× 10-core hybrid: 10 perf cores plus 10 efficiency cores, all coherent. ARM v9.2-a, SVE2 vectors, the latest Neoverse / Cortex generation circa 2024.

Cortex-X925 (perf cores)

  • 10 cores — the "big" cluster
  • 8-wide decode, 12-wide issue, deep OoO
  • 2× 128-bit SVE2 SIMD
  • ~1 MB L2 per core
  • Aggressive boost clocks ~3.6 GHz
  • Single-thread is M3 Pro-class

Cortex-A725 (efficiency cores)

  • 10 cores — the "little" cluster
  • 4-wide decode, in-order or shallow-OoO
  • 1× 128-bit SVE2 SIMD
  • ~512 KB L2 per core
  • Up to ~2.5 GHz
  • Background services, multi-threaded data prep

Why hybrid: a workstation needs single-thread responsiveness (X925) but also benefits from many threads for data preprocessing, JIT compilation, container builds (A725). Linux handles scheduling via EAS (Energy-Aware Scheduling) on Spark out-of-the-box.

04

The GPU — A Trimmed-Down Blackwell

The Spark GPU is the same Blackwell architecture as B200 / RTX 5090 but with much fewer SMs, sized for the GB10's 140 W envelope.

PropertyDGX Spark GPURTX 5090 (reference)B200 (reference)
SMs48 (6,144 CUDA cores)170148 (2 dies)
FP4 sparse TFLOPS~1000~335018000
FP8 dense TFLOPS~250~840 (FP16 acc.) / ~420 (FP32 acc.)4500
BF16 dense TFLOPS~125~210 (FP32 acc.)2250
FP64 TFLOPS~2~1.640
Memory BW273 GB/s (LPDDR5x)1792 GB/s (GDDR7)8 TB/s (HBM3e)

Same Blackwell tensor-core feature set as B200, so kernels targeting Blackwell features compile once and run on both. FP4 enabled (Blackwell introduces FP4 tensor cores; prior Ada/Hopper generations have no native FP4 support). The Spark GPU is best understood as "a quarter of an RTX 5090 with 4× the memory" — very different bandwidth/capacity profile.

05

NVLink-C2C In-Package

On rack-scale GB200, NVLink-C2C runs across the substrate between Grace and Blackwell dies at 900 GB/s. On GB10 the link is internal to the SoC — physically just a wide on-die fabric — so the headline number is moot. What matters is that CPU and GPU share the memory controller and address space; there's no copy across an external link.

Software model

The GPU sees system pages directly. cudaMallocManaged() just succeeds — no real migration cost because there's no other place to migrate to. cudaMemPrefetchAsync() is mostly a hint rather than a copy.

What it costs

The unified memory means CPU bursts and GPU streaming workloads contend for the same 273 GB/s bus. A heavy data-prep loop on the CPU can starve a decode-bound LLM. Pin tasks to NUMA-style core groups on Spark using numactl --cpunodebind.

06

128 GB Unified LPDDR5x — The Headline Feature

Spark's 128 GB is not the largest unified pool on a desk — the Apple Mac Studio M3 Ultra goes up to 512 GB at ~819 GB/s, and AMD Strix Halo mini-PCs also offer 128 GB — but it is the largest unified pool that runs the full CUDA stack with Blackwell FP4 tensor cores. For comparison: RTX 5090 has 32 GB at 1.79 TB/s; RTX PRO 6000 Blackwell has 96 GB at 1.79 TB/s but costs much more.

Memory characteristicSparkNotes
Capacity128 GB (single SKU at launch)Soldered LPDDR5x, not user-upgradable
Bandwidth~273 GB/s~15% of an RTX 5090, ~3.4% of a B200
Pin rate8.533 Gbps/pin (LPDDR5x-8533)256-bit interface
ECCOn-die ECC (LPDDR5x ECC mode)Side-band ECC not available on LPDDR
Voltage1.05 VSame spec as Grace 480 GB on GB200
Bandwidth reality check

273 GB/s is decode-bound for any LLM running locally. A 70 B model in MX-FP4 (35 GB weights) decodes at ~7 tok/s single-stream peak (273 / 35). A 405 B model in MX-FP4 (~200 GB — doesn't fit). A 70 B in FP8 (~70 GB) decodes at ~3.5 tok/s. The big-memory advantage is that the model fits, not that it's fast.

07

Storage, Display, and IO

08

200 GbE ConnectX-7 — The Pairing Port

The most unusual port on Spark is the ConnectX-7 200 Gb/s NIC with two QSFP ports, intended for two-Spark pairing. NVIDIA's reference scenario is direct attach: two Spark units connected by a 1-3 m DAC or AOC cable, behaving like a tiny 256 GB / 2-PFLOP cluster.

What 200 GbE buys you

  • ~25 GB/s/dir, RDMA-capable, GPUDirect supported
  • Roughly NVLink-3 era bandwidth between two chassis
  • Enough for pipeline parallelism across two GPUs
  • Not enough for tensor parallel of large layers (which want NVLink 5 = 1.8 TB/s)

Software

NVIDIA ships NCCL with RoCE/GPUDirect support out of the box on Spark. vLLM Pipeline Parallel is the recommended pattern for two-Spark serving 200 B+ models. NeMo / Megatron can train on a Spark pair with PP=2 + DP for tiny micro-batches.

Why not Ethernet RJ-45?

QSFP lets the same port carry InfiniBand, Ethernet, or RoCE depending on the cable plugged in. A single Spark connected to a Quantum-2 switch becomes a node in a small InfiniBand cluster. NVIDIA's intent: bottom of the same software stack as their datacenter products.

09

Power Delivery, Cooling, Form Factor

Compared to: a single RTX 5090 PC pulls 700-900 W at the wall under load — 3-4× Spark's 240 W maximum. For office or small-lab deployment, that's the difference between "just plug it in" and "dedicated 20 A circuit".

10

Tensor Core Throughput — What 1 PFLOP Buys

Spark's GPU has full Blackwell tensor core support. The 1 PFLOP figure is the headline (sparse FP4); dense FP4 is half that, FP8 is a quarter, BF16 is an eighth, FP32/TF32 is a sixteenth.

OperationDGX SparkWhat it enables
Sparse MX-FP4~1000 TFLOPSHeadline number; requires structured sparsity in the model.
Dense MX-FP4~500 TFLOPSRealistic for inference of MX-FP4-quantised LLMs.
Dense FP8 (E4M3 / E5M2)~250 TFLOPSFine-tune workloads; FP8 inference for models without FP4 quants.
Dense BF16~125 TFLOPSPretraining proxies; mixed-precision fine-tunes.
Dense TF32~62 TFLOPSDrop-in FP32 replacement; rarely used.
FP32 SIMT~30 TFLOPSNon-tensor-core kernels.
FP64 SIMT~2 TFLOPSHPC; low priority on a workstation.
11

Spark vs Other Workstations — Quick Look

Detailed comparison is in Deck 37; here's the headline:

MachineMemoryBWCompute$ bandWall power
DGX Spark128 GB unified273 GB/s~250 TF FP8$3-5k≤ 240 W
Mac Studio M3 Ultra96–512 GB unified819 GB/sno FP8 / FP4$4-10k270 W
RTX 5090 PC32 GB GDDR71792 GB/s~420–840 TF FP8$4-5k700 W
RTX PRO 6000 Blackwell PC96 GB GDDR71792 GB/s~1000 TF FP8$10-12k800 W
Cloud H100 1× / month80 GB HBM33350 GB/s~990 TF FP8$2-3k/mo700 W (theirs)

Spark's niche: the only desktop machine where 70 B-class models fit in unified memory at workstation prices, with the same software as datacenter NVIDIA. Bandwidth is its weakness; raw compute is decent; capacity is the win.

12

Interactive: Spark Capability Calculator

Weight bytes
—
Fits?
—
KV cache room (8k ctx)
—
Decode tok/s (single)
—