A low-level look at NVIDIA's personal AI workstation. The GB10 superchip pairs a 20-core ARM CPU with a Blackwell GPU through NVLink-C2C, exposes 128 GB of LPDDR5x as one coherent memory pool, and delivers ~1 PFLOP of FP4 in a 170 W desktop box. Block diagram, board layout, ports, cooling, and the silicon decisions that define what the box can do.
Announced as Project DIGITS at CES 2025, then renamed DGX Spark by mid-2025 and shipped in volume from Q3 2025. Spark is NVIDIA's first personal AI workstation — the same software stack as DGX H100/B200 in a desktop chassis the size of a Mac Mini, at consumer prices.
| Spec | Value |
|---|---|
| SoC | NVIDIA GB10 (Grace + Blackwell on one package) |
| CPU | 20 ARM cores: 10× Cortex-X925 (perf) + 10× Cortex-A725 (efficiency) |
| GPU | Blackwell-class with 5th-gen tensor cores supporting NVFP4 / MX-FP4 / MX-FP6 / FP8 |
| Memory | 128 GB unified LPDDR5x at ~273 GB/s |
| Peak compute | ~1 PFLOP sparse FP4 / ~500 TFLOPS dense FP4 / ~250 TFLOPS FP8 |
| Storage | 1 or 4 TB NVMe Gen5 SSD (slot, user-replaceable on some SKUs) |
| Networking | ConnectX-7 with 2× QSFP ports (200 Gb/s) + 1× 10 GbE RJ-45 + Wi-Fi 7 |
| Display | 1× HDMI 2.1a (no DisplayPort) |
| USB | 4× USB Type-C (one is the power input) |
| System TDP | GB10 SoC 140 W; 240 W power supply |
| Chassis | ~150×150×50 mm (sub-1 L), ~1.2 kg |
| Price | ~$3,000–$4,500 depending on SSD |
| OS | DGX OS (Ubuntu-based ARM64) |
Unlike the rack-scale GB200 (1 Grace + 2 B200 GPUs as separate dies on one board), GB10 puts everything on one SoC: CPU, GPU, NVLink-C2C, NPU helper, and IO — in a single package on TSMC 4NP.
Heritage: GB10 is to DGX Spark what Apple's M-series is to a Mac Mini — a CPU+GPU+memory SoC purpose-built for the form factor. NVIDIA partnered with MediaTek on the SoC platform; the CPU complex is licensed ARM IP, while the GPU is NVIDIA-designed Blackwell.
GB10's CPU complex is a 2× 10-core hybrid: 10 perf cores plus 10 efficiency cores, all coherent. ARM v9.2-a, SVE2 vectors, the latest Neoverse / Cortex generation circa 2024.
Why hybrid: a workstation needs single-thread responsiveness (X925) but also benefits from many threads for data preprocessing, JIT compilation, container builds (A725). Linux handles scheduling via EAS (Energy-Aware Scheduling) on Spark out-of-the-box.
The Spark GPU is the same Blackwell architecture as B200 / RTX 5090 but with much fewer SMs, sized for the GB10's 140 W envelope.
| Property | DGX Spark GPU | RTX 5090 (reference) | B200 (reference) |
|---|---|---|---|
| SMs | 48 (6,144 CUDA cores) | 170 | 148 (2 dies) |
| FP4 sparse TFLOPS | ~1000 | ~3350 | 18000 |
| FP8 dense TFLOPS | ~250 | ~840 (FP16 acc.) / ~420 (FP32 acc.) | 4500 |
| BF16 dense TFLOPS | ~125 | ~210 (FP32 acc.) | 2250 |
| FP64 TFLOPS | ~2 | ~1.6 | 40 |
| Memory BW | 273 GB/s (LPDDR5x) | 1792 GB/s (GDDR7) | 8 TB/s (HBM3e) |
Same Blackwell tensor-core feature set as B200, so kernels targeting Blackwell features compile once and run on both. FP4 enabled (Blackwell introduces FP4 tensor cores; prior Ada/Hopper generations have no native FP4 support). The Spark GPU is best understood as "a quarter of an RTX 5090 with 4× the memory" — very different bandwidth/capacity profile.
On rack-scale GB200, NVLink-C2C runs across the substrate between Grace and Blackwell dies at 900 GB/s. On GB10 the link is internal to the SoC — physically just a wide on-die fabric — so the headline number is moot. What matters is that CPU and GPU share the memory controller and address space; there's no copy across an external link.
The GPU sees system pages directly. cudaMallocManaged() just succeeds — no real migration cost because there's no other place to migrate to. cudaMemPrefetchAsync() is mostly a hint rather than a copy.
The unified memory means CPU bursts and GPU streaming workloads contend for the same 273 GB/s bus. A heavy data-prep loop on the CPU can starve a decode-bound LLM. Pin tasks to NUMA-style core groups on Spark using numactl --cpunodebind.
Spark's 128 GB is not the largest unified pool on a desk — the Apple Mac Studio M3 Ultra goes up to 512 GB at ~819 GB/s, and AMD Strix Halo mini-PCs also offer 128 GB — but it is the largest unified pool that runs the full CUDA stack with Blackwell FP4 tensor cores. For comparison: RTX 5090 has 32 GB at 1.79 TB/s; RTX PRO 6000 Blackwell has 96 GB at 1.79 TB/s but costs much more.
| Memory characteristic | Spark | Notes |
|---|---|---|
| Capacity | 128 GB (single SKU at launch) | Soldered LPDDR5x, not user-upgradable |
| Bandwidth | ~273 GB/s | ~15% of an RTX 5090, ~3.4% of a B200 |
| Pin rate | 8.533 Gbps/pin (LPDDR5x-8533) | 256-bit interface |
| ECC | On-die ECC (LPDDR5x ECC mode) | Side-band ECC not available on LPDDR |
| Voltage | 1.05 V | Same spec as Grace 480 GB on GB200 |
273 GB/s is decode-bound for any LLM running locally. A 70 B model in MX-FP4 (35 GB weights) decodes at ~7 tok/s single-stream peak (273 / 35). A 405 B model in MX-FP4 (~200 GB — doesn't fit). A 70 B in FP8 (~70 GB) decodes at ~3.5 tok/s. The big-memory advantage is that the model fits, not that it's fast.
The most unusual port on Spark is the ConnectX-7 200 Gb/s NIC with two QSFP ports, intended for two-Spark pairing. NVIDIA's reference scenario is direct attach: two Spark units connected by a 1-3 m DAC or AOC cable, behaving like a tiny 256 GB / 2-PFLOP cluster.
NVIDIA ships NCCL with RoCE/GPUDirect support out of the box on Spark. vLLM Pipeline Parallel is the recommended pattern for two-Spark serving 200 B+ models. NeMo / Megatron can train on a Spark pair with PP=2 + DP for tiny micro-batches.
QSFP lets the same port carry InfiniBand, Ethernet, or RoCE depending on the cable plugged in. A single Spark connected to a Quantum-2 switch becomes a node in a small InfiniBand cluster. NVIDIA's intent: bottom of the same software stack as their datacenter products.
Compared to: a single RTX 5090 PC pulls 700-900 W at the wall under load — 3-4× Spark's 240 W maximum. For office or small-lab deployment, that's the difference between "just plug it in" and "dedicated 20 A circuit".
Spark's GPU has full Blackwell tensor core support. The 1 PFLOP figure is the headline (sparse FP4); dense FP4 is half that, FP8 is a quarter, BF16 is an eighth, FP32/TF32 is a sixteenth.
| Operation | DGX Spark | What it enables |
|---|---|---|
| Sparse MX-FP4 | ~1000 TFLOPS | Headline number; requires structured sparsity in the model. |
| Dense MX-FP4 | ~500 TFLOPS | Realistic for inference of MX-FP4-quantised LLMs. |
| Dense FP8 (E4M3 / E5M2) | ~250 TFLOPS | Fine-tune workloads; FP8 inference for models without FP4 quants. |
| Dense BF16 | ~125 TFLOPS | Pretraining proxies; mixed-precision fine-tunes. |
| Dense TF32 | ~62 TFLOPS | Drop-in FP32 replacement; rarely used. |
| FP32 SIMT | ~30 TFLOPS | Non-tensor-core kernels. |
| FP64 SIMT | ~2 TFLOPS | HPC; low priority on a workstation. |
Detailed comparison is in Deck 37; here's the headline:
| Machine | Memory | BW | Compute | $ band | Wall power |
|---|---|---|---|---|---|
| DGX Spark | 128 GB unified | 273 GB/s | ~250 TF FP8 | $3-5k | ≤ 240 W |
| Mac Studio M3 Ultra | 96–512 GB unified | 819 GB/s | no FP8 / FP4 | $4-10k | 270 W |
| RTX 5090 PC | 32 GB GDDR7 | 1792 GB/s | ~420–840 TF FP8 | $4-5k | 700 W |
| RTX PRO 6000 Blackwell PC | 96 GB GDDR7 | 1792 GB/s | ~1000 TF FP8 | $10-12k | 800 W |
| Cloud H100 1× / month | 80 GB HBM3 | 3350 GB/s | ~990 TF FP8 | $2-3k/mo | 700 W (theirs) |
Spark's niche: the only desktop machine where 70 B-class models fit in unified memory at workstation prices, with the same software as datacenter NVIDIA. Bandwidth is its weakness; raw compute is decent; capacity is the win.