A practical guide to putting vLLM on a GB10 Grace-Blackwell workstation — unified memory, 128 GB fit, arm64 quirks, and two-Spark clustering.
The DGX Spark is a small-form-factor workstation with one GB10 Grace-Blackwell Superchip. Ubuntu pre-installed, full NVIDIA software stack, ~$4,700. Physically it looks like a Mac Studio; functionally it is a tiny GB200 DGX.
| Spec | DGX Spark | Context |
|---|---|---|
| Chip | GB10 Grace-Blackwell | 20-core Arm (10× Cortex-X925 + 10× Cortex-A725) + Blackwell GPU on one package |
| Memory | 128 GB LPDDR5x, unified CPU/GPU | ~5× an RTX 4090's VRAM |
| Mem BW | ~273 GB/s | below HBM3 (3 TB/s) but plenty for inference |
| Tensor perf | 1 PFLOP FP4 sparse | Blackwell tensor cores |
| Interconnect | NVLink-C2C between Grace & Blackwell | 900 GB/s, coherent |
| Networking | ConnectX-7 SmartNIC, 2× QSFP56 | 200 GbE / NDR IB each |
| OS | Ubuntu 24.04 (Grace) + DGX OS stack | CUDA 13.0, Triton, TensorRT preinstalled |
| Power | ≤ 240 W | desktop power budget |
On a discrete GPU, weights must be copied CPU→GPU once per load, and any spillover sits in slow system RAM across PCIe. On Spark the GPU directly addresses all 128 GB.
A 70B FP8 model streams in about 40 s from NVMe into VRAM on an H100 via PCIe. On Spark, the weights are simply mmap'd into the unified pool. Cold start can drop to 10–15 s for the same model.
Because weights and cache share the same pool, --gpu-memory-utilization effectively selects how much of 128 GB to give vLLM. A 32B FP8 model takes 32 GB; the other ~90 GB is KV pool. That's hundreds of concurrent 8k-token sessions.
You can keep tokeniser, router, vector DB, retrieval service, and the LLM all on the same chip without paying PCIe latency to cross between them. Useful for agent loops where small pieces of code thread between model calls.
20 Arm cores (10 Cortex-X925 + 10 Cortex-A725) with sane memory bandwidth run prefix hashing, tokenisation, JSON schema validation, and structured-output grammars quickly enough that they're not the bottleneck they are on a budget x86.
uname -m # aarch64 — arm64
nvidia-smi # shows one GB10 device
nvcc --version # CUDA 13.0 or later
docker --version # pre-installed
nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Most common containers are multi-arch now but a few aren't. Always pull with --platform linux/arm64 to be explicit:
docker pull --platform=linux/arm64 vllm/vllm-openai:v0.7.0
# If you see "exec format error" → you pulled an amd64 image. Re-pull with --platform.
bitsandbytes releases lack arm64 wheelsdocker run --rm -d --name vllm-70b \
--gpus all --ipc=host \
--ulimit memlock=-1 --ulimit stack=67108864 \
-p 8000:8000 \
-v ~/hf-cache:/root/.cache/huggingface \
-e HF_TOKEN=$HF_TOKEN \
vllm/vllm-openai:v0.7.0 \
--model neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8 \
--dtype auto \
--kv-cache-dtype fp8 \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--max-num-seqs 128
--model mistralai/Mistral-Small-24B-Instruct \
--dtype bfloat16 \
--max-model-len 32768 \
--gpu-memory-utilization 0.85 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-seqs 64
Because the 128 GB is shared with the OS and any side-car services, don't push --gpu-memory-utilization above 0.90 on a Spark. Leave ~12 GB for Grace-side processes, or you'll OOM your own router.
Rough guide. Assumes --gpu-memory-utilization 0.88, so ~112 GB usable, and FP8 KV.
| Model | Precision | Weights | KV/tok | Usable ctx @ 32 seqs |
|---|---|---|---|---|
| Llama-3.1-8B | BF16 | 16 GB | 0.07 MB | ~46k per session |
| Qwen2.5-14B | BF16 | 28 GB | 0.10 MB | ~27k |
| Mistral-Small-24B | BF16 | 48 GB | 0.08 MB | ~24k |
| Llama-3.3-70B | FP8 | 70 GB | 0.16 MB (FP8 KV) | ~8k |
| Llama-3.3-70B | AWQ INT4 | 40 GB | 0.33 MB (FP16 KV) | ~7k |
| DeepSeek-R1-Distill-32B | BF16 | 64 GB | 0.13 MB | ~11k |
| Mixtral-8x7B (MoE) | BF16 | 92 GB | 0.07 MB | ~9k |
A single Spark can't serve Llama-3-405B (405 GB FP8 weights, ~203 GB even at INT4) or DeepSeek-V3 (671 GB FP8, ~335 GB at INT4) at any precision. 405B at INT4 needs a second Spark (slide 07); DeepSeek-V3 doesn't fit even a pair.
Representative Llama-3.1-8B FP16 / FP8, prompt 256, output 128, with vLLM 0.7 and CUDA graphs.
Spark is not an H100. Single-user it's well behind a 4090 (LPDDR5x bandwidth is the ceiling: 273 GB/s vs 1008 GB/s). LMSYS measured Llama-3.1-8B FP8 on Spark at 20.5 tok/s at batch 1 and 368 tok/s at batch 32 (SGLang). The 4090 bar is the bandwidth ceiling (1008 GB/s ÷ 16 GB), not a measurement. What the giant unified pool buys is models and concurrency that don't fit a 24 GB card at all. It's a very good "local inference server" shape and a poor "single-user benchmarking rig" shape.
Each Spark has two QSFP56 ports. A direct QSFP56 DAC cable between the two delivers ~25 GB/s (200 Gb/s) between GPUs with RoCEv2. Two options for using it:
Layers 0..39 on Spark-0, 40..79 on Spark-1. Point-to-point send/recv only. Handles weights up to ~240 GB (e.g. Llama-3-405B at INT4, ~203 GB; 405B FP8 at 405 GB and DeepSeek-V3 INT4 at ~335 GB are too big).
ray start --head --port=6379
vllm serve hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4 \
--pipeline-parallel-size 2 \
--tensor-parallel-size 1 \
--distributed-executor-backend rayEach Spark runs an independent vLLM; a router shares traffic. Clean, linear scaling to ~2× throughput. Use if each model fits in 128 GB by itself.
docker run ... vllm/vllm-openai:v0.7.0 ...
# LiteLLM / Envoy LB out frontTP across 200 GbE works but costs you. The all-reduce traffic per layer is high relative to compute. PP and DP are both friendlier to this fabric. Reserve cross-node TP for proper IB NDR links.
Spark occupies the niche of "workstation-class local inference". Nothing else quite does: laptops don't have enough memory, consumer GPUs don't have 128 GB, H100 boxes are rack-scale. If your project maps onto that niche, Spark is genuinely good. If it doesn't, a 4090 or a cloud H100 will both serve you better.