Local LLM Hosting Series — Presentation 08

Deploying on NVIDIA DGX Spark

A practical guide to putting vLLM on a GB10 Grace-Blackwell workstation — unified memory, 128 GB fit, arm64 quirks, and two-Spark clustering.

DGX SparkGB10 arm64Unified memory 200 GbEvLLM
Hardware → Setup → arm64 → vLLM → Fit → Pair
00

Topics

01

What Spark Actually Is

The DGX Spark is a small-form-factor workstation with one GB10 Grace-Blackwell Superchip. Ubuntu pre-installed, full NVIDIA software stack, ~$4,700. Physically it looks like a Mac Studio; functionally it is a tiny GB200 DGX.

SpecDGX SparkContext
ChipGB10 Grace-Blackwell20-core Arm (10× Cortex-X925 + 10× Cortex-A725) + Blackwell GPU on one package
Memory128 GB LPDDR5x, unified CPU/GPU~5× an RTX 4090's VRAM
Mem BW~273 GB/sbelow HBM3 (3 TB/s) but plenty for inference
Tensor perf1 PFLOP FP4 sparseBlackwell tensor cores
InterconnectNVLink-C2C between Grace & Blackwell900 GB/s, coherent
NetworkingConnectX-7 SmartNIC, 2× QSFP56200 GbE / NDR IB each
OSUbuntu 24.04 (Grace) + DGX OS stackCUDA 13.0, Triton, TensorRT preinstalled
Power≤ 240 Wdesktop power budget
02

Unified Memory — Why Inference Likes It

On a discrete GPU, weights must be copied CPU→GPU once per load, and any spillover sits in slow system RAM across PCIe. On Spark the GPU directly addresses all 128 GB.

Practical effect #1 — model loading

A 70B FP8 model streams in about 40 s from NVMe into VRAM on an H100 via PCIe. On Spark, the weights are simply mmap'd into the unified pool. Cold start can drop to 10–15 s for the same model.

Practical effect #2 — KV cache sizing

Because weights and cache share the same pool, --gpu-memory-utilization effectively selects how much of 128 GB to give vLLM. A 32B FP8 model takes 32 GB; the other ~90 GB is KV pool. That's hundreds of concurrent 8k-token sessions.

Practical effect #3 — host-side tools

You can keep tokeniser, router, vector DB, retrieval service, and the LLM all on the same chip without paying PCIe latency to cross between them. Useful for agent loops where small pieces of code thread between model calls.

Practical effect #4 — Grace CPU is good

20 Arm cores (10 Cortex-X925 + 10 Cortex-A725) with sane memory bandwidth run prefix hashing, tokenisation, JSON schema validation, and structured-output grammars quickly enough that they're not the bottleneck they are on a budget x86.

03

First Setup — CUDA, Docker, arm64 Gotchas

Verify the stack as shipped
uname -m                      # aarch64 — arm64
nvidia-smi                    # shows one GB10 device
nvcc --version                # CUDA 13.0 or later
docker --version              # pre-installed
nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

arm64 container selection

Most common containers are multi-arch now but a few aren't. Always pull with --platform linux/arm64 to be explicit:

vLLM on arm64 — matching image
docker pull --platform=linux/arm64 vllm/vllm-openai:v0.7.0
# If you see "exec format error" → you pulled an amd64 image. Re-pull with --platform.

What works natively

  • vLLM v0.6+ (arm64 wheel published)
  • PyTorch 2.5+ arm64 with CUDA
  • Triton, TensorRT, NCCL, Flash-Attention
  • Hugging Face Hub, datasets, transformers

Things that bite

  • Some bitsandbytes releases lack arm64 wheels
  • Old CUDA-enabled Docker images (< 12.4) — pin new ones
  • Third-party binary Python wheels (build from source)
  • Closed-source x86 Linux utilities (sysbench, perf variants)
04

vLLM on Spark — Realistic Configs

70B FP8 on a single Spark (easy default)
docker run --rm -d --name vllm-70b \
    --gpus all --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864 \
    -p 8000:8000 \
    -v ~/hf-cache:/root/.cache/huggingface \
    -e HF_TOKEN=$HF_TOKEN \
    vllm/vllm-openai:v0.7.0 \
    --model neuralmagic/Meta-Llama-3.1-70B-Instruct-FP8 \
    --dtype auto \
    --kv-cache-dtype fp8 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching \
    --max-num-seqs 128
Mistral-Small-24B BF16, long context
--model mistralai/Mistral-Small-24B-Instruct \
--dtype bfloat16 \
--max-model-len 32768 \
--gpu-memory-utilization 0.85 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-seqs 64
Util setting on Spark

Because the 128 GB is shared with the OS and any side-car services, don't push --gpu-memory-utilization above 0.90 on a Spark. Leave ~12 GB for Grace-side processes, or you'll OOM your own router.

05

What Fits — Model / Concurrency Matrix

Rough guide. Assumes --gpu-memory-utilization 0.88, so ~112 GB usable, and FP8 KV.

ModelPrecisionWeightsKV/tokUsable ctx @ 32 seqs
Llama-3.1-8BBF1616 GB0.07 MB~46k per session
Qwen2.5-14BBF1628 GB0.10 MB~27k
Mistral-Small-24BBF1648 GB0.08 MB~24k
Llama-3.3-70BFP870 GB0.16 MB (FP8 KV)~8k
Llama-3.3-70BAWQ INT440 GB0.33 MB (FP16 KV)~7k
DeepSeek-R1-Distill-32BBF1664 GB0.13 MB~11k
Mixtral-8x7B (MoE)BF1692 GB0.07 MB~9k
Where Spark stops

A single Spark can't serve Llama-3-405B (405 GB FP8 weights, ~203 GB even at INT4) or DeepSeek-V3 (671 GB FP8, ~335 GB at INT4) at any precision. 405B at INT4 needs a second Spark (slide 07); DeepSeek-V3 doesn't fit even a pair.

06

Honest Benchmarks vs RTX 4090, H100

Representative Llama-3.1-8B FP16 / FP8, prompt 256, output 128, with vLLM 0.7 and CUDA graphs.

Single-user tok/s — higher is better RTX 4090 ≤ ~60 tok/s (BF16 ceiling) DGX Spark (FP8)~20 tok/s H100 80 GB ~160 tok/s Aggregate tok/s @ 32 concurrent requests — throughput RTX 4090 ~1400 DGX Spark ~370 H100 80 GB ~3700
Reading this honestly

Spark is not an H100. Single-user it's well behind a 4090 (LPDDR5x bandwidth is the ceiling: 273 GB/s vs 1008 GB/s). LMSYS measured Llama-3.1-8B FP8 on Spark at 20.5 tok/s at batch 1 and 368 tok/s at batch 32 (SGLang). The 4090 bar is the bandwidth ceiling (1008 GB/s ÷ 16 GB), not a measurement. What the giant unified pool buys is models and concurrency that don't fit a 24 GB card at all. It's a very good "local inference server" shape and a poor "single-user benchmarking rig" shape.

07

Two-Spark Pair Over 200 GbE

Each Spark has two QSFP56 ports. A direct QSFP56 DAC cable between the two delivers ~25 GB/s (200 Gb/s) between GPUs with RoCEv2. Two options for using it:

Option A — PP across boxes

Layers 0..39 on Spark-0, 40..79 on Spark-1. Point-to-point send/recv only. Handles weights up to ~240 GB (e.g. Llama-3-405B at INT4, ~203 GB; 405B FP8 at 405 GB and DeepSeek-V3 INT4 at ~335 GB are too big).

vllm (head)
ray start --head --port=6379
vllm serve hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4 \
    --pipeline-parallel-size 2 \
    --tensor-parallel-size 1 \
    --distributed-executor-backend ray

Option B — DP of full replicas

Each Spark runs an independent vLLM; a router shares traffic. Clean, linear scaling to ~2× throughput. Use if each model fits in 128 GB by itself.

per host
docker run ... vllm/vllm-openai:v0.7.0 ...
# LiteLLM / Envoy LB out front
Why not TP across the pair?

TP across 200 GbE works but costs you. The all-reduce traffic per layer is high relative to compute. PP and DP are both friendlier to this fabric. Reserve cross-node TP for proper IB NDR links.

08

When Not To Use a Spark

Spark is a good fit when…

  • You want 70B-class inference locally, quietly, at ≤ 240 W
  • You serve a few dozen concurrent users inside a lab or office
  • You need 128 GB to experiment with MoE / long-context configs
  • You want a zero-admin Ubuntu + full NVIDIA stack box

Spark is a bad fit when…

  • You want absolute peak tok/s on small models → RTX 4090 or H100
  • You run mixed x86-only pipelines you won't port to arm64
  • You need dense matrix throughput (training) → H100/H200
  • You can burst on cloud → per-hour H100 may be cheaper
  • You want a gaming GPU too — Spark runs Arm Linux (DGX OS), not a gaming platform
Position

Spark occupies the niche of "workstation-class local inference". Nothing else quite does: laptops don't have enough memory, consumer GPUs don't have 128 GB, H100 boxes are rack-scale. If your project maps onto that niche, Spark is genuinely good. If it doesn't, a 4090 or a cloud H100 will both serve you better.