NVIDIA GPU Architectures Series — Presentation 35

LLM Inference on DGX Spark — Practical Numbers and Patterns

What 273 GB/s and 128 GB unified actually deliver. Realistic tok/s on every model size from 7 B to 405 B; framework choice (Ollama, vLLM, NIM, TRT-LLM); precision selection; KV cache placement on unified memory; batching; and the patterns that get the most out of the hardware.

SparkvLLMOllama NIMTRT-LLM NVFP4MX-FP4FP8AWQ-INT4 KV cacheBatch
Pick model → Quantise → Pick framework → Tune batch → Profile → Serve
00

Topics We'll Cover

01

The Bandwidth-Bound Reality

Spark's GB10 GPU peaks at ~1 PFLOPS sparse FP4 (~500 TFLOPS dense FP4, ~250 TFLOPS dense FP8). Its memory bandwidth is ~273 GB/s. The arithmetic intensity ratio (FLOPS / byte) at FP8 is roughly 250000 / 273 = 916, meaning the GPU computes ~916 FLOPs per byte fetched at peak. LLM decode is a single matrix-vector multiply per layer; arithmetic intensity is ~2 (one FLOP per weight, two operands per result). You will be memory-bound by a factor of 400× on decode.

Implication

Single-stream decode tok/s ≈ memory_bandwidth / weight_bytes. Compute is not your bottleneck on a Spark. Choose the smallest precision the model tolerates; everything else (kernel choice, batch size, etc.) is secondary.

Prefill is different: prefill multiplies a (seq_len, hidden) matrix by all the layer weights, so arithmetic intensity rises with seq_len. For long prompts, Spark can run compute-bound prefill at full FP8 throughput.

02

Realistic tok/s — Model × Precision Matrix

Single-stream decode estimates assuming 273 GB/s × 0.85 efficiency = 232 GB/s sustained:

ModelBF16FP8MX-FP4AWQ-INT4
Llama-3.2 1 B~115 tok/s~230 tok/s~420 tok/s~390 tok/s
Llama-3.2 3 B~38 tok/s~75 tok/s~140 tok/s~130 tok/s
Llama-3.1 8 B~14 tok/s~29 tok/s~52 tok/s~48 tok/s
Mistral 7 B~16 tok/s~33 tok/s~60 tok/s~55 tok/s
Llama-2 13 B (older)~9 tok/s~18 tok/s~32 tok/s~30 tok/s
Mixtral 8×7 B (active 13 B)~9 tok/s~18 tok/s~32 tok/s~30 tok/s
Llama-3.3 70 BOOM~3.3 tok/s~6.7 tok/s~6 tok/s
Llama-3.1-Nemotron 70 B + toolOOM~3.0 tok/s~6.0 tok/s~5.5 tok/s
Mixtral 8×22 B (active 39 B)OOMOOM (141 GB)~11 tok/s~10 tok/s
DeepSeek-V3 671 B (active 37 B)OOMOOM (671 GB)OOM (~370 GB)OOM (full weights too big)
Llama-3.1 405 BOOMOOMOOM (single Spark)OOM

Two-Spark pair adds 128 GB more memory and extends fit, but pipeline-parallel adds latency hops; estimate ~70–80% of single-Spark per-GPU rates. Sweet spot for solo Spark: 7–13 B at FP8 / MX-FP4 (interactive 30–60 tok/s) or 70 B at MX-FP4 (5–7 tok/s, "reads slowly").

03

Framework Choice — Ollama vs vLLM vs NIM vs TRT-LLM

Ollama

Easiest. ARM64 build native. Pulls GGUF quants from ollama.com. Good for: solo desktop chat, hobby projects. Limitations: no real concurrency (sequential generation), no FP4 yet (uses GGUF q4 instead), llama.cpp backend doesn't use Blackwell tensor cores fully.

vLLM

Best mix of speed + flexibility. ARM64 image (vllm/vllm-openai) since v0.6. Supports MX-FP4 from v0.7. Use this for any multi-user or batched workload. Paged attention shines on Spark's tight memory.

NIM

Pre-tuned, one-line deploy, OpenAI-compatible. Engine pre-built per Spark's GPU class — no trtllm-build step. Catalog: Llama-3.x, Mistral, Mixtral, NeMo Embeddings, Riva ASR/TTS, Vision NIMs. Best if you want zero tuning.

TensorRT-LLM

Highest perf if you build an engine, but build time on Spark is 15–30 min per (model, precision, max_batch). Real value emerges once you have a stable workload. NIM is TRT-LLM under the hood.

Recommendation flow: start with Ollama if you're exploring → switch to NIM when you have a chosen model → vLLM for custom quants / fine-tunes → raw TRT-LLM only if you're optimising a production endpoint.

04

Picking a Quantisation Format

FormatBytes/paramSpark TC supportQuality costWhen to use
BF162.0native0% (reference)Anything ≤ ~50 B that fits.
FP8 (E4M3)1.0native~0.5% perplexity uplift typicalDefault for 60–120 B.
NVFP4~0.56native~0.5–1.5% perplexity upliftDefault in TensorRT-LLM / Transformer Engine. 16-element blocks, E4M3 scale + per-tensor FP32 scale; more accurate than MX-FP4.
MX-FP4~0.55native~1.5–3% perplexity upliftOpen OCP standard, 32-element blocks, E8M0 scale. Default for 70 B+ on Spark when an NVFP4 checkpoint isn't available.
AWQ-INT4~0.6via Marlin kernel (not TC native, but fast)~1–4%If MX-FP4 not available; ubiquitous on Hugging Face.
GPTQ-INT4~0.6via vLLM~1–5%Older quant; superseded by AWQ for new models.
GGUF Q5_K_M~0.7llama.cpp only~0.5%Ollama users; high-quality K-quant.
GGUF Q4_K_M~0.55llama.cpp only~1%Best space/quality tradeoff in GGUF.
FP6 (MX-FP6)~0.8native~0.7%Niche; FP4 sufficient for most LLMs.
Spark default

For any 70 B-class model, start with NVFP4 via TensorRT-LLM / NIM if a checkpoint exists; otherwise MX-FP4 (Llama-3.3-70B, DeepSeek-V3, Mistral-Large all ship FP4 since 2025). Fall back to AWQ-INT4 if neither is available. FP8 is fine for 30–50 B that fit comfortably. Skip BF16 unless debugging.

05

Unified Memory & KV Cache Placement

Unlike a discrete-VRAM GPU where weights and KV cache compete for the same fixed VRAM pool, Spark has one pool. The OS, your shell, your container, the model weights, the KV cache, and any other apps all live in the same 128 GB. Plan accordingly.

vLLM tuned for Spark + 70 B MX-FP4
docker run --gpus all -p 8000:8000 \
    -v /opt/models:/models \
    vllm/vllm-openai:v0.7.0 \
    --model /models/llama-3.3-70b-mxfp4 \
    --quantization mxfp4 \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.85 \
    --max-model-len 32768 \
    --max-num-seqs 8 \
    --enable-chunked-prefill
06

Batching Strategy on a Bandwidth-Bound Box

On a discrete-VRAM datacenter card, bigger batches yield more aggregate tok/s because compute scales while bandwidth stays the same. Spark inverts this for two reasons:

  1. The decode-bound bottleneck is bandwidth, and batched decode reads the same weights once per step regardless of batch size — so throughput does scale with batch.
  2. But KV cache per request scales with batch and easily eats the unified-memory headroom.

Empirical sweet spots on Spark with 70 B MX-FP4:

Batch (active streams)Per-stream tok/sAggregate tok/sKV cache used
1~77~2 GB
4~520~8 GB
8~3.528~16 GB
16~2.032~32 GB — tight
32~1.032OOM unless KV-cache FP8

The plateau at batch 16 reflects the bandwidth ceiling. For interactive single-user use, batch 1 is preferable (lowest latency). For multi-user serving on Spark, batch 4–8 is the realistic max.

07

Long Context & KV Quantisation

KV cache size at full FP16: 2 (k+v) × n_layers × n_kv_heads × head_dim × seq_len × batch × 2 bytes. For Llama-3.3-70B at 32k context, batch 1: 2 × 80 × 8 × 128 × 32768 × 1 × 2 = ~10.5 GB. At batch 8: 84 GB — doesn't fit.

Cure: FP8 KV cache (halves it) or INT4 KV (quarters it, with small quality loss). vLLM flag: --kv-cache-dtype fp8. Trades ~0.3% perplexity for 2× the context budget.

Practical reach

With FP8 weights + FP8 KV on 70 B, Spark routinely serves 32k-context, batch 4–6 chats. With MX-FP4 weights + FP8 KV, the same hardware reaches 64k-context, batch 6–8. Worth it for any agentic workload.

08

Multimodal Models — Vision, ASR, TTS

Spark runs multimodal models well because they're typically smaller than frontier text LLMs and the unified memory makes preprocessing pipelines (audio, image decode) efficient.

Stacking: an ASR + LLM + TTS chained pipeline (voice agent) fits comfortably in Spark's 128 GB with room for ~30k-token context.

09

Two-Spark Inference — Pipeline Parallel

The 200 Gb/s ConnectX-7 NIC gives Spark pairs ~25 GB/s/dir RDMA. That's enough for pipeline parallel (PP=2) which only ships hidden states between stages, not full weight matrices.

What works

  • PP=2: each Spark holds half the layers
  • Effective memory: 256 GB
  • 120 B model in MX-FP4 (~66 GB) fits easily
  • 200 B in MX-FP4 (~110 GB) fits with KV cache room
  • Llama-3.1-405B in MX-FP4 (~220 GB) fits but tight

What doesn't

  • TP=2 (tensor parallel): wants NVLink-class bandwidth (~1 TB/s); 200 GbE is 40× slower — layers stall
  • Latency: ~3 ms per cross-Spark hop — cuts decode tok/s ~30%
  • Throughput rarely beats single Spark unless model wouldn't fit otherwise
vLLM PP=2 across a Spark pair
# head Spark (10.0.0.1)
docker run --gpus all --network host \
    vllm/vllm-openai:v0.7.0 \
    --model /models/llama-3.1-405b-mxfp4 \
    --quantization mxfp4 \
    --pipeline-parallel-size 2 \
    --distributed-executor-backend ray \
    --master-addr 10.0.0.1 --master-port 29500
10

Monitoring & Profiling on Spark

Production observability: ship DCGM metrics to Prometheus via nvidia/dcgm-exporter container, scrape from your home Grafana. ~20 lines of docker-compose.

11

Production Patterns for a Single Spark

Solo developer

One vLLM container, one model, default port 8000. SSH from your laptop, hit it from VS Code's REST client or cline. Set up systemd unit so it survives reboots.

Small team

One vLLM container fronted by a LiteLLM proxy on the same Spark; basic auth or OAuth at LiteLLM. Multiple users share batch capacity. Real-world: 5–10 light users on one Spark with 70 B MX-FP4.

Home lab + agents

One Spark serving a 7–13 B for chat / coding (high tok/s), one cloud H100 reservation for 70 B+ batch jobs, one tiny Jetson Orin for ASR/TTS at the edge. Mix-and-match where each model lives.

Edge agent

Spark hidden in a closet with Wi-Fi 7. Tunnel via Tailscale; access from anywhere on your phone/laptop. Power draw < a small fridge. The original DIGITS pitch.

12

Interactive: Spark Inference Estimator

Weights
—
Fits + KV
—
Per-stream tok/s
—
Aggregate tok/s
—