What 273 GB/s and 128 GB unified actually deliver. Realistic tok/s on every model size from 7 B to 405 B; framework choice (Ollama, vLLM, NIM, TRT-LLM); precision selection; KV cache placement on unified memory; batching; and the patterns that get the most out of the hardware.
Spark's GB10 GPU peaks at ~1 PFLOPS sparse FP4 (~500 TFLOPS dense FP4, ~250 TFLOPS dense FP8). Its memory bandwidth is ~273 GB/s. The arithmetic intensity ratio (FLOPS / byte) at FP8 is roughly 250000 / 273 = 916, meaning the GPU computes ~916 FLOPs per byte fetched at peak. LLM decode is a single matrix-vector multiply per layer; arithmetic intensity is ~2 (one FLOP per weight, two operands per result). You will be memory-bound by a factor of 400× on decode.
Single-stream decode tok/s ≈ memory_bandwidth / weight_bytes. Compute is not your bottleneck on a Spark. Choose the smallest precision the model tolerates; everything else (kernel choice, batch size, etc.) is secondary.
Prefill is different: prefill multiplies a (seq_len, hidden) matrix by all the layer weights, so arithmetic intensity rises with seq_len. For long prompts, Spark can run compute-bound prefill at full FP8 throughput.
Single-stream decode estimates assuming 273 GB/s × 0.85 efficiency = 232 GB/s sustained:
| Model | BF16 | FP8 | MX-FP4 | AWQ-INT4 |
|---|---|---|---|---|
| Llama-3.2 1 B | ~115 tok/s | ~230 tok/s | ~420 tok/s | ~390 tok/s |
| Llama-3.2 3 B | ~38 tok/s | ~75 tok/s | ~140 tok/s | ~130 tok/s |
| Llama-3.1 8 B | ~14 tok/s | ~29 tok/s | ~52 tok/s | ~48 tok/s |
| Mistral 7 B | ~16 tok/s | ~33 tok/s | ~60 tok/s | ~55 tok/s |
| Llama-2 13 B (older) | ~9 tok/s | ~18 tok/s | ~32 tok/s | ~30 tok/s |
| Mixtral 8×7 B (active 13 B) | ~9 tok/s | ~18 tok/s | ~32 tok/s | ~30 tok/s |
| Llama-3.3 70 B | OOM | ~3.3 tok/s | ~6.7 tok/s | ~6 tok/s |
| Llama-3.1-Nemotron 70 B + tool | OOM | ~3.0 tok/s | ~6.0 tok/s | ~5.5 tok/s |
| Mixtral 8×22 B (active 39 B) | OOM | OOM (141 GB) | ~11 tok/s | ~10 tok/s |
| DeepSeek-V3 671 B (active 37 B) | OOM | OOM (671 GB) | OOM (~370 GB) | OOM (full weights too big) |
| Llama-3.1 405 B | OOM | OOM | OOM (single Spark) | OOM |
Two-Spark pair adds 128 GB more memory and extends fit, but pipeline-parallel adds latency hops; estimate ~70–80% of single-Spark per-GPU rates. Sweet spot for solo Spark: 7–13 B at FP8 / MX-FP4 (interactive 30–60 tok/s) or 70 B at MX-FP4 (5–7 tok/s, "reads slowly").
Easiest. ARM64 build native. Pulls GGUF quants from ollama.com. Good for: solo desktop chat, hobby projects. Limitations: no real concurrency (sequential generation), no FP4 yet (uses GGUF q4 instead), llama.cpp backend doesn't use Blackwell tensor cores fully.
Best mix of speed + flexibility. ARM64 image (vllm/vllm-openai) since v0.6. Supports MX-FP4 from v0.7. Use this for any multi-user or batched workload. Paged attention shines on Spark's tight memory.
Pre-tuned, one-line deploy, OpenAI-compatible. Engine pre-built per Spark's GPU class — no trtllm-build step. Catalog: Llama-3.x, Mistral, Mixtral, NeMo Embeddings, Riva ASR/TTS, Vision NIMs. Best if you want zero tuning.
Highest perf if you build an engine, but build time on Spark is 15–30 min per (model, precision, max_batch). Real value emerges once you have a stable workload. NIM is TRT-LLM under the hood.
Recommendation flow: start with Ollama if you're exploring → switch to NIM when you have a chosen model → vLLM for custom quants / fine-tunes → raw TRT-LLM only if you're optimising a production endpoint.
| Format | Bytes/param | Spark TC support | Quality cost | When to use |
|---|---|---|---|---|
| BF16 | 2.0 | native | 0% (reference) | Anything ≤ ~50 B that fits. |
| FP8 (E4M3) | 1.0 | native | ~0.5% perplexity uplift typical | Default for 60–120 B. |
| NVFP4 | ~0.56 | native | ~0.5–1.5% perplexity uplift | Default in TensorRT-LLM / Transformer Engine. 16-element blocks, E4M3 scale + per-tensor FP32 scale; more accurate than MX-FP4. |
| MX-FP4 | ~0.55 | native | ~1.5–3% perplexity uplift | Open OCP standard, 32-element blocks, E8M0 scale. Default for 70 B+ on Spark when an NVFP4 checkpoint isn't available. |
| AWQ-INT4 | ~0.6 | via Marlin kernel (not TC native, but fast) | ~1–4% | If MX-FP4 not available; ubiquitous on Hugging Face. |
| GPTQ-INT4 | ~0.6 | via vLLM | ~1–5% | Older quant; superseded by AWQ for new models. |
| GGUF Q5_K_M | ~0.7 | llama.cpp only | ~0.5% | Ollama users; high-quality K-quant. |
| GGUF Q4_K_M | ~0.55 | llama.cpp only | ~1% | Best space/quality tradeoff in GGUF. |
| FP6 (MX-FP6) | ~0.8 | native | ~0.7% | Niche; FP4 sufficient for most LLMs. |
For any 70 B-class model, start with NVFP4 via TensorRT-LLM / NIM if a checkpoint exists; otherwise MX-FP4 (Llama-3.3-70B, DeepSeek-V3, Mistral-Large all ship FP4 since 2025). Fall back to AWQ-INT4 if neither is available. FP8 is fine for 30–50 B that fit comfortably. Skip BF16 unless debugging.
Unlike a discrete-VRAM GPU where weights and KV cache compete for the same fixed VRAM pool, Spark has one pool. The OS, your shell, your container, the model weights, the KV cache, and any other apps all live in the same 128 GB. Plan accordingly.
--gpu-memory-utilization 0.85 or set --max-num-seqs explicitly.docker run --gpus all -p 8000:8000 \
-v /opt/models:/models \
vllm/vllm-openai:v0.7.0 \
--model /models/llama-3.3-70b-mxfp4 \
--quantization mxfp4 \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.85 \
--max-model-len 32768 \
--max-num-seqs 8 \
--enable-chunked-prefill
On a discrete-VRAM datacenter card, bigger batches yield more aggregate tok/s because compute scales while bandwidth stays the same. Spark inverts this for two reasons:
Empirical sweet spots on Spark with 70 B MX-FP4:
| Batch (active streams) | Per-stream tok/s | Aggregate tok/s | KV cache used |
|---|---|---|---|
| 1 | ~7 | 7 | ~2 GB |
| 4 | ~5 | 20 | ~8 GB |
| 8 | ~3.5 | 28 | ~16 GB |
| 16 | ~2.0 | 32 | ~32 GB — tight |
| 32 | ~1.0 | 32 | OOM unless KV-cache FP8 |
The plateau at batch 16 reflects the bandwidth ceiling. For interactive single-user use, batch 1 is preferable (lowest latency). For multi-user serving on Spark, batch 4–8 is the realistic max.
KV cache size at full FP16: 2 (k+v) × n_layers × n_kv_heads × head_dim × seq_len × batch × 2 bytes. For Llama-3.3-70B at 32k context, batch 1: 2 × 80 × 8 × 128 × 32768 × 1 × 2 = ~10.5 GB. At batch 8: 84 GB — doesn't fit.
Cure: FP8 KV cache (halves it) or INT4 KV (quarters it, with small quality loss). vLLM flag: --kv-cache-dtype fp8. Trades ~0.3% perplexity for 2× the context budget.
With FP8 weights + FP8 KV on 70 B, Spark routinely serves 32k-context, batch 4–6 chats. With MX-FP4 weights + FP8 KV, the same hardware reaches 64k-context, batch 6–8. Worth it for any agentic workload.
Spark runs multimodal models well because they're typically smaller than frontier text LLMs and the unified memory makes preprocessing pipelines (audio, image decode) efficient.
Stacking: an ASR + LLM + TTS chained pipeline (voice agent) fits comfortably in Spark's 128 GB with room for ~30k-token context.
The 200 Gb/s ConnectX-7 NIC gives Spark pairs ~25 GB/s/dir RDMA. That's enough for pipeline parallel (PP=2) which only ships hidden states between stages, not full weight matrices.
# head Spark (10.0.0.1)
docker run --gpus all --network host \
vllm/vllm-openai:v0.7.0 \
--model /models/llama-3.1-405b-mxfp4 \
--quantization mxfp4 \
--pipeline-parallel-size 2 \
--distributed-executor-backend ray \
--master-addr 10.0.0.1 --master-port 29500
Production observability: ship DCGM metrics to Prometheus via nvidia/dcgm-exporter container, scrape from your home Grafana. ~20 lines of docker-compose.
One vLLM container, one model, default port 8000. SSH from your laptop, hit it from VS Code's REST client or cline. Set up systemd unit so it survives reboots.
One vLLM container fronted by a LiteLLM proxy on the same Spark; basic auth or OAuth at LiteLLM. Multiple users share batch capacity. Real-world: 5–10 light users on one Spark with 70 B MX-FP4.
One Spark serving a 7–13 B for chat / coding (high tok/s), one cloud H100 reservation for 70 B+ batch jobs, one tiny Jetson Orin for ASR/TTS at the edge. Mix-and-match where each model lives.
Spark hidden in a closet with Wi-Fi 7. Tunnel via Tailscale; access from anywhere on your phone/laptop. Power draw < a small fridge. The original DIGITS pitch.