A walkthrough of bringing a Spark online. Unbox, OS install, drivers, NGC login, container runtime, networking, and the first inference run. The intent: a checklist you can read once and follow on the day, with the right command for each step and the gotchas surfaced in advance.
Before powering on: have a monitor + cable ready, a Wi-Fi network credential ready (or a wired Ethernet cable for the 10 GbE RJ-45 port), and an NVIDIA Developer account. Without an NVIDIA Developer account you can't pull NGC images, which is most of the value of the device.
Press the power button on the front. NVIDIA logo, ~30 seconds, then the OOBE wizard:
DGX OS is Ubuntu 24.04 LTS for ARM64 with NVIDIA's stack pre-installed and pinned: CUDA 13.0, NVIDIA driver 580, the NVIDIA Container Toolkit, DGX-validated kernel, NCCL, and a pre-configured docker-engine. Same DGX OS family as you'd find on a DGX H100 / B200; the apt repository is shared.
~/.bashrc sources /etc/profile.d/nvidia.shapt upgrade it casuallycuda-toolkit-13-0, nvidia-container-toolkit, docker.io, nvidia-fabricmanager, nvidia-dgxconfig, htop, iperf3, git, vim, tmux, nvtopWhat's not installed: PyTorch, TensorFlow, JupyterHub, vLLM. Those come in via NGC containers — Spark expects you to live in containers, not on the host.
All pre-installed. Verify after first boot:
nvidia-smi # shows GB10 GPU + 128 GB unified
nvcc --version # CUDA 13.0
docker info | grep -i runtime # nvidia runtime present
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
docker run --rm --gpus all nvcr.io/nvidia/cuda:13.0.1-base-ubuntu24.04 nvidia-smi
The last line should reproduce the same nvidia-smi output from inside a container — that's the "everything works" smoke test. If nvidia-smi inside the container fails, recheck nvidia-ctk output in /etc/docker/daemon.json.
DGX OS pins the driver to a Production Branch. Letting apt upgrade it onto a New Feature Branch driver can break the kernel-module signing chain. Use sudo dgx-release-upgrade for major version moves; let unattended-upgrades handle security patches only.
SSH is on by default after first user creation. To use Spark headless from your laptop:
# on the Spark, find its IP
ip -4 addr | grep inet
# on your laptop:
ssh-copy-id brendan@spark.local # avahi mDNS works on Wi-Fi 7
ssh brendan@spark.local nvidia-smi
sudo nvidia-smi -pm 1 # persistence on, run on the Spark
sudo systemctl enable nvidia-persistenced
For long-running services (vLLM serving), persistence mode + auto-login + a tmux session in a systemd unit is the canonical pattern. Restart-safe vLLM:
[Unit]
Description=vLLM serving llama-3.3-70b
After=docker.service
[Service]
Restart=always
ExecStart=/usr/bin/docker run --rm --gpus all -p 8000:8000 \
-v /opt/models:/models \
vllm/vllm-openai:v0.7.0 \
--model /models/llama-3.3-70b-mxfp4 --quantization mxfp4
[Install]
WantedBy=multi-user.target
# identify the device
ip link | grep enp # typical name enp1s0f0
# bring up at 200 GbE
sudo ip link set enp1s0f0 up
sudo ip link set enp1s0f0 mtu 9000 # jumbo frames for RDMA
sudo ip addr add 10.0.0.1/24 dev enp1s0f0 # direct attach to second Spark at 10.0.0.2
# verify
sudo mlxlink -d mlx5_0 --show_advanced # speed, link state, FEC
ibstat # InfiniBand mode if QSFP cable is IB
NVIDIA NGC (nvcr.io) is the registry for NVIDIA-validated containers and models. Spark needs an NGC API key for two reasons:
nvcr.io/nim/...).# get your API key from ngc.nvidia.com → Setup → Generate API Key
docker login nvcr.io
# Username: $oauthtoken
# Password: <paste API key>
# or with the ngc CLI (pre-installed)
ngc config set
# prompts for API key, org, team; ~/.ngc/config saved
# sanity check
ngc registry image list --column name --format_type csv | head
The free tier of NGC gives access to most CUDA / PyTorch / TensorRT / TRT-LLM / vLLM / Riva / NeMo containers. The premium NIM catalog (Llama-3, Mistral, Mixtral, Gemma, etc.) is included with AI Enterprise; standalone NIM access can be arranged via the Build Pricing tier.
Three idiomatic patterns:
One docker run = ready OpenAI-compatible endpoint. Engine pre-built for Spark's GPU class.
docker run --gpus all -p 8000:8000 nvcr.io/nim/meta/llama-3.3-70b-instruct:latest
ARM64 build available natively. Pulls quants from ollama.com.
curl -fsSL https://ollama.com/install.sh | sh
ollama run llama3.3:70b
Most flexibility, your choice of quant.
docker run --gpus all -p 8000:8000 \
vllm/vllm-openai:v0.7.0 \
--model meta-llama/Llama-3.3-70B-Instruct \
--quantization mxfp4
With NIM running on port 8000, prove end-to-end:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta/llama-3.3-70b-instruct",
"messages": [{"role": "user", "content": "hello, Spark"}],
"max_tokens": 32
}'
Expected on Spark with Llama-3.3-70B MX-FP4: TTFT ~2–4 s (cold cache), then ~6–8 tok/s sustained. If you see < 1 tok/s, check: GPU busy in nvidia-smi? Persistence mode on? Quant matches GPU support?
For a real benchmark, run the genai-perf tool from NGC:
docker run --rm --network host nvcr.io/nvidia/tritonserver:24.10-py3-sdk \
genai-perf -m meta/llama-3.3-70b-instruct \
--service-kind openai --endpoint-type chat \
--num-prompts 100 --concurrency 4 \
--url http://localhost:8000
sudo apt update && sudo apt upgrade — safe for OS / userland; do not upgrade the NVIDIA driver this way.sudo dgx-release-upgrade — pulls a tested PB driver and a vetted CUDA toolkit. Reboot required.:latest on production. Pin to vllm-openai:v0.7.0 (or whatever).boot from USB → recover → restore from cloud snapshot).snapper snapshots before risky changes. snapper create -d "before-vllm-upgrade".docker manifest inspect for arch: arm64.cat /etc/dgx-release.nvcr.io/nim/meta/llama-3.3-70b-instruct:latest changes silently; pin to a digest with @sha256:... for production.Click each step as you complete it.