NVIDIA GPU Architectures Series — Presentation 34

Setting Up DGX Spark — From Box to First Token

A walkthrough of bringing a Spark online. Unbox, OS install, drivers, NGC login, container runtime, networking, and the first inference run. The intent: a checklist you can read once and follow on the day, with the right command for each step and the gotchas surfaced in advance.

DGX OSUbuntu ARM64 JetPack-styleCUDA 13.0 NGCDocker Wi-Fi 7SSH 200 GbEvLLM
Unbox → OOBE → Driver / CUDA → NGC → Container → Pull model → First token
00

Topics We'll Cover

01

What's in the Box

Standard contents

  • 1× DGX Spark unit (~1.2 kg)
  • 1× 240 W power adapter (USB-C PD on most SKUs)
  • 1× AC cord (region-specific)
  • 1× quick-start card
  • 1× recovery USB-C drive (label: "DGX OS recovery")

What's NOT in the box

  • No HDMI / DP cable — bring your own
  • No keyboard / mouse — USB or Bluetooth
  • No QSFP cable — needed only if you're pairing two Sparks (DAC ~$80, AOC ~$150)
  • No printed manual; everything's online at docs.nvidia.com
Workspace pre-flight

Before powering on: have a monitor + cable ready, a Wi-Fi network credential ready (or a wired Ethernet cable for the 10 GbE RJ-45 port), and an NVIDIA Developer account. Without an NVIDIA Developer account you can't pull NGC images, which is most of the value of the device.

02

First Boot — Out-of-Box Experience

Press the power button on the front. NVIDIA logo, ~30 seconds, then the OOBE wizard:

  1. Language & keyboard.
  2. Wi-Fi or wired connection. Wi-Fi 7 is mostly painless on launch firmware; wired 10 GbE on the RJ-45 port is more reliable for headless.
  3. Time zone.
  4. Create user. You'll be the first sudoer; pick a strong password — the box exposes SSH by default if joined to a network.
  5. EULA acceptance for DGX OS and the bundled NVIDIA AI Enterprise components.
  6. Telemetry choice. Off by default in v1; opt-in only.
  7. First boot finalises. ~5 minutes to install initial packages and run the first kernel module compile for the GPU driver. Don't unplug.
03

DGX OS — What It Is

DGX OS is Ubuntu 24.04 LTS for ARM64 with NVIDIA's stack pre-installed and pinned: CUDA 13.0, NVIDIA driver 580, the NVIDIA Container Toolkit, DGX-validated kernel, NCCL, and a pre-configured docker-engine. Same DGX OS family as you'd find on a DGX H100 / B200; the apt repository is shared.

What's not installed: PyTorch, TensorFlow, JupyterHub, vLLM. Those come in via NGC containers — Spark expects you to live in containers, not on the host.

04

Drivers, CUDA Toolkit, Container Toolkit

All pre-installed. Verify after first boot:

verify the stack
nvidia-smi                                 # shows GB10 GPU + 128 GB unified
nvcc --version                             # CUDA 13.0
docker info | grep -i runtime              # nvidia runtime present
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
docker run --rm --gpus all nvcr.io/nvidia/cuda:13.0.1-base-ubuntu24.04 nvidia-smi

The last line should reproduce the same nvidia-smi output from inside a container — that's the "everything works" smoke test. If nvidia-smi inside the container fails, recheck nvidia-ctk output in /etc/docker/daemon.json.

Don't "apt full-upgrade" the driver

DGX OS pins the driver to a Production Branch. Letting apt upgrade it onto a New Feature Branch driver can break the kernel-module signing chain. Use sudo dgx-release-upgrade for major version moves; let unattended-upgrades handle security patches only.

05

SSH, Headless Mode, Persistence

SSH is on by default after first user creation. To use Spark headless from your laptop:

headless setup
# on the Spark, find its IP
ip -4 addr | grep inet

# on your laptop:
ssh-copy-id brendan@spark.local           # avahi mDNS works on Wi-Fi 7
ssh brendan@spark.local nvidia-smi
sudo nvidia-smi -pm 1                       # persistence on, run on the Spark
sudo systemctl enable nvidia-persistenced

For long-running services (vLLM serving), persistence mode + auto-login + a tmux session in a systemd unit is the canonical pattern. Restart-safe vLLM:

/etc/systemd/system/vllm.service
[Unit]
Description=vLLM serving llama-3.3-70b
After=docker.service

[Service]
Restart=always
ExecStart=/usr/bin/docker run --rm --gpus all -p 8000:8000 \
    -v /opt/models:/models \
    vllm/vllm-openai:v0.7.0 \
    --model /models/llama-3.3-70b-mxfp4 --quantization mxfp4

[Install]
WantedBy=multi-user.target
06

Networking — Wi-Fi 7 vs 200 GbE

enable the 200 GbE port
# identify the device
ip link | grep enp                         # typical name enp1s0f0

# bring up at 200 GbE
sudo ip link set enp1s0f0 up
sudo ip link set enp1s0f0 mtu 9000           # jumbo frames for RDMA
sudo ip addr add 10.0.0.1/24 dev enp1s0f0    # direct attach to second Spark at 10.0.0.2

# verify
sudo mlxlink -d mlx5_0 --show_advanced     # speed, link state, FEC
ibstat                                     # InfiniBand mode if QSFP cable is IB
07

NGC Login & The Catalog

NVIDIA NGC (nvcr.io) is the registry for NVIDIA-validated containers and models. Spark needs an NGC API key for two reasons:

  1. Pulling premium NIM containers (nvcr.io/nim/...).
  2. Validating your AI Enterprise entitlement, if you've bought one with the box.
NGC login
# get your API key from ngc.nvidia.com → Setup → Generate API Key
docker login nvcr.io
# Username: $oauthtoken
# Password: <paste API key>

# or with the ngc CLI (pre-installed)
ngc config set
# prompts for API key, org, team; ~/.ngc/config saved

# sanity check
ngc registry image list --column name --format_type csv | head

The free tier of NGC gives access to most CUDA / PyTorch / TensorRT / TRT-LLM / vLLM / Riva / NeMo containers. The premium NIM catalog (Llama-3, Mistral, Mixtral, Gemma, etc.) is included with AI Enterprise; standalone NIM access can be arranged via the Build Pricing tier.

08

Pulling Your First Model — NIM, Ollama, or HF

Three idiomatic patterns:

1. NIM (one container)

One docker run = ready OpenAI-compatible endpoint. Engine pre-built for Spark's GPU class.

docker run --gpus all -p 8000:8000 nvcr.io/nim/meta/llama-3.3-70b-instruct:latest

2. Ollama (zero friction)

ARM64 build available natively. Pulls quants from ollama.com.

curl -fsSL https://ollama.com/install.sh | sh
ollama run llama3.3:70b

3. vLLM + HF model

Most flexibility, your choice of quant.

docker run --gpus all -p 8000:8000 \
    vllm/vllm-openai:v0.7.0 \
    --model meta-llama/Llama-3.3-70B-Instruct \
    --quantization mxfp4
09

First Inference — The Smoke Test

With NIM running on port 8000, prove end-to-end:

first request
curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "meta/llama-3.3-70b-instruct",
        "messages": [{"role": "user", "content": "hello, Spark"}],
        "max_tokens": 32
    }'

Expected on Spark with Llama-3.3-70B MX-FP4: TTFT ~2–4 s (cold cache), then ~6–8 tok/s sustained. If you see < 1 tok/s, check: GPU busy in nvidia-smi? Persistence mode on? Quant matches GPU support?

For a real benchmark, run the genai-perf tool from NGC:

benchmark with genai-perf
docker run --rm --network host nvcr.io/nvidia/tritonserver:24.10-py3-sdk \
    genai-perf -m meta/llama-3.3-70b-instruct \
    --service-kind openai --endpoint-type chat \
    --num-prompts 100 --concurrency 4 \
    --url http://localhost:8000
10

Updates & Rollback

11

Common Gotchas the First Day

12

Interactive: Setup Step Tracker

Click each step as you complete it.

Done
0
Remaining
12
% complete
0%