From blank Linux box to an OpenAI-compatible vLLM endpoint: NVIDIA Container Toolkit, the vllm/vllm-openai image, docker compose, health checks, resource limits, and the mistakes that eat your evenings.
Docker can't touch a GPU on its own — it needs the NVIDIA Container Toolkit, which installs a runtime that mounts the driver and exposes the GPU device nodes into the container.
nvidia-ctk)nvidia runtime)# 1. NVIDIA driver (pick a version compatible with your CUDA)
sudo apt install -y nvidia-driver-550-server
# reboot, then check:
nvidia-smi
# 2. NVIDIA Container Toolkit repo + install
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
sudo gpg --dearmor -o /etc/apt/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb #deb [signed-by=/etc/apt/keyrings/nvidia-container-toolkit-keyring.gpg] #' \
| sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update && sudo apt install -y nvidia-container-toolkit
# 3. Register the runtime with Docker
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
# Expected output: same GPU listing as the host.
# If it errors "Failed to initialize NVML" → re-run `nvidia-ctk runtime configure`
# and restart Docker. If it says "no GPU" → check /etc/docker/daemon.json
| Flag | Meaning |
|---|---|
--gpus all | Every GPU on the host |
--gpus 1 | First available GPU |
--gpus '"device=0,2"' | Pin specific GPU ordinals (tensor-parallel needs adjacent indices) |
--gpus '"device=GPU-abc123…"' | Pin by UUID (stable across reboots) |
--runtime=nvidia | Legacy form, still works |
--shm-size=16g | Critical — vLLM uses POSIX shm for worker IPC with TP |
--ipc=host | Alternative to --shm-size, simpler |
--ulimit memlock=-1 | Required by NCCL for pinned memory |
--ulimit stack=67108864 | Avoids NCCL thread stack exhaustion |
The official image at vllm/vllm-openai is a CUDA-ready container with vLLM preinstalled and an OpenAI-compatible HTTP server as its entrypoint. Tags follow vLLM versions: v0.6.3, v0.7.0, latest. Never blindly chase :latest in production; pin a version.
python3 -m vllm.entrypoints.openai.api_serverrocm/vllmvllm/vllm-cpu| Host | Container | Why |
|---|---|---|
~/hf-cache | /root/.cache/huggingface | Persistent model cache across restarts |
/mnt/models | /models | Pre-downloaded weights / private fine-tunes |
~/vllm-logs | /var/log/vllm | Optional; stdout is fine for most |
docker run --rm -d \
--name vllm-llama3 \
--gpus all --ipc=host \
--ulimit memlock=-1 --ulimit stack=67108864 \
-p 8000:8000 \
-v ~/hf-cache:/root/.cache/huggingface \
-e HF_TOKEN=$HF_TOKEN \
vllm/vllm-openai:v0.7.0 \
--model meta-llama/Meta-Llama-3.1-8B-Instruct \
--dtype bfloat16 \
--max-model-len 16384 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enforce-eager=false
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Meta-Llama-3.1-8B-Instruct",
"messages": [{"role":"user","content":"Say hi"}]
}'
curl http://localhost:8000/v1/models # list
curl http://localhost:8000/health # liveness
curl http://localhost:8000/metrics # Prometheus
--gpu-memory-utilization 0.90: vLLM reserves this fraction of VRAM for weights + KV pool. Lower = safer; higher = more concurrent sessions. --max-model-len: caps context. --enforce-eager=false enables CUDA graphs for ~10–20% throughput. --enable-prefix-caching is almost always free throughput.
A single compose file that you can systemctl enable gives you a restartable, logged, monitored serve.
services:
vllm:
image: vllm/vllm-openai:v0.7.0
restart: unless-stopped
shm_size: '16gb'
ulimits:
memlock: -1
stack: 67108864
deploy:
resources:
reservations:
devices:
- capabilities: [gpu]
driver: nvidia
count: all
ports:
- "8000:8000"
environment:
HF_TOKEN: ${HF_TOKEN}
VLLM_LOGGING_LEVEL: INFO
NCCL_DEBUG: WARN
volumes:
- hf-cache:/root/.cache/huggingface
- /mnt/models:/models:ro
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 15s
timeout: 3s
retries: 6
start_period: 240s # model load takes time
command: >
--model meta-llama/Meta-Llama-3.1-8B-Instruct
--dtype bfloat16
--max-model-len 16384
--gpu-memory-utilization 0.90
--enable-prefix-caching
--max-num-seqs 128
--disable-log-requests
prometheus:
image: prom/prometheus:latest
volumes: [ "./prom.yml:/etc/prometheus/prometheus.yml:ro" ]
ports: [ "9090:9090" ]
volumes:
hf-cache: {}
deploy:
resources:
reservations:
devices:
- capabilities: [gpu]
driver: nvidia
device_ids: ["0", "1"]
command: >
--model meta-llama/Meta-Llama-3.1-70B-Instruct-AWQ
--tensor-parallel-size 2
--dtype half
--max-model-len 8192
--gpu-memory-utilization 0.92
Tensor-parallel requires a working NCCL path between GPUs. PCIe works; NVLink or NVSwitch helps. Deck 05 covers the details.
Pick your model and hardware; the builder prints a ready-to-run docker run line.
INFO Scheduled ... | running: N waiting: M — queue depthINFO gpu_cache_blocks=<k> — KV pool capacityWARN No available block groups — KV saturation, requests are being preemptedWARN nccl ... — TP/P2P issues; check NVLink / PCIe/metricsvllm:num_requests_runningvllm:num_requests_waitingvllm:gpu_cache_usage_percvllm:time_to_first_token_secondsvllm:time_per_output_token_secondsvllm:request_prompt_tokens / generation_tokensDon't autoscale on GPU %; it's always ~100% under any real load. Scale on num_requests_waiting > X for N minutes, or TTFT p95 > your SLO. That is the number your users actually feel.
Model + KV pool + NCCL scratch won't fit. Lower --gpu-memory-utilization, or shorten --max-model-len, or use a smaller dtype / quant.
Usually missing --ipc=host or --shm-size; for multi-GPU add --ulimit memlock=-1. On Spark / GH200 set NCCL_TOPO_FILE or rely on autodetect.
Weights are streaming from HuggingFace into the cache. Pre-pull onto a local NVMe and mount it read-only; or use huggingface-cli download before first run.
Some models or flash-attn versions don't graph cleanly. Fall back with --enforce-eager; you lose ~15% throughput but startup stabilises.
Continuous batching reorders work between submit and complete. If your client depends on strict sequential request ordering, you have a bug. Deck 09 covers non-determinism in depth.
TP across many GPUs uses a lot of POSIX shm. Bump shm_size to 32 GB for TP ≥ 4. Better: use --ipc=host and forget it.
Drop the compose file in /etc/vllm/, create a systemd unit ExecStart=/usr/bin/docker compose up, ExecStop=docker compose down, Restart=on-failure. You now have a GPU-backed inference server that survives reboots with no shell sessions open.