Local LLM Hosting Series — Presentation 04

vLLM in Docker on NVIDIA GPUs

From blank Linux box to an OpenAI-compatible vLLM endpoint: NVIDIA Container Toolkit, the vllm/vllm-openai image, docker compose, health checks, resource limits, and the mistakes that eat your evenings.

Docker NVIDIA Container Toolkit vllm/vllm-openai docker compose systemd Prometheus
Driver → Toolkit → Image → Run → Compose → Operate
00

Topics We'll Cover

01

Host Prerequisites

Docker can't touch a GPU on its own — it needs the NVIDIA Container Toolkit, which installs a runtime that mounts the driver and exposes the GPU device nodes into the container.

Linux kernel
↓
NVIDIA kernel driver (535 / 550 / 560 / …)
↓
NVIDIA Container Toolkit (nvidia-ctk)
↓
Docker Engine (uses nvidia runtime)
↓
vllm/vllm-openai:latest container
Ubuntu 24.04 — one-time host install
# 1. NVIDIA driver (pick a version compatible with your CUDA)
sudo apt install -y nvidia-driver-550-server
# reboot, then check:
nvidia-smi

# 2. NVIDIA Container Toolkit repo + install
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
  sudo gpg --dearmor -o /etc/apt/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
  sed 's#deb #deb [signed-by=/etc/apt/keyrings/nvidia-container-toolkit-keyring.gpg] #' \
  | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

sudo apt update && sudo apt install -y nvidia-container-toolkit

# 3. Register the runtime with Docker
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
02

Verifying GPU Access in Docker

Smoke test
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

# Expected output: same GPU listing as the host.
# If it errors "Failed to initialize NVML" → re-run `nvidia-ctk runtime configure`
# and restart Docker. If it says "no GPU" → check /etc/docker/daemon.json

Docker flags that expose GPUs

FlagMeaning
--gpus allEvery GPU on the host
--gpus 1First available GPU
--gpus '"device=0,2"'Pin specific GPU ordinals (tensor-parallel needs adjacent indices)
--gpus '"device=GPU-abc123…"'Pin by UUID (stable across reboots)
--runtime=nvidiaLegacy form, still works
--shm-size=16gCritical — vLLM uses POSIX shm for worker IPC with TP
--ipc=hostAlternative to --shm-size, simpler
--ulimit memlock=-1Required by NCCL for pinned memory
--ulimit stack=67108864Avoids NCCL thread stack exhaustion
03

The vllm/vllm-openai Image

The official image at vllm/vllm-openai is a CUDA-ready container with vLLM preinstalled and an OpenAI-compatible HTTP server as its entrypoint. Tags follow vLLM versions: v0.6.3, v0.7.0, latest. Never blindly chase :latest in production; pin a version.

What's baked in

  • Python 3.11 + vLLM + dependencies
  • PyTorch matching vLLM's supported CUDA
  • NCCL, Flash-Attention 3, Triton
  • Entrypoint: python3 -m vllm.entrypoints.openai.api_server

Architectures

  • amd64 — Ampere, Ada, Hopper, Blackwell
  • arm64 — Grace-Blackwell (DGX Spark, GH200)
  • ROCm variant: rocm/vllm
  • CPU-only variant for laptops/testing: vllm/vllm-cpu

Volumes you'll want

HostContainerWhy
~/hf-cache/root/.cache/huggingfacePersistent model cache across restarts
/mnt/models/modelsPre-downloaded weights / private fine-tunes
~/vllm-logs/var/log/vllmOptional; stdout is fine for most
04

One-liner — Serve Llama-3.1-8B

Minimal, working, single-GPU
docker run --rm -d \
    --name vllm-llama3 \
    --gpus all --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864 \
    -p 8000:8000 \
    -v ~/hf-cache:/root/.cache/huggingface \
    -e HF_TOKEN=$HF_TOKEN \
    vllm/vllm-openai:v0.7.0 \
    --model meta-llama/Meta-Llama-3.1-8B-Instruct \
    --dtype bfloat16 \
    --max-model-len 16384 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching \
    --enforce-eager=false

Smoke test

Query it
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3.1-8B-Instruct",
    "messages": [{"role":"user","content":"Say hi"}]
  }'

curl http://localhost:8000/v1/models       # list
curl http://localhost:8000/health          # liveness
curl http://localhost:8000/metrics         # Prometheus
Flag essentials

--gpu-memory-utilization 0.90: vLLM reserves this fraction of VRAM for weights + KV pool. Lower = safer; higher = more concurrent sessions. --max-model-len: caps context. --enforce-eager=false enables CUDA graphs for ~10–20% throughput. --enable-prefix-caching is almost always free throughput.

05

docker compose for Serious Use

A single compose file that you can systemctl enable gives you a restartable, logged, monitored serve.

docker-compose.yml
services:
  vllm:
    image: vllm/vllm-openai:v0.7.0
    restart: unless-stopped
    shm_size: '16gb'
    ulimits:
      memlock: -1
      stack: 67108864
    deploy:
      resources:
        reservations:
          devices:
            - capabilities: [gpu]
              driver: nvidia
              count: all
    ports:
      - "8000:8000"
    environment:
      HF_TOKEN: ${HF_TOKEN}
      VLLM_LOGGING_LEVEL: INFO
      NCCL_DEBUG: WARN
    volumes:
      - hf-cache:/root/.cache/huggingface
      - /mnt/models:/models:ro
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 15s
      timeout: 3s
      retries: 6
      start_period: 240s      # model load takes time
    command: >
      --model meta-llama/Meta-Llama-3.1-8B-Instruct
      --dtype bfloat16
      --max-model-len 16384
      --gpu-memory-utilization 0.90
      --enable-prefix-caching
      --max-num-seqs 128
      --disable-log-requests

  prometheus:
    image: prom/prometheus:latest
    volumes: [ "./prom.yml:/etc/prometheus/prometheus.yml:ro" ]
    ports: [ "9090:9090" ]

volumes:
  hf-cache: {}

Tensor-parallel on two GPUs in compose

compose fragment — 2× RTX 4090 TP=2
deploy:
  resources:
    reservations:
      devices:
        - capabilities: [gpu]
          driver: nvidia
          device_ids: ["0", "1"]
command: >
  --model meta-llama/Meta-Llama-3.1-70B-Instruct-AWQ
  --tensor-parallel-size 2
  --dtype half
  --max-model-len 8192
  --gpu-memory-utilization 0.92

Tensor-parallel requires a working NCCL path between GPUs. PCIe works; NVLink or NVSwitch helps. Deck 05 covers the details.

06

Interactive: Command Builder

Pick your model and hardware; the builder prints a ready-to-run docker run line.

1
0.90
128
07

Operating It — Logs, Metrics, Health

Logs worth watching

  • INFO Scheduled ... | running: N waiting: M — queue depth
  • INFO gpu_cache_blocks=<k> — KV pool capacity
  • WARN No available block groups — KV saturation, requests are being preempted
  • WARN nccl ... — TP/P2P issues; check NVLink / PCIe

Prometheus /metrics

  • vllm:num_requests_running
  • vllm:num_requests_waiting
  • vllm:gpu_cache_usage_perc
  • vllm:time_to_first_token_seconds
  • vllm:time_per_output_token_seconds
  • vllm:request_prompt_tokens / generation_tokens

A Grafana dashboard with four panels is enough

  1. TTFT p50/p95 (latency budget for tool calls, UIs)
  2. TPOT (time per output token) — generation speed
  3. KV usage % (tells you when to add concurrency caps)
  4. Requests running vs waiting (saturation indicator)
Autoscaling pragma

Don't autoscale on GPU %; it's always ~100% under any real load. Scale on num_requests_waiting > X for N minutes, or TTFT p95 > your SLO. That is the number your users actually feel.

08

Gotchas & How to Avoid Them

OOM at startup

Model + KV pool + NCCL scratch won't fit. Lower --gpu-memory-utilization, or shorten --max-model-len, or use a smaller dtype / quant.

"Failed to initialize NCCL"

Usually missing --ipc=host or --shm-size; for multi-GPU add --ulimit memlock=-1. On Spark / GH200 set NCCL_TOPO_FILE or rely on autodetect.

Startup takes forever

Weights are streaming from HuggingFace into the cache. Pre-pull onto a local NVMe and mount it read-only; or use huggingface-cli download before first run.

CUDA graph compile crash

Some models or flash-attn versions don't graph cleanly. Fall back with --enforce-eager; you lose ~15% throughput but startup stabilises.

Requests silently reordered

Continuous batching reorders work between submit and complete. If your client depends on strict sequential request ordering, you have a bug. Deck 09 covers non-determinism in depth.

Out of shared memory

TP across many GPUs uses a lot of POSIX shm. Bump shm_size to 32 GB for TP ≥ 4. Better: use --ipc=host and forget it.

systemd-ify it

Drop the compose file in /etc/vllm/, create a systemd unit ExecStart=/usr/bin/docker compose up, ExecStop=docker compose down, Restart=on-failure. You now have a GPU-backed inference server that survives reboots with no shell sessions open.