A grounded tour of the 2026 NVIDIA stack for LLM serving: which class of card for which model, what the memory type actually does to tok/s, how to gang GPUs together (including heterogeneous pairs), MIG, and the Ollama/vLLM-specific gotchas per architecture.
A practical, opinionated guide to picking, pairing, and operating NVIDIA GPUs for local LLM hosting — with the nuances that bite at 2 AM.
Five architecture families are relevant to local LLM serving today. Understanding which generation a card belongs to is the single biggest predictor of software behaviour.
| Arch | CC | Year | Examples | Tensor core native | LLM notes |
|---|---|---|---|---|---|
| Ampere | 8.0 / 8.6 | 2020–21 | A100, RTX 30xx, A40, A6000 | FP16, BF16, TF32 | No FP8. Workhorse for INT4 serving. |
| Ada Lovelace | 8.9 | 2022–23 | RTX 40xx, L40, L4, L40S | FP16, BF16, FP8 (all Ada) | Consumer king; no NVLink on 4090. |
| Hopper | 9.0 | 2023–24 | H100, H200, GH200 | FP16, BF16, FP8 (E4M3/E5M2) | First real FP8. Transformer Engine. |
| Blackwell | 10.0 / 12.0 | 2025–26 | B100, B200, RTX 50xx, GB200 | FP16, BF16, FP8, FP4 (MX) | Native FP4 doubles inference again. |
| Grace-Blackwell | 10.0 / 12.1 | 2025–26 | GB10 (Spark), GB200 NVL72 | same as Blackwell | Arm64 host, unified memory, C2C link. |
RTX 30xx / 40xx / 50xx. GDDR6 / 6X / 7. No ECC. No SR-IOV. Driver licence forbids datacenter deployment. On 40-series no NVLink. You're on your own for multi-GPU scaling.
RTX A6000, A6000 Ada, RTX PRO 6000 Blackwell (96 GB). GDDR6/7 ECC, proper driver, sometimes NVLink bridge. Datacenter-licensed. Sweet spot for serious lab work.
A100, H100, H200, B100, B200, L40S, L4. HBM or GDDR6 ECC, SR-IOV, MIG on A100/H100/B200, NVLink 3/4, NVSwitch optional, proper persistence and error recovery. What the frontier APIs actually run on.
NVIDIA's GeForce EULA restricts datacenter deployment of consumer cards. "Datacenter" is loosely defined but colocation, multi-tenant cloud, and large-scale commercial hosting are clearly out. For home / office / single-tenant internal use it's fine. If you're standing up a commercial service, use L4/L40S/A100/H100/B200 class hardware.
Decode tok/s is bounded by memory bandwidth, not FLOPS. Every token requires reading all weights once. The memory type therefore decides the upper bound on single-stream tok/s more than anything else.
| Type | Where | Characteristics |
|---|---|---|
| GDDR6 | A6000, 3060, 4060/70, L4 | Cheap per GB, ~400–800 GB/s, no ECC on consumer |
| GDDR6X | 3090 / 3090 Ti, 4080/90 | PAM4 signalling, higher data rate; same capacity tiers |
| GDDR7 | 5090, RTX PRO 6000 Blackwell | PAM3, ~2× GDDR6X; up to 96 GB on RTX PRO 6000 |
| HBM2e | A100 | On-package stacks; 80 GB, ~2 TB/s (the 40 GB A100 is HBM2, ~1.6 TB/s) |
| HBM3 | H100 | 80 GB, ~3.4 TB/s |
| HBM3e | H200, B100, B200, GB200 | 96–192 GB, 4.8–8 TB/s |
| LPDDR5x (unified) | GB10 Spark, GH200 | Huge capacity (128 GB+), moderate bandwidth (~0.3–0.5 TB/s); coherent with CPU via NVLink-C2C |
| EGM (Extended GPU Memory) | GH200, GB200 | GPU sees host LPDDR5x as extended VRAM over the C2C link — transparent to the kernel |
Single-stream tok/s on a 70B model is roughly min(bandwidth / weight_bytes, compute_limit). With 70B FP8 (70 GB weights), an H200 (4.8 TB/s) is theoretically capped at ~68 tok/s per stream; a pair of RTX 4090s layer-splitting a 40 GB INT4 quant (it doesn't fit one 24 GB card) peaks around ~25 tok/s. If you need raw tok/s on a single stream, spend on bandwidth; if you need concurrency, spend on capacity.
Two cards with the same VRAM can run very different software stacks because their compute capability (CC) determines which tensor-core formats, PTX intrinsics, and CUDA features the kernels can use.
| Feature | Ampere (8.x) | Ada (8.9) | Hopper (9.0) | Blackwell (10.0/12.0) |
|---|---|---|---|---|
| TF32 / BF16 / FP16 tensor cores | ✓ | ✓ | ✓ | ✓ |
| FP8 E4M3 / E5M2 tensor cores | — | ✓ | ✓ | ✓ |
| FP4 (MX) tensor cores | — | — | — | ✓ |
| 2:4 structured sparsity | ✓ | ✓ | ✓ | ✓ |
| Transformer Engine (auto FP8) | — | partial | ✓ | ✓ |
| Thread-Block Clusters | — | — | ✓ | ✓ |
| Distributed Shared Memory (DSMEM) | — | — | ✓ | ✓ |
| TMA (Tensor Memory Accelerator) | — | — | ✓ | ✓ |
| MIG (Multi-Instance GPU) | A100 only | — | ✓ | B200 only |
| NVLink 3/4/5 | A100 (3) | no on 4090 | ✓ (4) | ✓ (5) |
| Confidential Computing | — | — | ✓ | ✓ |
Every serious 2025–2026 model family ships FP8 checkpoints because Hopper onwards runs them at full tensor-core speed. On Ampere (A100, 3090, 4090's predecessor) FP8 is emulated — you pay the memory win but not the compute win. For most decoding that's still a good deal (decoding is bandwidth-bound), but prefill throughput lags Hopper by a lot.
Ollama wraps llama.cpp. The CUDA backend is built with GGML_CUDA=1 and uses the CUDA runtime + cuBLAS + llama.cpp's own kernels. Behaviour varies by arch.
| Card class | What actually happens | Practical notes |
|---|---|---|
| GTX 10xx / 16xx (Pascal, Turing) | Works via CUDA backend; no tensor cores → pure FP32 SMs carry matmul. | GGUF q4 on a 1080 Ti still runs 7B at ~15 tok/s. Fine for demos, not much else. |
| RTX 20xx / 30xx (Turing / Ampere) | Tensor cores used for FP16/BF16 in matmul. OLLAMA_FLASH_ATTENTION=1 switches to flash-attn kernel. |
Pin the model on NVMe; 3090's 24 GB runs 13B q5_K_M at 35–45 tok/s. |
| RTX 40xx (Ada) | Full BF16 tensor path. No NVLink on 4090 — multi-GPU is PCIe P2P only. | Sweet spot for solo Ollama on 7–32B. q4_K_M quant + FA on. |
| RTX 50xx (Blackwell consumer) | GDDR7 — memory bandwidth jumps ~1.8× vs 40xx. No FP4 in llama.cpp yet (2026), so you don't get the Blackwell inference speedup via Ollama. | 5090's 32 GB VRAM fits 32B q4 or 14B BF16 comfortably. |
| A100 / A40 / A6000 | Full BF16, ECC on; Ollama treats it like any CUDA GPU. A100 MIG instances show up as separate GPUs — Ollama will happily bind to one. | Datacenter cards are wasted on single-user Ollama. Still fine for RAG side-cars. |
| H100 / H200 | BF16 path; no FP8 in llama.cpp through Ollama (2026). Hopper-specific kernels largely unused. | Throws a lot of silicon away. Use vLLM if you've paid for an H100. |
| DGX Spark (GB10) | Arm64 build; unified memory means model load is extremely fast. FP4 unused. | Good dev target; 128 GB means 70B q4 comfortably. |
| GB200 / B200 | Overkill; llama.cpp/Ollama uses almost none of Blackwell's new tensor capabilities. | If you have one, run vLLM or TensorRT-LLM, not Ollama. |
CUDA_VISIBLE_DEVICES=0,1 — pin specific GPUs; useful when multi-tenantOLLAMA_SCHED_SPREAD=1 — spread a single model across all visible GPUs (experimental layer-splitting)OLLAMA_FLASH_ATTENTION=1 — flash-attention kernel (CC 8.0+)OLLAMA_KV_CACHE_TYPE=q8_0 — KV-quant, doubles effective context without VRAM costOLLAMA_NUM_PARALLEL=N — N concurrent sessions against the same loaded model (simulated batching)OLLAMA_MAX_LOADED_MODELS=N — how many models to keep resident; for a multi-GPU box default to number of GPUsGGML_CUDA_ENABLE_UNIFIED_MEMORY=1 — let CUDA page weights via unified memory; huge help on unified-memory hosts (Spark) but slower on discrete GPUsOllama has basic layer-splitting across multiple GPUs (ngl / tensor_split) but it is not tensor-parallel in the vLLM sense. It splits layers across cards and pipelines them — which means slower per-token than a single big card would be, and it does not help concurrent users at all. For real multi-GPU local serving, use vLLM.
vLLM's performance story is much closer to "what the hardware can actually do". It does ship arch-specific kernels and it's picky about them.
| Card class | What vLLM uses | Practical notes |
|---|---|---|
| Turing (RTX 20xx) | FP16 path only. Flash-attn v2 supported; v3 is Hopper-only. | Marginal; 20xx cards lack BF16 tensor cores, cuts throughput. |
| Ampere (A100, 3090, 3090 Ti, A40, A6000) | BF16 native, FP8 emulated (compute no faster than BF16; bandwidth win still applies). Flash-attn v2. PagedAttention works fully. | AWQ / GPTQ INT4 is the best point on this arch. Multi-GPU via NVLink on A100 / A6000, PCIe on 3090. |
| Ada (RTX 40xx, L4, L40, L40S) | BF16 + FP8 (all Ada, CC 8.9). No NVLink on 40-series. CUDA graphs fully supported. | Consumer Ada is great for single-GPU vLLM; multi-GPU scaling is rough without NVLink. |
| Hopper (H100 / H200) | Native FP8 tensor cores → 2× throughput vs BF16 on prefill. Flash-attn v3. Thread-block clusters. Transformer Engine via --quantization fp8. |
The default production choice in 2026. Use --kv-cache-dtype fp8 too. |
| Blackwell (B100 / B200 / GB200) | Adds native FP4 / MX-FP4 → another 2× on compute. vLLM 0.7+ and TensorRT-LLM 0.12+ required. | Highest ceiling; cheapest per token if utilisation is high. |
| DGX Spark (GB10) | Blackwell GPU + 128 GB unified. vLLM arm64 image. FP8/FP4 supported. Bandwidth-bound at LPDDR5x speeds. | Local workstation; see deck 08. |
| Flag | Ampere | Ada | Hopper | Blackwell |
|---|---|---|---|---|
FP8 checkpoint (auto-detected; --dtype has no fp8) | W8A16 via Marlin; bandwidth win only | native | native | native |
--quantization fp8 (TE) | — | native | native | native |
--quantization awq | ✓ | ✓ | ✓ | ✓ |
--quantization fp4 / mxfp4 | — | — | — | native |
--kv-cache-dtype fp8 | emulated | native | native | native |
--enforce-eager | safe | safe | safe | sometimes needed on early B200 kernels |
--enable-chunked-prefill | ✓ | ✓ | big win | big win |
Deck 05 covered the software axes (TP/PP/DP/EP). Here we cover what the physical link between GPUs buys you and where it goes wrong.
A high-speed, coherent point-to-point GPU↔GPU link. Three variants you'll meet:
Used for TP all-reduces and P2P memcpy. Not present on any 40-series; NVIDIA removed the bridge. 5090 reportedly also without.
A packet-switched fabric so that every GPU in a node talks to every other at full NVLink speed. DGX boxes and HGX baseboards. Turns a group of GPUs into one NVLink domain. TP across 8 H100s works because of NVSwitch.
Fallback on any NVLink-less pair. Per-lane bandwidth is bus-limited: PCIe 4 x16 ≈ 32 GB/s, PCIe 5 x16 ≈ 64 GB/s. Works but kills TP scaling.
Enablement:
iommu=pt or intel_iommu=off)lspci -vv | grep -E 'IOMMU|ACS'Cross-node path. NDR 400 = ~50 GB/s per GPU. GPUDirect RDMA bypasses the host so tensors move GPU→GPU without a PCIe-to-host-to-PCIe hop. Required for any useful cross-node TP.
nvidia-smi topo -m # matrix: NV#, PIX, PXB, PHB, SYS
nvidia-smi nvlink -s # link state per card
nvidia-smi -q -d TOPOLOGY | head -60
# Verify P2P actually works (simpleP2P from cuda-samples, or):
CUDA_VISIBLE_DEVICES=0,1 python -c "import torch;\
a=torch.randn(1,device='cuda:0');\
b=a.to('cuda:1');\
print('p2p ok',b.device)"
# Raw bandwidth (GPU 0 → GPU 1, H2H / D2D):
./bandwidthTest --device=0 --dtod
If nvidia-smi topo -m shows SYS between your two GPUs, they're separated by a CPU socket. All P2P traffic hops through the CPU's UPI / Infinity Fabric — catastrophic for TP. Physically move a card, or pin the workload to the GPUs that are both on the same socket with CUDA_VISIBLE_DEVICES.
The question everyone asks: "I have a 4090 and a 3090. Can I gang them?" The answer depends entirely on how you're "ganging" them.
| Scenario | Feasibility | Why |
|---|---|---|
| Run one model on each card (DP) | easy | Independent processes, independent CUDA contexts. Completely fine to pair any NVIDIA GPUs this way. Use CUDA_VISIBLE_DEVICES. |
Ollama OLLAMA_SCHED_SPREAD across mixed cards |
works | llama.cpp's layer split is indifferent to matched cards. Layer placement respects reported VRAM sizes. Slow card sets pace. |
| vLLM Tensor Parallel across mismatched cards | no | TP requires identical tensor shapes and bit-exact synchronised matmuls every layer. NCCL + Flash-Attention + CUTLASS all assume homogeneous GPUs. vLLM checks and refuses. |
| vLLM Pipeline Parallel across mismatched cards | technically possible, painful | Stages only exchange hidden states, so in principle stages can be on different GPUs. In practice vLLM/Ray still want matched ranks; you must set --pipeline-parallel-size and assign stages manually. Slowest stage sets throughput; bubble recovery is harder. |
| vLLM DP across mismatched cards behind a router | works | Each replica is its own vLLM process with its own flags. Different quants per replica are fine; router can even shape traffic towards the faster one. |
| Mix Ampere + Hopper (e.g. 3090 + H100) | no TP | Different compute capabilities → different kernel binaries. vLLM will load one arch's wheel, not the other's. |
| Mix 4090 + 4080 (same arch, different VRAM) | DP yes, TP no | Same CC but different VRAM sizes → unequal KV pools. TP wants identical shards per GPU; DP doesn't care. |
| A6000 48 GB + A6000 Ada 48 GB | no TP | Same memory but different arch (Ampere vs Ada). NCCL can still talk but vLLM TP will refuse because the kernels differ. |
Heterogeneous = DP-only. If you have mixed GPUs, run one vLLM per GPU (sized appropriately for that card's VRAM), put a load balancer in front, and call it a day. It's simpler, it actually works, and you lose very little vs a theoretical "ideal" unified deployment.
# Replica A on the 4090 — Llama-3.1-8B FP16 for fastest TTFT
CUDA_VISIBLE_DEVICES=0 docker run -d --name vllm-fast \
--gpus '"device=0"' -p 8001:8000 ... \
--model meta-llama/Meta-Llama-3.1-8B-Instruct --dtype bfloat16
# Replicas B/C on the 3090s — larger model, INT4
CUDA_VISIBLE_DEVICES=1 docker run -d --name vllm-big-a \
--gpus '"device=1"' -p 8002:8000 ... \
--model casperhansen/mistral-small-24b-awq --quantization awq
CUDA_VISIBLE_DEVICES=2 docker run -d --name vllm-big-b \
--gpus '"device=2"' -p 8003:8000 ... \
--model casperhansen/mistral-small-24b-awq --quantization awq
# LiteLLM router routes by model name in front of the three ports
MIG (Multi-Instance GPU) slices one physical GPU into up to 7 independent GPUs, each with dedicated SMs, L2, and memory slices. Unlike vGPU it's hardware-partitioned: no noisy-neighbour contention.
sudo nvidia-smi -mig 1 # enable MIG mode
sudo nvidia-smi mig -cgi 9,14,19,19 # create GIs (profile IDs vary per GPU)
sudo nvidia-smi mig -cci # create CIs
nvidia-smi -L # lists each slice with MIG UUID
# Launch a container bound to one slice
docker run --gpus '"device=MIG-abcd1234-..."' vllm/vllm-openai:v0.7.0 ...
No NVLink between MIG instances, no P2P across them. So MIG + TP is not a thing. MIG is for many small workloads, not one big sharded one. Also: only A100, A30, H100, H200, B200 and RTX PRO 6000 Blackwell (up to 4 instances) support MIG (not A6000, L40S, 4090).
Pick a GPU. See what model sizes and framework features are practical.
Pick two GPUs and a framework. The planner tells you which ganging modes actually work.
sudo nvidia-smi -pm 1. Without it the driver unloads between processes — adds ~30 s to every container start. Datacenter cards have it on by default; consumer cards don't.
BIOS feature; lets the CPU address the GPU's full VRAM as one PCIe aperture. Without it, large tensor transfers are chunked through BAR1 (a ~256 MB window) — cold loads slow by 2–5×. Always enable on modern boards.
Datacenter cards have ECC on by default; it costs ~6% of VRAM. sudo nvidia-smi -e 0 reclaims that capacity (reboot required). Do it for inference-only hosts where you don't care about single-bit flips; leave it on for training.
sudo nvidia-smi -pl 350 (watts). A 4090 capped at 350 W loses ~5% throughput but ~15% power. Useful in dense servers. Check nvidia-smi -q -d POWER for envelopes.
IOMMU groups determine which devices can do P2P. On consumer boards multiple GPUs often end up in the same group (fine) or blocked by PCIe switches. dmesg | grep iommu, lspci -vv | grep ACSCtl. Don't patch ACS on production hardware.
On multi-socket CPUs, GPU-to-CPU traffic hops sockets unless you pin. numactl --cpunodebind=0 --membind=0 vllm serve ... for the GPU on NUMA 0. Worth 5–10% on prefill throughput.
WSL2 works for dev but loses ~10–15% to the translation layer. No MIG, no SR-IOV. For serving, use bare Linux.
Match driver to CUDA. Blackwell needs driver ≥ 570 (CUDA 12.8); Hopper FP8 works from 525+; some flash-attn 3 kernels require ≥ 560. Pin the driver on serving hosts.
vLLM captures a graph of the decode step; can save 10–20% on small batch sizes. Sometimes crashes at load on bleeding-edge CUDA; fall back with --enforce-eager.
Not a joke: a 4090 at 85°C throttles. Check nvidia-smi -q -d TEMPERATURE. In a 1U rack a passive-cooled A100/H100 must have front-to-back airflow; blower-style consumer cards may not fit at all.
| I want to… | Buy this | Framework | Quant |
|---|---|---|---|
| Run a 7B locally, no fuss | Used RTX 3060 12 GB | Ollama | q4_K_M |
| Run a 13–30B locally, fast | RTX 4090 / 5090 / A6000 Ada | Ollama or vLLM (DP) | q5_K_M or AWQ-INT4 |
| Serve 10–50 users on 70B | H100 / H200 / RTX PRO 6000 Blackwell | vLLM | FP8 + FP8 KV |
| Local 70B at home | DGX Spark 128 GB | vLLM arm64 | FP8 / AWQ-INT4 |
| Serve 400B / MoE on-prem | 8× B200 (HGX) / GB200 NVL72 | vLLM TP=8 or TRT-LLM | FP8 weights + FP8 KV, MX-FP4 if Blackwell |
| Internal dev GPU pool | H100 80 GB + MIG | vLLM per slice | AWQ-INT4 for small slices, FP8 for big |
| Cheap inference service | L40S 48 GB × N | vLLM (DP), LB | FP8 |
| Long-context embeddings | RTX 3090 24 GB | Ollama or vLLM | FP16 weights + q8_0 KV |
Single-user → Ollama on whatever consumer GPU has enough VRAM. Multi-user → vLLM on a datacenter or workstation card with HBM or GDDR7 and NVLink if you plan to use more than one. Mixing different cards → DP-only, one replica per GPU, router in front. And measure before you optimise — the honest bottleneck on local hardware is almost always memory bandwidth.