Local LLM Hosting Series — Presentation 11

Deploying on NVIDIA GPUs — Architectures, Memory, Multi-GPU

A grounded tour of the 2026 NVIDIA stack for LLM serving: which class of card for which model, what the memory type actually does to tok/s, how to gang GPUs together (including heterogeneous pairs), MIG, and the Ollama/vLLM-specific gotchas per architecture.

AmpereAdaHopper BlackwellGDDR7HBM3e NVLinkMIG Heterogeneous
Arch → Memory → Capability → Ollama → vLLM → Multi-GPU → Hetero → MIG → Ops
00

Topics We'll Cover

A practical, opinionated guide to picking, pairing, and operating NVIDIA GPUs for local LLM hosting — with the nuances that bite at 2 AM.

01

The 2026 NVIDIA Stack — A Field Guide

Five architecture families are relevant to local LLM serving today. Understanding which generation a card belongs to is the single biggest predictor of software behaviour.

ArchCCYearExamplesTensor core nativeLLM notes
Ampere8.0 / 8.62020–21A100, RTX 30xx, A40, A6000FP16, BF16, TF32No FP8. Workhorse for INT4 serving.
Ada Lovelace8.92022–23RTX 40xx, L40, L4, L40SFP16, BF16, FP8 (all Ada)Consumer king; no NVLink on 4090.
Hopper9.02023–24H100, H200, GH200FP16, BF16, FP8 (E4M3/E5M2)First real FP8. Transformer Engine.
Blackwell10.0 / 12.02025–26B100, B200, RTX 50xx, GB200FP16, BF16, FP8, FP4 (MX)Native FP4 doubles inference again.
Grace-Blackwell10.0 / 12.12025–26GB10 (Spark), GB200 NVL72same as BlackwellArm64 host, unified memory, C2C link.

Product classes, same arch, very different machine

Consumer (GeForce)

RTX 30xx / 40xx / 50xx. GDDR6 / 6X / 7. No ECC. No SR-IOV. Driver licence forbids datacenter deployment. On 40-series no NVLink. You're on your own for multi-GPU scaling.

Prosumer / Workstation

RTX A6000, A6000 Ada, RTX PRO 6000 Blackwell (96 GB). GDDR6/7 ECC, proper driver, sometimes NVLink bridge. Datacenter-licensed. Sweet spot for serious lab work.

Datacenter

A100, H100, H200, B100, B200, L40S, L4. HBM or GDDR6 ECC, SR-IOV, MIG on A100/H100/B200, NVLink 3/4, NVSwitch optional, proper persistence and error recovery. What the frontier APIs actually run on.

Licence warning

NVIDIA's GeForce EULA restricts datacenter deployment of consumer cards. "Datacenter" is loosely defined but colocation, multi-tenant cloud, and large-scale commercial hosting are clearly out. For home / office / single-tenant internal use it's fine. If you're standing up a commercial service, use L4/L40S/A100/H100/B200 class hardware.

02

Memory Is the Bill — GDDR, HBM, LPDDR5x, EGM

Decode tok/s is bounded by memory bandwidth, not FLOPS. Every token requires reading all weights once. The memory type therefore decides the upper bound on single-stream tok/s more than anything else.

Memory bandwidth per card — GB/s, higher is better LPDDR5x (Spark) ~273 GB/s GDDR6 (A6000) ~768 GB/s GDDR6X (4090) ~1008 GB/s GDDR7 (5090 / RTX PRO 6000) ~1792 GB/s HBM2e (A100 80GB) ~2000 GB/s HBM3 (H100 SXM) ~3350 GB/s HBM3e (H200) ~4800 GB/s HBM3e (B200) ~8000 GB/s

Memory types, intuition

TypeWhereCharacteristics
GDDR6A6000, 3060, 4060/70, L4Cheap per GB, ~400–800 GB/s, no ECC on consumer
GDDR6X3090 / 3090 Ti, 4080/90PAM4 signalling, higher data rate; same capacity tiers
GDDR75090, RTX PRO 6000 BlackwellPAM3, ~2× GDDR6X; up to 96 GB on RTX PRO 6000
HBM2eA100On-package stacks; 80 GB, ~2 TB/s (the 40 GB A100 is HBM2, ~1.6 TB/s)
HBM3H10080 GB, ~3.4 TB/s
HBM3eH200, B100, B200, GB20096–192 GB, 4.8–8 TB/s
LPDDR5x (unified)GB10 Spark, GH200Huge capacity (128 GB+), moderate bandwidth (~0.3–0.5 TB/s); coherent with CPU via NVLink-C2C
EGM (Extended GPU Memory)GH200, GB200GPU sees host LPDDR5x as extended VRAM over the C2C link — transparent to the kernel
The practical consequence

Single-stream tok/s on a 70B model is roughly min(bandwidth / weight_bytes, compute_limit). With 70B FP8 (70 GB weights), an H200 (4.8 TB/s) is theoretically capped at ~68 tok/s per stream; a pair of RTX 4090s layer-splitting a 40 GB INT4 quant (it doesn't fit one 24 GB card) peaks around ~25 tok/s. If you need raw tok/s on a single stream, spend on bandwidth; if you need concurrency, spend on capacity.

03

Compute Capability & Feature Matrix

Two cards with the same VRAM can run very different software stacks because their compute capability (CC) determines which tensor-core formats, PTX intrinsics, and CUDA features the kernels can use.

FeatureAmpere (8.x)Ada (8.9)Hopper (9.0)Blackwell (10.0/12.0)
TF32 / BF16 / FP16 tensor cores✓✓✓✓
FP8 E4M3 / E5M2 tensor cores—✓✓✓
FP4 (MX) tensor cores———✓
2:4 structured sparsity✓✓✓✓
Transformer Engine (auto FP8)—partial✓✓
Thread-Block Clusters——✓✓
Distributed Shared Memory (DSMEM)——✓✓
TMA (Tensor Memory Accelerator)——✓✓
MIG (Multi-Instance GPU)A100 only—✓B200 only
NVLink 3/4/5A100 (3)no on 4090✓ (4)✓ (5)
Confidential Computing——✓✓
The FP8 cliff

Every serious 2025–2026 model family ships FP8 checkpoints because Hopper onwards runs them at full tensor-core speed. On Ampere (A100, 3090, 4090's predecessor) FP8 is emulated — you pay the memory win but not the compute win. For most decoding that's still a good deal (decoding is bandwidth-bound), but prefill throughput lags Hopper by a lot.

04

Ollama on NVIDIA — Card-by-Card Nuances

Ollama wraps llama.cpp. The CUDA backend is built with GGML_CUDA=1 and uses the CUDA runtime + cuBLAS + llama.cpp's own kernels. Behaviour varies by arch.

Card classWhat actually happensPractical notes
GTX 10xx / 16xx (Pascal, Turing) Works via CUDA backend; no tensor cores → pure FP32 SMs carry matmul. GGUF q4 on a 1080 Ti still runs 7B at ~15 tok/s. Fine for demos, not much else.
RTX 20xx / 30xx (Turing / Ampere) Tensor cores used for FP16/BF16 in matmul. OLLAMA_FLASH_ATTENTION=1 switches to flash-attn kernel. Pin the model on NVMe; 3090's 24 GB runs 13B q5_K_M at 35–45 tok/s.
RTX 40xx (Ada) Full BF16 tensor path. No NVLink on 4090 — multi-GPU is PCIe P2P only. Sweet spot for solo Ollama on 7–32B. q4_K_M quant + FA on.
RTX 50xx (Blackwell consumer) GDDR7 — memory bandwidth jumps ~1.8× vs 40xx. No FP4 in llama.cpp yet (2026), so you don't get the Blackwell inference speedup via Ollama. 5090's 32 GB VRAM fits 32B q4 or 14B BF16 comfortably.
A100 / A40 / A6000 Full BF16, ECC on; Ollama treats it like any CUDA GPU. A100 MIG instances show up as separate GPUs — Ollama will happily bind to one. Datacenter cards are wasted on single-user Ollama. Still fine for RAG side-cars.
H100 / H200 BF16 path; no FP8 in llama.cpp through Ollama (2026). Hopper-specific kernels largely unused. Throws a lot of silicon away. Use vLLM if you've paid for an H100.
DGX Spark (GB10) Arm64 build; unified memory means model load is extremely fast. FP4 unused. Good dev target; 128 GB means 70B q4 comfortably.
GB200 / B200 Overkill; llama.cpp/Ollama uses almost none of Blackwell's new tensor capabilities. If you have one, run vLLM or TensorRT-LLM, not Ollama.

Ollama env vars that matter on NVIDIA

Ollama's multi-GPU honestly

Ollama has basic layer-splitting across multiple GPUs (ngl / tensor_split) but it is not tensor-parallel in the vLLM sense. It splits layers across cards and pipelines them — which means slower per-token than a single big card would be, and it does not help concurrent users at all. For real multi-GPU local serving, use vLLM.

05

vLLM on NVIDIA — Card-by-Card Nuances

vLLM's performance story is much closer to "what the hardware can actually do". It does ship arch-specific kernels and it's picky about them.

Card classWhat vLLM usesPractical notes
Turing (RTX 20xx) FP16 path only. Flash-attn v2 supported; v3 is Hopper-only. Marginal; 20xx cards lack BF16 tensor cores, cuts throughput.
Ampere (A100, 3090, 3090 Ti, A40, A6000) BF16 native, FP8 emulated (compute no faster than BF16; bandwidth win still applies). Flash-attn v2. PagedAttention works fully. AWQ / GPTQ INT4 is the best point on this arch. Multi-GPU via NVLink on A100 / A6000, PCIe on 3090.
Ada (RTX 40xx, L4, L40, L40S) BF16 + FP8 (all Ada, CC 8.9). No NVLink on 40-series. CUDA graphs fully supported. Consumer Ada is great for single-GPU vLLM; multi-GPU scaling is rough without NVLink.
Hopper (H100 / H200) Native FP8 tensor cores → 2× throughput vs BF16 on prefill. Flash-attn v3. Thread-block clusters. Transformer Engine via --quantization fp8. The default production choice in 2026. Use --kv-cache-dtype fp8 too.
Blackwell (B100 / B200 / GB200) Adds native FP4 / MX-FP4 → another 2× on compute. vLLM 0.7+ and TensorRT-LLM 0.12+ required. Highest ceiling; cheapest per token if utilisation is high.
DGX Spark (GB10) Blackwell GPU + 128 GB unified. vLLM arm64 image. FP8/FP4 supported. Bandwidth-bound at LPDDR5x speeds. Local workstation; see deck 08.

vLLM flags that behave differently per arch

FlagAmpereAdaHopperBlackwell
FP8 checkpoint (auto-detected; --dtype has no fp8)W8A16 via Marlin; bandwidth win onlynativenativenative
--quantization fp8 (TE)—nativenativenative
--quantization awq✓✓✓✓
--quantization fp4 / mxfp4———native
--kv-cache-dtype fp8emulatednativenativenative
--enforce-eagersafesafesafesometimes needed on early B200 kernels
--enable-chunked-prefill✓✓big winbig win
06

Ganging GPUs — NVLink, NVSwitch, PCIe P2P, IOMMU

Deck 05 covered the software axes (TP/PP/DP/EP). Here we cover what the physical link between GPUs buys you and where it goes wrong.

NVLink — what it actually is

A high-speed, coherent point-to-point GPU↔GPU link. Three variants you'll meet:

  • NVLink 3: A100 (600 GB/s), A40/A6000 bridge (112 GB/s two-card).
  • NVLink 4: H100 (900 GB/s via NVSwitch).
  • NVLink 5: B200 / GB200 (1.8 TB/s, NVL72 domains up to 72 GPUs).

Used for TP all-reduces and P2P memcpy. Not present on any 40-series; NVIDIA removed the bridge. 5090 reportedly also without.

NVSwitch

A packet-switched fabric so that every GPU in a node talks to every other at full NVLink speed. DGX boxes and HGX baseboards. Turns a group of GPUs into one NVLink domain. TP across 8 H100s works because of NVSwitch.

PCIe P2P

Fallback on any NVLink-less pair. Per-lane bandwidth is bus-limited: PCIe 4 x16 ≈ 32 GB/s, PCIe 5 x16 ≈ 64 GB/s. Works but kills TP scaling.

Enablement:

  • BIOS: enable Above 4G Decoding and Resizable BAR
  • Kernel: IOMMU passthrough or off (iommu=pt or intel_iommu=off)
  • Same IOMMU group? Check with lspci -vv | grep -E 'IOMMU|ACS'
  • ACS override patch on consumer boards to split groups (risky)

InfiniBand / RoCE

Cross-node path. NDR 400 = ~50 GB/s per GPU. GPUDirect RDMA bypasses the host so tensors move GPU→GPU without a PCIe-to-host-to-PCIe hop. Required for any useful cross-node TP.

Diagnose the fabric before you deploy
nvidia-smi topo -m           # matrix: NV#, PIX, PXB, PHB, SYS
nvidia-smi nvlink -s         # link state per card
nvidia-smi -q -d TOPOLOGY | head -60

# Verify P2P actually works (simpleP2P from cuda-samples, or):
CUDA_VISIBLE_DEVICES=0,1 python -c "import torch;\
 a=torch.randn(1,device='cuda:0');\
 b=a.to('cuda:1');\
 print('p2p ok',b.device)"

# Raw bandwidth (GPU 0 → GPU 1, H2H / D2D):
./bandwidthTest --device=0 --dtod
Nuance: SYS in the matrix

If nvidia-smi topo -m shows SYS between your two GPUs, they're separated by a CPU socket. All P2P traffic hops through the CPU's UPI / Infinity Fabric — catastrophic for TP. Physically move a card, or pin the workload to the GPUs that are both on the same socket with CUDA_VISIBLE_DEVICES.

07

Heterogeneous GPUs — Can You Actually Mix?

The question everyone asks: "I have a 4090 and a 3090. Can I gang them?" The answer depends entirely on how you're "ganging" them.

ScenarioFeasibilityWhy
Run one model on each card (DP) easy Independent processes, independent CUDA contexts. Completely fine to pair any NVIDIA GPUs this way. Use CUDA_VISIBLE_DEVICES.
Ollama OLLAMA_SCHED_SPREAD across mixed cards works llama.cpp's layer split is indifferent to matched cards. Layer placement respects reported VRAM sizes. Slow card sets pace.
vLLM Tensor Parallel across mismatched cards no TP requires identical tensor shapes and bit-exact synchronised matmuls every layer. NCCL + Flash-Attention + CUTLASS all assume homogeneous GPUs. vLLM checks and refuses.
vLLM Pipeline Parallel across mismatched cards technically possible, painful Stages only exchange hidden states, so in principle stages can be on different GPUs. In practice vLLM/Ray still want matched ranks; you must set --pipeline-parallel-size and assign stages manually. Slowest stage sets throughput; bubble recovery is harder.
vLLM DP across mismatched cards behind a router works Each replica is its own vLLM process with its own flags. Different quants per replica are fine; router can even shape traffic towards the faster one.
Mix Ampere + Hopper (e.g. 3090 + H100) no TP Different compute capabilities → different kernel binaries. vLLM will load one arch's wheel, not the other's.
Mix 4090 + 4080 (same arch, different VRAM) DP yes, TP no Same CC but different VRAM sizes → unequal KV pools. TP wants identical shards per GPU; DP doesn't care.
A6000 48 GB + A6000 Ada 48 GB no TP Same memory but different arch (Ampere vs Ada). NCCL can still talk but vLLM TP will refuse because the kernels differ.
Rule of thumb

Heterogeneous = DP-only. If you have mixed GPUs, run one vLLM per GPU (sized appropriately for that card's VRAM), put a load balancer in front, and call it a day. It's simpler, it actually works, and you lose very little vs a theoretical "ideal" unified deployment.

A worked mixed-box example

One 4090 (24 GB) + two 3090s (24 GB each)
# Replica A on the 4090 — Llama-3.1-8B FP16 for fastest TTFT
CUDA_VISIBLE_DEVICES=0 docker run -d --name vllm-fast \
    --gpus '"device=0"' -p 8001:8000 ...  \
    --model meta-llama/Meta-Llama-3.1-8B-Instruct --dtype bfloat16

# Replicas B/C on the 3090s — larger model, INT4
CUDA_VISIBLE_DEVICES=1 docker run -d --name vllm-big-a \
    --gpus '"device=1"' -p 8002:8000 ... \
    --model casperhansen/mistral-small-24b-awq --quantization awq
CUDA_VISIBLE_DEVICES=2 docker run -d --name vllm-big-b \
    --gpus '"device=2"' -p 8003:8000 ... \
    --model casperhansen/mistral-small-24b-awq --quantization awq

# LiteLLM router routes by model name in front of the three ports
08

MIG — Slicing H100/H200/B200

MIG (Multi-Instance GPU) slices one physical GPU into up to 7 independent GPUs, each with dedicated SMs, L2, and memory slices. Unlike vGPU it's hardware-partitioned: no noisy-neighbour contention.

H100 80 GB MIG profiles — one possible split 3g.40gb (3× SMs / 40 GB)→ one big model per slice 2g.20gb(mid-size) 1g.10gb(small) 1g.10gb Each slice appears as a separate CUDA device with its own UUID.

Why MIG matters for LLM hosting

Configure MIG on an H100
sudo nvidia-smi -mig 1                      # enable MIG mode
sudo nvidia-smi mig -cgi 9,14,19,19      # create GIs (profile IDs vary per GPU)
sudo nvidia-smi mig -cci                    # create CIs
nvidia-smi -L                               # lists each slice with MIG UUID

# Launch a container bound to one slice
docker run --gpus '"device=MIG-abcd1234-..."' vllm/vllm-openai:v0.7.0 ...
MIG limits

No NVLink between MIG instances, no P2P across them. So MIG + TP is not a thing. MIG is for many small workloads, not one big sharded one. Also: only A100, A30, H100, H200, B200 and RTX PRO 6000 Blackwell (up to 4 instances) support MIG (not A6000, L40S, 4090).

09

Interactive: GPU Capability Picker

Pick a GPU. See what model sizes and framework features are practical.

Arch
—
VRAM
—
BW (GB/s)
—
70B FP8 tok/s (est.)
—
10

Interactive: Can I Gang These Two?

Pick two GPUs and a framework. The planner tells you which ganging modes actually work.

11

Operational Nuances That Burn Hours

Persistence mode

sudo nvidia-smi -pm 1. Without it the driver unloads between processes — adds ~30 s to every container start. Datacenter cards have it on by default; consumer cards don't.

Resizable BAR (ReBAR) / Above 4G Decoding

BIOS feature; lets the CPU address the GPU's full VRAM as one PCIe aperture. Without it, large tensor transfers are chunked through BAR1 (a ~256 MB window) — cold loads slow by 2–5×. Always enable on modern boards.

ECC on / off

Datacenter cards have ECC on by default; it costs ~6% of VRAM. sudo nvidia-smi -e 0 reclaims that capacity (reboot required). Do it for inference-only hosts where you don't care about single-bit flips; leave it on for training.

Power limits

sudo nvidia-smi -pl 350 (watts). A 4090 capped at 350 W loses ~5% throughput but ~15% power. Useful in dense servers. Check nvidia-smi -q -d POWER for envelopes.

IOMMU groups & ACS

IOMMU groups determine which devices can do P2P. On consumer boards multiple GPUs often end up in the same group (fine) or blocked by PCIe switches. dmesg | grep iommu, lspci -vv | grep ACSCtl. Don't patch ACS on production hardware.

NUMA pinning

On multi-socket CPUs, GPU-to-CPU traffic hops sockets unless you pin. numactl --cpunodebind=0 --membind=0 vllm serve ... for the GPU on NUMA 0. Worth 5–10% on prefill throughput.

WSL2 / Windows

WSL2 works for dev but loses ~10–15% to the translation layer. No MIG, no SR-IOV. For serving, use bare Linux.

Driver lineage

Match driver to CUDA. Blackwell needs driver ≥ 570 (CUDA 12.8); Hopper FP8 works from 525+; some flash-attn 3 kernels require ≥ 560. Pin the driver on serving hosts.

CUDA graphs

vLLM captures a graph of the decode step; can save 10–20% on small batch sizes. Sometimes crashes at load on bleeding-edge CUDA; fall back with --enforce-eager.

Cooling

Not a joke: a 4090 at 85°C throttles. Check nvidia-smi -q -d TEMPERATURE. In a 1U rack a passive-cooled A100/H100 must have front-to-back airflow; blower-style consumer cards may not fit at all.

12

Decision Matrix — A Cheat Sheet

I want to…Buy thisFrameworkQuant
Run a 7B locally, no fussUsed RTX 3060 12 GBOllamaq4_K_M
Run a 13–30B locally, fastRTX 4090 / 5090 / A6000 AdaOllama or vLLM (DP)q5_K_M or AWQ-INT4
Serve 10–50 users on 70BH100 / H200 / RTX PRO 6000 BlackwellvLLMFP8 + FP8 KV
Local 70B at homeDGX Spark 128 GBvLLM arm64FP8 / AWQ-INT4
Serve 400B / MoE on-prem8× B200 (HGX) / GB200 NVL72vLLM TP=8 or TRT-LLMFP8 weights + FP8 KV, MX-FP4 if Blackwell
Internal dev GPU poolH100 80 GB + MIGvLLM per sliceAWQ-INT4 for small slices, FP8 for big
Cheap inference serviceL40S 48 GB × NvLLM (DP), LBFP8
Long-context embeddingsRTX 3090 24 GBOllama or vLLMFP16 weights + q8_0 KV
The shortest correct answer

Single-user → Ollama on whatever consumer GPU has enough VRAM. Multi-user → vLLM on a datacenter or workstation card with HBM or GDDR7 and NVLink if you plan to use more than one. Mixing different cards → DP-only, one replica per GPU, router in front. And measure before you optimise — the honest bottleneck on local hardware is almost always memory bandwidth.