Drivers, tradeoffs, and a map of the local-inference ecosystem: Ollama, vLLM, llama.cpp, TGI, SGLang, TensorRT-LLM, LMDeploy, MLC, and the hardware they run on.
This first deck sets the scene. Before picking a framework or buying a GPU, it's worth being deliberate about why you're hosting an LLM locally and what you're trading off against the managed APIs.
People run LLMs on their own hardware for one of four reasons — often more than one, but almost always at least one. Being honest about which one is driving you keeps the stack simple.
Regulated data (patient records, legal discovery, on-prem code) cannot leave the network. The frontier-model APIs retain logs; an air-gapped local deployment side-steps that entirely. This is the single most common reason enterprises stand up vLLM or TensorRT-LLM internally.
The API break-even is somewhere around 2–5 million tokens/day for a 7–14B model. Below that, APIs win; above it, depreciating a £1,500 GPU over two years is cheaper per token. Batching is what makes the local side work.
A local model on an RTX 4090 can serve a short prompt in ~40 ms TTFT vs ~300–600 ms round-trip to a hosted API. For voice agents, IDE autocomplete, and sub-second user-facing loops this matters. No rate limits, no retry storms.
LoRA fine-tunes, quantised experts, speculative-decoding draft models, domain-specific tokenisers, and odd context-window configurations are all trivial locally and mostly impossible on closed APIs.
Local inference is not a free-tier optimisation. If "I want to save money on my weekend chatbot" is the only driver, you will almost certainly spend more in electricity and time than the £20/month you saved. Drivers 1 and 3 justify the effort far more cleanly than driver 2 alone.
The honest comparison is cost per million tokens, amortised over the life of the hardware. Every factor below matters.
| Cost Component | Typical Range (home / SMB) | Notes |
|---|---|---|
| GPU (CapEx) | £250–£4,000 | RTX 3060 → DGX Spark |
| Host (CPU / RAM / PSU) | £400–£1,200 | Needs 2× GPU TDP headroom in the PSU |
| Electricity @ £0.25/kWh | ~£0.10–£0.25 per hour under load | 4090 pulls 450 W, Spark ~240 W |
| Admin time | Highly variable | Usually the dominant hidden cost |
| Model / weights | £0 (most open weights) or licence fees | Llama 3.x: permissive, Mistral: Apache 2.0, DeepSeek: MIT |
Running a 13B model on a single RTX 4090 at 100 tok/s continuous, assuming £0.25/kWh and a 2-year hardware amortisation:
At ~£0.55 / M tokens you beat every frontier API on marginal cost — but only if you actually saturate the GPU. A 4090 sitting idle 22 hours a day is about £1.95/hr per utilised hour. Utilisation is the local-inference KPI. This is why continuous batching and request routing (decks 03, 08) matter so much.
Seven stacks matter in 2026. Each solves a different shape of problem. Deck 05 does a full head-to-head; this is the map.
| Framework | Language | Sweet Spot | Backend |
|---|---|---|---|
| Ollama | Go wrapper | Single-user desktop, laptops, dev loops | llama.cpp + GGUF |
| llama.cpp | C++ | CPU inference, Apple Silicon, embedded | own GGML/GGUF |
| vLLM | Python + CUDA | High-throughput OpenAI-compatible serving | PagedAttention, CUDA/ROCm |
| TGI (HF) | Rust + Python | HuggingFace-integrated production | Flash-Attention, TP |
| SGLang | Python | Structured output, multi-turn, RadixAttention | Triton kernels |
| TensorRT-LLM | C++/Python | Max-performance NVIDIA-only | TensorRT engines |
| LMDeploy / MLC | Python / TVM | Mobile, Vulkan, exotic targets | AWQ, TVM |
The best intuition: Ollama and llama.cpp are inference clients. vLLM, TGI, SGLang, TensorRT-LLM are inference servers. If more than one person is hitting the model, you want a server.
The architectural gulf between a desktop chat UI and a production serving cluster is wide. Knowing which side you're on dictates everything downstream.
Indicative numbers, prompt ~256 tok, output ~128 tok, FP16. vLLM's aggregate throughput scales with concurrency up to KV-cache saturation; Ollama's does not. This chart is the whole reason decks 03, 04 and 08 exist.
Local hosting spans from a £50 Raspberry-Pi-class CPU to a two-Spark cluster. Deck 07 goes into DGX Spark in depth; this is the shape of the space.
The fundamental VRAM relationship worth memorising — every other decision falls out of this.
VRAM ≈ params × bytes_per_param + KV_cache(batch, ctx)
— FP16 → 2 B/param, FP8 → 1 B, INT4 (AWQ/GGUF q4) → 0.5 B.
— KV cache per token ≈ 2 × layers × kv_heads × head_dim × bytes. For Llama-3 8B in FP16 (GQA, 8 KV heads × 128): ~0.13 MB per token per sequence.
| Model | FP16 weights | INT4 weights | KV @ 4k ctx · 1 seq | Fits on… |
|---|---|---|---|---|
| Llama-3.2-3B | 6 GB | 1.8 GB | ~0.45 GB | Any 8 GB GPU, CPU, Pi 5 |
| Llama-3.1-8B | 16 GB | 4.5 GB | ~0.5 GB | 12 GB GPU (q4) / 24 GB GPU (FP16) |
| Mistral-Small-24B | 48 GB | 13 GB | ~0.65 GB | 24 GB GPU (q4) / 48 GB GPU (FP16) |
| Llama-3.3-70B | 140 GB | 40 GB | ~1.3 GB | 48 GB (q4) / Spark (FP8) / H100 |
| DeepSeek-V3 (MoE 671B) | 1,342 GB | ~335 GB | ~0.3 GB (MLA) | Multi-GPU H100/H200 node (too big for 2× Spark) |
Deck 06 goes into quantisation formats and perplexity tradeoffs; deck 07 goes into DGX Spark unified memory specifically.
Pick your driver, concurrency, and hardware. The picker below applies the logic from slides 01–06.
The series index on GitHub is the canonical, up-to-date list — new decks are added there as they're written.
Each deck stands alone. Broadly: the Ollama and vLLM-architecture decks are the two stack deep-dives — pick the one matching your situation. The Docker, multi-GPU parallelism, DGX Spark and NVIDIA-GPUs decks are pragmatic deployment guides. The frameworks and quantisation decks are reference material. The determinism and production-patterns decks are the ones that matter once you're running something real.