Local LLM Hosting Series — Presentation 01

Why Host LLMs Locally — The Landscape

Drivers, tradeoffs, and a map of the local-inference ecosystem: Ollama, vLLM, llama.cpp, TGI, SGLang, TensorRT-LLM, LMDeploy, MLC, and the hardware they run on.

Local Inference vLLM Ollama llama.cpp NVIDIA DGX Spark
Why Local → Cost Model → Ecosystem → Hardware → Picker → Roadmap
00

Topics We'll Cover

This first deck sets the scene. Before picking a framework or buying a GPU, it's worth being deliberate about why you're hosting an LLM locally and what you're trading off against the managed APIs.

01

Four Drivers of Local Hosting

People run LLMs on their own hardware for one of four reasons — often more than one, but almost always at least one. Being honest about which one is driving you keeps the stack simple.

1 · Data Residency & Privacy

Regulated data (patient records, legal discovery, on-prem code) cannot leave the network. The frontier-model APIs retain logs; an air-gapped local deployment side-steps that entirely. This is the single most common reason enterprises stand up vLLM or TensorRT-LLM internally.

2 · Cost at Scale

The API break-even is somewhere around 2–5 million tokens/day for a 7–14B model. Below that, APIs win; above it, depreciating a £1,500 GPU over two years is cheaper per token. Batching is what makes the local side work.

3 · Latency & Determinism

A local model on an RTX 4090 can serve a short prompt in ~40 ms TTFT vs ~300–600 ms round-trip to a hosted API. For voice agents, IDE autocomplete, and sub-second user-facing loops this matters. No rate limits, no retry storms.

4 · Customisation

LoRA fine-tunes, quantised experts, speculative-decoding draft models, domain-specific tokenisers, and odd context-window configurations are all trivial locally and mostly impossible on closed APIs.

Honest Caveat

Local inference is not a free-tier optimisation. If "I want to save money on my weekend chatbot" is the only driver, you will almost certainly spend more in electricity and time than the £20/month you saved. Drivers 1 and 3 justify the effort far more cleanly than driver 2 alone.

02

The Total Cost Model

The honest comparison is cost per million tokens, amortised over the life of the hardware. Every factor below matters.

Cost ComponentTypical Range (home / SMB)Notes
GPU (CapEx)£250–£4,000RTX 3060 → DGX Spark
Host (CPU / RAM / PSU)£400–£1,200Needs 2× GPU TDP headroom in the PSU
Electricity @ £0.25/kWh~£0.10–£0.25 per hour under load4090 pulls 450 W, Spark ~240 W
Admin timeHighly variableUsually the dominant hidden cost
Model / weights£0 (most open weights) or licence feesLlama 3.x: permissive, Mistral: Apache 2.0, DeepSeek: MIT

Where the API Break-Even Actually Sits

Running a 13B model on a single RTX 4090 at 100 tok/s continuous, assuming £0.25/kWh and a 2-year hardware amortisation:

Hardware: £1,500 / (2 yr × 8,760 h) = £0.086/hr
+
Power: 0.45 kW × £0.25 = £0.11/hr
=
~£0.20/hr to run continuously
↓
At 100 tok/s → 360 k tok/hr → ~£0.55 per million tokens
Take-away

At ~£0.55 / M tokens you beat every frontier API on marginal cost — but only if you actually saturate the GPU. A 4090 sitting idle 22 hours a day is about £1.95/hr per utilised hour. Utilisation is the local-inference KPI. This is why continuous batching and request routing (decks 03, 08) matter so much.

03

The Framework Landscape

Seven stacks matter in 2026. Each solves a different shape of problem. Deck 05 does a full head-to-head; this is the map.

FrameworkLanguageSweet SpotBackend
OllamaGo wrapperSingle-user desktop, laptops, dev loopsllama.cpp + GGUF
llama.cppC++CPU inference, Apple Silicon, embeddedown GGML/GGUF
vLLMPython + CUDAHigh-throughput OpenAI-compatible servingPagedAttention, CUDA/ROCm
TGI (HF)Rust + PythonHuggingFace-integrated productionFlash-Attention, TP
SGLangPythonStructured output, multi-turn, RadixAttentionTriton kernels
TensorRT-LLMC++/PythonMax-performance NVIDIA-onlyTensorRT engines
LMDeploy / MLCPython / TVMMobile, Vulkan, exotic targetsAWQ, TVM

One-Axis Summary

Ease
Ollama  →  llama.cpp  →  vLLM  →  TGI  →  TensorRT-LLM
 
Throughput
Ollama  →  llama.cpp  →  vLLM  →  SGLang  →  TensorRT-LLM
 

The best intuition: Ollama and llama.cpp are inference clients. vLLM, TGI, SGLang, TensorRT-LLM are inference servers. If more than one person is hitting the model, you want a server.

04

Serving vs Single-User Stacks

The architectural gulf between a desktop chat UI and a production serving cluster is wide. Knowing which side you're on dictates everything downstream.

Single-user / desktop

  • One request at a time
  • KV cache per session, blown away on exit
  • GGUF quantised weights (usually 4-bit)
  • Optimised for latency not throughput
  • Ollama, llama.cpp, LM Studio, Jan

Multi-user / server

  • Continuous batching across many concurrent requests
  • PagedAttention / RadixAttention for KV re-use
  • FP16/BF16 or calibrated FP8/AWQ
  • Optimised for throughput — 5–20× higher tok/s
  • vLLM, TGI, SGLang, TensorRT-LLM

The Same Model, Two Wildly Different Numbers

Llama-3.1-8B on RTX 4090 — tokens/second (aggregate) Ollama (1 user) ~95 tok/s vLLM (1 user) ~110 tok/s vLLM (8 users) ~720 tok/s aggregate vLLM (32 users) ~1450 tok/s

Indicative numbers, prompt ~256 tok, output ~128 tok, FP16. vLLM's aggregate throughput scales with concurrency up to KV-cache saturation; Ollama's does not. This chart is the whole reason decks 03, 04 and 08 exist.

05

Hardware Tiers at a Glance

Local hosting spans from a £50 Raspberry-Pi-class CPU to a two-Spark cluster. Deck 07 goes into DGX Spark in depth; this is the shape of the space.

Laptop CPU
llama.cpp 3B q4
free
Apple Silicon
M2/M3/M4 — MLX, Ollama
£1.5k+
RTX 3060 12 GB
7B q8 / 13B q4
~£250
RTX 4090 24 GB
8B FP16 / 32B q4
~£1,500
DGX Spark 128 GB
70B FP8 / 200B quant
~£3,800
2× Spark
200B FP8 / 400B quant
~£7,600
H100 cloud
any weight, any precision
£2–6/hr
06

Model Sizes, Formats, and Fit

The fundamental VRAM relationship worth memorising — every other decision falls out of this.

Rule of thumb

VRAM ≈ params × bytes_per_param + KV_cache(batch, ctx)
— FP16 → 2 B/param, FP8 → 1 B, INT4 (AWQ/GGUF q4) → 0.5 B.
— KV cache per token ≈ 2 × layers × kv_heads × head_dim × bytes. For Llama-3 8B in FP16 (GQA, 8 KV heads × 128): ~0.13 MB per token per sequence.

Worked Examples

ModelFP16 weightsINT4 weightsKV @ 4k ctx · 1 seqFits on…
Llama-3.2-3B6 GB1.8 GB~0.45 GBAny 8 GB GPU, CPU, Pi 5
Llama-3.1-8B16 GB4.5 GB~0.5 GB12 GB GPU (q4) / 24 GB GPU (FP16)
Mistral-Small-24B48 GB13 GB~0.65 GB24 GB GPU (q4) / 48 GB GPU (FP16)
Llama-3.3-70B140 GB40 GB~1.3 GB48 GB (q4) / Spark (FP8) / H100
DeepSeek-V3 (MoE 671B)1,342 GB~335 GB~0.3 GB (MLA)Multi-GPU H100/H200 node (too big for 2× Spark)

Deck 06 goes into quantisation formats and perplexity tradeoffs; deck 07 goes into DGX Spark unified memory specifically.

07

Interactive: What Should I Run?

Pick your driver, concurrency, and hardware. The picker below applies the logic from slides 01–06.

1
8k
08

Rest of the Series

The series index on GitHub is the canonical, up-to-date list — new decks are added there as they're written.

How to read the series

Each deck stands alone. Broadly: the Ollama and vLLM-architecture decks are the two stack deep-dives — pick the one matching your situation. The Docker, multi-GPU parallelism, DGX Spark and NVIDIA-GPUs decks are pragmatic deployment guides. The frameworks and quantisation decks are reference material. The determinism and production-patterns decks are the ones that matter once you're running something real.