Local LLM Hosting Series — Presentation 10

Production Patterns for Local LLM Serving

Routing, caching, speculative decoding, observability, cost control, and the unglamorous ops work that turns a working vLLM into a stable service.

RoutingPrefix caching Speculative decodingPrometheus SLOsCost
Route → Cache → Spec → Metrics → SLO → Cost
00

Topics

01

A Reference Architecture

One diagram's worth of moving parts clients / apps gatewayauth, rate limit, routing semantic cacheRedis / Postgres pgvector KV-aware routerprefix-hash stickiness LoRA registryadapter → base map vLLM replica #0TP=2, prefix-cache on vLLM replica #1TP=2 vLLM replica #2TP=2 Prometheus + Grafana OpenTelemetry / traces log / audit store (S3 / Loki)

Not every project needs all of this. The gateway + one replica + Prometheus is 80% of the value. Add the rest incrementally as SLOs demand.

02

Gateway / Router — LiteLLM, Envoy, Custom

LiteLLM Proxy

Open-source OpenAI-compat gateway written in Python. Auth, per-key budgets, model routing, Prometheus, callbacks for Langfuse / Helicone. Great starter; deployable in minutes. Bottleneck is the Python hot path for very high QPS.

Envoy + ext_proc

Industrial-grade L7 proxy. Use ext_proc or Wasm filters for auth and routing logic; keep the data path in C++. Harder to author, much higher ceiling. Right choice when you're past ~10k QPS or need mTLS / ambient mesh.

Purpose-built routers

The vLLM router, SGLang Router, and NVIDIA Triton's ensemble scheduler all know about prefix cache and can route based on cache hit likelihood. Useful when you have many replicas and a stable prompt structure.

KV-aware stickiness

The only routing policy that matters for multi-turn: hash the first ~256 tokens, mod the number of replicas, and stick sessions to the matching replica. Deck 05 covered why.

03

Prefix & Semantic Caching Layers

Inside the GPU — prefix cache

vLLM's --enable-prefix-caching (deck 03). Hashes each 16-token block of prompt; reuses K/V across requests with a shared prefix. System prompts, tool schemas, few-shot examples get this for free.

Outside the GPU — semantic cache

Embed incoming prompts; if cosine similarity to a past prompt is > τ, return the cached response without touching the GPU. Redis Stack, Postgres pgvector, or Weaviate. Most useful for autocomplete-style workloads with repeated phrasings.

sketch
emb = embed(prompt)
hit = vstore.query(emb, k=1, threshold=0.93)
if hit: return hit.response

When semantic cache bites you

04

Speculative Decoding in Production

Deck 03 covered the mechanism; production has two extra concerns.

Pick the draft model with data

The acceptance rate depends on your workload. Benchmark three drafts (e.g. 1B, 3B, n-gram) on a representative log replay, measure tokens/second and p95 latency. Use the one that wins on your actual traffic, not what a paper says.

Budget the VRAM

The draft model eats a chunk of VRAM too. Llama-3.2-1B paired with a 70B base takes ~2 GB. On a 24 GB card that's a real tradeoff vs KV pool size. Usually worth it on H100/Spark; often not on a 4090.

Beware the determinism cost

Speculative decoding introduces a second path through sampling (deck 09). For eval or audit traffic, route around the speculative pool. Cheapest: keep two vLLM backends — one with --speculative-model, one without — and route based on a header.

05

Observability — the Four Signals

Four RED-style signals cover 95% of production questions:

SignalPrometheus metricAlert when…
TTFT p95vllm:time_to_first_token_seconds> SLO for 5 min
TPOT p95vllm:time_per_output_token_seconds> 2× baseline for 5 min
KV usagevllm:gpu_cache_usage_perc> 90% sustained → preemption
Queue depthvllm:num_requests_waiting> N for 2 min → scale out

Beyond RED — tracing

Wire OpenTelemetry around the gateway: tag each span with model, tp_size, prompt_tokens, completion_tokens, cache_hit, replica_id. Then you can answer "why was this request slow" instead of "why is p95 slow". Langfuse / Phoenix plug straight in.

Log policy

Don't log full prompt + completion in plaintext by default. Some data falls under GDPR / HIPAA / PCI. Log hashes + metrics; gate full body logging behind a role and an audit trail.

06

SLOs & Autoscaling Signals

Sensible starting SLOs

  • TTFT p95 < 400 ms (chat)
  • TPOT p95 < 50 ms (chat); < 20 ms for voice
  • Availability 99.5% monthly (rolling GPU restarts hurt this)
  • Token error rate < 0.1% (non-200 replies)

Autoscaling signals that work

  • num_requests_waiting > 8 for 2 min → scale up
  • Queue empty + TTFT stable 10 min → scale down
  • KV usage > 92% sustained → preempt alerts, maybe add a replica

Signals that don't work

07

Cost Accounting per Team / Feature

The killer question for internal hosting: "how much did feature X cost this month?" The inference cluster can answer it if you tag requests at the gateway.

Gateway stamps x-team, x-feature, x-env on every request
↓
vLLM returns prompt_tokens, completion_tokens in usage
↓
Aggregate by (team, feature, day) in a warehouse
↓
Multiply by your internal £/M-tokens rate — chargeback or showback
Rate-setting

Internal price should be stable, not GPU-utilisation-dependent. Compute monthly cost (hardware amortisation + power + admin) ÷ monthly tokens served; that's your £/M-token rate. Revisit quarterly. Charging teams for variable GPU cost creates incentives nobody wants.

08

Interactive: SLO Planner

20
3×
09

Rollout, Rollback, Canaries

Deploy cadence

  • Pin vLLM image + model revision in the compose file
  • Change one thing at a time — either model or vLLM or flags, not all three
  • Run the eval suite against a canary replica before promoting
  • Promote by flipping router weights (90/10 → 50/50 → 100/0) over an hour

Rollback

  • Keep the previous image + weights on disk, tagged
  • Rollback = re-launch old compose file; ~90 s
  • Do not delete old weights until a full week on the new version
  • Pin HF model commit SHA, not :latest

Draining a replica safely

stop accepting, finish in-flight, then exit
curl -X POST http://host:8000/v1/load  -d '{"drain": true}'
# router routes new traffic elsewhere; existing streams finish
# wait for vllm:num_requests_running == 0
docker stop --time=120 vllm

In practice: let the load balancer health-check the replica's /health, flip the replica's liveness to "drain", and stop it after in-flight traffic clears. This is the same playbook as any HTTP server.

10

Recap — What Moves the Needle

Where this fits

This deck is the one you want when moving from "works on my laptop" to "serves real users". The rest of the series covers orientation, Ollama and vLLM internals, Docker deployment, multi-GPU parallelism, frameworks, quantisation, hardware, and determinism — see the series index for the current full list.