Routing, caching, speculative decoding, observability, cost control, and the unglamorous ops work that turns a working vLLM into a stable service.
Not every project needs all of this. The gateway + one replica + Prometheus is 80% of the value. Add the rest incrementally as SLOs demand.
Open-source OpenAI-compat gateway written in Python. Auth, per-key budgets, model routing, Prometheus, callbacks for Langfuse / Helicone. Great starter; deployable in minutes. Bottleneck is the Python hot path for very high QPS.
Industrial-grade L7 proxy. Use ext_proc or Wasm filters for auth and routing logic; keep the data path in C++. Harder to author, much higher ceiling. Right choice when you're past ~10k QPS or need mTLS / ambient mesh.
The vLLM router, SGLang Router, and NVIDIA Triton's ensemble scheduler all know about prefix cache and can route based on cache hit likelihood. Useful when you have many replicas and a stable prompt structure.
The only routing policy that matters for multi-turn: hash the first ~256 tokens, mod the number of replicas, and stick sessions to the matching replica. Deck 05 covered why.
vLLM's --enable-prefix-caching (deck 03). Hashes each 16-token block of prompt; reuses K/V across requests with a shared prefix. System prompts, tool schemas, few-shot examples get this for free.
Embed incoming prompts; if cosine similarity to a past prompt is > τ, return the cached response without touching the GPU. Redis Stack, Postgres pgvector, or Weaviate. Most useful for autocomplete-style workloads with repeated phrasings.
emb = embed(prompt)
hit = vstore.query(emb, k=1, threshold=0.93)
if hit: return hit.responseDeck 03 covered the mechanism; production has two extra concerns.
The acceptance rate depends on your workload. Benchmark three drafts (e.g. 1B, 3B, n-gram) on a representative log replay, measure tokens/second and p95 latency. Use the one that wins on your actual traffic, not what a paper says.
The draft model eats a chunk of VRAM too. Llama-3.2-1B paired with a 70B base takes ~2 GB. On a 24 GB card that's a real tradeoff vs KV pool size. Usually worth it on H100/Spark; often not on a 4090.
Speculative decoding introduces a second path through sampling (deck 09). For eval or audit traffic, route around the speculative pool. Cheapest: keep two vLLM backends — one with --speculative-model, one without — and route based on a header.
Four RED-style signals cover 95% of production questions:
| Signal | Prometheus metric | Alert when… |
|---|---|---|
| TTFT p95 | vllm:time_to_first_token_seconds | > SLO for 5 min |
| TPOT p95 | vllm:time_per_output_token_seconds | > 2× baseline for 5 min |
| KV usage | vllm:gpu_cache_usage_perc | > 90% sustained → preemption |
| Queue depth | vllm:num_requests_waiting | > N for 2 min → scale out |
Wire OpenTelemetry around the gateway: tag each span with model, tp_size, prompt_tokens, completion_tokens, cache_hit, replica_id. Then you can answer "why was this request slow" instead of "why is p95 slow". Langfuse / Phoenix plug straight in.
Don't log full prompt + completion in plaintext by default. Some data falls under GDPR / HIPAA / PCI. Log hashes + metrics; gate full body logging behind a role and an audit trail.
num_requests_waiting > 8 for 2 min → scale upThe killer question for internal hosting: "how much did feature X cost this month?" The inference cluster can answer it if you tag requests at the gateway.
x-team, x-feature, x-env on every requestprompt_tokens, completion_tokens in usageInternal price should be stable, not GPU-utilisation-dependent. Compute monthly cost (hardware amortisation + power + admin) ÷ monthly tokens served; that's your £/M-token rate. Revisit quarterly. Charging teams for variable GPU cost creates incentives nobody wants.
:latestcurl -X POST http://host:8000/v1/load -d '{"drain": true}'
# router routes new traffic elsewhere; existing streams finish
# wait for vllm:num_requests_running == 0
docker stop --time=120 vllm
In practice: let the load balancer health-check the replica's /health, flip the replica's liveness to "drain", and stop it after in-flight traffic clears. This is the same playbook as any HTTP server.
This deck is the one you want when moving from "works on my laptop" to "serves real users". The rest of the series covers orientation, Ollama and vLLM internals, Docker deployment, multi-GPU parallelism, frameworks, quantisation, hardware, and determinism — see the series index for the current full list.