Why vLLM is 5–20× faster than a naive server: fragmentation, paged KV, continuous batching, prefix caching, chunked prefill, speculative decoding, and multi-LoRA.
During decoding each new token needs attention over every prior token. Without caching you'd re-compute K and V for the whole prefix at every step — quadratic. The KV cache stores them once. But for a server hosting many sessions this cache is the memory bill.
Llama-3-70B, 80 layers, hidden 8192 with GQA-8: KV per token ≈ 2 · 80 · 1024 · 2 B ≈ 320 KB. A single 8k-token session takes 2.5 GB. 32 concurrent 8k sessions = 80 GB — bigger than the weights themselves.
Static. Loaded once. 70B-FP16 = 140 GB, 70B-FP8 = 70 GB.
Transient per forward pass. Freed every step. Small relative to the others.
Grows with ( batch × sequence ). Unbounded from the engine's point of view. This is what PagedAttention is fighting.
Pre-vLLM engines reserved one contiguous KV region per request, sized to the maximum context. Then:
vLLM borrows the textbook OS trick: break the KV cache into fixed-size blocks (typically 16 tokens each), maintain a per-sequence block table mapping logical positions to physical blocks, and implement a custom attention kernel that can gather K/V from the scattered physical blocks.
page_id[0..n_blocks].paged_attention_v2) dereferences the block table per head.Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023. The PagedAttention CUDA kernel is in vllm/csrc/attention/paged_attention*.cu — worth reading if you've done any CUDA.
Block size is a knob: --block-size 16 is the default. Larger blocks reduce indirection overhead but increase fragmentation. 16 is the sweet spot on modern GPUs.
Dial up concurrent sessions and mean context, and watch the waste: naive versus paged.
Numbers use 300 KB/token — representative of Llama-70B FP16 with GQA-8. The point is the ratio, not the absolutes.
In a static batch, every request in a batch runs until the slowest finishes. That's fine for training but catastrophic for serving: a 2000-token generation holds up a 20-token one.
vLLM schedules at the iteration level. Every decode step the scheduler picks which sequences run in this step. Completed sequences are evicted; new ones joined in. Batches are re-formed step-by-step.
On the same hardware, continuous batching commonly delivers 5–10× the aggregate throughput of static batching for mixed workloads. It's the reason vLLM beats a naive HuggingFace model.generate loop.
--enable-prefix-caching)vLLM hashes each page's token content. Two requests with the same first N tokens share those N/block_size physical pages — no recomputation of K/V. For systems with a fixed system prompt or tool schema, TTFT (time-to-first-token) drops from hundreds of ms to ~10 ms on the cached prefix.
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct \
--enable-prefix-caching \
--block-size 16
--enable-chunked-prefill)Prefill is compute-bound; decode is memory-bound. If a long-prompt prefill is running, every concurrent decode stalls behind it. Chunked prefill breaks the prefill into batches of, say, 512 tokens and interleaves those chunks with decodes from other sessions. p99 decode latency drops dramatically without hurting prefill throughput much.
Decoding is memory-bandwidth bound: the bottleneck is reading the 70B weights once per token. Speculative decoding amortises that read by guessing several tokens at once with a cheap draft model and validating them in a single forward pass of the big model.
| Approach | Flag | Notes |
|---|---|---|
| External draft model | --speculative-model <small> | Pair Llama-3.2-1B with Llama-3.3-70B |
| Medusa heads | --speculative-model medusa | Extra LM heads predict next-next-… tokens |
| EAGLE / EAGLE-2 | --speculative-model eagle | Single-forward draft from hidden states |
| n-gram | --speculative-model [ngram] | Draft from the prompt itself, almost free |
Speculative decoding only wins when the acceptance rate is high. For open-ended creative writing you may accept 2/5 tokens; for code completion with tight distributions you can hit 4/5. Always measure — a low acceptance rate can actually make things slower.
LoRA adapters are tiny — a few hundred MB each. vLLM's --enable-lora keeps many adapters hot in VRAM and picks the right one per request via the model field. One base model, dozens of personalities.
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct \
--enable-lora \
--lora-modules sql-expert=/models/sql-lora \
code-review=/models/cr-lora \
legal-review=/models/legal-lora \
--max-lora-rank 32
curl http://host:8000/v1/chat/completions -d '{
"model": "sql-expert", # ← picks the SQL adapter
"messages":[{"role":"user","content":"explain this join plan"}]
}'
You can serve 30 fine-tuned variants on one 24 GB GPU without duplicating the base model. Economics flip completely compared to full-model fine-tunes. This is the feature that makes vLLM viable for internal "model per team" patterns.
Deck 04 shows the Docker image and docker run line that puts all of this behind an OpenAI-compatible port on your own NVIDIA GPU. Deck 05 adds multi-GPU parallelism.