Local LLM Hosting Series — Presentation 03

Inside vLLM — PagedAttention & Continuous Batching

Why vLLM is 5–20× faster than a naive server: fragmentation, paged KV, continuous batching, prefix caching, chunked prefill, speculative decoding, and multi-LoRA.

vLLM PagedAttention Continuous Batching KV Cache Speculative Decoding LoRA
KV Cache → Fragmentation → Paging → Batching → Prefix Cache → Specdec
00

Topics We'll Cover

01

Why the KV Cache Dominates Serving

During decoding each new token needs attention over every prior token. Without caching you'd re-compute K and V for the whole prefix at every step — quadratic. The KV cache stores them once. But for a server hosting many sessions this cache is the memory bill.

Scale it

Llama-3-70B, 80 layers, hidden 8192 with GQA-8: KV per token ≈ 2 · 80 · 1024 · 2 B ≈ 320 KB. A single 8k-token session takes 2.5 GB. 32 concurrent 8k sessions = 80 GB — bigger than the weights themselves.

The Three Memory Pools in a Serving Engine

Weights

Static. Loaded once. 70B-FP16 = 140 GB, 70B-FP8 = 70 GB.

Activations

Transient per forward pass. Freed every step. Small relative to the others.

KV Cache

Grows with ( batch × sequence ). Unbounded from the engine's point of view. This is what PagedAttention is fighting.

02

Internal Fragmentation — The Core Problem

Pre-vLLM engines reserved one contiguous KV region per request, sized to the maximum context. Then:

Naive contiguous KV allocation — 4 sessions, max_ctx=2048, actual use varies sess 1 · used 1800 / 2048 sess 2 · used 500 → wasted 1548 sess 3 · used 1024 sess 4 · used 200 → wasted 1848 >> effective utilisation ~40% >> fewer concurrent sessions than VRAM suggests
03

PagedAttention — OS Virtual Memory for KV

vLLM borrows the textbook OS trick: break the KV cache into fixed-size blocks (typically 16 tokens each), maintain a per-sequence block table mapping logical positions to physical blocks, and implement a custom attention kernel that can gather K/V from the scattered physical blocks.

What changes

  • KV memory = a pool of fixed 16-token pages, allocated on demand.
  • Each sequence holds a block table: page_id[0..n_blocks].
  • Attention kernel (paged_attention_v2) dereferences the block table per head.
  • Two sequences sharing a prefix point to the same physical pages for that prefix.
  • Freeing a session returns pages to the pool; no compaction, no holes.

The reference

Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023. The PagedAttention CUDA kernel is in vllm/csrc/attention/paged_attention*.cu — worth reading if you've done any CUDA.

Block size is a knob: --block-size 16 is the default. Larger blocks reduce indirection overhead but increase fragmentation. 16 is the sweet spot on modern GPUs.

Block Table — Schematic

Two sessions share a system prompt prefix; diverge into private pages. physical pages (16 tok each): p0 p1 p2 p3 p4 p5 p6 p7 p8 p9 p10 p11 A.blocks = [p0, p1, p2, p3, p4] shared prefix p0..p2 + private p3,p4 B.blocks = [p0, p1, p2, p6] shared prefix p0..p2 + private p6 Prefix pages (green) are ref-counted; freed only when all sessions release. Copy-on-write if a session overwrites.
04

Interactive: Page Allocation in Action

Dial up concurrent sessions and mean context, and watch the waste: naive versus paged.

8
1024
8192
Naive reserved
—
Paged used
—
Waste saved
—
Extra sessions that now fit
—

Numbers use 300 KB/token — representative of Llama-70B FP16 with GQA-8. The point is the ratio, not the absolutes.

05

Continuous Batching — Iteration-Level Scheduling

In a static batch, every request in a batch runs until the slowest finishes. That's fine for training but catastrophic for serving: a 2000-token generation holds up a 20-token one.

vLLM schedules at the iteration level. Every decode step the scheduler picks which sequences run in this step. Completed sequences are evicted; new ones joined in. Batches are re-formed step-by-step.

Static batching — one bar = one step, four sessions. Continuous batching — short sessions exit, new ones join mid-batch. grey = idle compute; batch shrinks as sessions end new sessions slot into vacated compute → GPU stays saturated
Real impact

On the same hardware, continuous batching commonly delivers 5–10× the aggregate throughput of static batching for mixed workloads. It's the reason vLLM beats a naive HuggingFace model.generate loop.

06

Prefix Caching & Chunked Prefill

Automatic Prefix Caching (--enable-prefix-caching)

vLLM hashes each page's token content. Two requests with the same first N tokens share those N/block_size physical pages — no recomputation of K/V. For systems with a fixed system prompt or tool schema, TTFT (time-to-first-token) drops from hundreds of ms to ~10 ms on the cached prefix.

Turn it on
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct \
    --enable-prefix-caching \
    --block-size 16

Chunked Prefill (--enable-chunked-prefill)

Prefill is compute-bound; decode is memory-bound. If a long-prompt prefill is running, every concurrent decode stalls behind it. Chunked prefill breaks the prefill into batches of, say, 512 tokens and interleaves those chunks with decodes from other sessions. p99 decode latency drops dramatically without hurting prefill throughput much.

Without chunked prefill

  • One long prompt (8k tokens) owns the GPU for ~400 ms
  • Every live decode session stalls for that window
  • Tail latency explodes

With chunked prefill

  • Prefill issued in 512-token chunks
  • Interleaved with ongoing decodes
  • Prefill completes slightly slower; decodes keep their latency budget
07

Speculative Decoding & Medusa

Decoding is memory-bandwidth bound: the bottleneck is reading the 70B weights once per token. Speculative decoding amortises that read by guessing several tokens at once with a cheap draft model and validating them in a single forward pass of the big model.

Draft 5 tokens, verify in parallel, accept the longest matching prefix. draft model — 5 guesses: A B C D E big model — one parallel forward pass, compares logits accepted 3 of 5 → 3 tokens in the time of one big-model step

vLLM options

ApproachFlagNotes
External draft model--speculative-model <small>Pair Llama-3.2-1B with Llama-3.3-70B
Medusa heads--speculative-model medusaExtra LM heads predict next-next-… tokens
EAGLE / EAGLE-2--speculative-model eagleSingle-forward draft from hidden states
n-gram--speculative-model [ngram]Draft from the prompt itself, almost free
Reality check

Speculative decoding only wins when the acceptance rate is high. For open-ended creative writing you may accept 2/5 tokens; for code completion with tight distributions you can hit 4/5. Always measure — a low acceptance rate can actually make things slower.

08

Multi-LoRA Serving

LoRA adapters are tiny — a few hundred MB each. vLLM's --enable-lora keeps many adapters hot in VRAM and picks the right one per request via the model field. One base model, dozens of personalities.

Serve a base model with three adapters
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct \
    --enable-lora \
    --lora-modules sql-expert=/models/sql-lora \
                   code-review=/models/cr-lora \
                   legal-review=/models/legal-lora \
    --max-lora-rank 32
Routing request
curl http://host:8000/v1/chat/completions -d '{
  "model": "sql-expert",          # ← picks the SQL adapter
  "messages":[{"role":"user","content":"explain this join plan"}]
}'
Why this matters

You can serve 30 fine-tuned variants on one 24 GB GPU without duplicating the base model. Economics flip completely compared to full-model fine-tunes. This is the feature that makes vLLM viable for internal "model per team" patterns.

09

What to Take Away

Next

Deck 04 shows the Docker image and docker run line that puts all of this behind an OpenAI-compatible port on your own NVIDIA GPU. Deck 05 adds multi-GPU parallelism.