vLLM vs Ollama vs llama.cpp vs TGI vs SGLang vs TensorRT-LLM vs LMDeploy vs MLC — what each is for, what each costs to run, and when to switch.
PagedAttention + continuous batching; the default for production NVIDIA serving. OpenAI-compat, multi-LoRA, CUDA + ROCm.
Rust+Python server with tight HF Hub integration. Enterprise support, solid on Hopper/H100; lags vLLM on cutting-edge kernels.
RadixAttention (tree-structured KV cache) + structured-output DSL. Best for agents / tool-heavy workloads.
NVIDIA's compiled-engine path. Highest peak tok/s on Hopper/Blackwell. Build step is painful; only NVIDIA GPUs.
llama.cpp + registry + REST. Single-user desktop king. Not for multi-tenant serving.
Pure C++. CPU, Metal, CUDA, Vulkan, ROCm, HIP, even Android. GGUF format.
InternLM's server. Strong AWQ path, good on Chinese-market models (Qwen, InternLM, DeepSeek).
TVM-based cross-compile. Browser (WebGPU), iOS, Android, Vulkan. Only one that runs an LLM in Chrome.
| Feature | vLLM | TGI | SGLang | TRT-LLM | Ollama | llama.cpp | LMDeploy | MLC |
|---|---|---|---|---|---|---|---|---|
| PagedAttention / KV re-use | ✓ | ✓ | ✓ Radix | ✓ | — | partial | ✓ | — |
| Continuous batching | ✓ | ✓ | ✓ | ✓ | — | — | ✓ | — |
| Tensor parallel | ✓ | ✓ | ✓ | ✓ | — | limited | ✓ | — |
| Pipeline parallel | ✓ | ✓ | ✓ | ✓ | — | — | ✓ | — |
| Expert parallel (MoE) | ✓ | ✓ | ✓ | ✓ | — | — | ✓ | — |
| Multi-LoRA hot swap | ✓ | ✓ | — | ✓ | — | — | ✓ | — |
| Speculative decoding | ✓ | ✓ | ✓ | ✓ | — | ✓ | ✓ | — |
| Prefix caching | ✓ | ✓ | ✓ (tree) | ✓ | — | basic | ✓ | — |
| Structured output (JSON/grammar) | outlines/xgrammar | ✓ | native | ✓ | format=json | grammar | ✓ | — |
| FP8 / FP4 | ✓ | FP8 | FP8 | ✓ full | — | — | FP8 | — |
| AWQ / GPTQ / INT4 | ✓ | ✓ | ✓ | ✓ | — | GGUF only | AWQ native | ✓ |
| CPU-only | — | — | — | — | ✓ | ✓ | — | ✓ |
| Apple Silicon (Metal) | — | — | — | — | ✓ | ✓ | — | ✓ |
| AMD ROCm | ✓ | ✓ | ✓ | — | ✓ | ✓ | — | ✓ |
| Browser / WebGPU | — | — | — | — | — | — | — | ✓ |
| Prometheus metrics | ✓ | ✓ | ✓ | Triton | — | — | ✓ | — |
| OpenAI-compat REST | ✓ | ✓ | ✓ | via Triton | ✓ | via wrapper | ✓ | via wrapper |
Representative Llama-3.1-8B FP16 on a single H100 80 GB, 256-token prompt, 128-token output, 2026-era software. Your mileage will vary by ±30%; the shape is the lesson.
Tick the features you need; shortlist updates.
SIGTERM (finish in-flight decodes)For peak NVIDIA performance in 2026, the production pattern is Triton Inference Server with a TensorRT-LLM backend. Triton handles batching, queueing, metrics, and model-repository management; TRT-LLM does the math. It's more moving parts than vLLM but it's how most NVIDIA reference architectures look.
All of these speak OpenAI-compat HTTP. Moving from vLLM to TGI to SGLang is usually one docker run change and a retest. Don't agonise — pick vLLM, measure, switch only if you've proven the alternative wins on your workload.