Local LLM Hosting Series — Presentation 06

Framework Shootout

vLLM vs Ollama vs llama.cpp vs TGI vs SGLang vs TensorRT-LLM vs LMDeploy vs MLC — what each is for, what each costs to run, and when to switch.

vLLMTGISGLang TensorRT-LLMOllamallama.cpp LMDeployMLC
Feature matrix → Throughput → Latency → Ops → Pick
00

Topics

01

The Eight Stacks, One-line Each

vLLM

PagedAttention + continuous batching; the default for production NVIDIA serving. OpenAI-compat, multi-LoRA, CUDA + ROCm.

TGI (HuggingFace)

Rust+Python server with tight HF Hub integration. Enterprise support, solid on Hopper/H100; lags vLLM on cutting-edge kernels.

SGLang

RadixAttention (tree-structured KV cache) + structured-output DSL. Best for agents / tool-heavy workloads.

TensorRT-LLM

NVIDIA's compiled-engine path. Highest peak tok/s on Hopper/Blackwell. Build step is painful; only NVIDIA GPUs.

Ollama

llama.cpp + registry + REST. Single-user desktop king. Not for multi-tenant serving.

llama.cpp

Pure C++. CPU, Metal, CUDA, Vulkan, ROCm, HIP, even Android. GGUF format.

LMDeploy

InternLM's server. Strong AWQ path, good on Chinese-market models (Qwen, InternLM, DeepSeek).

MLC LLM

TVM-based cross-compile. Browser (WebGPU), iOS, Android, Vulkan. Only one that runs an LLM in Chrome.

02

Feature Matrix

FeaturevLLMTGISGLangTRT-LLMOllamallama.cppLMDeployMLC
PagedAttention / KV re-use✓✓✓ Radix✓—partial✓—
Continuous batching✓✓✓✓——✓—
Tensor parallel✓✓✓✓—limited✓—
Pipeline parallel✓✓✓✓——✓—
Expert parallel (MoE)✓✓✓✓——✓—
Multi-LoRA hot swap✓✓—✓——✓—
Speculative decoding✓✓✓✓—✓✓—
Prefix caching✓✓✓ (tree)✓—basic✓—
Structured output (JSON/grammar)outlines/xgrammar✓native✓format=jsongrammar✓—
FP8 / FP4✓FP8FP8✓ full——FP8—
AWQ / GPTQ / INT4✓✓✓✓—GGUF onlyAWQ native✓
CPU-only————✓✓—✓
Apple Silicon (Metal)————✓✓—✓
AMD ROCm✓✓✓—✓✓—✓
Browser / WebGPU———————✓
Prometheus metrics✓✓✓Triton——✓—
OpenAI-compat REST✓✓✓via Triton✓via wrapper✓via wrapper
03

Throughput & Latency — Indicative Numbers

Representative Llama-3.1-8B FP16 on a single H100 80 GB, 256-token prompt, 128-token output, 2026-era software. Your mileage will vary by ±30%; the shape is the lesson.

Aggregate throughput @ 64 concurrent requests (tok/s) TensorRT-LLM ~4900 vLLM ~4400 SGLang ~4300 TGI ~3800 LMDeploy ~3600 Ollama ~1300 (single-stream scaling) llama.cpp ~1300
p50 TTFT @ single user (ms; lower is better) Ollama (llama.cpp) ~30 ms TensorRT-LLM ~33 ms vLLM (eager=false) ~36 ms SGLang ~37 ms TGI ~45 ms
04

Interactive: Side-by-side Filter

Tick the features you need; shortlist updates.

05

Ops & Observability

Things production cares about

  • Prometheus / OpenMetrics endpoint
  • Structured request / response logs with request IDs
  • Graceful drain on SIGTERM (finish in-flight decodes)
  • Rolling restart without dropping streaming connections
  • Per-API-key token accounting
  • Safe OOM recovery (vLLM pre-empts, restarts its KV pool)

Where each stack lands

  • vLLM: strong on all six
  • TGI: strong on all six; nicer default logs than vLLM
  • SGLang: good; request accounting still maturing
  • TensorRT-LLM: production depends on Triton Inference Server wrapper
  • Ollama / llama.cpp: minimal; no real metering
  • LMDeploy: parity with vLLM on metrics
Triton + TensorRT-LLM

For peak NVIDIA performance in 2026, the production pattern is Triton Inference Server with a TensorRT-LLM backend. Triton handles batching, queueing, metrics, and model-repository management; TRT-LLM does the math. It's more moving parts than vLLM but it's how most NVIDIA reference architectures look.

06

Decision Tree

Do you have any NVIDIA GPU?
↓ yes  /  no → Ollama / llama.cpp / MLC
Is it a personal machine, single-user?
↓ yes → Ollama   /   no ↓
Do you need browser / WebGPU / iOS?
↓ yes → MLC LLM   /   no ↓
Need heavy structured output or agent trees?
↓ yes → SGLang   /   no ↓
Need absolute peak tok/s on H100/B200?
↓ yes → TensorRT-LLM (via Triton)   /   no ↓
Default: vLLM
Switching cost

All of these speak OpenAI-compat HTTP. Moving from vLLM to TGI to SGLang is usually one docker run change and a retest. Don't agonise — pick vLLM, measure, switch only if you've proven the alternative wins on your workload.