Presentations in This Series
- Why Host Locally — The Landscape →Drivers, cost model, ecosystem map — Ollama, vLLM, llama.cpp, TGI, SGLang, TensorRT-LLM, LMDeploy, MLC. Interactive picker.
- Ollama — Zero-Friction Local LLMs →Architecture, Modelfiles, REST and OpenAI-compat APIs, VRAM tuning, the KV-cache-quant trick, sizing calculator.
- Inside vLLM — PagedAttention & Continuous Batching →Why vLLM is 5–20× faster — KV fragmentation, paged attention, continuous batching, prefix caching, chunked prefill, specdec, multi-LoRA.
- vLLM in Docker on NVIDIA GPUs →From a blank Linux box to an OpenAI-compatible vLLM endpoint — NVIDIA Container Toolkit, the vllm/vllm-openai image, compose, gotchas.
- Multi-GPU Parallelism for Serving →Tensor, pipeline, data and expert parallel; NVLink, NVSwitch, NCCL, InfiniBand; multi-node patterns; interactive TP/PP/DP/EP planner.
- Framework Shootout →Feature matrix and honest throughput/latency numbers across vLLM, Ollama, llama.cpp, TGI, SGLang, TensorRT-LLM, LMDeploy, MLC.
- Quantization for Local Hosting →GGUF, AWQ, GPTQ, FP8, INT4, MX-FP4 — bit layouts, perplexity impact, framework support, format picker.
- Deploying on NVIDIA DGX Spark →Practical vLLM on a GB10 Grace-Blackwell workstation — unified memory, arm64 gotchas, model/concurrency matrix, two-Spark pairing.
- Sources of Non-Determinism in LLM Serving →Why the same prompt gives different tokens twice — sampling, non-associative FP, batch-size dependence, scheduler drift.
- Production Patterns for Local LLM Serving →Routing, prefix & semantic caching, speculative decoding, observability, SLOs, rollouts, cost accounting, SLO planner.
- Deploying on NVIDIA GPUs →Architectures, memory, multi-GPU, ganging; NVLink/NVSwitch/PCIe P2P/IOMMU; MIG; Ollama/vLLM nuances per GPU class.