LLM Hub — Local LLM Hosting

Local LLM Hosting

Self-hosting LLMs in 2025/2026 — Ollama, vLLM, llama.cpp, TGI, SGLang and friends, plus Docker, multi-GPU, quantisation, determinism, and production patterns.

vLLMOllamallama.cppMulti-GPUQuantisationDGX Spark

Presentations in This Series

  1. Why Host Locally — The Landscape →
    Drivers, cost model, ecosystem map — Ollama, vLLM, llama.cpp, TGI, SGLang, TensorRT-LLM, LMDeploy, MLC. Interactive picker.
    live
  2. Ollama — Zero-Friction Local LLMs →
    Architecture, Modelfiles, REST and OpenAI-compat APIs, VRAM tuning, the KV-cache-quant trick, sizing calculator.
    live
  3. Inside vLLM — PagedAttention & Continuous Batching →
    Why vLLM is 5–20× faster — KV fragmentation, paged attention, continuous batching, prefix caching, chunked prefill, specdec, multi-LoRA.
    live
  4. vLLM in Docker on NVIDIA GPUs →
    From a blank Linux box to an OpenAI-compatible vLLM endpoint — NVIDIA Container Toolkit, the vllm/vllm-openai image, compose, gotchas.
    live
  5. Multi-GPU Parallelism for Serving →
    Tensor, pipeline, data and expert parallel; NVLink, NVSwitch, NCCL, InfiniBand; multi-node patterns; interactive TP/PP/DP/EP planner.
    live
  6. Framework Shootout →
    Feature matrix and honest throughput/latency numbers across vLLM, Ollama, llama.cpp, TGI, SGLang, TensorRT-LLM, LMDeploy, MLC.
    live
  7. Quantization for Local Hosting →
    GGUF, AWQ, GPTQ, FP8, INT4, MX-FP4 — bit layouts, perplexity impact, framework support, format picker.
    live
  8. Deploying on NVIDIA DGX Spark →
    Practical vLLM on a GB10 Grace-Blackwell workstation — unified memory, arm64 gotchas, model/concurrency matrix, two-Spark pairing.
    live
  9. Sources of Non-Determinism in LLM Serving →
    Why the same prompt gives different tokens twice — sampling, non-associative FP, batch-size dependence, scheduler drift.
    live
  10. Production Patterns for Local LLM Serving →
    Routing, prefix & semantic caching, speculative decoding, observability, SLOs, rollouts, cost accounting, SLO planner.
    live
  11. Deploying on NVIDIA GPUs →
    Architectures, memory, multi-GPU, ganging; NVLink/NVSwitch/PCIe P2P/IOMMU; MIG; Ollama/vLLM nuances per GPU class.
    live