Local LLM Hosting Series — Presentation 02

Ollama — Zero-Friction Local LLMs

Architecture, Modelfiles, the REST API, VRAM tuning, multi-model routing — everything needed to use Ollama as a serious tool, not just a demo.

Ollama llama.cpp GGUF Modelfile REST API OpenAI-compatible
Install → Pull → Modelfile → API → Tuning → Limits
00

Topics We'll Cover

01

What Ollama Actually Is

Ollama is a Go-language daemon wrapping llama.cpp, plus a model registry (ollama.com/library), plus a CLI, plus a REST API on localhost:11434. It's not a new inference engine — llama.cpp does the heavy lifting. The value Ollama adds is packaging and ergonomics.

Ollama stack — runtime view CLI ( ollama run ) SSE streaming, user stdin HTTP client curl / LiteLLM / LangChain Open WebUI / Obsidian / IDE any OpenAI-compatible client ollama serve · :11434 · REST + OpenAI-compat routes /api/generate, /api/chat, /v1/chat/completions llama.cpp (GGML / GGUF, CUDA · Metal · ROCm · Vulkan)
Mental model

Ollama = Docker-like Hub for models + daemon that keeps one loaded in VRAM. Under the hood every token is still produced by llama.cpp.

02

Install & First Run

Linux / macOS install
# One-line install — installs the daemon as a systemd service on Linux
curl -fsSL https://ollama.com/install.sh | sh

# Check the daemon is up
systemctl status ollama       # linux
curl http://localhost:11434   # should return "Ollama is running"
First chat
ollama pull llama3.1:8b       # ~4.7 GB download, q4_K_M
ollama run  llama3.1:8b       # drops into interactive chat

# Or one-shot:
ollama run llama3.1:8b "Explain PagedAttention in one paragraph."

Where Things Live

PathWhat
~/.ollama/models/ (Linux / macOS)GGUF weights + manifests
/usr/share/ollama/.ollama/modelssystem install (systemd)
OLLAMA_MODELS=/mnt/nvmeoverride (put weights on a fast SSD)
OLLAMA_HOST=0.0.0.0:11434bind to LAN (careful — unauthenticated!)
OLLAMA_KEEP_ALIVE=24hhow long to keep a loaded model in VRAM
03

The Model Library & Pulling Weights

Every image in the library is a GGUF bundle with a default quantisation, system prompt, and stop tokens. The tag after : is either a parameter size or a quant variant.

Pull variants
# Default (usually q4_K_M — a good middle ground)
ollama pull llama3.1:8b

# Explicit quantisation
ollama pull llama3.1:8b-instruct-q8_0    # near-FP16 quality, 8 GB
ollama pull llama3.1:8b-instruct-q4_K_M  # standard, 4.7 GB
ollama pull llama3.1:8b-instruct-q2_K    # tiny, lossy, 3.2 GB

# Useful families (2026 landscape)
ollama pull qwen2.5:14b        # strong all-rounder
ollama pull deepseek-r1:14b    # reasoning distills
ollama pull mistral-small:24b  # Apache licence, solid IF
ollama pull phi3:mini          # 3.8B, laptop-friendly
ollama pull nomic-embed-text   # embeddings for RAG

Storage Layout

Models are content-addressed blobs under ~/.ollama/models/blobs/ with manifests under manifests/registry.ollama.ai/library/. Because identical quantised layers de-dup by SHA-256, pulling ten variants of Llama-3.1 is not ten times the disk.

Practical tip

Put ~/.ollama on NVMe. Cold-load of a 40 GB q4 weight file off a spinning disk adds 30–60 s to first-token latency. On NVMe it's 2–5 s. OLLAMA_KEEP_ALIVE=-1 pins models forever — use this on a serving box.

04

Modelfiles — Your Custom Builds

A Modelfile is a Dockerfile-style spec that baked-in parameters, a system prompt, and optional LoRA adapters on top of a base image. Build one, tag it, and it behaves like any other pulled model.

Modelfile — code-assistant variant
FROM qwen2.5-coder:14b-instruct-q5_K_M

# Determinism for tests
PARAMETER temperature   0.1
PARAMETER top_p         0.95
PARAMETER num_ctx       16384
PARAMETER num_predict   2048
PARAMETER stop          "<|im_end|>"

SYSTEM """
You are a senior Python engineer. Reply with code blocks only unless
a short prose explanation is strictly necessary. Prefer the standard
library. Never invent APIs.
"""

# Optional: bake in a LoRA adapter exported as a GGUF file
# ADAPTER ./my-codebase-lora.gguf

TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}
{{ range .Messages }}<|im_start|>{{ .Role }}
{{ .Content }}<|im_end|>
{{ end }}<|im_start|>assistant
"""
Build & use
ollama create myco-coder -f Modelfile
ollama run    myco-coder "Refactor utils/paths.py to use pathlib"

# Inspect what got baked
ollama show myco-coder --modelfile
ollama show myco-coder --parameters
When to use Modelfiles

Anytime you'd otherwise paste the same system prompt 50 times — repo-specific coder, RAG-aware retriever, tool-using agent, redaction filter for PII. Modelfiles also commit nicely to your repo next to the code that talks to the model.

05

The REST API & OpenAI Compatibility

Two API surfaces, same daemon. Use /api/chat for Ollama-native features, use /v1/chat/completions for drop-in compatibility with any OpenAI SDK.

Native chat — streaming NDJSON
curl -s http://localhost:11434/api/chat -d '{
  "model": "llama3.1:8b",
  "messages": [{"role":"user","content":"Explain KV cache"}],
  "options": { "num_ctx": 8192, "temperature": 0.2 },
  "stream": true
}'
OpenAI-compatible — any OpenAI SDK just works
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
r = client.chat.completions.create(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "Hi"}],
    stream=True,
)
for chunk in r:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Endpoints Cheat-Sheet

EndpointPurpose
POST /api/generateSingle-prompt completion
POST /api/chatChat with message history + tool calls
POST /api/embeddingsVectors (nomic-embed-text etc.)
POST /v1/chat/completionsOpenAI-compat chat, incl. function calling
GET /api/tagsList local models
GET /api/psWhat's currently loaded in VRAM
POST /api/pullPull a model by API (useful in CI)
Tool calling

Ollama supports tools on the OpenAI endpoint for models trained on tool-calling (Llama 3.1/3.2, Qwen 2.5, Mistral). The daemon parses the model's JSON tool-call output and returns it in the OpenAI-standard shape — so LangChain and the OpenAI SDK's function-calling helpers just work.

06

VRAM & Performance Tuning

What the knobs actually do

ParameterEffectGood defaults
num_ctxContext window. Larger → more VRAM for KV cache.4096–16384
num_gpuHow many layers to push to GPU. -1 = all.-1 if it fits, else a number
num_batchBatch size for prefill. More throughput, more VRAM.512
num_threadCPU threads for CPU-offloaded layers.physical cores
OLLAMA_FLASH_ATTENTION=1Enable flash-attention kernels (huge win on long ctx).on
OLLAMA_KV_CACHE_TYPE=q8_0Quantise the KV cache itself → ~2× longer ctx in same VRAMq8_0 or q4_0

Split-offload: when the model doesn't fit

If the model is larger than VRAM, Ollama offloads the top N layers to CPU. Throughput drops roughly linearly with the fraction offloaded — one layer on CPU often halves tok/s. Prefer a smaller quant over offloading.

Force full GPU, see what actually happened
OLLAMA_VERBOSE=1 ollama run llama3.1:70b-q4_K_M
# Look for "offloaded N/M layers to GPU" in the logs.

# Explicit via options
curl -s http://localhost:11434/api/generate -d '{
  "model":"llama3.1:70b-q4_K_M",
  "prompt":"ping",
  "options":{"num_gpu": 60}
}'
The KV-cache quant trick

Setting OLLAMA_KV_CACHE_TYPE=q8_0 (or q4_0) halves the KV cache memory with near-zero quality loss on short-to-medium contexts. On a 24 GB card this is often the difference between 16k and 32k of context for a 7B model.

07

Interactive: Sizing Calculator

Pick a model size, quant, and context length. The calculator applies weights × B/param + KV and reports whether it fits on common GPUs.

8 B
8192
Weights
—
KV cache
—
Total VRAM
—
08

Integrations — LiteLLM, LangChain, Open WebUI

Open WebUI

Self-hosted ChatGPT-alike. Points at http://ollama:11434, handles multi-user auth, chat history, image uploads, RAG, tool plugins. The usual "chatbot for the house".

docker compose
docker run -d -p 3000:8080 \
  -e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
  --name open-webui ghcr.io/open-webui/open-webui:main

LiteLLM

A drop-in OpenAI-compat proxy that can route to Ollama, Claude, GPT, Bedrock, etc. Lets you swap Ollama for a hosted model with one env var.

config.yaml
model_list:
  - model_name: llama3
    litellm_params:
      model: ollama/llama3.1:8b
      api_base: http://localhost:11434

LangChain / LangGraph

langchain_ollama.ChatOllama is a first-class chat model. Function calling and streaming work; useful for prototyping agent graphs before running them against a paid API.

IDE / editor

Continue.dev, Zed, Cursor-local, Neovim's Avante — all support a generic OpenAI endpoint. Point them at http://localhost:11434/v1 and set api_key=ollama.

09

Limits — When to Graduate to vLLM

Ollama is excellent at what it does. It is not a multi-tenant inference server. The limits below are where you should start looking at vLLM (deck 03).

Where Ollama wins

  • One human + one model, or a few humans sharing casually
  • Desktop / laptop / single-GPU home server
  • Apple Silicon (MLX isn't yet as polished)
  • Fast model-swapping: it unloads and loads in seconds
  • Privacy-by-default — no cloud round-trips, no telemetry of prompts

Where Ollama falls over

  • No continuous batching. 8 concurrent users → 8× the latency, not 8× the throughput.
  • No tensor parallel. Can't split a 70B FP16 across two GPUs.
  • No multi-LoRA serving. One adapter at a time.
  • GGUF-only. No native AWQ / GPTQ / FP8; you get llama.cpp's quants only.
  • No production telemetry. No Prometheus, no token-usage accounting, no per-API-key metering.
Rule of thumb

If you are serving > 2 concurrent users, or if throughput per GPU matters, move to vLLM (deck 03). If you're doing mobile / browser / Vulkan targets, look at MLC (deck 05). Otherwise — Ollama is the right tool and you should stop reading.