Architecture, Modelfiles, the REST API, VRAM tuning, multi-model routing — everything needed to use Ollama as a serious tool, not just a demo.
Ollama is a Go-language daemon wrapping llama.cpp, plus a model registry (ollama.com/library), plus a CLI, plus a REST API on localhost:11434. It's not a new inference engine — llama.cpp does the heavy lifting. The value Ollama adds is packaging and ergonomics.
Ollama = Docker-like Hub for models + daemon that keeps one loaded in VRAM. Under the hood every token is still produced by llama.cpp.
# One-line install — installs the daemon as a systemd service on Linux
curl -fsSL https://ollama.com/install.sh | sh
# Check the daemon is up
systemctl status ollama # linux
curl http://localhost:11434 # should return "Ollama is running"
ollama pull llama3.1:8b # ~4.7 GB download, q4_K_M
ollama run llama3.1:8b # drops into interactive chat
# Or one-shot:
ollama run llama3.1:8b "Explain PagedAttention in one paragraph."
| Path | What |
|---|---|
~/.ollama/models/ (Linux / macOS) | GGUF weights + manifests |
/usr/share/ollama/.ollama/models | system install (systemd) |
OLLAMA_MODELS=/mnt/nvme | override (put weights on a fast SSD) |
OLLAMA_HOST=0.0.0.0:11434 | bind to LAN (careful — unauthenticated!) |
OLLAMA_KEEP_ALIVE=24h | how long to keep a loaded model in VRAM |
Every image in the library is a GGUF bundle with a default quantisation, system prompt, and stop tokens. The tag after : is either a parameter size or a quant variant.
# Default (usually q4_K_M — a good middle ground)
ollama pull llama3.1:8b
# Explicit quantisation
ollama pull llama3.1:8b-instruct-q8_0 # near-FP16 quality, 8 GB
ollama pull llama3.1:8b-instruct-q4_K_M # standard, 4.7 GB
ollama pull llama3.1:8b-instruct-q2_K # tiny, lossy, 3.2 GB
# Useful families (2026 landscape)
ollama pull qwen2.5:14b # strong all-rounder
ollama pull deepseek-r1:14b # reasoning distills
ollama pull mistral-small:24b # Apache licence, solid IF
ollama pull phi3:mini # 3.8B, laptop-friendly
ollama pull nomic-embed-text # embeddings for RAG
Models are content-addressed blobs under ~/.ollama/models/blobs/ with manifests under manifests/registry.ollama.ai/library/. Because identical quantised layers de-dup by SHA-256, pulling ten variants of Llama-3.1 is not ten times the disk.
Put ~/.ollama on NVMe. Cold-load of a 40 GB q4 weight file off a spinning disk adds 30–60 s to first-token latency. On NVMe it's 2–5 s. OLLAMA_KEEP_ALIVE=-1 pins models forever — use this on a serving box.
A Modelfile is a Dockerfile-style spec that baked-in parameters, a system prompt, and optional LoRA adapters on top of a base image. Build one, tag it, and it behaves like any other pulled model.
FROM qwen2.5-coder:14b-instruct-q5_K_M
# Determinism for tests
PARAMETER temperature 0.1
PARAMETER top_p 0.95
PARAMETER num_ctx 16384
PARAMETER num_predict 2048
PARAMETER stop "<|im_end|>"
SYSTEM """
You are a senior Python engineer. Reply with code blocks only unless
a short prose explanation is strictly necessary. Prefer the standard
library. Never invent APIs.
"""
# Optional: bake in a LoRA adapter exported as a GGUF file
# ADAPTER ./my-codebase-lora.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}
{{ range .Messages }}<|im_start|>{{ .Role }}
{{ .Content }}<|im_end|>
{{ end }}<|im_start|>assistant
"""
ollama create myco-coder -f Modelfile
ollama run myco-coder "Refactor utils/paths.py to use pathlib"
# Inspect what got baked
ollama show myco-coder --modelfile
ollama show myco-coder --parameters
Anytime you'd otherwise paste the same system prompt 50 times — repo-specific coder, RAG-aware retriever, tool-using agent, redaction filter for PII. Modelfiles also commit nicely to your repo next to the code that talks to the model.
Two API surfaces, same daemon. Use /api/chat for Ollama-native features, use /v1/chat/completions for drop-in compatibility with any OpenAI SDK.
curl -s http://localhost:11434/api/chat -d '{
"model": "llama3.1:8b",
"messages": [{"role":"user","content":"Explain KV cache"}],
"options": { "num_ctx": 8192, "temperature": 0.2 },
"stream": true
}'
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
r = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Hi"}],
stream=True,
)
for chunk in r:
print(chunk.choices[0].delta.content or "", end="", flush=True)
| Endpoint | Purpose |
|---|---|
POST /api/generate | Single-prompt completion |
POST /api/chat | Chat with message history + tool calls |
POST /api/embeddings | Vectors (nomic-embed-text etc.) |
POST /v1/chat/completions | OpenAI-compat chat, incl. function calling |
GET /api/tags | List local models |
GET /api/ps | What's currently loaded in VRAM |
POST /api/pull | Pull a model by API (useful in CI) |
Ollama supports tools on the OpenAI endpoint for models trained on tool-calling (Llama 3.1/3.2, Qwen 2.5, Mistral). The daemon parses the model's JSON tool-call output and returns it in the OpenAI-standard shape — so LangChain and the OpenAI SDK's function-calling helpers just work.
| Parameter | Effect | Good defaults |
|---|---|---|
num_ctx | Context window. Larger → more VRAM for KV cache. | 4096–16384 |
num_gpu | How many layers to push to GPU. -1 = all. | -1 if it fits, else a number |
num_batch | Batch size for prefill. More throughput, more VRAM. | 512 |
num_thread | CPU threads for CPU-offloaded layers. | physical cores |
OLLAMA_FLASH_ATTENTION=1 | Enable flash-attention kernels (huge win on long ctx). | on |
OLLAMA_KV_CACHE_TYPE=q8_0 | Quantise the KV cache itself → ~2× longer ctx in same VRAM | q8_0 or q4_0 |
If the model is larger than VRAM, Ollama offloads the top N layers to CPU. Throughput drops roughly linearly with the fraction offloaded — one layer on CPU often halves tok/s. Prefer a smaller quant over offloading.
OLLAMA_VERBOSE=1 ollama run llama3.1:70b-q4_K_M
# Look for "offloaded N/M layers to GPU" in the logs.
# Explicit via options
curl -s http://localhost:11434/api/generate -d '{
"model":"llama3.1:70b-q4_K_M",
"prompt":"ping",
"options":{"num_gpu": 60}
}'
Setting OLLAMA_KV_CACHE_TYPE=q8_0 (or q4_0) halves the KV cache memory with near-zero quality loss on short-to-medium contexts. On a 24 GB card this is often the difference between 16k and 32k of context for a 7B model.
Pick a model size, quant, and context length. The calculator applies weights × B/param + KV and reports whether it fits on common GPUs.
Self-hosted ChatGPT-alike. Points at http://ollama:11434, handles multi-user auth, chat history, image uploads, RAG, tool plugins. The usual "chatbot for the house".
docker run -d -p 3000:8080 \
-e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
--name open-webui ghcr.io/open-webui/open-webui:mainA drop-in OpenAI-compat proxy that can route to Ollama, Claude, GPT, Bedrock, etc. Lets you swap Ollama for a hosted model with one env var.
model_list:
- model_name: llama3
litellm_params:
model: ollama/llama3.1:8b
api_base: http://localhost:11434langchain_ollama.ChatOllama is a first-class chat model. Function calling and streaming work; useful for prototyping agent graphs before running them against a paid API.
Continue.dev, Zed, Cursor-local, Neovim's Avante — all support a generic OpenAI endpoint. Point them at http://localhost:11434/v1 and set api_key=ollama.
Ollama is excellent at what it does. It is not a multi-tenant inference server. The limits below are where you should start looking at vLLM (deck 03).
If you are serving > 2 concurrent users, or if throughput per GPU matters, move to vLLM (deck 03). If you're doing mobile / browser / Vulkan targets, look at MLC (deck 05). Otherwise — Ollama is the right tool and you should stop reading.