Local LLM Hosting Series — Presentation 09

Sources of Non-Determinism in LLM Serving

Why the same prompt gives different tokens twice, where determinism really comes from, and how to get it back when you need reproducible evals.

DeterminismBatch invariance FP16Sampling Evals
Why we care → Sampling → FP math → Batch → Scheduler → Mitigate
00

Topics

01

Why Reproducibility Matters — And When It Doesn't

Matters

  • Evals — comparing two models or two prompt variants is impossible if outputs drift
  • Regression tests in CI — "does this prompt still produce valid JSON"
  • Audit / compliance — reproducing what a model said last Tuesday
  • Research — ablation studies that require identical conditions

Does not matter

  • Production chat / agents — some variability is harmless and often desirable
  • Creative writing — you want diversity
  • Multi-turn search — you care about final answer quality, not token-by-token identity
Starting premise

Most engineers think "temperature=0 means deterministic". It does not. Below we walk through seven independent reasons the same prompt can produce different tokens even at temperature 0.

02

Sampling — The Obvious Source

Temperature-based sampling draws from the softmax distribution. With temperature > 0, different RNG seeds give different tokens. Fixing this is easy if you know where the seed lives:

Seed the sampler (vLLM)
curl http://host:8000/v1/chat/completions -d '{
  "model": "...",
  "messages": [...],
  "temperature": 0.0,
  "seed": 42,
  "top_p": 1.0,
  "top_k": -1
}'

The ties problem

Even with temperature=0, the argmax can tie — two tokens with the exact same logit. Which one wins depends on the CUDA kernel's tie-break rule and sometimes on the warp scheduling. Rare in practice but it's one of the reasons temperature=0 isn't a strict guarantee.

Speculative decoding silently re-rolls

With speculative decoding enabled, the draft model proposes N tokens and the big model accepts a prefix. When the draft is rejected, the big model's sample replaces it. Different draft rejection counts → different interleavings → different sampler calls. Numbers identical over many runs only with temperature=0 AND a deterministic draft model (beware batch-dependent kernels).

03

Non-Associative Floating-Point — The Quiet One

Floating-point addition is not associative: (a + b) + c ≠ a + (b + c) in general. Every reduction the GPU does — softmax, layer-norm, row-sums in matmul — depends on order. Change the order, change the bits.

Live demo you can paste in Python
>>> import numpy as np
>>> a = np.array([1e20, 1.0, -1e20], dtype=np.float32)
>>> a.sum()           # left-to-right reduction
0.0
>>> a[::-1].sum()      # reversed order
1.0

Where this bites in transformers

The argmax cliff

Individually these are tiny — ulp-level differences in logits. But the argmax operation is discontinuous: if the difference straddles a tie-breaking boundary, the output token flips. Two "logit-identical" runs on different hardware can produce different text. You'll see it most on the 3rd or 4th token, then it cascades.

04

Batch-Size-Dependent Outputs

The single most common surprise: run a prompt alone → token X. Run it in a batch of 8 → token Y. Same weights, same prompt, same seed, different tokens.

Why

One sequence, three different batch neighbours each step step t mine neighbour neighbour step t+1 mine (new arrival) (finished) step t+2 mine neighbour neighbour Effective batch shape changes each step → kernels pick different tiles → reductions re-order → logits differ at ulp.
05

Scheduler-Induced Drift

vLLM, TGI, SGLang, TRT-LLM all have iteration-level schedulers. Three sources of drift in there:

Preemption

When KV pool fills up, vLLM preempts the lowest-priority sequence and recomputes its prefix later. Recomputation can land at a different batch size → different logits.

Prefix caching

If someone else already populated the pages for your system prompt, you reuse those K/V tensors. The bits depend on the batch size that originally filled them. So re-run order matters.

Chunked prefill interleaving

Chunked prefill reorders prompt tiles. For the big matmul, reordering rows matters to bitwise output but not to semantics.

LoRA switching

When multi-LoRA is enabled, the scheduler may group same-adapter requests together. The batch composition therefore depends on which adapters arrive, not just which sequences.

06

Quantisation & Hardware Differences

Precision

FP16, BF16, FP8, INT4 each round to different grids. Two deployments identical except precision will disagree on ~1-3% of tokens, nearly always on ambiguous next-token choices.

KV cache precision

--kv-cache-dtype fp8 rounds every cached K/V entry into FP8. Long after the prompt was processed, those tokens look different in attention. Quality impact is tiny, determinism impact vs an FP16 KV run is real.

Kernel variant

Ampere vs Ada vs Hopper vs Blackwell pick different CUTLASS schedules for the same matmul. The same library at the same version will not produce the same bytes across architectures.

TP world size

TP=2 and TP=4 reduce over different groups of heads. All-reduce sums happen in different tree orders. Determinism is preserved within a fixed TP size, not across.

07

Library & Kernel Version Drift

Pinning the model is the obvious part. Easy to forget: kernels ship with libraries, and libraries change.

LayerWhat can change under you
NVIDIA driverPTX JIT version; affects compiled kernels
CUDA toolkitcuBLAS / cuDNN pick different algorithms per version
cuBLASLt auto-heuristicsPicks an algorithm at run-time based on probe timings
Flash-Attentionv2 vs v3 vs Hopper-tuned pick different tile sizes
vLLM / SGLang / TRT-LLMVersion bumps add/remove fused kernels
NCCLAlgorithm choice (ring / tree) can flip with topology cache
TokeniserDifferent tokeniser versions can produce different prompt IDs. Lock this especially.
The one pin nobody sets

CUBLASLT_WORKSPACE_CONFIG and CUBLAS_WORKSPACE_CONFIG=:4096:8 force deterministic algorithms in cuBLAS. Without that, cuBLAS is free to pick a faster non-deterministic kernel and sometimes does.

08

Interactive: What Would Differ?

Flip the bits and see which determinism guarantees survive.

09

Mitigations — Practical Recipes

For evals & CI

  • temperature=0, top_p=1, top_k=-1, fixed seed
  • Turn off speculative decoding
  • Turn off prefix caching (--no-enable-prefix-caching)
  • Set --max-num-seqs 1 for eval, or pad to a fixed batch size
  • Pin image digest (vllm/vllm-openai@sha256:…), not a tag
  • Pin NVIDIA driver, CUDA, vLLM, flash-attn versions
  • Run on the same GPU arch (don't mix 4090 and H100 for the same eval)

For production (just to reduce surprise)

  • Log seed, temperature, model revision, and server build with every request
  • Lock model digest (HF revision SHA, not tag)
  • Lock tokeniser; use the same tokenizer.json everywhere
  • Avoid upgrading vLLM mid-sprint; tie upgrades to eval re-runs
  • Don't compare outputs across different TP sizes; use a per-TP baseline
Deterministic-enough vLLM config for evals
docker run --rm \
    --gpus all --ipc=host \
    -e CUBLAS_WORKSPACE_CONFIG=:4096:8 \
    -e VLLM_USE_V1=0 \
    vllm/vllm-openai:v0.7.0 \
    --model <pinned-revision> \
    --dtype bfloat16 \
    --max-num-seqs 1 \
    --no-enable-prefix-caching \
    --disable-log-requests
Honest bottom line

Bitwise reproducibility across hardware and versions is not a thing you get for free. What you can get, reliably, is reproducibility within a pinned deployment: same image digest, same GPU arch, same flags, same batch policy, fixed seed, no speculative decoding, no prefix cache. That's enough for evals and CI. Anything else requires hardware-level reproducibility guarantees most kernels don't provide.