Why the same prompt gives different tokens twice, where determinism really comes from, and how to get it back when you need reproducible evals.
Most engineers think "temperature=0 means deterministic". It does not. Below we walk through seven independent reasons the same prompt can produce different tokens even at temperature 0.
Temperature-based sampling draws from the softmax distribution. With temperature > 0, different RNG seeds give different tokens. Fixing this is easy if you know where the seed lives:
curl http://host:8000/v1/chat/completions -d '{
"model": "...",
"messages": [...],
"temperature": 0.0,
"seed": 42,
"top_p": 1.0,
"top_k": -1
}'
Even with temperature=0, the argmax can tie — two tokens with the exact same logit. Which one wins depends on the CUDA kernel's tie-break rule and sometimes on the warp scheduling. Rare in practice but it's one of the reasons temperature=0 isn't a strict guarantee.
With speculative decoding enabled, the draft model proposes N tokens and the big model accepts a prefix. When the draft is rejected, the big model's sample replaces it. Different draft rejection counts → different interleavings → different sampler calls. Numbers identical over many runs only with temperature=0 AND a deterministic draft model (beware batch-dependent kernels).
Floating-point addition is not associative: (a + b) + c ≠ a + (b + c) in general. Every reduction the GPU does — softmax, layer-norm, row-sums in matmul — depends on order. Change the order, change the bits.
>>> import numpy as np
>>> a = np.array([1e20, 1.0, -1e20], dtype=np.float32)
>>> a.sum() # left-to-right reduction
0.0
>>> a[::-1].sum() # reversed order
1.0
Individually these are tiny — ulp-level differences in logits. But the argmax operation is discontinuous: if the difference straddles a tie-breaking boundary, the output token flips. Two "logit-identical" runs on different hardware can produce different text. You'll see it most on the 3rd or 4th token, then it cascades.
The single most common surprise: run a prompt alone → token X. Run it in a batch of 8 → token Y. Same weights, same prompt, same seed, different tokens.
vLLM, TGI, SGLang, TRT-LLM all have iteration-level schedulers. Three sources of drift in there:
When KV pool fills up, vLLM preempts the lowest-priority sequence and recomputes its prefix later. Recomputation can land at a different batch size → different logits.
If someone else already populated the pages for your system prompt, you reuse those K/V tensors. The bits depend on the batch size that originally filled them. So re-run order matters.
Chunked prefill reorders prompt tiles. For the big matmul, reordering rows matters to bitwise output but not to semantics.
When multi-LoRA is enabled, the scheduler may group same-adapter requests together. The batch composition therefore depends on which adapters arrive, not just which sequences.
FP16, BF16, FP8, INT4 each round to different grids. Two deployments identical except precision will disagree on ~1-3% of tokens, nearly always on ambiguous next-token choices.
--kv-cache-dtype fp8 rounds every cached K/V entry into FP8. Long after the prompt was processed, those tokens look different in attention. Quality impact is tiny, determinism impact vs an FP16 KV run is real.
Ampere vs Ada vs Hopper vs Blackwell pick different CUTLASS schedules for the same matmul. The same library at the same version will not produce the same bytes across architectures.
TP=2 and TP=4 reduce over different groups of heads. All-reduce sums happen in different tree orders. Determinism is preserved within a fixed TP size, not across.
Pinning the model is the obvious part. Easy to forget: kernels ship with libraries, and libraries change.
| Layer | What can change under you |
|---|---|
| NVIDIA driver | PTX JIT version; affects compiled kernels |
| CUDA toolkit | cuBLAS / cuDNN pick different algorithms per version |
| cuBLASLt auto-heuristics | Picks an algorithm at run-time based on probe timings |
| Flash-Attention | v2 vs v3 vs Hopper-tuned pick different tile sizes |
| vLLM / SGLang / TRT-LLM | Version bumps add/remove fused kernels |
| NCCL | Algorithm choice (ring / tree) can flip with topology cache |
| Tokeniser | Different tokeniser versions can produce different prompt IDs. Lock this especially. |
CUBLASLT_WORKSPACE_CONFIG and CUBLAS_WORKSPACE_CONFIG=:4096:8 force deterministic algorithms in cuBLAS. Without that, cuBLAS is free to pick a faster non-deterministic kernel and sometimes does.
Flip the bits and see which determinism guarantees survive.
temperature=0, top_p=1, top_k=-1, fixed seed--no-enable-prefix-caching)--max-num-seqs 1 for eval, or pad to a fixed batch sizevllm/vllm-openai@sha256:…), not a tagseed, temperature, model revision, and server build with every requesttokenizer.json everywheredocker run --rm \
--gpus all --ipc=host \
-e CUBLAS_WORKSPACE_CONFIG=:4096:8 \
-e VLLM_USE_V1=0 \
vllm/vllm-openai:v0.7.0 \
--model <pinned-revision> \
--dtype bfloat16 \
--max-num-seqs 1 \
--no-enable-prefix-caching \
--disable-log-requests
Bitwise reproducibility across hardware and versions is not a thing you get for free. What you can get, reliably, is reproducibility within a pinned deployment: same image digest, same GPU arch, same flags, same batch policy, fixed seed, no speculative decoding, no prefix cache. That's enough for evals and CI. Anything else requires hardware-level reproducibility guarantees most kernels don't provide.