Turning a simulation into evidence: latency distributions, utilisation, MFU/MBU, stage breakdowns and hot-spot attribution; traces in Perfetto; statistics that survive review; and the verification ladder — unit, invariant, analytic, behavioural, correlation — wired into CI.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.
A simulator's output is not a number; it is an argument to an architect, a programme manager or a customer that a design will or won't meet its targets. Metrics are the evidence in that argument, so they need the same engineering care as the model. Four families:
How long things take, as distributions: TTFT, TPOT, ITL, end-to-end; per stage and per request class.
How hard each resource works: busy fraction, MFU, MBU, mean batch, queue lengths, buffer high-water marks.
Why: which stage dominates, which resource owns it, and what changes if that resource were faster.
What it costs: average and peak power, joules per token, the static/compute/memory/link split, time spent power-capped (deck 07).
| Probe | Records | Cost | Example in the companion simulator |
|---|---|---|---|
| Counter / accumulator | Sums: busy time, bytes, FLOPs | O(1) per event | Instance.busy, .flops, .bytes |
| Time-weighted integral | ∫ N(t) dt for mean queue or population | O(1) per change | Simulation._population, used for Little's law |
| Per-entity timestamps | Every milestone of every request | O(entities) | Request.prefill_start, … .finish |
| Trace / sampler | Timeline of steps; periodic snapshots | O(events), can dominate | Tracer, sampler() |
def percentile(xs, p):
"""Linear-interpolated percentile (numpy's default), so results can be cross-checked."""
s = sorted(xs); k = (len(s) - 1) * p / 100
lo, hi = math.floor(k), math.ceil(k)
return s[lo] + (s[hi] - s[lo]) * (k - lo)
"Utilisation" hides two different questions: how often is it doing something, and how much of its capability is it using while it does.
| Resource (defaults, 4 req/s) | Busy | MFU | MBU | Mean batch | Reading |
|---|---|---|---|---|---|
| prefill-0 | 51% | 28% | 3% | 1.2 | Compute-bound when busy; idle half the time: headroom for TTFT |
| decode-0 | 100% | 4% | 77% | 14 | Always busy (continuous batching) but memory-bound; the "100%" is not saturation |
| kv-link | 5% | — | — | — | Plenty of headroom on IB NDR |
A continuous-batching decode engine is "busy" whenever any sequence is running, even at batch size 1. The useful saturation signals are the decode queue (requests waiting for admission), KV occupancy, and how close the batch is to the bandwidth roofline. This is why the simulator attributes hot-spots by where requests wait, not by busy time.
A hot-spot is the place where making things faster would most improve the metric you care about. Three ways to find it, in increasing order of rigour:
Cheap, and often wrong: see the decode instance at "100%". Good only for resources with no internal batching.
Split each request's life into stages that sum exactly to its latency (queue, prefill, KV wait, transfer, decode queue, decode). The largest waiting stage, mapped to the resource that owns it, is the hot-spot. This is what the simulator reports.
Speed up one resource by 10% in the model and measure the change in goodput or p99. This is causal profiling (the idea behind the Coz profiler) and only a simulator can do it cheaply. It also catches bottlenecks that shift as soon as you relieve them.
utilisation prefill-0 47% prefill-1 13% decode-0 100% kv-link 96%
where time goes prefill 1% kv_wait 81% kv_transfer 1% decode 18%
hot-spot stage=kv_wait -> kv-link (busy 96%) busy>90%: decode-0, kv-link
The stage breakdown carries a tested invariant (stages sum to end-to-end latency for every request), so the attribution cannot silently lose time.
Rerun after the cost-model correction of 2026-10-03 (decode steps had been charged the whole embedding table; deck 05, slide 08): only kv_wait moved, from 80% to 81%, and the utilisation table on the next slide is unchanged.
Aggregate metrics tell you that something is wrong; a timeline tells you why. The Chrome trace-event JSON format is the lingua franca: Perfetto (ui.perfetto.dev), chrome://tracing and many tools read it, and it is easy to emit.
disagg-sim --trace (first events, rounded){"traceEvents": [
{"ph":"M", "name":"process_name", "pid":2, "args":{"name":"prefill-0"}},
{"ph":"X", "name":"prefill n=1 tok=1236", "pid":2, "ts":36072.8, "dur":80384.3},
{"ph":"X", "name":"kv r0 405 MB", "pid":3, "ts":116457.1, "dur":8110.2},
{"ph":"X", "name":"decode b=1", "pid":4, "ts":124567.3, "dur":13700.6},
{"ph":"C", "name":"population", "pid":1, "ts":0, "args":{"in_system":0, "link_queue":0}}
], "displayTimeUnit": "ms"}
X = complete slice (start + duration, microseconds); C = counter track; M = metadata (names).A stochastic simulation produces a random answer. Treat each run as one experiment and report uncertainty, or someone in the review will.
| Technique | Problem it solves | How |
|---|---|---|
| Warm-up deletion | The system starts empty, so early requests see no queue | Discard an initial period, chosen by eye with Welch's graphical method or by a rule; the simulator drops the first 10% of requests |
| Independent replications | One run gives no error bar | Repeat with different seeds; mean ± tn−1 · s/√n |
| Batch means | Replications each pay warm-up | One long run cut into batches long enough to be roughly independent |
| Common random numbers | Comparing two designs with independent noise needs many runs | Feed both designs the same arrival stream; the noise largely cancels in the difference |
| Run to drain | Long requests censored at the end | Stop when all measured requests finish, as the simulator does |
A p99 from 800 requests rests on about eight observations beyond it. Expect it to move noticeably from seed to seed. Run more requests or more replications before quoting a p99 to two significant figures.
Run the deck 05 simulator n times with different seeds, for disaggregated and colocated, and compare the difference in TTFT p99. With common random numbers both designs see the same request stream in each replication; without them each design gets its own stream.
| Quantity | Mean | 95% CI half-width |
|---|
The CRN interval on the difference is usually much narrower than the independent one: the same number of runs gives a sharper answer to "which design is better, and by how much?". Every comparison in the companion simulator (--compare, the deck 05 Compare button) uses common random numbers.
"Is the simulator right?" splits into verification (does the code implement the model?) and validation (does the model represent reality?). Each rung catches errors the others miss. The companion simulator's 36 tests are organised this way.
| Rung | Checks | Example test |
|---|---|---|
| Unit | Components against hand calculation | Llama-3-70B has 70.6 B parameters and 320 KiB of KV per token; decode at batch 1 ≈ weights / bandwidth |
| Invariant | Properties true for every run | Every request completes with all its tokens; timestamps ordered; stages sum to E2E; KV released; same seed gives the same answer |
| Analytic | The engine against queueing theory | KV link matches M/D/1 (Pollaczek–Khinchine) within 8%; Little's law holds |
| Behavioural | Expected qualitative effects | Disaggregation cuts p99 ITL more than 4×; a slow link becomes the hot-spot; a second prefill instance cuts TTFT |
| Property-based | Invariants over generated configurations | Hypothesis draws random rates, pool sizes, modes and seeds and checks conservation for each |
| Differential | Two implementations agree | Fast path is bit-identical to baseline; the browser JavaScript port is bit-identical to Python |
| Correlation | Model against measurement | (The next step) compare with vLLM on real GPUs, Vidur, or RTL cycle counts; track the error |
For a performance model the deliverable is a correlation report: for a defined set of workloads, model against measurement, error per metric, and a target (say within ±10% on throughput and p50 latency). Track it over time; it is one of the most persuasive engineering metrics a simulation team has.
The default for Python. Fixtures for shared setup, @pytest.mark.parametrize to run the same check across modes and configurations, pytest.approx for floating tolerance, markers to split quick and nightly suites.
Generates hundreds of random inputs, including nasty edge cases, checks invariants, and shrinks any failure to a minimal example. Ideal for simulators, whose correctness is mostly invariants.
Store the metrics of reference runs; fail CI when they change unexpectedly, and require an explicit "re-bless" when a change is intended. This is how you stop silent model drift.
cocotb runs Python testbenches against RTL, so the simulator's models become scoreboards. UVM is the SystemVerilog standard. GoogleTest and Catch2 cover C++ cores; cargo test and proptest cover Rust.
from hypothesis import given, settings, strategies as st
@settings(max_examples=50, deadline=None)
@given(rate=st.floats(0.5, 8), n_prefill=st.integers(1, 3), n_decode=st.integers(1, 3),
seed=st.integers(0, 10_000), mode=st.sampled_from(["disagg", "colocated"]))
def test_conservation_for_any_config(rate, n_prefill, n_decode, seed, mode):
cfg = SimConfig(mode=mode, n_prefill=n_prefill, n_decode=n_decode)
wl = poisson_workload(rate, 100, LengthDist(1024, 1.0), LengthDist(64, 1.0), seed)
simulate(cfg, wl)
for r in wl:
assert r.tokens_out == r.output_len
assert abs(sum(r.stages().values()) - r.e2e) < 1e-9
A simulator used for sign-off needs the same pipeline discipline as production software. Jenkins remains common in hardware companies (on-premises agents, licences, large compute farms); GitHub Actions or GitLab CI are equivalent.
pipeline {
agent { label 'sim' }
triggers { cron('H 2 * * *') } // nightly
stages {
stage('Lint & unit') { steps { sh 'ruff check . && pytest -m "not slow" --junitxml=unit.xml' } }
stage('Golden metrics') { steps { sh 'pytest tests/golden --junitxml=golden.xml' } }
stage('Nightly sweep') {
when { triggeredBy 'TimerTrigger' }
steps { sh 'python sweeps/run_all.py --workers 32 --out results/' }
}
stage('Report') { steps { sh 'python reports/build.py results/ > report.html' } }
}
post { always { junit '*.xml'; archiveArtifacts 'report.html, results/**' } }
}
Three documents turn a simulator from a personal tool into a team asset. Each has a recognisable skeleton.
Lead with a one-paragraph answer a non-expert can act on; put the methods and caveats where the experts will look for them. The same evidence, layered.
Deck 07 adds the other half of the performance story: power and energy, modelled, measured and traded against latency.