LLM Inference Simulators — Presentation 06

Metrics, Hot-Spots & Validation

Turning a simulation into evidence: latency distributions, utilisation, MFU/MBU, stage breakdowns and hot-spot attribution; traces in Perfetto; statistics that survive review; and the verification ladder — unit, invariant, analytic, behavioural, correlation — wired into CI.

Percentiles Utilisation Hot-spots Perfetto Little's law CI / Jenkins
Probe → Aggregate → Attribute → Validate → Report
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.

01

From Events to Evidence

A simulator's output is not a number; it is an argument to an architect, a programme manager or a customer that a design will or won't meet its targets. Metrics are the evidence in that argument, so they need the same engineering care as the model. Four families:

Latency

How long things take, as distributions: TTFT, TPOT, ITL, end-to-end; per stage and per request class.

Utilisation

How hard each resource works: busy fraction, MFU, MBU, mean batch, queue lengths, buffer high-water marks.

Attribution

Why: which stage dominates, which resource owns it, and what changes if that resource were faster.

Power & energy

What it costs: average and peak power, joules per token, the static/compute/memory/link split, time spent power-capped (deck 07).

Four kinds of probe

ProbeRecordsCostExample in the companion simulator
Counter / accumulatorSums: busy time, bytes, FLOPsO(1) per eventInstance.busy, .flops, .bytes
Time-weighted integral∫ N(t) dt for mean queue or populationO(1) per changeSimulation._population, used for Little's law
Per-entity timestampsEvery milestone of every requestO(entities)Request.prefill_start, … .finish
Trace / samplerTimeline of steps; periodic snapshotsO(events), can dominateTracer, sampler()
02

Latency: Distributions, Not Averages

metrics.py: one percentile convention, used everywhere
def percentile(xs, p):
    """Linear-interpolated percentile (numpy's default), so results can be cross-checked."""
    s = sorted(xs); k = (len(s) - 1) * p / 100
    lo, hi = math.floor(k), math.ceil(k)
    return s[lo] + (s[hi] - s[lo]) * (k - lo)
03

Utilisation and Efficiency

"Utilisation" hides two different questions: how often is it doing something, and how much of its capability is it using while it does.

Resource (defaults, 4 req/s)BusyMFUMBUMean batchReading
prefill-051%28%3%1.2Compute-bound when busy; idle half the time: headroom for TTFT
decode-0100%4%77%14Always busy (continuous batching) but memory-bound; the "100%" is not saturation
kv-link5%———Plenty of headroom on IB NDR
Why busy fraction misleads

A continuous-batching decode engine is "busy" whenever any sequence is running, even at batch size 1. The useful saturation signals are the decode queue (requests waiting for admission), KV occupancy, and how close the batch is to the bandwidth roofline. This is why the simulator attributes hot-spots by where requests wait, not by busy time.

Operational laws: free consistency checks

04

Hot-Spot Attribution

A hot-spot is the place where making things faster would most improve the metric you care about. Three ways to find it, in increasing order of rigour:

1. Highest utilisation

Cheap, and often wrong: see the decode instance at "100%". Good only for resources with no internal batching.

2. Stage breakdown

Split each request's life into stages that sum exactly to its latency (queue, prefill, KV wait, transfer, decode queue, decode). The largest waiting stage, mapped to the resource that owns it, is the hot-spot. This is what the simulator reports.

3. Sensitivity (what-if)

Speed up one resource by 10% in the model and measure the change in goodput or p99. This is causal profiling (the idea behind the Coz profiler) and only a simulator can do it cheaply. It also catches bottlenecks that shift as soon as you relieve them.

The simulator's report for 2 prefill, 6 req/s, 25 GbE
utilisation  prefill-0 47%  prefill-1 13%  decode-0 100%  kv-link 96%
where time goes  prefill 1%  kv_wait 81%  kv_transfer 1%  decode 18%
hot-spot     stage=kv_wait -> kv-link (busy 96%)   busy>90%: decode-0, kv-link

The stage breakdown carries a tested invariant (stages sum to end-to-end latency for every request), so the attribution cannot silently lose time.

Rerun after the cost-model correction of 2026-10-03 (decode steps had been charged the whole embedding table; deck 05, slide 08): only kv_wait moved, from 80% to 81%, and the utilisation table on the next slide is unchanged.

05

Traces: Seeing the Timeline in Perfetto

Aggregate metrics tell you that something is wrong; a timeline tells you why. The Chrome trace-event JSON format is the lingua franca: Perfetto (ui.perfetto.dev), chrome://tracing and many tools read it, and it is easy to emit.

trace.py output from disagg-sim --trace (first events, rounded)
{"traceEvents": [
  {"ph":"M", "name":"process_name", "pid":2, "args":{"name":"prefill-0"}},
  {"ph":"X", "name":"prefill n=1 tok=1236", "pid":2, "ts":36072.8, "dur":80384.3},
  {"ph":"X", "name":"kv r0 405 MB", "pid":3, "ts":116457.1, "dur":8110.2},
  {"ph":"X", "name":"decode b=1", "pid":4, "ts":124567.3, "dur":13700.6},
  {"ph":"C", "name":"population", "pid":1, "ts":0, "args":{"in_system":0, "link_queue":0}}
], "displayTimeUnit": "ms"}
06

Statistics That Survive Review

A stochastic simulation produces a random answer. Treat each run as one experiment and report uncertainty, or someone in the review will.

TechniqueProblem it solvesHow
Warm-up deletionThe system starts empty, so early requests see no queueDiscard an initial period, chosen by eye with Welch's graphical method or by a rule; the simulator drops the first 10% of requests
Independent replicationsOne run gives no error barRepeat with different seeds; mean ± tn−1 · s/√n
Batch meansReplications each pay warm-upOne long run cut into batches long enough to be roughly independent
Common random numbersComparing two designs with independent noise needs many runsFeed both designs the same arrival stream; the noise largely cancels in the difference
Run to drainLong requests censored at the endStop when all measured requests finish, as the simulator does
Tail percentiles need many samples

A p99 from 800 requests rests on about eight observations beyond it. Expect it to move noticeably from seed to seed. Run more requests or more replications before quoting a p99 to two significant figures.

07

Interactive: Replications, Confidence Intervals and Common Random Numbers

Run the deck 05 simulator n times with different seeds, for disaggregated and colocated, and compare the difference in TTFT p99. With common random numbers both designs see the same request stream in each replication; without them each design gets its own stream.

8
400
4
QuantityMean95% CI half-width

The CRN interval on the difference is usually much narrower than the independent one: the same number of runs gives a sharper answer to "which design is better, and by how much?". Every comparison in the companion simulator (--compare, the deck 05 Compare button) uses common random numbers.

08

The Verification Ladder

"Is the simulator right?" splits into verification (does the code implement the model?) and validation (does the model represent reality?). Each rung catches errors the others miss. The companion simulator's 36 tests are organised this way.

RungChecksExample test
UnitComponents against hand calculationLlama-3-70B has 70.6 B parameters and 320 KiB of KV per token; decode at batch 1 ≈ weights / bandwidth
InvariantProperties true for every runEvery request completes with all its tokens; timestamps ordered; stages sum to E2E; KV released; same seed gives the same answer
AnalyticThe engine against queueing theoryKV link matches M/D/1 (Pollaczek–Khinchine) within 8%; Little's law holds
BehaviouralExpected qualitative effectsDisaggregation cuts p99 ITL more than 4×; a slow link becomes the hot-spot; a second prefill instance cuts TTFT
Property-basedInvariants over generated configurationsHypothesis draws random rates, pool sizes, modes and seeds and checks conservation for each
DifferentialTwo implementations agreeFast path is bit-identical to baseline; the browser JavaScript port is bit-identical to Python
CorrelationModel against measurement(The next step) compare with vLLM on real GPUs, Vidur, or RTL cycle counts; track the error
Validation is a number, not a feeling

For a performance model the deliverable is a correlation report: for a defined set of workloads, model against measurement, error per metric, and a target (say within ±10% on throughput and p50 latency). Track it over time; it is one of the most persuasive engineering metrics a simulation team has.

09

Testing Frameworks

pytest

The default for Python. Fixtures for shared setup, @pytest.mark.parametrize to run the same check across modes and configurations, pytest.approx for floating tolerance, markers to split quick and nightly suites.

Hypothesis (property-based)

Generates hundreds of random inputs, including nasty edge cases, checks invariants, and shrinks any failure to a minimal example. Ideal for simulators, whose correctness is mostly invariants.

Golden / approval tests

Store the metrics of reference runs; fail CI when they change unexpectedly, and require an explicit "re-bless" when a change is intended. This is how you stop silent model drift.

cocotb, UVM, GoogleTest, cargo test

cocotb runs Python testbenches against RTL, so the simulator's models become scoreboards. UVM is the SystemVerilog standard. GoogleTest and Catch2 cover C++ cores; cargo test and proptest cover Rust.

A Hypothesis property test for the simulator
from hypothesis import given, settings, strategies as st

@settings(max_examples=50, deadline=None)
@given(rate=st.floats(0.5, 8), n_prefill=st.integers(1, 3), n_decode=st.integers(1, 3),
       seed=st.integers(0, 10_000), mode=st.sampled_from(["disagg", "colocated"]))
def test_conservation_for_any_config(rate, n_prefill, n_decode, seed, mode):
    cfg = SimConfig(mode=mode, n_prefill=n_prefill, n_decode=n_decode)
    wl = poisson_workload(rate, 100, LengthDist(1024, 1.0), LengthDist(64, 1.0), seed)
    simulate(cfg, wl)
    for r in wl:
        assert r.tokens_out == r.output_len
        assert abs(sum(r.stages().values()) - r.e2e) < 1e-9
10

CI/CD and Engineering Metrics

A simulator used for sign-off needs the same pipeline discipline as production software. Jenkins remains common in hardware companies (on-premises agents, licences, large compute farms); GitHub Actions or GitLab CI are equivalent.

Jenkinsfile (declarative): fast checks and golden metrics on every push, sweeps nightly
pipeline {
  agent { label 'sim' }
  triggers { cron('H 2 * * *') }                       // nightly
  stages {
    stage('Lint & unit')   { steps { sh 'ruff check . && pytest -m "not slow" --junitxml=unit.xml' } }
    stage('Golden metrics') { steps { sh 'pytest tests/golden --junitxml=golden.xml' } }
    stage('Nightly sweep')  {
      when { triggeredBy 'TimerTrigger' }
      steps { sh 'python sweeps/run_all.py --workers 32 --out results/' }
    }
    stage('Report') { steps { sh 'python reports/build.py results/ > report.html' } }
  }
  post { always { junit '*.xml'; archiveArtifacts 'report.html, results/**' } }
}

Engineering metrics worth tracking (and reporting)

Model quality

  • Correlation error against hardware or RTL, per workload
  • Operator and feature coverage (fraction of target models that run end to end)
  • Open model-accuracy issues by severity

Engineering health

  • Test pass rate and code coverage
  • Simulator speed (simulated seconds per wall-clock second) as a tracked number
  • Time from architecture question to answer
  • Milestones in Jira linked to spec sections
11

Specifications, Test Plans and Reports

Three documents turn a simulator from a personal tool into a team asset. Each has a recognisable skeleton.

Simulator specification

  • Purpose and questions in scope
  • Abstraction level per component
  • Modelling assumptions (FCFS or shared links, KV reservation policy)
  • Inputs, configuration schema, outputs and metric definitions
  • Interfaces to frameworks and RTL
  • Known limitations

Test plan

  • Features → tests traceability matrix
  • Levels: unit, invariant, analytic, behavioural, differential, correlation
  • Pass criteria and tolerances
  • Workload set and seeds
  • Regression cadence (per push, nightly, release)
  • Coverage goals

Performance report

  • Question and answer, first
  • Configuration and workload, reproducible (hash, SHA, seeds)
  • Results with confidence intervals
  • Hot-spots and sensitivities
  • Assumptions that could change the answer
  • Recommendation
Writing for two audiences

Lead with a one-paragraph answer a non-expert can act on; put the methods and caveats where the experts will look for them. The same evidence, layered.

12

What to Take Away

Next

Deck 07 adds the other half of the performance story: power and energy, modelled, measured and traded against latency.