A test strategy for a simulator and the frameworks that implement it: pytest in depth (fixtures, parametrize, markers, conftest, plugins, xdist), Hypothesis including stateful tests, golden tests and re-blessing, mutation testing, coverage, GoogleTest, cargo test and proptest, cocotb, and putting it all in CI. Every example runs.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.
A simulator's output is a number nobody can check by eye, so its tests carry more weight than usual. LLM Inference Simulators 06 introduced the verification ladder; this deck is about the frameworks that implement each rung, and about testing the tests. Every example here runs: the outputs quoted are recorded in snippets/RESULTS.md or in a code repository's examples/results.md.
Not "what fraction of lines run?" but "if I put a plausible bug in this model, would a test fail?". Slides 07 to 09 measure exactly that.
A fixture is a function that provides something a test needs; a test asks for it by naming it as a parameter. Fixtures in conftest.py are visible to every test in that directory, with no import.
@pytest.fixture(scope="session")
def baseline_run():
"""One 300-request run, built once and shared by every test that asks for it."""
wl = poisson_workload(4.0, 300, LengthDist(2048, 0.5), LengthDist(128, 0.5), seed=1)
return simulate(SimConfig(), wl)
@pytest.fixture
def small_workload():
"""A fresh workload per test: simulate() mutates requests, so never share one."""
return poisson_workload(2.0, 50, LengthDist(512, 0.3), LengthDist(32, 0.3), seed=7)
| Scope | Built | Use for |
|---|---|---|
function (default) | Once per test | Anything a test might change: workloads, temporary files |
module / class | Once per file or class | A model shared by one group of tests |
session | Once per run | An expensive reference run that tests only read |
yield; code after the yield runs when the scope ends, even if the test failed.tmp_path (a fresh directory), monkeypatch (patch attributes and environment variables), capsys (captured output), request (which test is asking; used by the golden fixture on slide 06).simulate() stamps timestamps onto the request objects it is given. A shared workload, or a shallow copy of one (dataclasses.replace still shares each request's list of inter-token latencies), carries one run's results into the next. This bit the Rust_DES_Kernel differential suite (deck 02, slide 08); build a fresh one per test.@pytest.mark.parametrize("cfg", [
SimConfig(),
SimConfig(mode="colocated", n_colocated=2),
SimConfig(n_prefill=2, n_decode=2),
], ids=["1P1D", "colocated", "2P2D"])
def test_kv_is_released(cfg, small_workload):
res = simulate(cfg, small_workload)
assert all(i.kv_used == 0 for i in res.instances)
@pytest.mark.slow
def test_long_run_is_stable():
from disagg_sim.workload import LengthDist, poisson_workload
wl = poisson_workload(3.0, 3000, LengthDist(2048, 0.5), LengthDist(128, 0.5), seed=2)
m = summarise(simulate(SimConfig(fast_forward=True), wl))
assert m["throughput"]["slo_attainment"] > 0.9
@pytest.mark.xfail(strict=True, reason="known optimism: tensor-parallel all-reduce is not modelled")
[tool.pytest.ini_options]
markers = ["slow: long runs, excluded from the quick suite (run with -m slow)"]
addopts = "-ra -q -m 'not slow'"
xfail_strict = true
test_sim.py::test_every_request_finishes
test_sim.py::test_stages_sum_to_end_to_end
test_sim.py::test_kv_is_released[1P1D]
test_sim.py::test_kv_is_released[colocated]
test_sim.py::test_kv_is_released[2P2D]
test_sim.py::test_long_run_is_stable
test_sim.py::test_eight_devices_cost_more_than_four_in_communication
test_sim.py::test_summary_matches_goldenids say which case failed.slow, gpu, nightly); -m selects them. Register markers so a typo is an error, not a silently unselected test.pytest-cov (coverage), pytest-xdist (-n auto: tests in parallel processes), Hypothesis's own plugin, --junitxml for CI. Measured on Disaggregated_Inference_Sim's suite:| Run | Tests | Wall time (including start-up) | Speed-up |
|---|---|---|---|
pytest | 32 | 7.4 s | 1.0x |
pytest -n auto (8 workers) | 32 | 3.5 s | 2.1x |
Source: snippets/RESULTS.md in _simeng_build
The speed-up is well short of the worker count: start-up costs, and a suite dominated by a few long tests, cap it. xdist pays off most on large suites of independent tests; tests must not share files or global state to run this way.
A property-based test states something that must hold for every input and lets Hypothesis search for a counterexample. When it finds one, it shrinks it: it tries simpler inputs until it has the simplest failing case it can find. Here is a deliberately false claim about the simulator:
@settings(max_examples=40, deadline=None, derandomize=True)
@given(rate=st.floats(0.5, 20.0))
def test_ttft_p99_under_one_second_at_any_rate(rate):
wl = poisson_workload(rate, 200, LengthDist(2048, 0.5), LengthDist(64, 0.5), seed=0)
m = summarise(simulate(SimConfig(fast_forward=True), wl))
assert m["latency_s"]["ttft"]["p99"] < 1.0
| Property | Result | Counterexample after shrinking |
|---|---|---|
| p99 TTFT < 1 s at any rate in [0.5, 20] | fails | rate=5.0 (p99 TTFT 1.088 s) |
Source: snippets/RESULTS.md in _simeng_build
max_examples: more search, longer runs; raise it in a nightly job.deadline=None for simulations, which are legitimately slow.derandomize=True for reproducible CI; or keep the example database, which replays past failures first.@example(...) pins a past failure as a permanent regression test.Some bugs need a sequence of operations to appear. A stateful test describes the operations as rules; Hypothesis generates random programs from them and, after every step, compares the system with a model simple enough to be obviously right. For an event queue, the model is a list and min():
class EventQueueMachine(RuleBasedStateMachine):
def __init__(self):
super().__init__()
self.q = EventQueue(fifo_ties=FIFO)
self.model = [] # (time, priority, insertion order, name): sorted() is the spec
self.n = 0
@rule(dt=st.integers(0, 3), urgent=st.booleans())
def schedule(self, dt, urgent):
prio = URGENT if urgent else NORMAL
name = f"e{self.n}"
self.q.schedule(self.q.now + dt, name, prio)
self.model.append((self.q.now + dt, prio, self.n, name))
self.n += 1
@precondition(lambda self: self.model)
@rule()
def pop(self):
want = min(self.model)
self.model.remove(want)
assert self.q.pop() == (want[0], want[3])
@invariant()
def sizes_agree(self):
assert len(self.q) == len(self.model)
TestEventQueue = EventQueueMachine.TestCase
The queue under test can be switched to a buggy tie-break (newest first among equal times). Hypothesis finds a failing program and shrinks it to this (recorded):
state = EventQueueMachine()
state.schedule(dt=0, urgent=False)
state.schedule(dt=0, urgent=False)
state.pop()
AssertionError: assert (0.0, 'e1') == (0.0, 'e0')Read it as a minimal reproduction: two events at the same time, popped in the wrong order. Shrinking is a heuristic, so the reported program is short but not always the shortest possible. The same idea tests a cache, a scheduler or a DRAM bank state machine against a reference model.
A golden (approval) test stores the output of a reference run and fails when today's output differs. It catches silent drift: the change nobody meant to make to a number nobody was watching.
@pytest.fixture
def small_workload():
"""A fresh workload per test: simulate() mutates requests, so never share one."""
return poisson_workload(2.0, 50, LengthDist(512, 0.3), LengthDist(32, 0.3), seed=7)
@pytest.fixture
def golden(request, pytestconfig):
"""Compare a result with tests/golden/<test name>.json, or rewrite it with --bless."""
path = request.path.parent / "golden" / f"{request.node.name}.json"
def check(result: dict):
if pytestconfig.getoption("bless") or not path.exists():
path.parent.mkdir(exist_ok=True)
path.write_text(json.dumps(result, indent=1, sort_keys=True))
pytest.skip(f"blessed {path.name}")
assert result == json.loads(path.read_text())
return check
pytest --bless rewrites the files; the diff of the golden file goes in the same commit as the model change, and a reviewer reads it.Rust_DES_Kernel's tests/golden.rs replays fourteen Python runs stored as JSON and demands identical timestamps and summaries, so the Rust build checks parity without Python installed. pytests/make_golden.py is its --bless: it regenerates the file from the reference simulator.
#[test]
fn matches_the_python_simulator_exactly() {
let cases: Vec<Value> = serde_json::from_str(include_str!("fixtures/golden.json")).unwrap();
assert!(cases.len() >= 5);
for case in &cases {
let name = case["name"].as_str().unwrap();
let spec: ConfigSpec = serde_json::from_value(case["config"].clone()).unwrap();
let rows = case["rows"].as_array().unwrap();
let wl = rows
.iter()
.enumerate()
.map(|(i, r)| {
Request::new(
i,
r[0].as_f64().unwrap(),
r[1].as_i64().unwrap(),
r[2].as_i64().unwrap(),
)
})
.collect();
let res = simulate(spec.build().unwrap(), wl).unwrap();
for (r, want) in res.requests.iter().zip(case["stamps"].as_array().unwrap()) {
let got = [
r.prefill_start,
r.first_token,
r.kv_start,
r.kv_ready,
r.decode_start,
r.finish,
];
let want: Vec<Option<f64>> =
want.as_array().unwrap().iter().map(Value::as_f64).collect();
assert_eq!(got.to_vec(), want, "{name}: request {}", r.rid);
}
assert_eq!(summarise(&res), case["summary"], "{name}: summary differs");
}
}
Mutation testing measures the tests rather than the code. A tool makes small deliberate bugs (mutants): >= becomes >, + becomes -, a return value is replaced. It runs the suite against each. A failing suite kills the mutant; a passing one lets it survive, and each survivor is a bug the tests would not notice. Tools: mutmut for Python, cargo-mutants for Rust.
A toy, run for real: the roofline step model, tested by three suites of increasing strength.
def step_time(flops, nbytes, peak_flops, bandwidth, overhead=0.5e-3):
"""Seconds for one step, and which roof bounds it."""
tc = flops / peak_flops
tm = nbytes / bandwidth
if tc >= tm:
return tc + overhead, "compute"
return tm + overhead, "memory"
| Test suite | Tests | Line + branch coverage | Mutants killed | Survived | Mutation score |
|---|---|---|---|---|---|
| weak | 1 | 75% | 5 | 7 | 42% |
| medium | 2 | 100% | 8 | 4 | 67% |
| strong | 5 | 100% | 12 | 0 | 100% |
Source: snippets/RESULTS.md in _simeng_build
-def step_time(flops, nbytes, peak_flops, bandwidth, overhead=0.5e-3):
+def step_time(flops, nbytes, peak_flops, bandwidth, overhead=1.0005):
- if tc >= tm:
+ if tc > tm:
- return tc + overhead, "compute"
+ return tc - overhead, "compute"
- return tm + overhead, "memory"
+ return tm - overhead, "memory"The medium suite runs every line and both branches, yet it only checks which roof bounds the step, never how long the step takes, so four bugs in the timing survive. Applied to Rust_DES_Kernel, cargo-mutants found the same pattern in places that matter (slide 09).
The same function and the same twelve mutants that mutmut generated (slide 07). Tick tests to build a suite; the page runs every mutant against it, live, and reports coverage alongside the mutation score. The presets reproduce the three recorded suites.
Things to try: the weak suite covers 5 of 6 lines and kills 5 mutants; add the two strong timing tests and watch the score jump while coverage barely moves; find the single test that kills mutant #6 (the tie).
Coverage measures what the tests execute: lines, branches, or regions (cargo llvm-cov's unit). It is cheap, and a low number is a real warning. A high number is not proof, because executing a line is not checking its result. Rust_DES_Kernel's final numbers:
| File | Lines | Regions | Functions |
|---|---|---|---|
src/bin/disagg-rs.rs | 89.8% | 69.1% | 69.2% |
src/disagg/engine.rs | 97.9% | 96.7% | 92.3% |
src/disagg/hardware.rs | 99.2% | 99.4% | 100.0% |
src/disagg/metrics.rs | 97.6% | 97.0% | 97.3% |
src/disagg/workload.rs | 100.0% | 100.0% | 100.0% |
src/kernel.rs | 97.2% | 98.3% | 95.0% |
src/pymath.rs | 97.6% | 97.1% | 100.0% |
src/pyrand.rs | 100.0% | 100.0% | 100.0% |
src/queueing.rs | 100.0% | 100.0% | 100.0% |
| Total | 97.6% | 94.9% | 94.5% |
Source: examples/results.md in Rust_DES_Kernel
Getting there took three rounds of cargo-mutants, each time writing tests for the survivors:
| Round | What the survivors showed | Tests added |
|---|---|---|
| 1 (four files) | fsum's final rounding correction and floor division's sign fix-ups ran under test but no result was checked: 42 survivors in pymath.rs alone | 1,416 cases computed by CPython itself; the kernel's equality test; a tight-KV golden case with rejections |
| 2 (whole crate) | The A100, optical and 8B presets, three links and several power-cap branches were only exercised by the Python differential suite; the Rust workload generator was never checked from Rust; the M/D/1 test used a service time of 1 s, so * and / by it were indistinguishable | Golden cases for every preset, link and power-cap mode; four Python-generated workloads; M/D/1 at 2.5 s |
| 3 | A mutant that made the M/D/1 theory negative passed: the test divided by the theory, and a negative denominator makes every relative error look small | Assert the theory is positive |
| File | Mutants | Caught | Missed | Timeout | Unviable | Score (caught / (caught + missed)) |
|---|---|---|---|---|---|---|
src/bin/disagg-rs.rs | 4 | 3 | 0 | 0 | 1 | 100% |
src/disagg/engine.rs | 183 | 161 | 14 | 6 | 2 | 92% |
src/disagg/hardware.rs | 313 | 265 | 13 | 0 | 35 | 95% |
src/disagg/metrics.rs | 86 | 83 | 3 | 0 | 0 | 97% |
src/disagg/workload.rs | 50 | 49 | 0 | 0 | 1 | 100% |
src/kernel.rs | 26 | 23 | 0 | 0 | 3 | 100% |
src/pymath.rs | 99 | 88 | 10 | 1 | 0 | 90% |
src/pyrand.rs | 129 | 106 | 1 | 18 | 4 | 99% |
src/queueing.rs | 38 | 35 | 3 | 0 | 0 | 92% |
| Total | 928 | 813 | 44 | 25 | 46 | 95% |
Source: examples/results.md in Rust_DES_Kernel
The score rose from 89% after the first full run to 95%; the per-round tables are in examples/results.md. The final run was repeated on 2026-10-03, after the cost-model correction (deck 10): the new code added mutants, every one was caught, and the same 44 survive. Most of them are worth reading, not fixing:
| becomes ^ on two bit masks that never overlap; a / b becomes a * b where only the sign is used; the explicit stop event is deleted, but the event list drains at the same instant anyway. No test can kill these.> to >=) on float comparisons that never tie exactly in practice. Killing them means constructing exact ties; worth it only where a tie is plausible.When a model's core is C++ (SystemC, gem5-style simulators; deck 03 of this series, planned), GoogleTest is the usual framework. The same event-queue contract as deck 01 and slide 05, in C++:
// A fixture: SetUp() runs before every TEST_F that names it.
class TiedQueue : public ::testing::Test {
protected:
void SetUp() override {
q.schedule(1.0, "n1");
q.schedule(1.0, "n2");
q.schedule(1.0, "urgent", Priority::Urgent);
}
EventQueue q;
};
TEST_F(TiedQueue, UrgentBeatsNormalAtTheSameTime) {
EXPECT_EQ(q.pop().name, "urgent");
}
TEST_F(TiedQueue, EqualKeysAreFirstInFirstOut) {
q.pop();
ASSERT_EQ(q.size(), 2u); // ASSERT stops this test if it fails; EXPECT carries on
EXPECT_EQ(q.pop().name, "n1");
EXPECT_EQ(q.pop().name, "n2");
}
// A value-parameterised test: one body, many inputs.
class ClockNeverGoesBack : public ::testing::TestWithParam<std::tuple<double, double, double>> {};
TEST_P(ClockNeverGoesBack, AcrossInsertionOrders) {
auto [a, b, c] = GetParam();
EventQueue q;
for (double t : {a, b, c}) q.schedule(t, "e");
double last = -1.0;
while (q.size()) {
double t = q.pop().time;
EXPECT_GE(t, last);
last = t;
}
}
INSTANTIATE_TEST_SUITE_P(Orders, ClockNeverGoesBack,
// A death test: the assertion must abort the process (debug builds only).
TEST(EventQueueDeathTest, PopFromEmptyAborts) {
EventQueue q;
EXPECT_DEATH(q.pop(), "empty queue");
}
include(FetchContent)
FetchContent_Declare(googletest
URL https://github.com/google/googletest/archive/refs/tags/v1.18.0.tar.gz
DOWNLOAD_EXTRACT_TIMESTAMP TRUE)
FetchContent_MakeAvailable(googletest)
add_executable(test_event_queue test_event_queue.cpp)
target_compile_options(test_event_queue PRIVATE -Wall -Wextra -fsanitize=address,undefined -fno-omit-frame-pointer)
target_link_options(test_event_queue PRIVATE -fsanitize=address,undefined)
target_link_libraries(test_event_queue GTest::gtest_main)
include(GoogleTest)
gtest_discover_tests(test_event_queue)
Recorded result: [==========] 9 tests from 4 test suites ran. (1155 ms total) [ PASSED ] 9 tests. Building tests with AddressSanitizer and UndefinedBehaviorSanitizer turns memory errors and undefined behaviour into test failures, which is where most C++ model bugs hide. gtest_discover_tests registers each test with CTest, so ctest --output-junit feeds CI.
Rust's test harness is built in: #[test] functions in a module (they see private items) or in tests/ (they see the public API, as a user would). proptest brings Hypothesis-style generation and shrinking. From Rust_DES_Kernel:
fn spec() -> impl Strategy<Value = ConfigSpec> {
(
prop::bool::ANY,
1usize..4,
1usize..4,
prop::sample::select(vec!["ib-ndr", "eth-25g", "nvlink4", "pcie5", "eth-100g"]),
1usize..4,
prop::option::of(200.0f64..700.0),
prop::bool::ANY,
prop::sample::select(vec![8usize, 64, 256]),
)
.prop_map(|(colo, a, b, link, ch, cap, dvfs, maxb)| ConfigSpec {
mode: if colo { "colocated" } else { "disagg" }.into(),
n_prefill: a,
n_decode: b,
n_colocated: a,
link: link.into(),
link_channels: ch,
power_cap_w: cap,
dvfs,
max_decode_batch: maxb,
..Default::default()
})
}
// Every request either finishes or is rejected, never both.
let finished = res.requests.iter().filter(|r| r.finish.is_some()).count();
prop_assert_eq!(finished + res.rejected.len(), res.requests.len());
for r in res.requests.iter().filter(|r| r.finish.is_some()) {
// Timestamps never go backwards along a request's path.
let path: Vec<f64> = [Some(r.arrival), r.prefill_start, r.first_token, r.kv_start,
r.kv_ready, r.decode_start, r.finish].into_iter().flatten().collect();
prop_assert!(path.windows(2).all(|w| w[0] <= w[1]), "{:?}", path);
// One inter-token latency per token after the first, all positive.
prop_assert_eq!(r.itls.len() as i64, (r.output_len - 1).max(0));
prop_assert!(r.itls.iter().all(|&x| x > 0.0));
// Stages partition end-to-end latency.
let sum: f64 = r.stages().iter().sum();
prop_assert!((sum - r.e2e()).abs() <= 1e-9 * r.e2e().max(1.0));
proptest-regressions/; commit that directory and the case is replayed first, forever.prop_map, prop::option::of, prop::sample::select, ranges of floats and integers.| Test target | Tests | What it checks |
|---|---|---|
unittests src/lib.rs | 22 | Unit tests inside the modules |
tests/cli.rs | 5 | The disagg-rs binary end to end |
tests/golden.rs | 2 | Recorded Python runs replayed bit for bit |
tests/kernel.rs | 2 | M/D/1 against theory; ordering property |
tests/props.rs | 2 | proptest invariants over random configurations |
tests/pymath.rs | 2 | fsum and floor division against CPython |
pytests/ (pytest + Hypothesis) | 28 | Differential tests against the live Python simulator |
| Total | 63 |
Source: examples/results.md in Rust_DES_Kernel
cocotb drives an HDL simulator from Python coroutines, so the simulator's Python model can serve as the RTL's scoreboard. This is the bridge between a performance model and pre-tape-out verification (deck 05 of this series, planned, builds it out for an NTT pipeline). A modular adder, the smallest piece of an NTT butterfly:
def golden(a: int, b: int) -> int:
"""The reference model: what the hardware must compute."""
return (a + b) % Q
def stimulus(rng: random.Random, n: int):
"""Constrained random: mostly uniform, plus the corners a uniform draw rarely hits."""
corners = [(0, 0), (Q - 1, 0), (Q - 1, 1), (Q - 1, Q - 1), (Q // 2, Q - Q // 2)]
yield from corners
for _ in range(n):
yield rng.randrange(Q), rng.randrange(Q)
@cocotb.test()
async def matches_golden_model_with_full_coverage(dut):
cocotb.start_soon(Clock(dut.clk, 10, unit="ns").start())
dut.rst_n.value, dut.in_valid.value = 0, 0
await RisingEdge(dut.clk)
dut.rst_n.value = 1
bins = {"no wrap": 0, "wrap": 0, "sum == Q (result 0)": 0, "max operands": 0}
for a, b in stimulus(random.Random(2026), 2000):
await FallingEdge(dut.clk) # drive away from the active edge
dut.a.value, dut.b.value, dut.in_valid.value = a, b, 1
await RisingEdge(dut.clk)
await ReadOnly() # sample after the register updates
assert dut.out_valid.value == 1
assert int(dut.r.value) == golden(a, b), f"{a} + {b}: got {int(dut.r.value)}"
bins["wrap" if a + b >= Q else "no wrap"] += 1
bins["sum == Q (result 0)"] += a + b == Q
bins["max operands"] += a == b == Q - 1
dut._log.info("coverage: %s", bins)
assert all(bins.values()), f"coverage hole: {bins}"
ReadOnly() following the rising edge, so the check sees the registered result.a + b == Q exactly, so the corners are listed explicitly; the coverage bins prove they were exercised.TESTS=1 PASS=1 FAIL=0 SKIP=0; coverage: {'no wrap': 1013, 'wrap': 992, 'sum == Q (result 0)': 2, 'max operands': 1}>= to >) is caught: Seeded bug (sum >= Q changed to sum > Q): TESTS=1 PASS=0 FAIL=1 SKIP=0, first mismatch 65520 + 1: got 65521Tests only protect a model if they run on every change. The frameworks above all speak the same two formats, which is what makes a mixed-language project manageable in one pipeline:
| Framework | JUnit XML | Coverage |
|---|---|---|
| pytest | --junitxml=report.xml | pytest-cov: --cov-report=xml (Cobertura) |
| cargo | cargo nextest with a [profile.ci.junit] section | cargo llvm-cov --cobertura |
| GoogleTest | --gtest_output=xml, or ctest --output-junit | gcov / llvm-cov, then gcovr |
| cocotb | results.xml, written by every run | The simulator's own coverage (Verilator, Questa) |
stage('Rust tests') {
steps {
sh(env.WITH_CARGO + 'cargo nextest run --release --profile ci')
}
}
stage('Python differential tests') {
steps {
sh(env.WITH_CARGO + '''
python3 -m venv .venv
.venv/bin/pip install -q maturin
.venv/bin/maturin develop --release -q -E dev
# The venv outlives builds, and pip keeps an installed git dependency whose version
# number has not changed, so fetch the Python reference's current commit every time.
.venv/bin/pip install -q --force-reinstall --no-deps "disagg-sim @ git+https://github.com/BrendanJamesLynskey/Disaggregated_Inference_Sim"
.venv/bin/pytest pytests --junitxml=pytest-junit.xml
''')
}
}
stage('Coverage') {
steps {
sh(env.WITH_CARGO + 'cargo llvm-cov --release --cobertura --output-path coverage.xml')
// cargo-llvm-cov lists each generic instantiation as its own method, and the
// Coverage plugin's Cobertura parser rejects duplicate names: skip them.
recordCoverage(tools: [[parser: 'COBERTURA', pattern: 'coverage.xml']],
ci/perf_gate.py).