Simulation Engineering Toolkit — Presentation 06

Testing Frameworks for Simulators

A test strategy for a simulator and the frameworks that implement it: pytest in depth (fixtures, parametrize, markers, conftest, plugins, xdist), Hypothesis including stateful tests, golden tests and re-blessing, mutation testing, coverage, GoogleTest, cargo test and proptest, cocotb, and putting it all in CI. Every example runs.

pytest Hypothesis Golden tests Mutation testing GoogleTest proptest cocotb
Unit → Property → Golden → Differential → Mutation → CI
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

01

A Test Strategy for a Simulator

A simulator's output is a number nobody can check by eye, so its tests carry more weight than usual. LLM Inference Simulators 06 introduced the verification ladder; this deck is about the frameworks that implement each rung, and about testing the tests. Every example here runs: the outputs quoted are recorded in snippets/RESULTS.md or in a code repository's examples/results.md.

Each rung catches what the others miss; each has a tool tool Unit: a component against a hand calculationpytest, cargo test, GoogleTest Analytic: the engine against queueing theorypytest / cargo test Property: invariants over generated inputsHypothesis, proptest Stateful: sequences of operations vs a modelHypothesis stateful Golden: today's answers vs blessed onesJSON + --bless Differential: two implementations agreepytest + PyO3, cocotb Mutation: do the tests notice a bug?mutmut, cargo-mutants Above them all: correlation with hardware, a validation question (LLM Inference Simulators 06, slide 08).
The question to ask of any suite

Not "what fraction of lines run?" but "if I put a plausible bug in this model, would a test fail?". Slides 07 to 09 measure exactly that.

02

pytest I: Fixtures, conftest and Scopes

A fixture is a function that provides something a test needs; a test asks for it by naming it as a parameter. Fixtures in conftest.py are visible to every test in that directory, with no import.

snippets/t06/pytest_demo/conftest.py source
@pytest.fixture(scope="session")
def baseline_run():
    """One 300-request run, built once and shared by every test that asks for it."""
    wl = poisson_workload(4.0, 300, LengthDist(2048, 0.5), LengthDist(128, 0.5), seed=1)
    return simulate(SimConfig(), wl)


@pytest.fixture
def small_workload():
    """A fresh workload per test: simulate() mutates requests, so never share one."""
    return poisson_workload(2.0, 50, LengthDist(512, 0.3), LengthDist(32, 0.3), seed=7)
ScopeBuiltUse for
function (default)Once per testAnything a test might change: workloads, temporary files
module / classOnce per file or classA model shared by one group of tests
sessionOnce per runAn expensive reference run that tests only read
03

pytest II: Parametrize, Markers, Plugins and xdist

snippets/t06/pytest_demo/test_sim.py: one body, three configurations; a slow marker; a strict xfail source
@pytest.mark.parametrize("cfg", [
    SimConfig(),
    SimConfig(mode="colocated", n_colocated=2),
    SimConfig(n_prefill=2, n_decode=2),
], ids=["1P1D", "colocated", "2P2D"])
def test_kv_is_released(cfg, small_workload):
    res = simulate(cfg, small_workload)
    assert all(i.kv_used == 0 for i in res.instances)


@pytest.mark.slow
def test_long_run_is_stable():
    from disagg_sim.workload import LengthDist, poisson_workload
    wl = poisson_workload(3.0, 3000, LengthDist(2048, 0.5), LengthDist(128, 0.5), seed=2)
    m = summarise(simulate(SimConfig(fast_forward=True), wl))
    assert m["throughput"]["slo_attainment"] > 0.9


@pytest.mark.xfail(strict=True, reason="known optimism: tensor-parallel all-reduce is not modelled")
pyproject.toml: register markers; quick suite by default source
[tool.pytest.ini_options]
markers = ["slow: long runs, excluded from the quick suite (run with -m slow)"]
addopts = "-ra -q -m 'not slow'"
xfail_strict = true
pytest --collect-only -q (recorded)
test_sim.py::test_every_request_finishes
test_sim.py::test_stages_sum_to_end_to_end
test_sim.py::test_kv_is_released[1P1D]
test_sim.py::test_kv_is_released[colocated]
test_sim.py::test_kv_is_released[2P2D]
test_sim.py::test_long_run_is_stable
test_sim.py::test_eight_devices_cost_more_than_four_in_communication
test_sim.py::test_summary_matches_golden
RunTestsWall time (including start-up)Speed-up
pytest327.4 s1.0x
pytest -n auto (8 workers)323.5 s2.1x

Source: snippets/RESULTS.md in _simeng_build

The speed-up is well short of the worker count: start-up costs, and a suite dominated by a few long tests, cap it. xdist pays off most on large suites of independent tests; tests must not share files or global state to run this way.

04

Hypothesis: Properties and Shrinking

A property-based test states something that must hold for every input and lets Hypothesis search for a counterexample. When it finds one, it shrinks it: it tries simpler inputs until it has the simplest failing case it can find. Here is a deliberately false claim about the simulator:

snippets/t06/shrinking/test_shrinking.py source
@settings(max_examples=40, deadline=None, derandomize=True)
@given(rate=st.floats(0.5, 20.0))
def test_ttft_p99_under_one_second_at_any_rate(rate):
    wl = poisson_workload(rate, 200, LengthDist(2048, 0.5), LengthDist(64, 0.5), seed=0)
    m = summarise(simulate(SimConfig(fast_forward=True), wl))
    assert m["latency_s"]["ttft"]["p99"] < 1.0
PropertyResultCounterexample after shrinking
p99 TTFT < 1 s at any rate in [0.5, 20]failsrate=5.0 (p99 TTFT 1.088 s)

Source: snippets/RESULTS.md in _simeng_build

Writing good properties for a simulator

  • Conservation: every request finishes or is rejected; tokens out equal tokens requested.
  • Ordering: timestamps never go backwards; stages sum to end-to-end latency.
  • Limits: utilisation ≤ 1; power ≤ TDP; latency ≥ the analytic minimum.
  • Agreement: two implementations, or a fast path and a slow one, give the same answer.

Settings worth knowing

  • max_examples: more search, longer runs; raise it in a nightly job.
  • deadline=None for simulations, which are legitimately slow.
  • derandomize=True for reproducible CI; or keep the example database, which replays past failures first.
  • @example(...) pins a past failure as a permanent regression test.
05

Hypothesis: Stateful Testing

Some bugs need a sequence of operations to appear. A stateful test describes the operations as rules; Hypothesis generates random programs from them and, after every step, compares the system with a model simple enough to be obviously right. For an event queue, the model is a list and min():

snippets/t06/stateful/test_event_queue_stateful.py source
class EventQueueMachine(RuleBasedStateMachine):
    def __init__(self):
        super().__init__()
        self.q = EventQueue(fifo_ties=FIFO)
        self.model = []          # (time, priority, insertion order, name): sorted() is the spec
        self.n = 0

    @rule(dt=st.integers(0, 3), urgent=st.booleans())
    def schedule(self, dt, urgent):
        prio = URGENT if urgent else NORMAL
        name = f"e{self.n}"
        self.q.schedule(self.q.now + dt, name, prio)
        self.model.append((self.q.now + dt, prio, self.n, name))
        self.n += 1

    @precondition(lambda self: self.model)
    @rule()
    def pop(self):
        want = min(self.model)
        self.model.remove(want)
        assert self.q.pop() == (want[0], want[3])

    @invariant()
    def sizes_agree(self):
        assert len(self.q) == len(self.model)


TestEventQueue = EventQueueMachine.TestCase

The queue under test can be switched to a buggy tie-break (newest first among equal times). Hypothesis finds a failing program and shrinks it to this (recorded):

QUEUE_BUG=lifo pytest (recorded in snippets/RESULTS.md)
state = EventQueueMachine()
state.schedule(dt=0, urgent=False)
state.schedule(dt=0, urgent=False)
state.pop()
AssertionError: assert (0.0, 'e1') == (0.0, 'e0')

Read it as a minimal reproduction: two events at the same time, popped in the wrong order. Shrinking is a heuristic, so the reported program is short but not always the shortest possible. The same idea tests a cache, a scheduler or a DRAM bank state machine against a reference model.

06

Golden Tests and Re-Blessing

A golden (approval) test stores the output of a reference run and fails when today's output differs. It catches silent drift: the change nobody meant to make to a number nobody was watching.

snippets/t06/pytest_demo/conftest.py: a golden fixture with a --bless option source
@pytest.fixture
def small_workload():
    """A fresh workload per test: simulate() mutates requests, so never share one."""
    return poisson_workload(2.0, 50, LengthDist(512, 0.3), LengthDist(32, 0.3), seed=7)


@pytest.fixture
def golden(request, pytestconfig):
    """Compare a result with tests/golden/<test name>.json, or rewrite it with --bless."""
    path = request.path.parent / "golden" / f"{request.node.name}.json"

    def check(result: dict):
        if pytestconfig.getoption("bless") or not path.exists():
            path.parent.mkdir(exist_ok=True)
            path.write_text(json.dumps(result, indent=1, sort_keys=True))
            pytest.skip(f"blessed {path.name}")
        assert result == json.loads(path.read_text())

    return check

Rules that keep golden files honest

  • Re-bless deliberately. pytest --bless rewrites the files; the diff of the golden file goes in the same commit as the model change, and a reviewer reads it.
  • Store what matters, not everything: summary metrics and a sample of timestamps, so a diff is readable.
  • Decide the tolerance per quantity. Exact for a deterministic simulator; a tolerance only where the platform's maths library genuinely differs.

Golden files across languages

Rust_DES_Kernel's tests/golden.rs replays fourteen Python runs stored as JSON and demands identical timestamps and summaries, so the Rust build checks parity without Python installed. pytests/make_golden.py is its --bless: it regenerates the file from the reference simulator.

tests/golden.rs in Rust_DES_Kernel source
#[test]
fn matches_the_python_simulator_exactly() {
    let cases: Vec<Value> = serde_json::from_str(include_str!("fixtures/golden.json")).unwrap();
    assert!(cases.len() >= 5);
    for case in &cases {
        let name = case["name"].as_str().unwrap();
        let spec: ConfigSpec = serde_json::from_value(case["config"].clone()).unwrap();
        let rows = case["rows"].as_array().unwrap();
        let wl = rows
            .iter()
            .enumerate()
            .map(|(i, r)| {
                Request::new(
                    i,
                    r[0].as_f64().unwrap(),
                    r[1].as_i64().unwrap(),
                    r[2].as_i64().unwrap(),
                )
            })
            .collect();
        let res = simulate(spec.build().unwrap(), wl).unwrap();
        for (r, want) in res.requests.iter().zip(case["stamps"].as_array().unwrap()) {
            let got = [
                r.prefill_start,
                r.first_token,
                r.kv_start,
                r.kv_ready,
                r.decode_start,
                r.finish,
            ];
            let want: Vec<Option<f64>> =
                want.as_array().unwrap().iter().map(Value::as_f64).collect();
            assert_eq!(got.to_vec(), want, "{name}: request {}", r.rid);
        }
        assert_eq!(summarise(&res), case["summary"], "{name}: summary differs");
    }
}
07

Mutation Testing

Mutation testing measures the tests rather than the code. A tool makes small deliberate bugs (mutants): >= becomes >, + becomes -, a return value is replaced. It runs the suite against each. A failing suite kills the mutant; a passing one lets it survive, and each survivor is a bug the tests would not notice. Tools: mutmut for Python, cargo-mutants for Rust.

A toy, run for real: the roofline step model, tested by three suites of increasing strength.

snippets/t06/mutmut_*/src/roofline/__init__.py source
def step_time(flops, nbytes, peak_flops, bandwidth, overhead=0.5e-3):
    """Seconds for one step, and which roof bounds it."""
    tc = flops / peak_flops
    tm = nbytes / bandwidth
    if tc >= tm:
        return tc + overhead, "compute"
    return tm + overhead, "memory"
Test suiteTestsLine + branch coverageMutants killedSurvivedMutation score
weak175%5742%
medium2100%8467%
strong5100%120100%

Source: snippets/RESULTS.md in _simeng_build

The four mutants that survive the "medium" suite, which has 100% line and branch coverage
-def step_time(flops, nbytes, peak_flops, bandwidth, overhead=0.5e-3):
+def step_time(flops, nbytes, peak_flops, bandwidth, overhead=1.0005):

-    if tc >= tm:
+    if tc > tm:

-        return tc + overhead, "compute"
+        return tc - overhead, "compute"

-    return tm + overhead, "memory"
+    return tm - overhead, "memory"

The medium suite runs every line and both branches, yet it only checks which roof bounds the step, never how long the step takes, so four bugs in the timing survive. Applied to Rust_DES_Kernel, cargo-mutants found the same pattern in places that matter (slide 09).

08

Interactive: Mutation Score versus Coverage

The same function and the same twelve mutants that mutmut generated (slide 07). Tick tests to build a suite; the page runs every mutant against it, live, and reports coverage alongside the mutation score. The presets reproduce the three recorded suites.

Things to try: the weak suite covers 5 of 6 lines and kills 5 mutants; add the two strong timing tests and watch the score jump while coverage barely moves; find the single test that kills mutant #6 (the tie).

09

Coverage, and What It Does Not Tell You

Coverage measures what the tests execute: lines, branches, or regions (cargo llvm-cov's unit). It is cheap, and a low number is a real warning. A high number is not proof, because executing a line is not checking its result. Rust_DES_Kernel's final numbers:

FileLinesRegionsFunctions
src/bin/disagg-rs.rs89.8%69.1%69.2%
src/disagg/engine.rs97.9%96.7%92.3%
src/disagg/hardware.rs99.2%99.4%100.0%
src/disagg/metrics.rs97.6%97.0%97.3%
src/disagg/workload.rs100.0%100.0%100.0%
src/kernel.rs97.2%98.3%95.0%
src/pymath.rs97.6%97.1%100.0%
src/pyrand.rs100.0%100.0%100.0%
src/queueing.rs100.0%100.0%100.0%
Total97.6%94.9%94.5%

Source: examples/results.md in Rust_DES_Kernel

Getting there took three rounds of cargo-mutants, each time writing tests for the survivors:

RoundWhat the survivors showedTests added
1 (four files)fsum's final rounding correction and floor division's sign fix-ups ran under test but no result was checked: 42 survivors in pymath.rs alone1,416 cases computed by CPython itself; the kernel's equality test; a tight-KV golden case with rejections
2 (whole crate)The A100, optical and 8B presets, three links and several power-cap branches were only exercised by the Python differential suite; the Rust workload generator was never checked from Rust; the M/D/1 test used a service time of 1 s, so * and / by it were indistinguishableGolden cases for every preset, link and power-cap mode; four Python-generated workloads; M/D/1 at 2.5 s
3A mutant that made the M/D/1 theory negative passed: the test divided by the theory, and a negative denominator makes every relative error look smallAssert the theory is positive
FileMutantsCaughtMissedTimeoutUnviableScore (caught / (caught + missed))
src/bin/disagg-rs.rs43001100%
src/disagg/engine.rs183161146292%
src/disagg/hardware.rs3132651303595%
src/disagg/metrics.rs868330097%
src/disagg/workload.rs5049001100%
src/kernel.rs2623003100%
src/pymath.rs9988101090%
src/pyrand.rs129106118499%
src/queueing.rs383530092%
Total92881344254695%

Source: examples/results.md in Rust_DES_Kernel

The score rose from 89% after the first full run to 95%; the per-round tables are in examples/results.md. The final run was repeated on 2026-10-03, after the cost-model correction (deck 10): the new code added mutants, every one was caught, and the same 44 survive. Most of them are worth reading, not fixing:

10

GoogleTest for C++ Models

When a model's core is C++ (SystemC, gem5-style simulators; deck 03 of this series, planned), GoogleTest is the usual framework. The same event-queue contract as deck 01 and slide 05, in C++:

snippets/t06/gtest/test_event_queue.cpp: fixtures, ASSERT vs EXPECT, value-parameterised tests source
// A fixture: SetUp() runs before every TEST_F that names it.
class TiedQueue : public ::testing::Test {
protected:
    void SetUp() override {
        q.schedule(1.0, "n1");
        q.schedule(1.0, "n2");
        q.schedule(1.0, "urgent", Priority::Urgent);
    }
    EventQueue q;
};

TEST_F(TiedQueue, UrgentBeatsNormalAtTheSameTime) {
    EXPECT_EQ(q.pop().name, "urgent");
}

TEST_F(TiedQueue, EqualKeysAreFirstInFirstOut) {
    q.pop();
    ASSERT_EQ(q.size(), 2u);              // ASSERT stops this test if it fails; EXPECT carries on
    EXPECT_EQ(q.pop().name, "n1");
    EXPECT_EQ(q.pop().name, "n2");
}

// A value-parameterised test: one body, many inputs.
class ClockNeverGoesBack : public ::testing::TestWithParam<std::tuple<double, double, double>> {};

TEST_P(ClockNeverGoesBack, AcrossInsertionOrders) {
    auto [a, b, c] = GetParam();
    EventQueue q;
    for (double t : {a, b, c}) q.schedule(t, "e");
    double last = -1.0;
    while (q.size()) {
        double t = q.pop().time;
        EXPECT_GE(t, last);
        last = t;
    }
}

INSTANTIATE_TEST_SUITE_P(Orders, ClockNeverGoesBack,
death tests check that an invariant aborts source
// A death test: the assertion must abort the process (debug builds only).
TEST(EventQueueDeathTest, PopFromEmptyAborts) {
    EventQueue q;
    EXPECT_DEATH(q.pop(), "empty queue");
}
CMakeLists.txt: fetch GoogleTest; build with sanitizers source
include(FetchContent)
FetchContent_Declare(googletest
  URL https://github.com/google/googletest/archive/refs/tags/v1.18.0.tar.gz
  DOWNLOAD_EXTRACT_TIMESTAMP TRUE)
FetchContent_MakeAvailable(googletest)

add_executable(test_event_queue test_event_queue.cpp)
target_compile_options(test_event_queue PRIVATE -Wall -Wextra -fsanitize=address,undefined -fno-omit-frame-pointer)
target_link_options(test_event_queue PRIVATE -fsanitize=address,undefined)
target_link_libraries(test_event_queue GTest::gtest_main)

include(GoogleTest)
gtest_discover_tests(test_event_queue)

Recorded result: [==========] 9 tests from 4 test suites ran. (1155 ms total) [ PASSED ] 9 tests. Building tests with AddressSanitizer and UndefinedBehaviorSanitizer turns memory errors and undefined behaviour into test failures, which is where most C++ model bugs hide. gtest_discover_tests registers each test with CTest, so ctest --output-junit feeds CI.

11

cargo test and proptest

Rust's test harness is built in: #[test] functions in a module (they see private items) or in tests/ (they see the public API, as a user would). proptest brings Hypothesis-style generation and shrinking. From Rust_DES_Kernel:

tests/props.rs: a strategy that generates whole configurations source
fn spec() -> impl Strategy<Value = ConfigSpec> {
    (
        prop::bool::ANY,
        1usize..4,
        1usize..4,
        prop::sample::select(vec!["ib-ndr", "eth-25g", "nvlink4", "pcie5", "eth-100g"]),
        1usize..4,
        prop::option::of(200.0f64..700.0),
        prop::bool::ANY,
        prop::sample::select(vec![8usize, 64, 256]),
    )
        .prop_map(|(colo, a, b, link, ch, cap, dvfs, maxb)| ConfigSpec {
            mode: if colo { "colocated" } else { "disagg" }.into(),
            n_prefill: a,
            n_decode: b,
            n_colocated: a,
            link: link.into(),
            link_channels: ch,
            power_cap_w: cap,
            dvfs,
            max_decode_batch: maxb,
            ..Default::default()
        })
}
tests/props.rs: the properties source
// Every request either finishes or is rejected, never both.
let finished = res.requests.iter().filter(|r| r.finish.is_some()).count();
prop_assert_eq!(finished + res.rejected.len(), res.requests.len());

for r in res.requests.iter().filter(|r| r.finish.is_some()) {
    // Timestamps never go backwards along a request's path.
    let path: Vec<f64> = [Some(r.arrival), r.prefill_start, r.first_token, r.kv_start,
                          r.kv_ready, r.decode_start, r.finish].into_iter().flatten().collect();
    prop_assert!(path.windows(2).all(|w| w[0] <= w[1]), "{:?}", path);
    // One inter-token latency per token after the first, all positive.
    prop_assert_eq!(r.itls.len() as i64, (r.output_len - 1).max(0));
    prop_assert!(r.itls.iter().all(|&x| x > 0.0));
    // Stages partition end-to-end latency.
    let sum: f64 = r.stages().iter().sum();
    prop_assert!((sum - r.e2e()).abs() <= 1e-9 * r.e2e().max(1.0));
Test targetTestsWhat it checks
unittests src/lib.rs22Unit tests inside the modules
tests/cli.rs5The disagg-rs binary end to end
tests/golden.rs2Recorded Python runs replayed bit for bit
tests/kernel.rs2M/D/1 against theory; ordering property
tests/props.rs2proptest invariants over random configurations
tests/pymath.rs2fsum and floor division against CPython
pytests/ (pytest + Hypothesis)28Differential tests against the live Python simulator
Total63

Source: examples/results.md in Rust_DES_Kernel

12

cocotb: Python Testbenches for RTL

cocotb drives an HDL simulator from Python coroutines, so the simulator's Python model can serve as the RTL's scoreboard. This is the bridge between a performance model and pre-tape-out verification (deck 05 of this series, planned, builds it out for an NTT pipeline). A modular adder, the smallest piece of an NTT butterfly:

snippets/t06/cocotb/test_mod_add.py: golden model, constrained random stimulus, functional coverage source
def golden(a: int, b: int) -> int:
    """The reference model: what the hardware must compute."""
    return (a + b) % Q


def stimulus(rng: random.Random, n: int):
    """Constrained random: mostly uniform, plus the corners a uniform draw rarely hits."""
    corners = [(0, 0), (Q - 1, 0), (Q - 1, 1), (Q - 1, Q - 1), (Q // 2, Q - Q // 2)]
    yield from corners
    for _ in range(n):
        yield rng.randrange(Q), rng.randrange(Q)


@cocotb.test()
async def matches_golden_model_with_full_coverage(dut):
    cocotb.start_soon(Clock(dut.clk, 10, unit="ns").start())
    dut.rst_n.value, dut.in_valid.value = 0, 0
    await RisingEdge(dut.clk)
    dut.rst_n.value = 1

    bins = {"no wrap": 0, "wrap": 0, "sum == Q (result 0)": 0, "max operands": 0}
    for a, b in stimulus(random.Random(2026), 2000):
        await FallingEdge(dut.clk)                      # drive away from the active edge
        dut.a.value, dut.b.value, dut.in_valid.value = a, b, 1
        await RisingEdge(dut.clk)
        await ReadOnly()                                # sample after the register updates
        assert dut.out_valid.value == 1
        assert int(dut.r.value) == golden(a, b), f"{a} + {b}: got {int(dut.r.value)}"
        bins["wrap" if a + b >= Q else "no wrap"] += 1
        bins["sum == Q (result 0)"] += a + b == Q
        bins["max operands"] += a == b == Q - 1

    dut._log.info("coverage: %s", bins)
    assert all(bins.values()), f"coverage hole: {bins}"
13

Running It All in CI

Tests only protect a model if they run on every change. The frameworks above all speak the same two formats, which is what makes a mixed-language project manageable in one pipeline:

FrameworkJUnit XMLCoverage
pytest--junitxml=report.xmlpytest-cov: --cov-report=xml (Cobertura)
cargocargo nextest with a [profile.ci.junit] sectioncargo llvm-cov --cobertura
GoogleTest--gtest_output=xml, or ctest --output-junitgcov / llvm-cov, then gcovr
cocotbresults.xml, written by every runThe simulator's own coverage (Verilator, Questa)
Jenkinsfile in Rust_DES_Kernel: tests and coverage stages source
stage('Rust tests') {
    steps {
        sh(env.WITH_CARGO + 'cargo nextest run --release --profile ci')
    }
}

stage('Python differential tests') {
    steps {
        sh(env.WITH_CARGO + '''
            python3 -m venv .venv
            .venv/bin/pip install -q maturin
            .venv/bin/maturin develop --release -q -E dev
            # The venv outlives builds, and pip keeps an installed git dependency whose version
            # number has not changed, so fetch the Python reference's current commit every time.
            .venv/bin/pip install -q --force-reinstall --no-deps "disagg-sim @ git+https://github.com/BrendanJamesLynskey/Disaggregated_Inference_Sim"
            .venv/bin/pytest pytests --junitxml=pytest-junit.xml
        ''')
    }
}

stage('Coverage') {
    steps {
        sh(env.WITH_CARGO + 'cargo llvm-cov --release --cobertura --output-path coverage.xml')
        // cargo-llvm-cov lists each generic instantiation as its own method, and the
        // Coverage plugin's Cobertura parser rejects duplicate names: skip them.
        recordCoverage(tools: [[parser: 'COBERTURA', pattern: 'coverage.xml']],
14

What to Take Away