Simulation Engineering Toolkit — Presentation 02

Rust ↔ Python: Porting a Simulator Core with PyO3

Moving the core of a SimPy simulator into Rust without changing a single answer: when a rewrite pays and when it does not, PyO3 and maturin, what should cross the boundary, releasing the GIL, everything it took to be bit-exact (operation order, SimPy's event order, fsum, Python's random stream, a JSON-parsing trap), differential and golden tests, and measured speed-ups.

PyO3 maturin GIL Bit-exact Differential tests Speed-ups
Profile → Port → Bind → Diff-test → Measure
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

01

Should You Port at All?

A rewrite is the most expensive way to make a simulator faster. LLM Inference Simulators 08 puts it last for a reason: profile first, then try fewer events, cheaper events and parallel runs. Disaggregated_Inference_Sim's exact fast path already halves the run time in pure Python. Port when what remains is still too slow and the model has stopped changing every week.

Port the core when…Do not port yet when…
Profiling shows the event loop and cost model dominateTime goes on I/O, analysis or plotting
The model is validated and its interface is stableThe physics is still being argued about; every change would be made twice
Thousands of runs a night are needed (sweeps, search, CI)A few runs a day, and a coffee is a fine wait
A Python reference exists to test the port againstThere is nothing to check the port's answers with
Someone on the team will maintain the RustThe port would be one person's island
What this deck ports

The engine, roofline cost model (with DVFS and power caps), workload generator and metrics of Disaggregated_Inference_Sim, into Rust_DES_Kernel. Python stays in charge of configuration, the reference model and analysis. The requirement: identical answers, to the last bit.

02

The Shape of a Mixed Rust and Python Project

One repository, one wheel, two languages your notebook / sweep scriptrust_des.simulate(SimConfig(...), rows) python/rust_des/__init__.pythin wrapper: SimConfig to JSON, results to a dataclass src/python.rs -> rust_des._native (PyO3)converts plain data; releases the GIL; maps errors src/kernel.rs, src/disagg/*.rspure Rust: no Python types; cargo test and the CLI use it too Disaggregated_Inference_Simthe SimPy reference model(a dev dependency only) pytests/, tests/fixtures/differential tests (live) andgolden files (recorded)

The crucial line is the bottom box: the Rust core never sees a Python type. That keeps it testable with cargo test alone, usable from a Rust CLI, and free to run with the GIL released. maturin's "mixed" layout puts the Python wrapper next to the compiled module:

pyproject.toml: one setting tells maturin where the Python half lives source
[tool.maturin]
features = ["python"]
python-source = "python"
module-name = "rust_des._native"
03

PyO3 and maturin in Practice

PyO3 generates the CPython glue from attributes on ordinary Rust functions; maturin builds and installs the result as a wheel. Argument and return types are converted automatically: a Python list of tuples arrives as Vec<(f64, i64, i64)>.

src/python.rs: a #[pyfunction] with keyword defaults source
/// `poisson_workload` with Python's own random stream: the same rows as
/// `disagg_sim.workload.poisson_workload` for the same arguments.
#[pyfunction]
#[pyo3(signature = (rate, n, prompt_mean, prompt_cv, output_mean, output_cv, seed=0))]
fn poisson_workload(
    rate: f64,
    n: usize,
    prompt_mean: f64,
    prompt_cv: f64,
    output_mean: f64,
    output_cv: f64,
    seed: u64,
) -> Vec<Row> {
    disagg::poisson_workload(
        rate,
        n,
        LengthDist::new(prompt_mean, prompt_cv),
        LengthDist::new(output_mean, output_cv),
        seed,
    )
    .into_iter()
    .map(|r| (r.arrival, r.prompt_len, r.output_len))
    .collect()
}
src/python.rs: the module source
/// Rust core of rust_des: a bit-exact port of Disaggregated_Inference_Sim's engine.
#[pymodule]
fn _native(m: &Bound<'_, PyModule>) -> PyResult<()> {
    m.add_function(wrap_pyfunction!(simulate_json, m)?)?;
    m.add_function(wrap_pyfunction!(simulate_many_json, m)?)?;
    m.add_function(wrap_pyfunction!(poisson_workload, m)?)?;
    m.add("__version__", env!("CARGO_PKG_VERSION"))?;
    Ok(())
}
The development loop
pip install maturin
maturin develop --release -E dev    # compile, install into the venv, plus test extras
pytest pytests                      # Python-side tests import rust_des like any package
maturin build --release             # a wheel for distribution
Versions move: check the guide

This code uses PyO3 0.29. Names changed recently: Python::allow_threads became Python::detach and Python::with_gil became Python::attach in 0.26 (changelog). Older tutorials use the old names.

04

What Crosses the Boundary

Every value that crosses from Python to Rust is converted, and every call has a fixed cost. So the boundary should be coarse (one call per simulation, never one per event) and plain (numbers, tuples, strings; no Python objects held by Rust).

In

The configuration as a JSON string (names of presets, numbers), and the workload as a list of (arrival, prompt_len, output_len) tuples.

Out

One tuple of six timestamps per request, the summary as JSON, the event count, and the time spent in Rust.

Never

Callbacks into Python during a run, Python objects inside the engine, or per-event logging across the boundary.

The wrapper times each part of a call, so the cost of the boundary is measured, not guessed:

WorkloadRust simulateRust summariseBoundary (conversion)Python summarise()
1P1D, 1,000 requests at 4 req/s4.0 ms8.5 ms0.7 ms47 ms
Colocated x2, 1,000 requests at 4 req/s4.4 ms9.0 ms0.6 ms46 ms
2P2D, 10,000 requests at 8 req/s40.9 ms108.4 ms4.6 ms593 ms

Source: examples/results.md in Rust_DES_Kernel

Two lessons. The conversion itself is small (under a millisecond for 1,000 requests). But the first version left the summary in Python, and Python's summarise() takes far longer than the simulation it summarises: about 50 ms against 4 ms (about 100 ms before it learned to sort once; deck 11). Port what profiling says is slow, and that includes the code around the engine. The summary is the biggest Rust cost too, because it sorts and sums every inter-token latency: one per output token.

05

Releasing the GIL

CPython's global interpreter lock (GIL) lets one thread run Python bytecode at a time. Rust code that touches no Python objects does not need it, and can say so. Then other Python threads keep running, including other calls into the same simulator.

src/python.rs: py.detach releases the GIL for the whole run source
/// Run one simulation. Returns (per-request stamps, summary JSON, events processed,
/// seconds simulating, seconds summarising).
#[pyfunction]
#[pyo3(signature = (config_json, rows, with_summary=true))]
fn simulate_json(
    py: Python<'_>,
    config_json: &str,
    rows: Vec<Row>,
    with_summary: bool,
) -> PyResult<Output> {
    let spec = spec(config_json)?;
    // The GIL is released for the whole run: nothing in here touches Python.
    py.detach(|| run_one(&spec, &rows, with_summary))
        .map_err(PyValueError::new_err)
}
WhatRunsWall timeSpeed-up over sequential
Rust via PyO3, sequential8 x 4,000 requests0.430 s1.0x
Rust via PyO3, 8 Python threads (GIL released)8 x 4,000 requests0.144 s3.0x
Rust simulate_many (rayon, one call)8 x 4,000 requests0.131 s3.3x
Python SimPy, sequential4 x 1,000 requests1.11 s1.0x
Python SimPy, 4 threads (GIL held)4 x 1,000 requests2.40 s0.5x
Python SimPy, 4 processes4 x 1,000 requests0.30 s3.7x

Source: examples/results.md in Rust_DES_Kernel

SimPy holds the GIL, so four threads are slower than one (they contend for the lock); only processes help it. The Rust rows reach about 3× on this 4-core, 8-thread machine. On the free-threaded build of CPython there is no GIL to release; PyO3 0.28 and later declare modules safe for it by default (gil_used = true opts out), but an abi3 wheel does not cover that build, which has its own ABI (PyO3 guide).

06

Bit-Exact: The Parity Checklist

"Agrees to within a tolerance" hides bugs: a wrong tie-break or an off-by-one in batching changes results by less than the tolerance on most workloads. Identical answers rule out everything at once, and they are achievable, because IEEE 754 arithmetic is deterministic. What makes two programs differ is doing different operations. Everything that had to match:

ItemPython doesSo the Rust does
Operation ordera * b * c is (a * b) * cThe same expression, never "simplified"
Exact integersFLOP counts are unbounded ints until they meet a floati128, converted at the same point (round to nearest even, as Python does)
Event order at equal timesSimPy: priority, then scheduling orderThe kernel's (time, priority, sequence) rule (deck 01, slide 03)
Timeoutsenv.timeout(max(0, a - now)) fires at now + (a - now)The same sum, which is not always exactly a
Trigger onceA triggered event is not triggered againThe wake_triggered flag (deck 01, slide 07)
Meansstatistics.fmean sums with math.fsumA port of CPython's fsum
Rounding and divisionround() is half-to-even; float // is fmod-basedround_ties_even; a port of CPython's floor division
Random numbersMT19937 with Python's seeding and variatesA port (slide 07)
Transcendentalslog, exp, pow from the C libraryRust's ln, exp, powf call the same libm on Linux; tested equal there
07

Reproducing Python's Random Numbers

A port that cannot generate the same workload can only be tested on workloads passed in from Python. Reproducing random.Random removes that limit, and it is less work than it sounds: CPython uses the Mersenne Twister (MT19937), seeds it with init_by_array, and builds the variates in pure Python.

src/pyrand.rs: random() takes 27 + 26 bits from two 32-bit outputs source
/// `random()`: a float in [0, 1) with 53 random bits.
pub fn random(&mut self) -> f64 {
    let a = (self.next_u32() >> 5) as f64;
    let b = (self.next_u32() >> 6) as f64;
    (a * 67_108_864.0 + b) * (1.0 / 9_007_199_254_740_992.0)
}
src/pyrand.rs: Kinderman and Monahan, line for line from random.py source
/// `normalvariate(mu, sigma)`: Kinderman and Monahan's ratio-of-uniforms method.
pub fn normalvariate(&mut self, mu: f64, sigma: f64) -> f64 {
    let nv_magicconst = 4.0 * (-0.5f64).exp() / 2.0f64.sqrt();
    loop {
        let u1 = self.random();
        let u2 = 1.0 - self.random();
        let z = nv_magicconst * (u1 - 0.5) / u2;
        let zz = z * z / 4.0;
        if zz <= -u2.ln() {
            return mu + z * sigma;
        }
    }
}
08

Three Traps That Cost One ulp

Each of these produced, or would have produced, a difference in the last bit of some result. Each was found by a test, not by reading the code.

1. fmean is not a loop

statistics.fmean calls math.fsum, which is correctly rounded. A left-to-right loop over ten 0.1s gives 0.9999999999999999; fsum gives 1.0. Python 3.12's built-in sum() is compensated too. The port carries CPython's fsum algorithm, now pinned by 804 CPython-computed cases.

2. JSON parsing

serde_json's default float parser is fast but not always correctly rounded. The golden test (Python results stored as JSON) failed with timestamps one ulp apart, because stored arrival times had been read back one ulp off. The float_roundtrip feature fixes it.

3. Port the expression, not the intent

Python's power-cap solver writes a cube root as abs(x) ** (1 / 3). Rust's f64::cbrt is a different function. The port calls powf(1.0 / 3.0), the same libm pow.

snippets/t02/cbrt_vs_pow: how often the two disagree source
fn main() {
    let (mut differ, mut total) = (0u32, 0u32);
    let mut x = 1e-3_f64;
    while x < 1e6 {
        total += 1;
        if x.cbrt() != x.powf(1.0 / 3.0) {
            differ += 1;
        }
        x *= 1.000123; // a geometric sweep over nine decades
    }
    println!("{differ} of {total} values differ");
    println!(
        "0.001: cbrt {:?}, powf(1/3) {:?}",
        0.001f64.cbrt(),
        0.001f64.powf(1.0 / 3.0)
    );
}
Recorded output (snippets/RESULTS.md)
100783 of 168493 values differ
0.001: cbrt 0.1, powf(1/3) 0.10000000000000002
A fourth trap, in the tests

A differential test that reused a workload through dataclasses.replace failed on the colocated case. The port was right: replace makes a shallow copy, so the two Python runs shared each request's list of inter-token latencies. Build a fresh workload for every run (deck 06, slide 02).

09

Testing the Port: Golden and Differential

Two complementary checks, both demanding identical timestamps and an identical summary dictionary:

Golden files (tests/golden.rs)

Fourteen Python runs stored as JSON, replayed by cargo test, plus four Python-generated workloads the Rust generator must reproduce. No Python needed, so they run in every Rust build. They cover every hardware preset, link and power-cap mode, fixed-length requests (simultaneous events), bursty arrivals, two KV transfers landing together on an idle decoder, and a tight KV budget with rejections. Several of these were added because mutation testing showed no Rust test reached those paths (deck 06, slide 09)

Live differential (pytests/)

The Python simulator in the loop: ten named configurations, ties, bursts, rejections and single-token outputs, plus 25 Hypothesis-generated configurations per run (mode, pool sizes, link, channels, rate, seed, power cap, DVFS).

pytests/test_differential.py: "identical" means every stamp and the whole summary source
def assert_identical(py, rs):
    for j, name in enumerate(STAMPS):
        a = [getattr(r, name) for r in py.requests]
        b = [s[j] for s in rs.stamps]
        if a != b:
            k = next(i for i, (x, y) in enumerate(zip(a, b)) if x != y)
            pytest.fail(f"{name} differs first at request {k}: python {a[k]!r} rust {b[k]!r}")
    assert normalise(summarise(py)) == rs.summary
ConfigurationRequestsTimestamps comparedDifferingSummary dict identical
1P1D, InfiniBand NDR1,0006,0000yes
2P2D, 25 GbE, 2 channels1,0006,0000yes
Colocated, 2 instances1,0004,0000yes
1P1D, 350 W cap + DVFS1,0006,0000yes
Colocated, 300 W cap1,0004,0000yes
Fixed lengths (simultaneous events)1,0006,0000yes

Source: examples/results.md in Rust_DES_Kernel

When a difference appears, the message names the first request and stamp that differ, which usually points straight at the event-order rule involved.

10

Measured: Speed-Ups

WorkloadPython SimPyPython fast pathRust coreRust via PyO3 (total)Core speed-up vs SimPyvs fast pathRust events/sSimPy events/s
1P1D, 1,000 requests at 4 req/s0.285 s0.141 s4.0 ms13.2 ms72x36x5.17 M86 k
Colocated x2, 1,000 requests at 4 req/s0.324 sn/a4.4 ms13.9 ms74xn/a6.22 M85 k
2P2D, 10,000 requests at 8 req/s2.901 s1.460 s40.9 ms153.8 ms71x36x5.06 M85 k

Source: examples/results.md in Rust_DES_Kernel

11

Interactive: What Should You Port?

Amdahl's law with measured parts. Choose a workload and decide where each stage runs. Times come from examples/results.md; the boundary cost applies whenever Rust is involved. LLM Inference Simulators 08, slide 14 has the general version.

12

Packaging and CI

.github/workflows/ci.yml: the Python half of CI source
python:
  runs-on: ubuntu-latest
  steps:
    - uses: actions/checkout@v4
    - uses: dtolnay/rust-toolchain@stable
      with:
        components: clippy
    - uses: Swatinem/rust-cache@v2
    - uses: actions/setup-python@v5
      with:
        python-version: "3.12"
    - name: Clippy (PyO3 bindings)
      run: cargo clippy --features python -- -D warnings
    - name: Build the extension and install the Python reference simulator
      run: |
        python -m venv .venv
        .venv/bin/pip install maturin
        .venv/bin/maturin develop --release -E dev
    - name: Differential tests against Disaggregated_Inference_Sim
      run: .venv/bin/pytest pytests
Exactness is a platform property

Bit-exactness through ln, exp and pow holds when both languages use the same C library, which is the case on Linux with glibc (tested locally and in GitHub Actions). On another platform the Hypothesis-generated power-capped cases fall back to a relative tolerance of 10−12; everything else stays exact.

13

What to Take Away