Moving the core of a SimPy simulator into Rust without changing a single answer: when a rewrite pays and when it does not, PyO3 and maturin, what should cross the boundary, releasing the GIL, everything it took to be bit-exact (operation order, SimPy's event order, fsum, Python's random stream, a JSON-parsing trap), differential and golden tests, and measured speed-ups.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.
A rewrite is the most expensive way to make a simulator faster. LLM Inference Simulators 08 puts it last for a reason: profile first, then try fewer events, cheaper events and parallel runs. Disaggregated_Inference_Sim's exact fast path already halves the run time in pure Python. Port when what remains is still too slow and the model has stopped changing every week.
| Port the core when… | Do not port yet when… |
|---|---|
| Profiling shows the event loop and cost model dominate | Time goes on I/O, analysis or plotting |
| The model is validated and its interface is stable | The physics is still being argued about; every change would be made twice |
| Thousands of runs a night are needed (sweeps, search, CI) | A few runs a day, and a coffee is a fine wait |
| A Python reference exists to test the port against | There is nothing to check the port's answers with |
| Someone on the team will maintain the Rust | The port would be one person's island |
The engine, roofline cost model (with DVFS and power caps), workload generator and metrics of Disaggregated_Inference_Sim, into Rust_DES_Kernel. Python stays in charge of configuration, the reference model and analysis. The requirement: identical answers, to the last bit.
The crucial line is the bottom box: the Rust core never sees a Python type. That keeps it testable with cargo test alone, usable from a Rust CLI, and free to run with the GIL released. maturin's "mixed" layout puts the Python wrapper next to the compiled module:
[tool.maturin]
features = ["python"]
python-source = "python"
module-name = "rust_des._native"
PyO3 generates the CPython glue from attributes on ordinary Rust functions; maturin builds and installs the result as a wheel. Argument and return types are converted automatically: a Python list of tuples arrives as Vec<(f64, i64, i64)>.
/// `poisson_workload` with Python's own random stream: the same rows as
/// `disagg_sim.workload.poisson_workload` for the same arguments.
#[pyfunction]
#[pyo3(signature = (rate, n, prompt_mean, prompt_cv, output_mean, output_cv, seed=0))]
fn poisson_workload(
rate: f64,
n: usize,
prompt_mean: f64,
prompt_cv: f64,
output_mean: f64,
output_cv: f64,
seed: u64,
) -> Vec<Row> {
disagg::poisson_workload(
rate,
n,
LengthDist::new(prompt_mean, prompt_cv),
LengthDist::new(output_mean, output_cv),
seed,
)
.into_iter()
.map(|r| (r.arrival, r.prompt_len, r.output_len))
.collect()
}
/// Rust core of rust_des: a bit-exact port of Disaggregated_Inference_Sim's engine.
#[pymodule]
fn _native(m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_function(wrap_pyfunction!(simulate_json, m)?)?;
m.add_function(wrap_pyfunction!(simulate_many_json, m)?)?;
m.add_function(wrap_pyfunction!(poisson_workload, m)?)?;
m.add("__version__", env!("CARGO_PKG_VERSION"))?;
Ok(())
}
pip install maturin
maturin develop --release -E dev # compile, install into the venv, plus test extras
pytest pytests # Python-side tests import rust_des like any package
maturin build --release # a wheel for distributionThis code uses PyO3 0.29. Names changed recently: Python::allow_threads became Python::detach and Python::with_gil became Python::attach in 0.26 (changelog). Older tutorials use the old names.
Every value that crosses from Python to Rust is converted, and every call has a fixed cost. So the boundary should be coarse (one call per simulation, never one per event) and plain (numbers, tuples, strings; no Python objects held by Rust).
The configuration as a JSON string (names of presets, numbers), and the workload as a list of (arrival, prompt_len, output_len) tuples.
One tuple of six timestamps per request, the summary as JSON, the event count, and the time spent in Rust.
Callbacks into Python during a run, Python objects inside the engine, or per-event logging across the boundary.
The wrapper times each part of a call, so the cost of the boundary is measured, not guessed:
| Workload | Rust simulate | Rust summarise | Boundary (conversion) | Python summarise() |
|---|---|---|---|---|
| 1P1D, 1,000 requests at 4 req/s | 4.0 ms | 8.5 ms | 0.7 ms | 47 ms |
| Colocated x2, 1,000 requests at 4 req/s | 4.4 ms | 9.0 ms | 0.6 ms | 46 ms |
| 2P2D, 10,000 requests at 8 req/s | 40.9 ms | 108.4 ms | 4.6 ms | 593 ms |
Source: examples/results.md in Rust_DES_Kernel
Two lessons. The conversion itself is small (under a millisecond for 1,000 requests). But the first version left the summary in Python, and Python's summarise() takes far longer than the simulation it summarises: about 50 ms against 4 ms (about 100 ms before it learned to sort once; deck 11). Port what profiling says is slow, and that includes the code around the engine. The summary is the biggest Rust cost too, because it sorts and sums every inter-token latency: one per output token.
CPython's global interpreter lock (GIL) lets one thread run Python bytecode at a time. Rust code that touches no Python objects does not need it, and can say so. Then other Python threads keep running, including other calls into the same simulator.
/// Run one simulation. Returns (per-request stamps, summary JSON, events processed,
/// seconds simulating, seconds summarising).
#[pyfunction]
#[pyo3(signature = (config_json, rows, with_summary=true))]
fn simulate_json(
py: Python<'_>,
config_json: &str,
rows: Vec<Row>,
with_summary: bool,
) -> PyResult<Output> {
let spec = spec(config_json)?;
// The GIL is released for the whole run: nothing in here touches Python.
py.detach(|| run_one(&spec, &rows, with_summary))
.map_err(PyValueError::new_err)
}
detach must be Send and cannot use the Python token, so it cannot touch a Python object by mistake. That is why the inputs are converted to plain Rust data first.rust_des.simulate, or one call to simulate_many, which runs a batch on rayon's thread pool.| What | Runs | Wall time | Speed-up over sequential |
|---|---|---|---|
| Rust via PyO3, sequential | 8 x 4,000 requests | 0.430 s | 1.0x |
| Rust via PyO3, 8 Python threads (GIL released) | 8 x 4,000 requests | 0.144 s | 3.0x |
Rust simulate_many (rayon, one call) | 8 x 4,000 requests | 0.131 s | 3.3x |
| Python SimPy, sequential | 4 x 1,000 requests | 1.11 s | 1.0x |
| Python SimPy, 4 threads (GIL held) | 4 x 1,000 requests | 2.40 s | 0.5x |
| Python SimPy, 4 processes | 4 x 1,000 requests | 0.30 s | 3.7x |
Source: examples/results.md in Rust_DES_Kernel
SimPy holds the GIL, so four threads are slower than one (they contend for the lock); only processes help it. The Rust rows reach about 3× on this 4-core, 8-thread machine. On the free-threaded build of CPython there is no GIL to release; PyO3 0.28 and later declare modules safe for it by default (gil_used = true opts out), but an abi3 wheel does not cover that build, which has its own ABI (PyO3 guide).
"Agrees to within a tolerance" hides bugs: a wrong tie-break or an off-by-one in batching changes results by less than the tolerance on most workloads. Identical answers rule out everything at once, and they are achievable, because IEEE 754 arithmetic is deterministic. What makes two programs differ is doing different operations. Everything that had to match:
| Item | Python does | So the Rust does |
|---|---|---|
| Operation order | a * b * c is (a * b) * c | The same expression, never "simplified" |
| Exact integers | FLOP counts are unbounded ints until they meet a float | i128, converted at the same point (round to nearest even, as Python does) |
| Event order at equal times | SimPy: priority, then scheduling order | The kernel's (time, priority, sequence) rule (deck 01, slide 03) |
| Timeouts | env.timeout(max(0, a - now)) fires at now + (a - now) | The same sum, which is not always exactly a |
| Trigger once | A triggered event is not triggered again | The wake_triggered flag (deck 01, slide 07) |
| Means | statistics.fmean sums with math.fsum | A port of CPython's fsum |
| Rounding and division | round() is half-to-even; float // is fmod-based | round_ties_even; a port of CPython's floor division |
| Random numbers | MT19937 with Python's seeding and variates | A port (slide 07) |
| Transcendentals | log, exp, pow from the C library | Rust's ln, exp, powf call the same libm on Linux; tested equal there |
A port that cannot generate the same workload can only be tested on workloads passed in from Python. Reproducing random.Random removes that limit, and it is less work than it sounds: CPython uses the Mersenne Twister (MT19937), seeds it with init_by_array, and builds the variates in pure Python.
/// `random()`: a float in [0, 1) with 53 random bits.
pub fn random(&mut self) -> f64 {
let a = (self.next_u32() >> 5) as f64;
let b = (self.next_u32() >> 6) as f64;
(a * 67_108_864.0 + b) * (1.0 / 9_007_199_254_740_992.0)
}
/// `normalvariate(mu, sigma)`: Kinderman and Monahan's ratio-of-uniforms method.
pub fn normalvariate(&mut self, mu: f64, sigma: f64) -> f64 {
let nv_magicconst = 4.0 * (-0.5f64).exp() / 2.0f64.sqrt();
loop {
let u1 = self.random();
let u2 = 1.0 - self.random();
let z = nv_magicconst * (u1 - 0.5) / u2;
let zz = z * z / 4.0;
if zz <= -u2.ln() {
return mu + z * sigma;
}
}
}
init_by_array, which is why Random(5) and Random(2**40 + 7) both work.poisson_workload for four seeds and 500 requests each (test_workload_generator_reproduces_python_random).round() (half to even, not Rust's default round, which rounds halves away from zero) and Python's clamp order.Each of these produced, or would have produced, a difference in the last bit of some result. Each was found by a test, not by reading the code.
fmean is not a loopstatistics.fmean calls math.fsum, which is correctly rounded. A left-to-right loop over ten 0.1s gives 0.9999999999999999; fsum gives 1.0. Python 3.12's built-in sum() is compensated too. The port carries CPython's fsum algorithm, now pinned by 804 CPython-computed cases.
serde_json's default float parser is fast but not always correctly rounded. The golden test (Python results stored as JSON) failed with timestamps one ulp apart, because stored arrival times had been read back one ulp off. The float_roundtrip feature fixes it.
Python's power-cap solver writes a cube root as abs(x) ** (1 / 3). Rust's f64::cbrt is a different function. The port calls powf(1.0 / 3.0), the same libm pow.
fn main() {
let (mut differ, mut total) = (0u32, 0u32);
let mut x = 1e-3_f64;
while x < 1e6 {
total += 1;
if x.cbrt() != x.powf(1.0 / 3.0) {
differ += 1;
}
x *= 1.000123; // a geometric sweep over nine decades
}
println!("{differ} of {total} values differ");
println!(
"0.001: cbrt {:?}, powf(1/3) {:?}",
0.001f64.cbrt(),
0.001f64.powf(1.0 / 3.0)
);
}
100783 of 168493 values differ
0.001: cbrt 0.1, powf(1/3) 0.10000000000000002A differential test that reused a workload through dataclasses.replace failed on the colocated case. The port was right: replace makes a shallow copy, so the two Python runs shared each request's list of inter-token latencies. Build a fresh workload for every run (deck 06, slide 02).
Two complementary checks, both demanding identical timestamps and an identical summary dictionary:
tests/golden.rs)Fourteen Python runs stored as JSON, replayed by cargo test, plus four Python-generated workloads the Rust generator must reproduce. No Python needed, so they run in every Rust build. They cover every hardware preset, link and power-cap mode, fixed-length requests (simultaneous events), bursty arrivals, two KV transfers landing together on an idle decoder, and a tight KV budget with rejections. Several of these were added because mutation testing showed no Rust test reached those paths (deck 06, slide 09)
pytests/)The Python simulator in the loop: ten named configurations, ties, bursts, rejections and single-token outputs, plus 25 Hypothesis-generated configurations per run (mode, pool sizes, link, channels, rate, seed, power cap, DVFS).
def assert_identical(py, rs):
for j, name in enumerate(STAMPS):
a = [getattr(r, name) for r in py.requests]
b = [s[j] for s in rs.stamps]
if a != b:
k = next(i for i, (x, y) in enumerate(zip(a, b)) if x != y)
pytest.fail(f"{name} differs first at request {k}: python {a[k]!r} rust {b[k]!r}")
assert normalise(summarise(py)) == rs.summary
| Configuration | Requests | Timestamps compared | Differing | Summary dict identical |
|---|---|---|---|---|
| 1P1D, InfiniBand NDR | 1,000 | 6,000 | 0 | yes |
| 2P2D, 25 GbE, 2 channels | 1,000 | 6,000 | 0 | yes |
| Colocated, 2 instances | 1,000 | 4,000 | 0 | yes |
| 1P1D, 350 W cap + DVFS | 1,000 | 6,000 | 0 | yes |
| Colocated, 300 W cap | 1,000 | 4,000 | 0 | yes |
| Fixed lengths (simultaneous events) | 1,000 | 6,000 | 0 | yes |
Source: examples/results.md in Rust_DES_Kernel
When a difference appears, the message names the first request and stamp that differ, which usually points straight at the event-order rule involved.
| Workload | Python SimPy | Python fast path | Rust core | Rust via PyO3 (total) | Core speed-up vs SimPy | vs fast path | Rust events/s | SimPy events/s |
|---|---|---|---|---|---|---|---|---|
| 1P1D, 1,000 requests at 4 req/s | 0.285 s | 0.141 s | 4.0 ms | 13.2 ms | 72x | 36x | 5.17 M | 86 k |
| Colocated x2, 1,000 requests at 4 req/s | 0.324 s | n/a | 4.4 ms | 13.9 ms | 74x | n/a | 6.22 M | 85 k |
| 2P2D, 10,000 requests at 8 req/s | 2.901 s | 1.460 s | 40.9 ms | 153.8 ms | 71x | 36x | 5.06 M | 85 k |
Source: examples/results.md in Rust_DES_Kernel
summarise()).Amdahl's law with measured parts. Choose a workload and decide where each stage runs. Times come from examples/results.md; the boundary cost applies whenever Rust is involved. LLM Inference Simulators 08, slide 14 has the general version.
abi3-py310 feature builds against the stable ABI, so one wheel per platform serves CPython 3.10 and later.python feature that maturin switches on; cargo test and the CLI never link against Python.cargo test (including the golden files) and a CLI smoke test. The other: clippy on the bindings, maturin develop, and the differential tests with the Python reference installed from its repository.python:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
with:
components: clippy
- uses: Swatinem/rust-cache@v2
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Clippy (PyO3 bindings)
run: cargo clippy --features python -- -D warnings
- name: Build the extension and install the Python reference simulator
run: |
python -m venv .venv
.venv/bin/pip install maturin
.venv/bin/maturin develop --release -E dev
- name: Differential tests against Disaggregated_Inference_Sim
run: .venv/bin/pytest pytests
Bit-exactness through ln, exp and pow holds when both languages use the same C library, which is the case on Linux with glibc (tested locally and in GitHub Actions). On another platform the Hypothesis-generated power-capped cases fall back to a relative tolerance of 10−12; everything else stays exact.
py.detach (formerly allow_threads): Python threads then run simulations in parallel.fsum, rounding, the random stream and the exact expressions.examples/results.md.