Simulation Hub — Simulation Engineering Toolkit

Simulation Engineering Toolkit

The engineering around a simulator: the languages its fast paths are written in, the memory systems and RTL it must agree with, the tests, pipelines and trackers that keep it trustworthy, the specifications it is built to, the front ends that feed it real workloads, an accelerator model built end to end, the tools it is measured with, and the power, performance and area trade-offs it informs. Each deck is backed by a tested code repository, and every number in the decks comes from a recorded run.

RustPyO3C++ / SystemCMemory systemscocotbTesting frameworksJenkinsJiraSpecificationsPerformance analysisMeasurement toolsPPAAccelerator models

Presentations in This Series

  1. Rust for Simulation Engineers →
    Why Rust suits discrete-event simulators, taught through a real kernel: a BinaryHeap event list with SimPy's tie-breaking, events as enums, ownership in an event loop (arenas and indices, not references), SimPy processes as explicit state machines, traits for cost models, error handling, rayon sweeps, and the cargo, clippy, rustfmt, proptest and criterion toolchain.
    liveOwnership · BinaryHeap · Enums & traits · rayon · clippy
  2. Rust and Python: Porting a Simulator Core with PyO3 →
    Moving the core of a SimPy simulator into Rust without changing a single answer: when a rewrite pays and when it does not, PyO3 and maturin, what should cross the boundary, releasing the GIL, everything it took to be bit-exact (operation order, SimPy's event order, fsum, Python's random stream, a JSON-parsing trap), differential and golden tests, and measured speed-ups.
    livePyO3 · maturin · GIL · Bit-exact · Differential tests
  3. C++ Performance Models with SystemC TLM-2.0 →
    An FHE accelerator tile in C++17 and SystemC, driven by the SimPy simulator's own traces: modern C++ for models, the SystemC kernel and delta cycles, TLM-2.0 sockets and payloads, approximately- and loosely-timed coding, temporal decoupling and the quantum, why tie-breaks are part of the model, op-by-op agreement with SimPy, GoogleTest and sanitizers.
    liveSystemC · TLM-2.0 · C++17 · AT and LT · Temporal decoupling
  4. Modelling Memory Systems: DRAM and HBM →
    Why bandwidth times efficiency fails, and a command-level model that derives the efficiency instead: banks, rows and bank groups, the timing parameters, hits, misses and conflicts, FCFS and FR-FCFS, page policy, address mapping, refresh and tFAW, load and latency, validation four ways (including against DRAMsim3), and the model plugged into the FHE simulator.
    liveDRAM timing · HBM · FR-FCFS · Address mapping · Refresh
  5. Verification Bridge: cocotb, Verilator and Golden Models →
    A SystemVerilog NTT core verified in both directions: the FHE simulator's NTT as the golden model, a bit-accurate scoreboard, transactors, constrained-random stimulus, functional coverage and the crosses that make a bug observable, a seeded bug that uniform stimulus misses, Verilator and Icarus in CI, and RTL cycle counts calibrating the simulator.
    livecocotb · Verilator · Golden models · Constrained random · Functional coverage
  6. Testing Frameworks for Simulators →
    A test strategy for a simulator and the frameworks that implement it: pytest in depth (fixtures, parametrize, markers, conftest, plugins, xdist), Hypothesis including stateful tests, golden tests and re-blessing, mutation testing, coverage, GoogleTest, cargo test and proptest, cocotb, and putting it all in CI. Every example runs.
    livepytest · Hypothesis · Golden tests · Mutation testing · GoogleTest
  7. Jenkins for Hardware and Simulation Teams →
    Continuous integration for simulators, from real pipelines that ran: declarative and scripted Jenkinsfiles, agents and labels for tools and licences, matrix builds as parameter sweeps, shared libraries, JUnit and coverage reports, nightly regressions, exact, speed and drift gates, credentials, the declarative linter, and how Jenkins and GitHub Actions divide the work.
    liveJenkinsfile · Matrix builds · Shared libraries · JUnit & coverage · Performance gates
  8. Jira and Engineering Metrics →
    The tracker as a simulation team's data model: issue types, workflows and status categories, versions, sprints and Kanban, JQL with worked queries and its traps, Git and Jenkins integration, traceability from requirement to build, the metrics worth reporting, flow metrics and Little's law from a workflow history, and evidence for sign-off.
    liveIssue model · Workflows · JQL · Smart commits · Flow metrics
  9. Specifications, Requirements and Test Plans →
    Writing down what a simulator must do, and proving it does: requirement quality, EARS patterns with an interactive checker, non-functional requirements for simulators, requirements capture for mission-mode software, the V-model, a traceability matrix generated from the test run (which found a real gap), and worked templates for a specification, a test plan and a performance report.
    liveEARS · Requirement quality · Mission mode · V-model · Traceability
  10. From PyTorch and ONNX to an Accelerator Model →
    A working front end for an accelerator simulator: real model configurations (Llama-3-8B and -70B, Mistral, Qwen, GPT-2) traced without weights on the meta device and under fake tensors, torch.export, a torch.compile backend and an ONNX graph walk, all four agreeing exactly on the arithmetic, checked against closed forms and PyTorch's FLOP counter, costed on a roofline, with operator coverage counted three ways.
    liveMeta device · Fake tensors · torch.export · torch.compile · ONNX
  11. Performance Analysis of Simulators and Systems →
    Measuring before optimising, on the simulators in this series: the USE method and the profile-first loop, Linux perf and what to do when it is locked down, sampling profilers and interactive flame graphs from py-spy, instruction counts with cachegrind (Rust against Python), a hot spot found in the metrics code and fixed with identical answers, benchmark distributions, and regression gates designed from measured noise.
    liveUSE method · py-spy · Flame graphs · cachegrind · Benchmark statistics
  12. Measurement Tools and Methods →
    A catalogue of the tools and methods the simulation series measure with, grouped by what they measure, each with its principle, overhead, accuracy and pitfalls, when to use it, and where this GitHub uses it: profilers and flame graphs, benchmarks, simulated caches and hardware counters, GPU and CPU energy telemetry, trace viewers and framework profilers, RTL coverage and power flows, DRAM simulators, the CACTI, McPAT and Accelergy estimators (CACTI built and swept here), the statistics of simulation output, and agreement checks. Overheads measured on this series' own simulators.
    liveProfilers · Hardware counters · Energy telemetry · Traces · RTL coverage
  13. Power, Performance and Area: the Architect's Trade-offs →
    The three axes an architect trades: performance (latency and throughput), power (dynamic, static, DVFS, dark silicon) and area (SRAM, logic, PHYs and wires, estimated before layout), and why area is cost (dies per wafer, Poisson and Murphy yield, the reticle). Composite metrics (perf/W, perf/mm², EDP, TCO), Pareto fronts, a live explorer on FHE_Accelerator_Sim's new area model, a worked three-way trade-off, and the other trade-offs the simulation series make, named.
    livePPA · Dynamic and static power · Area and yield · perf/W, perf/mm², EDP · Pareto fronts
  14. An Accelerator Model in SimPy, End to End →
    A real graph from torch.export or ONNX run through an event-driven model of a tiled accelerator: off-chip memory, interconnect, DMA engines, an on-chip buffer with back-pressure, a compute array and a vector unit. Lowering to tiles, a cycle-approximate timing model, the roofline, Little's law and stall attribution, a timeline and hot-spot report, a cycle-stepped twin, how gem5 and SST are structured, a bit-identical C++ fast path with pybind11, ONNX Runtime execution providers, and an FHE-style NTT workload.
    liveSimPy · Back-pressure · Stall attribution · Cycle-based twin · gem5 and SST

Companion Code

  1. </>
    Rust_DES_Kernel →
    A discrete-event kernel in Rust and a bit-exact Rust port of Disaggregated_Inference_Sim's engine, cost model and metrics, exposed to Python with PyO3. Differential and golden tests against the Python simulator, proptest invariants, an M/D/1 check against theory, criterion benchmarks, measured speed-ups, GitHub Actions and a Jenkins pipeline with JUnit, coverage and a performance gate.
    codeRust · PyO3 · maturin · proptest · criterion · pytest · Hypothesis
  2. </>
    RTL_CoSim_NTT →
    A Barrett multiplier, NTT butterfly and P-lane NTT core in SystemVerilog, verified with cocotb on Verilator and Icarus against FHE_Accelerator_Sim's NTT as the golden model. Constrained-random stimulus with functional coverage crosses, a seeded bug that uniform stimulus misses, an exact cycle model, a toggle-count power proxy, and the measured NTT efficiency fed back into the FHE simulator.
    codeSystemVerilog · cocotb · Verilator · Icarus · pytest
  3. </>
    Memory_System_Sim →
    A command-level DRAM/HBM timing simulator: bank state machines, JEDEC-style timing constraints, address mapping with XOR hashing, FCFS and FR-FCFS, refresh. Checked by closed forms, an independent protocol checker on Hypothesis traces (which found a real bug) and DRAMsim3, and plugged into FHE_Accelerator_Sim as an optional HBM model.
    codePython · pytest · Hypothesis · DRAMsim3 cross-check
  4. </>
    SystemC_Accelerator_Model →
    A SystemC TLM-2.0 model of an FHE accelerator tile, approximately and loosely timed, driven by FHE_Accelerator_Sim's trace format. An explicit tie-break arbiter makes it agree with the SimPy model op by op; LT quantum trade-offs measured; GoogleTest with sanitizers.
    codeC++17 · SystemC 3 · TLM-2.0 · GoogleTest · ASan/UBSan
  5. </>
    Torch_Sim_Frontend →
    PyTorch and ONNX models turned into operator traces four ways (a dispatch trace on the meta device or under fake tensors, torch.export, a torch.compile backend, an ONNX graph walk) and costed on a roofline or offload model, or run end to end through a SimPy accelerator model (memory, interconnect, DMA, buffer with back-pressure, compute array, vector unit) with utilisation, stall and hot-spot reports, timelines, a bit-identical C++ fast path (pybind11), a cycle-stepped twin, an NTT workload and execution-provider partitioning. Runs Llama-3-70B without weights; the four routes agree exactly with each other, with closed forms and with PyTorch's FLOP counter. An EARS specification with a generated traceability matrix; GitHub Actions and Jenkins.
    codePython · PyTorch · torch.export · torch.compile · ONNX · pytest · Hypothesis

How this series relates to its sisters. LLM Inference Simulators teaches how to build, validate and accelerate a simulator; FHE Accelerator Simulators applies that to an FHE accelerator, and Fourier Optics for Inference extends the LLM simulator with heterogeneous pools and an optical transform engine. This series covers the engineering practice around them: the languages a simulator's fast core is written in, the hardware it must agree with, and the tests, pipelines, trackers and specifications that make its answers trustworthy. Interview-style questions on each topic are in the Interview_* repositories linked from every deck.

New to simulation? Start with Introduction to Simulation: the levels engineers simulate at, from field solvers and SPICE to RTL, architecture and system models, and the methods they share, with links into all three series.

Glossary: Concepts and Where They Are Explained

Every concept the decks rely on, with a short explanation and links to the slides that explain it in depth. Concepts that the sister series already explain are listed at the end with links to their glossary entries rather than repeated here. Each deck's contents slide links to the entries it uses.

Simulation kernels in Rust

BinaryHeap and Reverse
Rust's standard priority queue is a max-heap; wrapping entries in std::cmp::Reverse makes it the min-heap an event list needs, so the earliest event pops first. Push and pop are O(log n).Explained in: SimEng 01 · the event list
Total ordering of floats (total_cmp)
f64 is only partially ordered because NaN compares unequal to everything. f64::total_cmp implements IEEE 754 totalOrder, which lets float times key a heap.Explained in: SimEng 01 · ordering floats, breaking ties
Deterministic tie-breaking
Events at equal times are ordered by priority (SimPy's URGENT before NORMAL) and then by a sequence number (first scheduled, first served), so the same inputs always give the same event order and the same answers.Explained in: SimEng 01 · the ordering rule · SimEng 01 · interactive event queue · SimEng 03 · ties in SystemC · LLM Inference Simulators 02 · determinism pitfalls
Events as enums; exhaustive match
Events are plain data (an enum of small, copyable variants) and the model handles them in a match; the compiler rejects a match that forgets a variant.Explained in: SimEng 01 · events as enums, models as traits
Ownership and borrowing
Every value has one owner; references borrow it, either shared (many readers) or exclusive (one writer), never both at once. The compiler enforces this, which rules out dangling references and data races.Explained in: SimEng 01 · ownership in an event loop · The Rust Programming Language, chapter 4
Arenas and indices
The simulation owns all requests and instances in vectors; events and queues hold indices into them. Borrows last one statement, so the borrow checker is satisfied without Rc<RefCell>, and indices can be logged and replayed.Explained in: SimEng 01 · ownership in an event loop
Processes as state machines
A SimPy process is a generator that sleeps at each yield. Without stable generators, Rust names those sleep points as states and lets each event resume the component from where it slept.Explained in: SimEng 01 · from SimPy processes to state machines
Traits; static and dynamic dispatch
A trait is an interface. Generic code (<C: CostModel>) is compiled once per implementation and can inline calls; a trait object (Box<dyn CostModel>) chooses the implementation at run time through a vtable.Explained in: SimEng 01 · traits for pluggable cost models
Result, ? and panic
Expected failures (a bad configuration) return Result and propagate with ?; broken invariants (an event in the past) panic!, stopping the run at the point of the bug.Explained in: SimEng 01 · errors
rayon, Send and Sync
rayon turns iter() into par_iter() on a work-stealing thread pool. The Send and Sync traits let the compiler prove that what each thread touches is safe to share, so a data race is a compile error.Explained in: SimEng 01 · fearless parallel sweeps
cargo, clippy and rustfmt
Rust's build tool, linter and formatter. cargo clippy -- -D warnings fails the build on any lint; cargo fmt --check enforces one layout.Explained in: SimEng 01 · the toolchain
criterion benchmarks
A Rust benchmarking library: warm-up, many samples, confidence intervals and comparison with the previous run, so performance changes are measured rather than guessed.Explained in: SimEng 12 · criterion · SimEng 01 · the toolchain · SimEng 06 · a performance gate in CI

Rust and Python

Amdahl's law for a port
Speeding up part of a program speeds up the whole only in proportion to that part's share of the time. After porting a simulator's engine, the Python code around it can dominate.Explained in: SimEng 02 · interactive: what should you port? · SimEng 02 · what crosses the boundary · LLM Inference Simulators 08 · Amdahl's law for simulators
PyO3 and maturin
PyO3 generates the CPython bindings for Rust functions marked #[pyfunction]; maturin builds and installs them as a Python wheel, alongside any pure-Python code.Explained in: SimEng 02 · PyO3 and maturin in practice · SimEng 02 · project layout
The language boundary
Every call from Python into Rust converts its arguments and results. A good boundary is coarse (one call per simulation) and plain (numbers, tuples, strings), and its cost is measured.Explained in: SimEng 02 · what crosses the boundary
The GIL and releasing it
CPython's global interpreter lock lets one thread run Python at a time. Rust code that touches no Python objects can release it (py.detach, formerly allow_threads), so Python threads run simulations in parallel.Explained in: SimEng 02 · releasing the GIL
Bit-exact parity
Two implementations that produce identical floating-point results, bit for bit. It requires the same operations in the same order, the same event order and the same library functions, and it turns any discrepancy into a located bug.Explained in: SimEng 02 · the parity checklist · SimEng 02 · three traps that cost one ulp · SimEng 03 · the same order in C++
Correctly rounded summation (math.fsum)
math.fsum returns the exactly rounded sum of its inputs; statistics.fmean uses it. A plain loop accumulates rounding error, so a port must reproduce the algorithm, not just add.Explained in: SimEng 02 · three traps that cost one ulp
MT19937 and Python's random module
CPython's generator is the Mersenne Twister, seeded with init_by_array; random() combines two 32-bit outputs into a 53-bit float, and the variates are written in Python, so the whole stream can be reproduced exactly.Explained in: SimEng 02 · reproducing Python's random numbers
abi3 wheels
Extension modules built against CPython's stable ABI work on every later Python version, so one wheel per platform suffices. The free-threaded build has its own ABI and is not covered.Explained in: SimEng 02 · packaging and CI

C++ and SystemC models

Modern C++ for models
Plain C++17 value types, RAII and lambdas for the model's data and arithmetic, kept free of the simulation kernel so it can be unit-tested alone and ported from a reference in the same floating-point order.Explained in: SimEng 03 · modern C++ for models
The SystemC kernel and delta cycles
SystemC (IEEE 1666) runs processes in evaluate/update cycles. Zero-time notifications start another delta cycle at the same simulated time; time advances only when nothing more is runnable. Time is an integer count of a global resolution.Explained in: SimEng 03 · the SystemC kernel
Process order and explicit tie-breaks
IEEE 1666 does not specify the order in which processes runnable at the same instant execute, so simultaneous requests for a resource need a tie-break written into the model if results are to be reproducible across kernels.Explained in: SimEng 03 · ties are part of the model · SimEng 01 · ordering floats, breaking ties
TLM-2.0 generic payload and sockets
TLM-2.0 replaces pin-level signals with function calls through initiator and target sockets, carrying a generic payload (command, address, data pointer and length, response) and optional extensions, so models from different sources interoperate.Explained in: SimEng 03 · payloads, sockets and coding styles
Approximately timed (AT) and the base protocol
AT models use non-blocking transport with four phases (BEGIN_REQ, END_REQ, BEGIN_RESP, END_RESP), each at its own simulated time, so contention resolves in time order. Accurate, at the cost of a context switch per phase.Explained in: SimEng 03 · the four-phase protocol
Loosely timed (LT) and temporal decoupling
LT models use one blocking b_transport call whose delay argument carries timing. Processes run ahead of simulated time by up to a global quantum and synchronise rarely: faster, but resources get booked out of time order.Explained in: SimEng 03 · temporal decoupling and the quantum · SimEng 03 · interactive: the quantum trade-off
Global quantum
The longest a loosely-timed process may run ahead of the kernel's time before it must synchronise (tlm_quantumkeeper). Larger quanta mean fewer context switches and larger timing errors.Explained in: SimEng 03 · interactive: the quantum trade-off
One elaboration per process
A SystemC program elaborates its module hierarchy and calls sc_start once; tests that each run a simulation therefore fork a child per test.Explained in: SimEng 03 · GoogleTest and sanitizers

Memory systems

“Bandwidth × efficiency”
Modelling memory as peak bandwidth times a fixed derating factor. The achieved fraction of peak depends on access pattern, controller, address mapping and load, so a single factor is right for at most one of them.Explained in: SimEng 04 · why bandwidth times efficiency fails · SimEng 04 · plugging it into the FHE simulator
Channels, ranks, bank groups, banks and rows
A DRAM channel has ranks of chips; each rank has bank groups of banks; each bank is an array of rows read through one row buffer. Banks work in parallel; accesses within a bank group must be further apart than across groups.Explained in: SimEng 04 · how DRAM is organised
DRAM commands and timing parameters
ACT, RD/WR, PRE and REF, separated by minimum gaps such as tRCD, tRP, tCL, tRAS, tRC, tRRD, tCCD, tWTR, tWR and tRTP, defined per speed bin by the JEDEC standards.Explained in: SimEng 04 · commands and timing parameters
Row hits, misses and conflicts
A hit finds its row open (CL); a miss finds the bank closed (tRCD + CL); a conflict finds another row open (tRP + tRCD + CL). A controller's main job is turning conflicts into hits.Explained in: SimEng 04 · row hits, misses and conflicts
FCFS and FR-FCFS scheduling
FCFS serves requests in arrival order. FR-FCFS (Rixner et al., 2000) issues ready row hits first, then the oldest ready command, overlapping row opening in some banks with transfers in others.Explained in: SimEng 04 · scheduling and page policy
Open and closed page policy
Open page leaves a row open after an access, betting on a hit; closed page precharges at once (auto-precharge), betting on a miss rather than a conflict.Explained in: SimEng 04 · scheduling and page policy
Address mapping and XOR bank hashing
Which address bits select channel, bank group, bank, row and column decides which patterns spread across banks. XOR hashing mixes row bits into the bank index so row-sized strides stop hitting one bank.Explained in: SimEng 04 · address mapping
Refresh (tREFI, tRFC)
Every tREFI a rank closes its banks and refreshes for tRFC, costing a stream about tRFC/tREFI of its bandwidth and setting the latency tail.Explained in: SimEng 04 · refresh and the four-activate window · SimEng 04 · load and latency
The four-activate window (tFAW)
At most four ACTs per rank in any window of tFAW, a limit on peak current. It caps random traffic, which needs one ACT per access, at 4 (BL/2) / tFAW of peak.Explained in: SimEng 04 · refresh and the four-activate window
Independent protocol checker
A checker that replays a controller's command log and verifies every timing rule directly, sharing no code with the controller; with property-based traces it is the oracle that finds scheduling bugs.Explained in: SimEng 04 · how we know the model is right
Cross-checking against another simulator
Running identical traces and parameters through an independent simulator (here DRAMsim3). Agreement where the physics dominates and disagreement where the designs differ both carry information.Explained in: SimEng 04 · cross-checked against DRAMsim3

Verification bridge

Golden model and bit-accurate model
The golden model says what the hardware must compute (here the FHE simulator's NTT); a bit-accurate model mirrors the RTL's internal steps so a scoreboard can check them.Explained in: SimEng 05 · two golden models and a scoreboard
Scoreboard
The testbench component that predicts each expected output from the model and compares it with what the monitor observed from the design.Explained in: SimEng 05 · two golden models and a scoreboard
Transactors (driver and monitor)
Testbench components that turn transactions into pin activity and back. A ready/valid handshake must be sampled in the cycle it happens, not after the edge that commits it.Explained in: SimEng 05 · transactors · SimEng 05 · cocotb in one page
Constrained-random stimulus
Random stimulus biased towards corners (0, 1, q - 1, near-maximal operands) while still covering the whole range, so rare situations occur often enough to test.Explained in: SimEng 05 · constrained-random stimulus · SimEng 05 · interactive: coverage closure
Coverage crosses and observability
A cross counts combinations of bins. Reaching a corner is not enough if later logic masks a bug there; the cross with the condition that makes the bug visible is the bin that matters.Explained in: SimEng 05 · functional coverage and crosses · SimEng 05 · a seeded bug
Seeded bugs
Deliberately injected faults (here a missing Barrett correction behind a define) used to check that a testbench can catch them: mutation testing for hardware.Explained in: SimEng 05 · a seeded bug: caught or escaped
Verilator and Icarus Verilog
Verilator compiles synthesisable SystemVerilog to fast two-state C++ and lints it; Icarus is a four-state event-driven simulator. Running tests on both catches simulator-dependent behaviour.Explained in: SimEng 12 · Verilator coverage · SimEng 05 · Verilator, Icarus and CI for RTL
Cycle models from RTL
A closed-form cycle count (here log2 n (n/2P + 6)) fitted to and checked against every RTL measurement, so it can be trusted to extrapolate to sizes too large to simulate.Explained in: SimEng 05 · RTL cycle counts and a cycle model
Calibrating a simulator from RTL
Derating the architecture simulator's ideal unit throughput by the efficiency the RTL achieves; it can change which resource the simulator calls the bottleneck.Explained in: SimEng 05 · feeding the RTL back into the simulator
Switching activity (toggle counts)
The number of bit changes per cycle on a design's nets or registers; activity times capacitance times V squared times f gives dynamic power, so toggle counts are an early power proxy.Explained in: SimEng 12 · toggle counts and switching activity · SimEng 05 · switching activity as a power proxy

Test strategy and frameworks

pytest fixtures and scopes
Functions that provide what a test needs, requested by naming them as parameters, defined in conftest.py for sharing, and cached per function, module or session.Explained in: SimEng 06 · pytest I
parametrize and markers
@pytest.mark.parametrize runs one test body over many inputs; markers label tests (slow, nightly) so -m can select them; a strict xfail records a known limitation.Explained in: SimEng 06 · pytest II
pytest-xdist
A pytest plugin that runs tests in parallel worker processes (-n auto). Tests must be independent: no shared files or global state.Explained in: SimEng 06 · pytest II
Shrinking
When a property-based test fails, the framework searches for simpler inputs that still fail and reports the simplest it finds, which is usually a readable reproduction of the bug.Explained in: SimEng 06 · properties and shrinking · SimEng 06 · stateful testing
Stateful property testing
Hypothesis generates random sequences of operations (rules) and compares the system with a simple model after every step; failures shrink to a minimal sequence.Explained in: SimEng 06 · Hypothesis: stateful testing
proptest
Rust's property-based testing library: strategies generate inputs, failures shrink, and failing seeds are saved in proptest-regressions/ and replayed first.Explained in: SimEng 06 · cargo test and proptest · SimEng 01 · testing in Rust
Golden tests and re-blessing
Store the output of a reference run and fail when today's differs; re-bless (rewrite the stored output) deliberately, with the diff reviewed alongside the change that caused it.Explained in: SimEng 12 · golden, differential and correlation · SimEng 06 · golden tests and re-blessing · SimEng 02 · golden files across languages
GoogleTest
The common C++ test framework: TEST, fixtures with TEST_F, value-parameterised TEST_P, EXPECT_* (continue) and ASSERT_* (stop), death tests.Explained in: SimEng 06 · GoogleTest for C++ models · SimEng 03 · GoogleTest and sanitizers
Sanitizers (ASan, UBSan)
Compiler instrumentation that turns memory errors (AddressSanitizer) and undefined behaviour (UndefinedBehaviorSanitizer) into immediate failures with a stack trace. Build C++ tests with them.Explained in: SimEng 06 · GoogleTest for C++ models · SimEng 03 · GoogleTest and sanitizers
cocotb
A framework for writing RTL testbenches in Python coroutines that drive an HDL simulator (Icarus, Verilator and others), so a Python model can check the hardware.Explained in: SimEng 06 · cocotb: Python testbenches for RTL · SimEng 05 · cocotb in one page
Functional coverage
Bins describing the situations a testbench must exercise (a wrap-around, a boundary value), counted during the run; an empty bin is a hole in the verification, whatever the code coverage says.Explained in: SimEng 12 · functional and code coverage · SimEng 06 · cocotb · SimEng 05 · functional coverage and crosses

Test quality and CI

Mutation testing
Make small deliberate bugs (mutants) in the code and run the tests against each: a killed mutant was noticed, a surviving one was not. The mutation score measures how much the tests actually check.Explained in: SimEng 12 · mutation testing · SimEng 06 · mutation testing · SimEng 06 · interactive: mutation score versus coverage
Equivalent mutants
Mutants that change the code but not its behaviour, so no test can kill them; survivors must be triaged rather than all chased.Explained in: SimEng 06 · coverage, and what it does not tell you
Line, branch and region coverage
The fraction of lines, branch outcomes or code regions that the tests execute. Low coverage is a warning; high coverage does not show that results are checked.Explained in: SimEng 12 · code coverage tools · SimEng 06 · coverage, and what it does not tell you · SimEng 06 · interactive
MC/DC
Modified condition/decision coverage: each condition in a decision must be shown to change the outcome independently. DO-178C requires it for the most critical avionics software.Explained in: SimEng 06 · coverage, and what it does not tell you
JUnit XML and Cobertura reports
The de facto interchange formats for test results and coverage, written by pytest, cargo-nextest, GoogleTest and cocotb, and read by Jenkins and most CI systems.Explained in: SimEng 06 · running it all in CI · SimEng 07 · test and coverage reports
Performance-regression gate
A CI step that benchmarks the build and fails it if a benchmark is slower than a stored baseline by more than a set margin.Explained in: SimEng 06 · running it all in CI · SimEng 01 · criterion · SimEng 05 · an exact cycle-count gate · SimEng 07 · gating on performance and accuracy · SimEng 11 · regression detection in CI
Flaky tests
Tests that pass or fail without a code change. In a deterministic simulator they point to shared state, an unseeded random number or a wall-clock dependency, and should be fixed, not retried.Explained in: SimEng 06 · running it all in CI

Pipelines and Jenkins

Jenkins: controller, agents, plugins
A self-hosted automation server. The controller (web UI, build queue, credentials, plugins) schedules pipeline runs onto the executors of agents; almost every feature, Pipeline itself included, is a plugin.Explained in: Introduction to Jenkins · architecture: controller, agents, executors · Introduction to Jenkins · plugins, the update centre and JENKINS_HOME · SimEng 07 · agents, labels and the tools on them
GitHub Actions: workflows, jobs, runners
GitHub's built-in automation. An event (a push, a pull request, a schedule) starts a workflow, a YAML file in .github/workflows/; its jobs run in parallel on fresh runners, each step runs a shell command or a reusable action, and each job reports a check that a ruleset can require before a merge.Explained in: Introduction to GitHub Actions · events, workflows, jobs, steps and runners · Introduction to GitHub Actions · required checks, shown blocking · SimEng 07 · Jenkins and GitHub Actions
Declarative and scripted pipelines
A Jenkinsfile is either declarative (a fixed structure of agent, parameters, stages and post conditions that Jenkins validates and draws) or scripted (Groovy code inside node { }, with loops and try/catch).Explained in: SimEng 07 · declarative and scripted pipelines
Agents and labels
The controller schedules work; agents run it. A stage asks for an agent by label (tools, licences, capacity), so a pipeline states what it needs rather than which machine to use.Explained in: SimEng 07 · agents, labels and the tools on them
Matrix builds
A declarative matrix runs the same stages for every combination of its axes as parallel branches, minus exclusions: a parameter sweep in CI.Explained in: SimEng 07 · matrix builds
Shared libraries
A Git repository of Groovy that pipelines load with @Library; each file in vars/ becomes a pipeline step, so many repositories share one definition of their stages.Explained in: SimEng 07 · shared libraries
Nightly regressions and build parameters
Fast checks on every change, expensive sweeps nightly, selected by a build parameter. A parameter default computed from env becomes the text null, and once a job has parameters a plain /build request is rejected (HTTP 400).Explained in: SimEng 07 · nightly regressions and parameters
Exact, speed and drift gates
Three ways a CI gate compares a build with a baseline: exact behaviour (any change fails), speed (within a margin, best of several runs) and drift (reported, not failed).Explained in: SimEng 07 · gating on performance and accuracy
Credentials and masking
Secrets stored on the controller and referred to by ID; withCredentials binds one for a block and masks its value in the log.Explained in: SimEng 07 · credentials
The declarative linter
Jenkins validates a declarative Jenkinsfile without running it (/pipeline-model-converter/validate), catching structural mistakes such as a misspelt post condition.Explained in: SimEng 07 · the real runs

Tracking and engineering metrics

The issue model
Epics group stories, tasks and bugs, which may have sub-tasks; each item has fields (status, assignee, fix version, sprint, components, labels), links and a full change history.Explained in: SimEng 08 · the issue model
Workflows and status categories
A workflow is a state machine of statuses and allowed transitions; every status belongs to one of three categories (To Do, In Progress, Done) that queries and boards can rely on.Explained in: SimEng 08 · workflows, transitions and status categories
Versions, sprints and Kanban
A fix version names the release that will contain a change; Scrum plans work in fixed sprints, Kanban pulls it continuously under WIP limits.Explained in: SimEng 08 · versions, sprints and Kanban
JQL
Jira Query Language: clauses of field, operator and value joined by AND, OR and NOT, with functions and ORDER BY. != does not match empty fields; WAS and CHANGED search the history of six fields.Explained in: SimEng 08 · JQL: the query language · SimEng 08 · interactive: worked JQL queries
Smart commits and the development panel
Work-item keys in branch names, commit messages or pull-request titles link the development work to the item; smart commits can also comment, log time and run a transition.Explained in: SimEng 08 · Git and Jenkins integration
Cycle time, throughput and WIP
Flow metrics computed from the workflow history: how long items take from start to done (as percentiles), how many finish per week, how many are in progress, and which have been in progress too long.Explained in: SimEng 08 · interactive: flow metrics · SimEng 08 · the metrics to report
Cumulative flow diagram
The number of items in each status category on each day, stacked; widening bands show work accumulating in a stage.Explained in: SimEng 08 · interactive: flow metrics
Goodhart's law
A measure that becomes a target stops being a good measure; prefer metrics that tools compute from records, and report distributions rather than single numbers.Explained in: SimEng 08 · the metrics a simulation team should report

Specifications and requirements

Requirement quality
A good requirement is necessary, unambiguous, singular, feasible and verifiable, among the characteristics INCOSE's guide and ISO/IEC/IEEE 29148 describe; writing rules ban vague terms, escape and open-ended clauses.Explained in: SimEng 09 · what makes a good requirement · SimEng 09 · interactive: check a requirement
EARS
The Easy Approach to Requirements Syntax: ubiquitous, event-driven (When), state-driven (While), unwanted-behaviour (If ... then) and optional-feature (Where) patterns, and combinations of them.Explained in: SimEng 09 · EARS: five patterns · SimEng 09 · interactive
Verification methods
Each requirement names how it will be verified: test, analysis, inspection or demonstration. A requirement no method could fail is not verifiable.Explained in: SimEng 09 · what makes a good requirement
Non-functional requirements for simulators
How well a simulator must work: accuracy against a named reference and tolerance, determinism, speed, capacity and stated fidelity boundaries.Explained in: SimEng 09 · functional and non-functional requirements
Mission-mode software
Here: the software that runs on the product in its operational mode, as opposed to test, bring-up, calibration or diagnostic software. Its requirements must capture modes, fault handling and resource budgets.Explained in: SimEng 09 · requirements capture for mission-mode software
The V-model; verification and validation
Each level of specification is verified by the level of test opposite it. Verification asks whether it was built right; validation, whether the right thing was built.Explained in: SimEng 09 · the V-model
Traceability matrix
The links from each requirement to the tests and evidence that verify it, checked in both directions; generated from the test run, a missing link is a visible gap.Explained in: SimEng 09 · traceability, generated from the test run · SimEng 08 · requirement, work item, test, build
Test plans and test oracles
A test plan states scope, approach, pass/fail and exit criteria, environment, deliverables and risks; its core is the oracle each level uses to decide that a result is right.Explained in: SimEng 09 · template 2: a test plan
Generated performance reports
A report written by a script from recorded runs (question, environment, method, validation, results, limitations), so it can be regenerated after any change.Explained in: SimEng 09 · template 3: a performance report

Framework front ends

Operator trace
The operators a model executes, in order, with each tensor's shape, element size, identity and whether it is a weight: the input a cost model needs from a front end.Explained in: SimEng 10 · one trace format, four front ends
Fake tensors and traced code paths
Tensors with shapes and a claimed device but no data. What a trace shows depends on that device (fused or decomposed attention) and on whether the library detects tracing and takes another path.Explained in: SimEng 10 · meta or fake tensors
ONNX shape inference and constant folding
Walking an exported graph needs shape values propagated through Shape and Reshape chains (data_prop) and constant subgraphs folded as a runtime would; weights can be exported as typed inputs without their data.Explained in: SimEng 10 · front end 4: ONNX without the weights
Unfused and ideal-fusion bounds
Costing a trace per operator assumes every operator reads and writes memory (an upper bound); assuming only matmul, attention and gathers touch memory gives a lower bound. Real compilers fall between.Explained in: SimEng 10 · costing the trace
Operator coverage
Which operators a cost model has rules for, and which the accelerator runs: counted by operators, FLOPs and time, because a device can hold nearly all the FLOPs and little of the time.Explained in: SimEng 10 · interactive: operator coverage · SimEng 10 · operator coverage as a metric

Performance analysis

The USE method
For every resource (CPUs, memory, disks, network, locks), check utilisation, saturation (work queued because the resource is busy) and errors: a checklist that finds system bottlenecks without guessing.Explained in: SimEng 11 · USE for resources, profile first for code · SimEng 11 · USE in practice
The profile-first loop
Measure a baseline, profile, form one hypothesis, change one thing, re-measure, and check the answers did not change; repeat.Explained in: SimEng 11 · two methods · SimEng 11 · closing the loop: sort once · LLM Inference Simulators 08 · step zero: profile
Linux perf and perf_event_paranoid
The Linux profiler: hardware event counts (perf stat), sampled call stacks (perf record) and annotated code. Unprivileged access is governed by the kernel.perf_event_paranoid setting.Explained in: SimEng 12 · Linux perf · SimEng 11 · Linux perf: counters and call stacks
Sampling profilers (py-spy)
A profiler that records the call stack at a fixed rate; time spent shows up as samples. py-spy samples a Python process from outside, with no changes to it.Explained in: SimEng 12 · py-spy and sampling profilers · SimEng 11 · sampling profilers and flame graphs · SimEng 11 · reading a profile
Flame graphs
Sampled stacks merged by common prefix: a frame's width is its share of samples, the top edge of each tower is where time is spent, and the x-axis is not time.Explained in: SimEng 12 · flame graphs, on- and off-CPU · SimEng 11 · sampling profilers and flame graphs · SimEng 11 · interactive: flame graphs
Self time and total time
Self time counts samples where a function is running; total time counts samples where it is anywhere on the stack. Total locates the phase, self locates the code.Explained in: SimEng 11 · reading a profile: self and total time
Cachegrind and instruction counts
Valgrind's tool that runs a program on a simulated CPU, counting every instruction and modelling the caches: slow, but deterministic, so small differences can be compared without timing noise.Explained in: SimEng 12 · cachegrind and callgrind · SimEng 11 · counting instructions with cachegrind
Benchmark distributions
Repeated timings form a distribution, often with several clusters: report the median, the median absolute deviation and a bootstrap confidence interval rather than a mean and standard deviation.Explained in: SimEng 11 · benchmarks are distributions
A/A tests and regression-gate design
Running a timing gate on unchanged code measures its false-alarm rate; the number of runs per side and the margin then set what it catches. The margin must sit well below the regression to be caught.Explained in: SimEng 11 · interactive: designing a regression gate · SimEng 11 · regression detection in CI

Measurement tools and methods

cProfile and snakeviz (instrumenting profilers)
Python's deterministic profiler hooks every call and return: exact call counts and call trees, but the per-call cost inflates small, frequent functions. snakeviz draws the saved profile.Explained in: SimEng 12 · cProfile and snakeviz · LLM Inference Simulators 08 · step zero: profile
Hardware performance counters and multiplexing
The CPU's performance-monitoring unit counts events (cycles, instructions, cache and branch misses) at almost no cost; with more events than counters the kernel rotates them and scales the counts, which adds error.Explained in: SimEng 12 · hardware performance counters · SimEng 11 · perf stat
GNU time (/usr/bin/time -v)
Prints what the kernel accounted for a finished command: wall and CPU time, maximum resident set size, page faults and context switches. Free, and the first look in the USE method.Explained in: SimEng 12 · GNU time · SimEng 11 · USE in practice
Load generators and coordinated omission
A load generator issues timestamped requests to a system under test (MLPerf's LoadGen defines the standard scenarios). A closed-loop generator that waits for each response hides stalls from its own measurements: coordinated omission.Explained in: SimEng 12 · load generators
tracemalloc (Python heap profiling)
Records the allocating traceback of every live Python memory block, so snapshots show which lines hold memory; expensive, and blind to memory that native extensions allocate.Explained in: SimEng 12 · tracemalloc
nvidia-smi and NVML
NVIDIA's driver library for GPU monitoring and its command-line front end: utilisation, clocks, memory, power and an energy counter. Power readings are averaged over sensor windows that can miss most of a run.Explained in: SimEng 12 · nvidia-smi and NVML
NVIDIA DCGM
The Data Center GPU Manager samples numbered fields (power, energy, and profiling metrics such as SM and tensor-pipe activity) on every GPU, continuously; some profiling metrics cannot be collected together and are multiplexed.Explained in: SimEng 12 · DCGM
Zeus (ML.ENERGY)
A Python library measuring time and energy over named code windows from NVML's energy counter (and RAPL for CPUs); the measurement behind the ML.ENERGY leaderboard.Explained in: SimEng 12 · Zeus
RAPL (CPU energy counters)
Energy counters per CPU domain (package, cores, DRAM) in model-specific registers, read through Linux powercap or perf's power events. Package energy, not wall energy; root-only on current kernels.Explained in: SimEng 12 · RAPL
External power analysers and board telemetry
Meters at the wall or on a supply rail that measure voltage and current directly: the reference other power readings are calibrated against, including everything in the box.Explained in: SimEng 12 · power analysers
PyTorch profiler
Records an event per PyTorch operator (and GPU kernel), with shapes and memory on request; aggregates by operator or exports a Chrome trace.Explained in: SimEng 12 · PyTorch profiler
ONNX Runtime profiler
With profiling enabled, ONNX Runtime writes a Chrome-trace event per graph node per run, with operator type, execution provider and shapes, for the graph after its own optimisations.Explained in: SimEng 12 · ONNX Runtime profiler · LLM Inference Simulators 09 · ONNX and ONNX Runtime
NVIDIA Nsight Systems
A system-wide tracer of CPU threads, CUDA API calls, GPU kernels and copies on one timeline, for finding the idle gaps on a GPU.Explained in: SimEng 12 · Nsight Systems
DRAMsim3
A cycle-accurate DRAM controller and device simulator (DDR, LPDDR, GDDR, HBM) configured with JEDEC timing; used here as the reference that Memory_System_Sim was cross-checked against.Explained in: SimEng 12 · DRAMsim3 · SimEng 04 · cross-checked against DRAMsim3
Ramulator
A cycle-level DRAM simulator that describes each standard as a generic state machine, so new standards and mechanisms are easy to add.Explained in: SimEng 12 · Ramulator
McPAT
An analytical power, area and timing model of multicore processors built from their structure, driven by a performance simulator's activity counts; useful for relative comparisons, with documented sources of error.Explained in: SimEng 12 · McPAT
CACTI
An analytical model of SRAM and DRAM arrays and caches: from capacity, organisation and technology it gives access time, energy per access, leakage and area. Built and swept over capacity at 22 nm in SimEng 12.Explained in: SimEng 12 · CACTI · SimEng 12 · CACTI on this machine
Accelergy and Timeloop
Timeloop searches the mappings of a tensor workload onto an accelerator and counts the actions at every level; Accelergy prices those actions in energy and area from per-component models.Explained in: SimEng 12 · Accelergy and Timeloop
Analytic lower bounds
Bounds on run time from counts alone (the roofline, the busiest resource's busy time): optimistic by construction, so a simulated result below one is a bug.Explained in: SimEng 12 · analytic bounds
PyTorch's FLOP counter (FlopCounterMode)
A dispatch mode that adds up FLOPs from per-operator formulas as a model runs, even on the meta device; operators without a formula (fused CPU attention) are silently missed.Explained in: SimEng 12 · FlopCounterMode · SimEng 10 · operator coverage as a metric

Power, performance and area

PPA: power, performance and area
The three quantities a chip design is judged on, traded together at every level from architecture to layout: more SRAM saves traffic and costs area, a higher clock raises performance and power, a bigger die has longer wires.Explained in: SimEng 13 · what PPA is, and who trades it · SimEng 13 · the other trade-offs, named
Dynamic power (αCV²f) and leakage
Dynamic power is αCV²f: activity, switched capacitance, supply voltage squared, clock. Static power is V · Ileak, drawn whether busy or not. DVFS lowers V and f together, so energy per operation falls roughly as V².Explained in: SimEng 13 · dynamic, static, DVFS and dark silicon
Dark silicon
Since supply voltage stopped scaling with feature size, a chip has more transistors than its power budget can switch at once, so part of it must be idle or slowed at any moment; specialised accelerators are one response.Explained in: SimEng 13 · dark silicon
What sets die area
SRAM bitcells and their periphery, logic, I/O and PHYs (which shrink little with the node), and wiring with its repeaters and registers. In ARK's 7 nm breakdown the scratchpad is over half the die and most of the NTT and permutation units is wire.Explained in: SimEng 13 · area: what sets it · SimEng 13 · estimating area before layout
Dies per wafer
π(d/2)²/A − πd/√(2A) for wafer diameter d and die area A: the wafer area divided by the die area, minus the partial dies lost round the edge.Explained in: SimEng 13 · why area is cost
Poisson yield
Y = e−A·D0: the chance a die of area A has no defect when defects (density D0) land independently and uniformly. Pessimistic for large dies.Explained in: SimEng 13 · yield models, interactive
Murphy yield
Y = ((1 − e−A·D0)/(A·D0))²: yield when the defect density varies across the wafer. Agrees with Poisson for small dies and is higher for large ones.Explained in: SimEng 13 · yield models, interactive
Reticle limit and chiplets
A lithography field is about 26 × 33 mm (858 mm²), the largest die one exposure prints. Larger designs, or ones past the yield knee, are split into chiplets, which pay for die-to-die links, packaging and assembly.Explained in: SimEng 13 · the reticle · SimEng 13 · when to split the die
perf/W
Throughput per watt, which for a fixed task equals 1 / energy per task: it ranks designs by energy alone and ignores speed.Explained in: SimEng 13 · composite metrics
perf/mm² and perf/$
Throughput per square millimetre of silicon, or per dollar of manufacturing cost: what a design delivers for its die. Both ignore energy, and perf/mm² also ignores yield.Explained in: SimEng 13 · composite metrics · SimEng 13 · the scratchpad sweep
EDP and ED²P
Energy-delay product E·t weights energy and speed equally; E·t² weights speed more and is roughly independent of supply voltage under DVFS, so it compares designs rather than operating points.Explained in: SimEng 13 · composite metrics
Total cost of ownership (TCO)
Purchase cost amortised over the service life plus the energy bill (energy × price × datacentre overhead): the figure an operator optimises, and one that needs prices rarely known before silicon.Explained in: SimEng 13 · composite metrics
Pareto front and dominance
Design A dominates B if it is no worse on every axis and better on one; the Pareto front is the set nothing dominates. Adding an axis (area) can turn a single optimum into a front.Explained in: SimEng 13 · Pareto fronts · SimEng 13 · the live PPA explorer · FHE glossary · design sweeps and Pareto fronts
An area model calibrated from papers and CACTI
FHE_Accelerator_Sim's ppa.py: per-unit areas from ARK's published 7 nm breakdown, SRAM from a CACTI sweep scaled to ARK's scratchpad, HBM PHYs and uncore. Illustrative: good for ratios between designs, not absolute mm².Explained in: SimEng 13 · estimating area before layout · SimEng 13 · pitfalls

Accelerator models in SimPy

Lowering a graph to tiles
Turning each operator of a trace into tiles, the unit an accelerator loads, computes and stores: GEMMs are blocked so operands and result fit the on-chip buffer, other operators are cut into byte slices. Tiling sets both the extra memory traffic and how much load, compute and store overlap.Explained in: SimEng 14 · from a graph to tiles · SimEng 14 · a bigger buffer made it slower
im2col: a convolution as a GEMM
A convolution is a matrix product once each output position's input patch is laid out as a row: M = output positions, K = input channels per group × kernel size, N = output channels per group. It lets one matrix engine run both.Explained in: SimEng 14 · from a graph to tiles
Systolic array (output-stationary)
A grid of R × C multiply-accumulate cells that pass operands to their neighbours; in the output-stationary dataflow each cell accumulates one output. Its time per tile is about ⌈m/R⌉⌈n/C⌉k cycles plus a fill and drain, so PE efficiency drops when m or n is not a multiple of the array.Explained in: SimEng 14 · the timing model
Cycle-approximate and cycle-accurate models
A cycle-approximate model computes a block's cycle count from a formula (or a few state changes); a cycle-accurate one evaluates every register on every cycle and matches RTL exactly. The first is fast enough to explore designs; the second is worth building when the question lies inside the cycles, or to calibrate the first.Explained in: SimEng 14 · cycle-approximate against cycle-accurate · SimEng 14 · process-based against cycle-based
Stall attribution
Splitting a run's time into computing and the waits between tiles, each wait named by what the next load was waiting for (load bandwidth, buffer space, an earlier operator's results). The parts sum to the latency, and the largest names the hot-spot: a sharper tool than utilisation.Explained in: SimEng 14 · finding the stalls · SimEng 14 · where the time goes, interactive
Read-after-write through memory
When every operator writes its results to off-chip memory and the next reads them back, the next operator cannot start loading until the previous one is stored: the pipeline drains at every operator boundary. Fusion, keeping results on chip, removes the wait.Explained in: SimEng 14 · from a graph to tiles · SimEng 14 · dependency stalls
Process-based against cycle-based modelling
A process-based (event-driven) model jumps from one state change to the next; a cycle-based one evaluates every block on every clock edge, as RTL simulation does. On the same whole-cycle program they give identical answers; the cost scales with state changes for one and with cycles for the other.Explained in: SimEng 14 · process-based against cycle-based
From an event loop to a recurrence
When nothing in a model is left to arbitrate (no two engines ever contend), every event time follows from earlier ones by additions and maxima, so the event loop can be replaced by a straight-line recurrence: same answers, bit for bit, at a fraction of the cost.Explained in: SimEng 14 · the hot path in C++
pybind11
A header-only C++ library that exposes C++ functions and classes to Python, converting standard containers automatically; built here by setuptools with Pybind11Extension. Bound objects are owned through a holder, std::unique_ptr by default.Explained in: SimEng 14 · the hot path in C++ with pybind11 · SimEng 14 · modern C++ in the port
Templates, concepts, move semantics and smart pointers
Templates write code once for many types, and C++20 concepts state what a type must support; move semantics transfer a container's storage instead of copying it; smart pointers (unique_ptr, shared_ptr) free memory when its owner goes, so no delete is written.Explained in: SimEng 14 · modern C++ in the port · SimEng 03 · modern C++ for models
How gem5 and SST are structured
gem5: SimObjects configured from Python, connected by request and response ports that carry packets, with retries for flow control, and an event queue in ticks. SST: components connected by links with latencies, exchanging events, with clock handlers and conservative parallel simulation across ranks.Explained in: SimEng 14 · how gem5 and SST are structured
Execution-provider partitioning (GetCapability)
How ONNX Runtime hands part of a graph to a hardware backend: each execution provider says which nodes it can run, connected claimed nodes are fused into subgraphs, the rest falls back to the CPU provider, and tensors crossing between them cost transfers.Explained in: SimEng 14 · ONNX Runtime execution providers
Polynomial products via the NTT
Multiplying two polynomials modulo XN+1 by transforming both with the number-theoretic transform, multiplying pointwise and transforming back, after a twist by a 2N-th root of unity: O(N log N) instead of O(N2), and the core workload of lattice-based FHE.Explained in: SimEng 14 · an FHE workload
Computing in the data path
Doing work on data while it moves (in memory, in the network, in a link) instead of moving it to a processor first. It pays when the moved data needs little arithmetic per byte and the destination engine is the bottleneck; a stage slower than the stream throttles it.Explained in: SimEng 14 · computing in the data path (speculative)

Explained in the sister series

Discrete-event simulation (DES)
A simulator whose state changes only at instants (events). A clock, a state, a time-ordered event list and handlers; the clock jumps from one event to the next (next-event time advance), so idle time costs nothing.Explained in: the LLM Inference Simulators glossary
SimPy
A small, pure-Python DES library: Environment, processes, timeouts, events, Resource, Container, Store and interrupts. It has no built-in statistics or tracing; the probes are yours to design.Explained in: the LLM Inference Simulators glossary
Process interaction and coroutines
Writing each entity's life as one sequential function that suspends while it waits, instead of scattering it over event handlers. Python generators make this natural: yield an event, resume when it fires.Explained in: the LLM Inference Simulators glossary
Determinism, tie-breaking and RNG streams
Same configuration and seed must give bit-identical results: break ties between simultaneous events with a sequence number, prefer integer time for clocked models, and give each component its own random stream.Explained in: the LLM Inference Simulators glossary
Cost model
The part of a simulator that answers "how long does this take?" (and, with a power model, "how many joules?"). Kept separate from the event engine so a roofline can later be replaced by a measured table or an RTL-calibrated model.Explained in: the LLM Inference Simulators glossary
Queueing theory checks (M/M/1, M/D/1)
Closed-form results that a simulator's engine must reproduce: M/M/1 mean time in system 1/(μ−λ); M/D/1 mean wait λS²/(2(1−ρ)) (Pollaczek–Khinchine). The KV link is tested this way.Explained in: the LLM Inference Simulators glossary
Little's law and the operational laws
L = λW: mean population equals arrival rate times mean time in system, for any stable system. With the utilisation and forced-flow laws it gives free consistency checks on a simulator's probes.Explained in: the LLM Inference Simulators glossary
The verification ladder
Layers of evidence: unit tests against hand calculations, invariants, analytic checks, behavioural expectations, property-based tests, differential tests and, finally, correlation against measurement.Explained in: the LLM Inference Simulators glossary
Property-based testing (Hypothesis)
Generate many random configurations and check that invariants hold for all of them; Hypothesis shrinks any failure to a minimal example.Explained in: the LLM Inference Simulators glossary
Differential testing
Requiring two independent implementations to agree exactly on the same input: here the fast path against the baseline, and the browser JavaScript against the Python simulator.Explained in: the LLM Inference Simulators glossary
Profiling a simulator
Measure before optimising (cProfile, py-spy, perf). Time usually goes to model code that runs per entity per step, not to the event kernel.Explained in: the LLM Inference Simulators glossary
Parallel runs and Amdahl's law
Independent runs parallelise trivially (processes, job arrays), but each technique only speeds up the share of time it touches; serial overheads cap the gain.Explained in: the LLM Inference Simulators glossary
CI, golden metrics and engineering metrics
Run tests on every push and sweeps nightly (Jenkins or similar); fail the build if golden metrics drift; track correlation error, coverage, simulator speed and test health as engineering metrics.Explained in: the LLM Inference Simulators glossary
Transaction-level modelling (SystemC TLM-2.0)
Modelling bus reads and writes as function calls with annotated delays rather than signal wiggles. Loosely-timed TLM boots software; approximately-timed TLM models contention. The industry route to virtual platforms.Explained in: the LLM Inference Simulators glossary
RTL simulation, co-simulation and golden models
The software model is the executable specification and the golden reference (a scoreboard in the RTL testbench). RTL cycle counts calibrate the model, workload traces become directed tests, and co-simulation embeds RTL blocks in the system model.Explained in: the LLM Inference Simulators glossary
RTL power and thermal analysis
Switching activity from RTL simulation or emulation, run on real workload traces, gives per-block energies that calibrate the simulator; power maps feed thermal models that find thermal hot-spots.Explained in: the LLM Inference Simulators glossary
Percentiles and tail-latency CCDFs
Report p50, p90 and p99, not means, and plot the fraction of samples slower than x on log-log axes, where tails show. Say which population (all tokens or per request), and use streaming sketches when samples are many.Explained in: the LLM Inference Simulators glossary
torch.export and FX graphs
Capturing a PyTorch model ahead of time as one graph of ATen operators with fake-tensor shapes on every node; run_decompositions() lowers it to the smaller Core ATen set a cost model can cover.Explained in: the LLM Inference Simulators glossary
torch.compile backends
TorchDynamo captures graphs as Python runs and hands each to a backend callable, a clean hook for a simulator that estimates cost and then runs the graph eagerly.Explained in: the LLM Inference Simulators glossary
Dispatch interception and the meta device
A TorchDispatchMode sees every ATen call; on the meta device tensors have shapes but no storage, so a 70B model's exact operator trace can be recorded on a laptop.Explained in: the LLM Inference Simulators glossary
ONNX and ONNX Runtime execution providers
ONNX is a framework-neutral graph format; ONNX Runtime partitions graphs between execution providers, each claiming the nodes it can run, with fallback to the CPU.Explained in: the LLM Inference Simulators glossary
FLOP and byte counting
About 2 FLOPs per matmul parameter per token, plus attention that grows with context. A decode step must also read every weight and all of the KV cache. These two counts drive every timing and energy number in the series.Explained in: the LLM Inference Simulators glossary
Roofline, arithmetic intensity and ridge point
Attainable performance is the lower of peak compute and bandwidth × arithmetic intensity (FLOPs per byte). The ridge point is the intensity at which both limits meet; below it a kernel is memory-bound.Explained in: the LLM Inference Simulators glossary
Warm-up, replications and confidence intervals
A stochastic run is one experiment: drop the warm-up period, repeat with different seeds, and report mean ± a t-based confidence interval. Tail percentiles need many samples.Explained in: the LLM Inference Simulators glossary
Hot-spot attribution
Finding where speeding things up would help most. A stage breakdown that sums exactly to each request's latency, the largest waiting stage mapped to the resource that owns it, confirmed by what-if sensitivity runs.Explained in: the LLM Inference Simulators glossary
Chrome trace events and Perfetto
A JSON timeline format (complete slices, counters, metadata) that Perfetto and chrome://tracing display. The simulator, ONNX Runtime's profiler and RTL testbenches can all emit it, so timelines can be compared side by side.Explained in: the LLM Inference Simulators glossary
Calibration and correlation
Fitting a model's few free coefficients to measurements (or RTL), then tracking its error on workloads it was not fitted to. The correlation report is the deliverable that makes predictions trustworthy.Explained in: the LLM Inference Simulators glossary
Prefill and decode
Prefill processes the whole prompt in one pass (matrix-matrix, compute-bound) and produces the first token. Decode then generates one token per sequence per step (matrix-vector, memory-bound). They need different hardware balances.Explained in: the LLM Inference Simulators glossary
KV cache
The stored keys and values of every token of every live sequence, so attention need not recompute them. Its size (320 KiB per token for Llama-3-70B) limits batch size and is what disaggregation must move.Explained in: the LLM Inference Simulators glossary
Utilisation, MFU and MBU
Busy fraction says how often a resource works; model FLOPs and bandwidth utilisation say how much of its capability it uses. A continuous-batching decode engine is 100% busy at 4% MFU and about 77% MBU.Explained in: the LLM Inference Simulators glossary
Common random numbers
Feeding two designs the same random request stream, so noise cancels in their difference: a much narrower confidence interval on "which is better, and by how much" for the same number of runs.Explained in: the LLM Inference Simulators glossary
Batch means
One long run cut into batches long enough to be nearly independent, each giving one sample of the metric: warm-up is paid once instead of once per replication, at the risk of correlated batches.Explained in: the LLM Inference Simulators glossary
Static and dynamic power
Static power (leakage, clocks, always-on logic) is paid whenever a chip is on; dynamic power αCV²f is paid for switching. Energy per operation scales with V², which is why lowering voltage pays quadratically.Explained in: the LLM Inference Simulators glossary
DVFS and power caps (the third roof)
Lowering clock and voltage to cut power. A memory-bound step can slow its compute clock at no time cost; a power cap makes some steps power-bound, a third roof beside compute and memory.Explained in: the LLM Inference Simulators glossary
TDP and peak power
The board power limit a device throttles to. Steps that saturate compute and memory at once can exceed it, so the simulator enforces TDP by default and reports the peak power of each step.Explained in: the LLM Inference Simulators glossary
Energy per token and energy–delay trade-off
Joules per output token (or tokens per joule) is the efficiency headline. Capping power trades time for energy until static power over the longer run dominates, giving an energy-optimal operating point.Explained in: the LLM Inference Simulators glossary
Continuous batching
Re-forming the batch at every decode iteration: finished sequences leave and new ones join, so the GPU stays full. The scheduler loop at the heart of every serving engine and serving simulator.Explained in: the LLM Inference Simulators glossary
TTFT, TPOT and inter-token latency
Time to first token (prefill queueing and compute); time per output token after the first (DistServe's definition); and the gap between consecutive tokens, whose tail shows stalls the mean hides.Explained in: the LLM Inference Simulators glossary
Disaggregated (prefill/decode) serving
Running prefill and decode on separate machine pools (Splitwise, DistServe, Mooncake, NVIDIA Dynamo), so each phase is provisioned and power-managed for its own SLO. The cost is moving the KV cache.Explained in: the LLM Inference Simulators glossary
The fidelity ladder
The range of model detail from analytical formulas through discrete-event, transaction-level, cycle-level and RTL simulation to emulation, FPGA prototypes and silicon. Each rung is slower and more detailed; use the highest rung that answers the question.Explained in: the LLM Inference Simulators glossary
Resources, stores, containers and back-pressure
Resource: N identical servers held for a while (links, engines). Store: a buffer of items (FIFOs). Container: an amount (credits, bytes). A finite-capacity Store makes producers wait when it is full, which is back-pressure.Explained in: the LLM Inference Simulators glossary
Parallel discrete-event simulation (PDES)
Splitting one run across cores as logical processes: conservative (Chandy–Misra–Bryant, needs lookahead) or optimistic (Time Warp, rolls back). Worth it for large, loosely coupled models.Explained in: the LLM Inference Simulators glossary
Number-theoretic transform (NTT)
The FFT over a finite field Zq. It moves a limb between coefficient and evaluation form in (N/2) log N butterflies, so polynomial multiplication becomes element-wise. Every key switch and rescale runs NTTs.Explained in: the FHE Accelerator Simulators glossary
Modular reduction (Barrett, Montgomery)
Ways to compute a·b mod q without division, using precomputed constants. Every butterfly and multiply-add lane in an FHE accelerator contains one.Explained in: the FHE Accelerator Simulators glossary
HBM and key streaming
Off-chip high-bandwidth memory. Evaluation keys and DFT plaintexts are too big to keep on chip, so they stream from HBM, which is why bootstrapping is often memory-bound.Explained in: the FHE Accelerator Simulators glossary
NTT-, MAC-, memory- and power-bound
The simulator's verdict: the most-utilised resource sets the bound (HBM, NTT units, multiply-add lanes), or power when the power limit throttles a compute-bound design.Explained in: the FHE Accelerator Simulators glossary
Kernel-level trace
The contract between a scheme model or compiler and the hardware model: HE operations with their inputs, outputs, keys, plaintexts, levels and primitive kernels.Explained in: the FHE Accelerator Simulators glossary
Scratchpad, LRU and spills
The accelerator's on-chip SRAM, modelled as a least-recently-used store of ciphertexts, keys and plaintexts. A miss loads from HBM; evicting a live ciphertext writes it back (a spill).Explained in: the FHE Accelerator Simulators glossary
Calibration and out-of-sample prediction
Fit as few model coefficients as possible to measurements (here, one rate to an OpenFHE multiplication), then judge the model on quantities it was not fitted to.Explained in: the FHE Accelerator Simulators glossary
Key switching and evaluation keys
After a multiplication (relinearisation) or a rotation, part of the ciphertext is under the wrong key; key switching fixes it using an evaluation key (78–210 MiB for the N = 216 sets used here). It dominates the cost of FHE.Explained in: the FHE Accelerator Simulators glossary
Design sweeps and Pareto fronts
Simulate a grid of designs; a design is Pareto-optimal if no other is at least as good on every metric (latency, energy, and now area) and better on one. The front is the set of sensible choices; everything else is dominated.Explained in: the FHE Accelerator Simulators glossary
Min-KS, seeded keys and on-the-fly plaintexts
Algorithmic ways to cut key and plaintext traffic: rotate iteratively by one step so a BSGS loop reuses one key (Min-KS, from ARK), regenerate half of each key from a seed, and generate DFT plaintexts on chip.Explained in: the FHE Accelerator Simulators glossary
ENOB and exact rounding
Effective number of bits: the real precision of an analogue chain. Rounding to the exact integer needs the error below one half, which sets a minimum ENOB for each block size and digit width.Explained in: the FHE Accelerator Simulators glossary
Dynamic power manager against worst-case clocking
Worst-case clocking picks one clock at which everything at full rate fits the TDP; a dynamic manager gives each kernel the highest clock that fits the power actually free at that moment.Explained in: the FHE Accelerator Simulators glossary
Silicon area and the area model
Die area in mm² per component (NTT, MAC and permutation units, scratchpad, uncore, HBM PHYs). The simulator's model is at 7 nm, calibrated from ARK's published breakdown and a CACTI sweep, and illustrative; area never changes a simulated time or energy.Explained in: the FHE Accelerator Simulators glossary
Die yield and silicon cost
The fraction of dies with no fatal defect: Poisson e−A·D0 or Murphy's model, which is kinder to large dies. With dies per wafer it turns area into cost per good die, which grows faster than area.Explained in: the FHE Accelerator Simulators glossary
RNS polynomials and limbs
The ciphertext modulus Q is a product of word-sized primes, so each polynomial is stored as one row ('limb') of N residues per prime. A ciphertext at level ℓ has ℓ+1 limbs; all arithmetic stays in 64-bit words.Explained in: the FHE Accelerator Simulators glossary
Bootstrapping
Evaluating decryption homomorphically to turn a level-0 ciphertext into one at a high level: ModRaise, CoeffToSlot, EvalMod, SlotToCoeff. It costs hundreds of key switches and gigabytes of keys.Explained in: the FHE Accelerator Simulators glossary
Fourier optics and the 4f system
A lens Fourier-transforms the light field; two lenses with a mask between them multiply in the frequency domain, which is a convolution. The transform is free; the converters and lasers around it are not.Explained in: the FHE Accelerator Simulators glossary
HEIR
Google's MLIR-based FHE compiler: it chooses parameters, packing, rotations, levels and bootstrap placement, and emits code for libraries or hardware. Its ckks-dialect output drives this simulator.Explained in: the FHE Accelerator Simulators glossary