Presentations in This Series
- Rust for Simulation Engineers →Why Rust suits discrete-event simulators, taught through a real kernel: a BinaryHeap event list with SimPy's tie-breaking, events as enums, ownership in an event loop (arenas and indices, not references), SimPy processes as explicit state machines, traits for cost models, error handling, rayon sweeps, and the cargo, clippy, rustfmt, proptest and criterion toolchain.
- Rust and Python: Porting a Simulator Core with PyO3 →Moving the core of a SimPy simulator into Rust without changing a single answer: when a rewrite pays and when it does not, PyO3 and maturin, what should cross the boundary, releasing the GIL, everything it took to be bit-exact (operation order, SimPy's event order, fsum, Python's random stream, a JSON-parsing trap), differential and golden tests, and measured speed-ups.
- C++ Performance Models with SystemC TLM-2.0 →An FHE accelerator tile in C++17 and SystemC, driven by the SimPy simulator's own traces: modern C++ for models, the SystemC kernel and delta cycles, TLM-2.0 sockets and payloads, approximately- and loosely-timed coding, temporal decoupling and the quantum, why tie-breaks are part of the model, op-by-op agreement with SimPy, GoogleTest and sanitizers.
- Modelling Memory Systems: DRAM and HBM →Why bandwidth times efficiency fails, and a command-level model that derives the efficiency instead: banks, rows and bank groups, the timing parameters, hits, misses and conflicts, FCFS and FR-FCFS, page policy, address mapping, refresh and tFAW, load and latency, validation four ways (including against DRAMsim3), and the model plugged into the FHE simulator.
- Verification Bridge: cocotb, Verilator and Golden Models →A SystemVerilog NTT core verified in both directions: the FHE simulator's NTT as the golden model, a bit-accurate scoreboard, transactors, constrained-random stimulus, functional coverage and the crosses that make a bug observable, a seeded bug that uniform stimulus misses, Verilator and Icarus in CI, and RTL cycle counts calibrating the simulator.
- Testing Frameworks for Simulators →A test strategy for a simulator and the frameworks that implement it: pytest in depth (fixtures, parametrize, markers, conftest, plugins, xdist), Hypothesis including stateful tests, golden tests and re-blessing, mutation testing, coverage, GoogleTest, cargo test and proptest, cocotb, and putting it all in CI. Every example runs.
- Jenkins for Hardware and Simulation Teams →Continuous integration for simulators, from real pipelines that ran: declarative and scripted Jenkinsfiles, agents and labels for tools and licences, matrix builds as parameter sweeps, shared libraries, JUnit and coverage reports, nightly regressions, exact, speed and drift gates, credentials, the declarative linter, and how Jenkins and GitHub Actions divide the work.
- Jira and Engineering Metrics →The tracker as a simulation team's data model: issue types, workflows and status categories, versions, sprints and Kanban, JQL with worked queries and its traps, Git and Jenkins integration, traceability from requirement to build, the metrics worth reporting, flow metrics and Little's law from a workflow history, and evidence for sign-off.
- Specifications, Requirements and Test Plans →Writing down what a simulator must do, and proving it does: requirement quality, EARS patterns with an interactive checker, non-functional requirements for simulators, requirements capture for mission-mode software, the V-model, a traceability matrix generated from the test run (which found a real gap), and worked templates for a specification, a test plan and a performance report.
- From PyTorch and ONNX to an Accelerator Model →A working front end for an accelerator simulator: real model configurations (Llama-3-8B and -70B, Mistral, Qwen, GPT-2) traced without weights on the meta device and under fake tensors, torch.export, a torch.compile backend and an ONNX graph walk, all four agreeing exactly on the arithmetic, checked against closed forms and PyTorch's FLOP counter, costed on a roofline, with operator coverage counted three ways.
- Performance Analysis of Simulators and Systems →Measuring before optimising, on the simulators in this series: the USE method and the profile-first loop, Linux perf and what to do when it is locked down, sampling profilers and interactive flame graphs from py-spy, instruction counts with cachegrind (Rust against Python), a hot spot found in the metrics code and fixed with identical answers, benchmark distributions, and regression gates designed from measured noise.
- Measurement Tools and Methods →A catalogue of the tools and methods the simulation series measure with, grouped by what they measure, each with its principle, overhead, accuracy and pitfalls, when to use it, and where this GitHub uses it: profilers and flame graphs, benchmarks, simulated caches and hardware counters, GPU and CPU energy telemetry, trace viewers and framework profilers, RTL coverage and power flows, DRAM simulators, the CACTI, McPAT and Accelergy estimators (CACTI built and swept here), the statistics of simulation output, and agreement checks. Overheads measured on this series' own simulators.
- Power, Performance and Area: the Architect's Trade-offs →The three axes an architect trades: performance (latency and throughput), power (dynamic, static, DVFS, dark silicon) and area (SRAM, logic, PHYs and wires, estimated before layout), and why area is cost (dies per wafer, Poisson and Murphy yield, the reticle). Composite metrics (perf/W, perf/mm², EDP, TCO), Pareto fronts, a live explorer on FHE_Accelerator_Sim's new area model, a worked three-way trade-off, and the other trade-offs the simulation series make, named.
- An Accelerator Model in SimPy, End to End →A real graph from torch.export or ONNX run through an event-driven model of a tiled accelerator: off-chip memory, interconnect, DMA engines, an on-chip buffer with back-pressure, a compute array and a vector unit. Lowering to tiles, a cycle-approximate timing model, the roofline, Little's law and stall attribution, a timeline and hot-spot report, a cycle-stepped twin, how gem5 and SST are structured, a bit-identical C++ fast path with pybind11, ONNX Runtime execution providers, and an FHE-style NTT workload.
Companion Code
- </>Rust_DES_Kernel →A discrete-event kernel in Rust and a bit-exact Rust port of Disaggregated_Inference_Sim's engine, cost model and metrics, exposed to Python with PyO3. Differential and golden tests against the Python simulator, proptest invariants, an M/D/1 check against theory, criterion benchmarks, measured speed-ups, GitHub Actions and a Jenkins pipeline with JUnit, coverage and a performance gate.
- </>RTL_CoSim_NTT →A Barrett multiplier, NTT butterfly and P-lane NTT core in SystemVerilog, verified with cocotb on Verilator and Icarus against FHE_Accelerator_Sim's NTT as the golden model. Constrained-random stimulus with functional coverage crosses, a seeded bug that uniform stimulus misses, an exact cycle model, a toggle-count power proxy, and the measured NTT efficiency fed back into the FHE simulator.
- </>Memory_System_Sim →A command-level DRAM/HBM timing simulator: bank state machines, JEDEC-style timing constraints, address mapping with XOR hashing, FCFS and FR-FCFS, refresh. Checked by closed forms, an independent protocol checker on Hypothesis traces (which found a real bug) and DRAMsim3, and plugged into FHE_Accelerator_Sim as an optional HBM model.
- </>SystemC_Accelerator_Model →A SystemC TLM-2.0 model of an FHE accelerator tile, approximately and loosely timed, driven by FHE_Accelerator_Sim's trace format. An explicit tie-break arbiter makes it agree with the SimPy model op by op; LT quantum trade-offs measured; GoogleTest with sanitizers.
- </>Torch_Sim_Frontend →PyTorch and ONNX models turned into operator traces four ways (a dispatch trace on the meta device or under fake tensors, torch.export, a torch.compile backend, an ONNX graph walk) and costed on a roofline or offload model, or run end to end through a SimPy accelerator model (memory, interconnect, DMA, buffer with back-pressure, compute array, vector unit) with utilisation, stall and hot-spot reports, timelines, a bit-identical C++ fast path (pybind11), a cycle-stepped twin, an NTT workload and execution-provider partitioning. Runs Llama-3-70B without weights; the four routes agree exactly with each other, with closed forms and with PyTorch's FLOP counter. An EARS specification with a generated traceability matrix; GitHub Actions and Jenkins.
How this series relates to its sisters. LLM Inference Simulators teaches how to build, validate and accelerate a simulator; FHE Accelerator Simulators applies that to an FHE accelerator, and Fourier Optics for Inference extends the LLM simulator with heterogeneous pools and an optical transform engine. This series covers the engineering practice around them: the languages a simulator's fast core is written in, the hardware it must agree with, and the tests, pipelines, trackers and specifications that make its answers trustworthy. Interview-style questions on each topic are in the Interview_* repositories linked from every deck.
New to simulation? Start with Introduction to Simulation: the levels engineers simulate at, from field solvers and SPICE to RTL, architecture and system models, and the methods they share, with links into all three series.
Glossary: Concepts and Where They Are Explained
Every concept the decks rely on, with a short explanation and links to the slides that explain it in depth. Concepts that the sister series already explain are listed at the end with links to their glossary entries rather than repeated here. Each deck's contents slide links to the entries it uses.
Simulation kernels in Rust
- BinaryHeap and Reverse
- Rust's standard priority queue is a max-heap; wrapping entries in
std::cmp::Reversemakes it the min-heap an event list needs, so the earliest event pops first. Push and pop are O(log n).Explained in: SimEng 01 · the event list - Total ordering of floats (total_cmp)
f64is only partially ordered because NaN compares unequal to everything.f64::total_cmpimplements IEEE 754 totalOrder, which lets float times key a heap.Explained in: SimEng 01 · ordering floats, breaking ties- Deterministic tie-breaking
- Events at equal times are ordered by priority (SimPy's URGENT before NORMAL) and then by a sequence number (first scheduled, first served), so the same inputs always give the same event order and the same answers.Explained in: SimEng 01 · the ordering rule · SimEng 01 · interactive event queue · SimEng 03 · ties in SystemC · LLM Inference Simulators 02 · determinism pitfalls
- Events as enums; exhaustive match
- Events are plain data (an
enumof small, copyable variants) and the model handles them in amatch; the compiler rejects amatchthat forgets a variant.Explained in: SimEng 01 · events as enums, models as traits - Ownership and borrowing
- Every value has one owner; references borrow it, either shared (many readers) or exclusive (one writer), never both at once. The compiler enforces this, which rules out dangling references and data races.Explained in: SimEng 01 · ownership in an event loop · The Rust Programming Language, chapter 4
- Arenas and indices
- The simulation owns all requests and instances in vectors; events and queues hold indices into them. Borrows last one statement, so the borrow checker is satisfied without
Rc<RefCell>, and indices can be logged and replayed.Explained in: SimEng 01 · ownership in an event loop - Processes as state machines
- A SimPy process is a generator that sleeps at each
yield. Without stable generators, Rust names those sleep points as states and lets each event resume the component from where it slept.Explained in: SimEng 01 · from SimPy processes to state machines - Traits; static and dynamic dispatch
- A trait is an interface. Generic code (
<C: CostModel>) is compiled once per implementation and can inline calls; a trait object (Box<dyn CostModel>) chooses the implementation at run time through a vtable.Explained in: SimEng 01 · traits for pluggable cost models - Result, ? and panic
- Expected failures (a bad configuration) return
Resultand propagate with?; broken invariants (an event in the past)panic!, stopping the run at the point of the bug.Explained in: SimEng 01 · errors - rayon, Send and Sync
- rayon turns
iter()intopar_iter()on a work-stealing thread pool. TheSendandSynctraits let the compiler prove that what each thread touches is safe to share, so a data race is a compile error.Explained in: SimEng 01 · fearless parallel sweeps - cargo, clippy and rustfmt
- Rust's build tool, linter and formatter.
cargo clippy -- -D warningsfails the build on any lint;cargo fmt --checkenforces one layout.Explained in: SimEng 01 · the toolchain - criterion benchmarks
- A Rust benchmarking library: warm-up, many samples, confidence intervals and comparison with the previous run, so performance changes are measured rather than guessed.Explained in: SimEng 12 · criterion · SimEng 01 · the toolchain · SimEng 06 · a performance gate in CI
Rust and Python
- Amdahl's law for a port
- Speeding up part of a program speeds up the whole only in proportion to that part's share of the time. After porting a simulator's engine, the Python code around it can dominate.Explained in: SimEng 02 · interactive: what should you port? · SimEng 02 · what crosses the boundary · LLM Inference Simulators 08 · Amdahl's law for simulators
- PyO3 and maturin
- PyO3 generates the CPython bindings for Rust functions marked
#[pyfunction]; maturin builds and installs them as a Python wheel, alongside any pure-Python code.Explained in: SimEng 02 · PyO3 and maturin in practice · SimEng 02 · project layout - The language boundary
- Every call from Python into Rust converts its arguments and results. A good boundary is coarse (one call per simulation) and plain (numbers, tuples, strings), and its cost is measured.Explained in: SimEng 02 · what crosses the boundary
- The GIL and releasing it
- CPython's global interpreter lock lets one thread run Python at a time. Rust code that touches no Python objects can release it (
py.detach, formerlyallow_threads), so Python threads run simulations in parallel.Explained in: SimEng 02 · releasing the GIL - Bit-exact parity
- Two implementations that produce identical floating-point results, bit for bit. It requires the same operations in the same order, the same event order and the same library functions, and it turns any discrepancy into a located bug.Explained in: SimEng 02 · the parity checklist · SimEng 02 · three traps that cost one ulp · SimEng 03 · the same order in C++
- Correctly rounded summation (math.fsum)
math.fsumreturns the exactly rounded sum of its inputs;statistics.fmeanuses it. A plain loop accumulates rounding error, so a port must reproduce the algorithm, not just add.Explained in: SimEng 02 · three traps that cost one ulp- MT19937 and Python's random module
- CPython's generator is the Mersenne Twister, seeded with
init_by_array;random()combines two 32-bit outputs into a 53-bit float, and the variates are written in Python, so the whole stream can be reproduced exactly.Explained in: SimEng 02 · reproducing Python's random numbers - abi3 wheels
- Extension modules built against CPython's stable ABI work on every later Python version, so one wheel per platform suffices. The free-threaded build has its own ABI and is not covered.Explained in: SimEng 02 · packaging and CI
C++ and SystemC models
- Modern C++ for models
- Plain C++17 value types, RAII and lambdas for the model's data and arithmetic, kept free of the simulation kernel so it can be unit-tested alone and ported from a reference in the same floating-point order.Explained in: SimEng 03 · modern C++ for models
- The SystemC kernel and delta cycles
- SystemC (IEEE 1666) runs processes in evaluate/update cycles. Zero-time notifications start another delta cycle at the same simulated time; time advances only when nothing more is runnable. Time is an integer count of a global resolution.Explained in: SimEng 03 · the SystemC kernel
- Process order and explicit tie-breaks
- IEEE 1666 does not specify the order in which processes runnable at the same instant execute, so simultaneous requests for a resource need a tie-break written into the model if results are to be reproducible across kernels.Explained in: SimEng 03 · ties are part of the model · SimEng 01 · ordering floats, breaking ties
- TLM-2.0 generic payload and sockets
- TLM-2.0 replaces pin-level signals with function calls through initiator and target sockets, carrying a generic payload (command, address, data pointer and length, response) and optional extensions, so models from different sources interoperate.Explained in: SimEng 03 · payloads, sockets and coding styles
- Approximately timed (AT) and the base protocol
- AT models use non-blocking transport with four phases (BEGIN_REQ, END_REQ, BEGIN_RESP, END_RESP), each at its own simulated time, so contention resolves in time order. Accurate, at the cost of a context switch per phase.Explained in: SimEng 03 · the four-phase protocol
- Loosely timed (LT) and temporal decoupling
- LT models use one blocking b_transport call whose delay argument carries timing. Processes run ahead of simulated time by up to a global quantum and synchronise rarely: faster, but resources get booked out of time order.Explained in: SimEng 03 · temporal decoupling and the quantum · SimEng 03 · interactive: the quantum trade-off
- Global quantum
- The longest a loosely-timed process may run ahead of the kernel's time before it must synchronise (tlm_quantumkeeper). Larger quanta mean fewer context switches and larger timing errors.Explained in: SimEng 03 · interactive: the quantum trade-off
- One elaboration per process
- A SystemC program elaborates its module hierarchy and calls sc_start once; tests that each run a simulation therefore fork a child per test.Explained in: SimEng 03 · GoogleTest and sanitizers
Memory systems
- “Bandwidth × efficiency”
- Modelling memory as peak bandwidth times a fixed derating factor. The achieved fraction of peak depends on access pattern, controller, address mapping and load, so a single factor is right for at most one of them.Explained in: SimEng 04 · why bandwidth times efficiency fails · SimEng 04 · plugging it into the FHE simulator
- Channels, ranks, bank groups, banks and rows
- A DRAM channel has ranks of chips; each rank has bank groups of banks; each bank is an array of rows read through one row buffer. Banks work in parallel; accesses within a bank group must be further apart than across groups.Explained in: SimEng 04 · how DRAM is organised
- DRAM commands and timing parameters
- ACT, RD/WR, PRE and REF, separated by minimum gaps such as tRCD, tRP, tCL, tRAS, tRC, tRRD, tCCD, tWTR, tWR and tRTP, defined per speed bin by the JEDEC standards.Explained in: SimEng 04 · commands and timing parameters
- Row hits, misses and conflicts
- A hit finds its row open (CL); a miss finds the bank closed (tRCD + CL); a conflict finds another row open (tRP + tRCD + CL). A controller's main job is turning conflicts into hits.Explained in: SimEng 04 · row hits, misses and conflicts
- FCFS and FR-FCFS scheduling
- FCFS serves requests in arrival order. FR-FCFS (Rixner et al., 2000) issues ready row hits first, then the oldest ready command, overlapping row opening in some banks with transfers in others.Explained in: SimEng 04 · scheduling and page policy
- Open and closed page policy
- Open page leaves a row open after an access, betting on a hit; closed page precharges at once (auto-precharge), betting on a miss rather than a conflict.Explained in: SimEng 04 · scheduling and page policy
- Address mapping and XOR bank hashing
- Which address bits select channel, bank group, bank, row and column decides which patterns spread across banks. XOR hashing mixes row bits into the bank index so row-sized strides stop hitting one bank.Explained in: SimEng 04 · address mapping
- Refresh (tREFI, tRFC)
- Every tREFI a rank closes its banks and refreshes for tRFC, costing a stream about tRFC/tREFI of its bandwidth and setting the latency tail.Explained in: SimEng 04 · refresh and the four-activate window · SimEng 04 · load and latency
- The four-activate window (tFAW)
- At most four ACTs per rank in any window of tFAW, a limit on peak current. It caps random traffic, which needs one ACT per access, at 4 (BL/2) / tFAW of peak.Explained in: SimEng 04 · refresh and the four-activate window
- Independent protocol checker
- A checker that replays a controller's command log and verifies every timing rule directly, sharing no code with the controller; with property-based traces it is the oracle that finds scheduling bugs.Explained in: SimEng 04 · how we know the model is right
- Cross-checking against another simulator
- Running identical traces and parameters through an independent simulator (here DRAMsim3). Agreement where the physics dominates and disagreement where the designs differ both carry information.Explained in: SimEng 04 · cross-checked against DRAMsim3
Verification bridge
- Golden model and bit-accurate model
- The golden model says what the hardware must compute (here the FHE simulator's NTT); a bit-accurate model mirrors the RTL's internal steps so a scoreboard can check them.Explained in: SimEng 05 · two golden models and a scoreboard
- Scoreboard
- The testbench component that predicts each expected output from the model and compares it with what the monitor observed from the design.Explained in: SimEng 05 · two golden models and a scoreboard
- Transactors (driver and monitor)
- Testbench components that turn transactions into pin activity and back. A ready/valid handshake must be sampled in the cycle it happens, not after the edge that commits it.Explained in: SimEng 05 · transactors · SimEng 05 · cocotb in one page
- Constrained-random stimulus
- Random stimulus biased towards corners (0, 1, q - 1, near-maximal operands) while still covering the whole range, so rare situations occur often enough to test.Explained in: SimEng 05 · constrained-random stimulus · SimEng 05 · interactive: coverage closure
- Coverage crosses and observability
- A cross counts combinations of bins. Reaching a corner is not enough if later logic masks a bug there; the cross with the condition that makes the bug visible is the bin that matters.Explained in: SimEng 05 · functional coverage and crosses · SimEng 05 · a seeded bug
- Seeded bugs
- Deliberately injected faults (here a missing Barrett correction behind a define) used to check that a testbench can catch them: mutation testing for hardware.Explained in: SimEng 05 · a seeded bug: caught or escaped
- Verilator and Icarus Verilog
- Verilator compiles synthesisable SystemVerilog to fast two-state C++ and lints it; Icarus is a four-state event-driven simulator. Running tests on both catches simulator-dependent behaviour.Explained in: SimEng 12 · Verilator coverage · SimEng 05 · Verilator, Icarus and CI for RTL
- Cycle models from RTL
- A closed-form cycle count (here log2 n (n/2P + 6)) fitted to and checked against every RTL measurement, so it can be trusted to extrapolate to sizes too large to simulate.Explained in: SimEng 05 · RTL cycle counts and a cycle model
- Calibrating a simulator from RTL
- Derating the architecture simulator's ideal unit throughput by the efficiency the RTL achieves; it can change which resource the simulator calls the bottleneck.Explained in: SimEng 05 · feeding the RTL back into the simulator
- Switching activity (toggle counts)
- The number of bit changes per cycle on a design's nets or registers; activity times capacitance times V squared times f gives dynamic power, so toggle counts are an early power proxy.Explained in: SimEng 12 · toggle counts and switching activity · SimEng 05 · switching activity as a power proxy
Test strategy and frameworks
- pytest fixtures and scopes
- Functions that provide what a test needs, requested by naming them as parameters, defined in
conftest.pyfor sharing, and cached per function, module or session.Explained in: SimEng 06 · pytest I - parametrize and markers
@pytest.mark.parametrizeruns one test body over many inputs; markers label tests (slow, nightly) so-mcan select them; a strict xfail records a known limitation.Explained in: SimEng 06 · pytest II- pytest-xdist
- A pytest plugin that runs tests in parallel worker processes (
-n auto). Tests must be independent: no shared files or global state.Explained in: SimEng 06 · pytest II - Shrinking
- When a property-based test fails, the framework searches for simpler inputs that still fail and reports the simplest it finds, which is usually a readable reproduction of the bug.Explained in: SimEng 06 · properties and shrinking · SimEng 06 · stateful testing
- Stateful property testing
- Hypothesis generates random sequences of operations (rules) and compares the system with a simple model after every step; failures shrink to a minimal sequence.Explained in: SimEng 06 · Hypothesis: stateful testing
- proptest
- Rust's property-based testing library: strategies generate inputs, failures shrink, and failing seeds are saved in
proptest-regressions/and replayed first.Explained in: SimEng 06 · cargo test and proptest · SimEng 01 · testing in Rust - Golden tests and re-blessing
- Store the output of a reference run and fail when today's differs; re-bless (rewrite the stored output) deliberately, with the diff reviewed alongside the change that caused it.Explained in: SimEng 12 · golden, differential and correlation · SimEng 06 · golden tests and re-blessing · SimEng 02 · golden files across languages
- GoogleTest
- The common C++ test framework:
TEST, fixtures withTEST_F, value-parameterisedTEST_P,EXPECT_*(continue) andASSERT_*(stop), death tests.Explained in: SimEng 06 · GoogleTest for C++ models · SimEng 03 · GoogleTest and sanitizers - Sanitizers (ASan, UBSan)
- Compiler instrumentation that turns memory errors (AddressSanitizer) and undefined behaviour (UndefinedBehaviorSanitizer) into immediate failures with a stack trace. Build C++ tests with them.Explained in: SimEng 06 · GoogleTest for C++ models · SimEng 03 · GoogleTest and sanitizers
- cocotb
- A framework for writing RTL testbenches in Python coroutines that drive an HDL simulator (Icarus, Verilator and others), so a Python model can check the hardware.Explained in: SimEng 06 · cocotb: Python testbenches for RTL · SimEng 05 · cocotb in one page
- Functional coverage
- Bins describing the situations a testbench must exercise (a wrap-around, a boundary value), counted during the run; an empty bin is a hole in the verification, whatever the code coverage says.Explained in: SimEng 12 · functional and code coverage · SimEng 06 · cocotb · SimEng 05 · functional coverage and crosses
Test quality and CI
- Mutation testing
- Make small deliberate bugs (mutants) in the code and run the tests against each: a killed mutant was noticed, a surviving one was not. The mutation score measures how much the tests actually check.Explained in: SimEng 12 · mutation testing · SimEng 06 · mutation testing · SimEng 06 · interactive: mutation score versus coverage
- Equivalent mutants
- Mutants that change the code but not its behaviour, so no test can kill them; survivors must be triaged rather than all chased.Explained in: SimEng 06 · coverage, and what it does not tell you
- Line, branch and region coverage
- The fraction of lines, branch outcomes or code regions that the tests execute. Low coverage is a warning; high coverage does not show that results are checked.Explained in: SimEng 12 · code coverage tools · SimEng 06 · coverage, and what it does not tell you · SimEng 06 · interactive
- MC/DC
- Modified condition/decision coverage: each condition in a decision must be shown to change the outcome independently. DO-178C requires it for the most critical avionics software.Explained in: SimEng 06 · coverage, and what it does not tell you
- JUnit XML and Cobertura reports
- The de facto interchange formats for test results and coverage, written by pytest, cargo-nextest, GoogleTest and cocotb, and read by Jenkins and most CI systems.Explained in: SimEng 06 · running it all in CI · SimEng 07 · test and coverage reports
- Performance-regression gate
- A CI step that benchmarks the build and fails it if a benchmark is slower than a stored baseline by more than a set margin.Explained in: SimEng 06 · running it all in CI · SimEng 01 · criterion · SimEng 05 · an exact cycle-count gate · SimEng 07 · gating on performance and accuracy · SimEng 11 · regression detection in CI
- Flaky tests
- Tests that pass or fail without a code change. In a deterministic simulator they point to shared state, an unseeded random number or a wall-clock dependency, and should be fixed, not retried.Explained in: SimEng 06 · running it all in CI
Pipelines and Jenkins
- Jenkins: controller, agents, plugins
- A self-hosted automation server. The controller (web UI, build queue, credentials, plugins) schedules pipeline runs onto the executors of agents; almost every feature, Pipeline itself included, is a plugin.Explained in: Introduction to Jenkins · architecture: controller, agents, executors · Introduction to Jenkins · plugins, the update centre and JENKINS_HOME · SimEng 07 · agents, labels and the tools on them
- GitHub Actions: workflows, jobs, runners
- GitHub's built-in automation. An event (a push, a pull request, a schedule) starts a workflow, a YAML file in
.github/workflows/; its jobs run in parallel on fresh runners, each step runs a shell command or a reusable action, and each job reports a check that a ruleset can require before a merge.Explained in: Introduction to GitHub Actions · events, workflows, jobs, steps and runners · Introduction to GitHub Actions · required checks, shown blocking · SimEng 07 · Jenkins and GitHub Actions - Declarative and scripted pipelines
- A Jenkinsfile is either declarative (a fixed structure of agent, parameters, stages and post conditions that Jenkins validates and draws) or scripted (Groovy code inside
node { }, with loops and try/catch).Explained in: SimEng 07 · declarative and scripted pipelines - Agents and labels
- The controller schedules work; agents run it. A stage asks for an agent by label (tools, licences, capacity), so a pipeline states what it needs rather than which machine to use.Explained in: SimEng 07 · agents, labels and the tools on them
- Matrix builds
- A declarative
matrixruns the same stages for every combination of its axes as parallel branches, minus exclusions: a parameter sweep in CI.Explained in: SimEng 07 · matrix builds - A Git repository of Groovy that pipelines load with
@Library; each file invars/becomes a pipeline step, so many repositories share one definition of their stages.Explained in: SimEng 07 · shared libraries - Nightly regressions and build parameters
- Fast checks on every change, expensive sweeps nightly, selected by a build parameter. A parameter default computed from
envbecomes the textnull, and once a job has parameters a plain/buildrequest is rejected (HTTP 400).Explained in: SimEng 07 · nightly regressions and parameters - Exact, speed and drift gates
- Three ways a CI gate compares a build with a baseline: exact behaviour (any change fails), speed (within a margin, best of several runs) and drift (reported, not failed).Explained in: SimEng 07 · gating on performance and accuracy
- Credentials and masking
- Secrets stored on the controller and referred to by ID;
withCredentialsbinds one for a block and masks its value in the log.Explained in: SimEng 07 · credentials - The declarative linter
- Jenkins validates a declarative Jenkinsfile without running it (
/pipeline-model-converter/validate), catching structural mistakes such as a misspeltpostcondition.Explained in: SimEng 07 · the real runs
Tracking and engineering metrics
- The issue model
- Epics group stories, tasks and bugs, which may have sub-tasks; each item has fields (status, assignee, fix version, sprint, components, labels), links and a full change history.Explained in: SimEng 08 · the issue model
- Workflows and status categories
- A workflow is a state machine of statuses and allowed transitions; every status belongs to one of three categories (To Do, In Progress, Done) that queries and boards can rely on.Explained in: SimEng 08 · workflows, transitions and status categories
- Versions, sprints and Kanban
- A fix version names the release that will contain a change; Scrum plans work in fixed sprints, Kanban pulls it continuously under WIP limits.Explained in: SimEng 08 · versions, sprints and Kanban
- JQL
- Jira Query Language: clauses of field, operator and value joined by AND, OR and NOT, with functions and ORDER BY.
!=does not match empty fields; WAS and CHANGED search the history of six fields.Explained in: SimEng 08 · JQL: the query language · SimEng 08 · interactive: worked JQL queries - Smart commits and the development panel
- Work-item keys in branch names, commit messages or pull-request titles link the development work to the item; smart commits can also comment, log time and run a transition.Explained in: SimEng 08 · Git and Jenkins integration
- Cycle time, throughput and WIP
- Flow metrics computed from the workflow history: how long items take from start to done (as percentiles), how many finish per week, how many are in progress, and which have been in progress too long.Explained in: SimEng 08 · interactive: flow metrics · SimEng 08 · the metrics to report
- Cumulative flow diagram
- The number of items in each status category on each day, stacked; widening bands show work accumulating in a stage.Explained in: SimEng 08 · interactive: flow metrics
- Goodhart's law
- A measure that becomes a target stops being a good measure; prefer metrics that tools compute from records, and report distributions rather than single numbers.Explained in: SimEng 08 · the metrics a simulation team should report
Specifications and requirements
- Requirement quality
- A good requirement is necessary, unambiguous, singular, feasible and verifiable, among the characteristics INCOSE's guide and ISO/IEC/IEEE 29148 describe; writing rules ban vague terms, escape and open-ended clauses.Explained in: SimEng 09 · what makes a good requirement · SimEng 09 · interactive: check a requirement
- EARS
- The Easy Approach to Requirements Syntax: ubiquitous, event-driven (When), state-driven (While), unwanted-behaviour (If ... then) and optional-feature (Where) patterns, and combinations of them.Explained in: SimEng 09 · EARS: five patterns · SimEng 09 · interactive
- Verification methods
- Each requirement names how it will be verified: test, analysis, inspection or demonstration. A requirement no method could fail is not verifiable.Explained in: SimEng 09 · what makes a good requirement
- Non-functional requirements for simulators
- How well a simulator must work: accuracy against a named reference and tolerance, determinism, speed, capacity and stated fidelity boundaries.Explained in: SimEng 09 · functional and non-functional requirements
- Mission-mode software
- Here: the software that runs on the product in its operational mode, as opposed to test, bring-up, calibration or diagnostic software. Its requirements must capture modes, fault handling and resource budgets.Explained in: SimEng 09 · requirements capture for mission-mode software
- The V-model; verification and validation
- Each level of specification is verified by the level of test opposite it. Verification asks whether it was built right; validation, whether the right thing was built.Explained in: SimEng 09 · the V-model
- Traceability matrix
- The links from each requirement to the tests and evidence that verify it, checked in both directions; generated from the test run, a missing link is a visible gap.Explained in: SimEng 09 · traceability, generated from the test run · SimEng 08 · requirement, work item, test, build
- Test plans and test oracles
- A test plan states scope, approach, pass/fail and exit criteria, environment, deliverables and risks; its core is the oracle each level uses to decide that a result is right.Explained in: SimEng 09 · template 2: a test plan
- Generated performance reports
- A report written by a script from recorded runs (question, environment, method, validation, results, limitations), so it can be regenerated after any change.Explained in: SimEng 09 · template 3: a performance report
Framework front ends
- Operator trace
- The operators a model executes, in order, with each tensor's shape, element size, identity and whether it is a weight: the input a cost model needs from a front end.Explained in: SimEng 10 · one trace format, four front ends
- Fake tensors and traced code paths
- Tensors with shapes and a claimed device but no data. What a trace shows depends on that device (fused or decomposed attention) and on whether the library detects tracing and takes another path.Explained in: SimEng 10 · meta or fake tensors
- ONNX shape inference and constant folding
- Walking an exported graph needs shape values propagated through Shape and Reshape chains (
data_prop) and constant subgraphs folded as a runtime would; weights can be exported as typed inputs without their data.Explained in: SimEng 10 · front end 4: ONNX without the weights - Unfused and ideal-fusion bounds
- Costing a trace per operator assumes every operator reads and writes memory (an upper bound); assuming only matmul, attention and gathers touch memory gives a lower bound. Real compilers fall between.Explained in: SimEng 10 · costing the trace
- Operator coverage
- Which operators a cost model has rules for, and which the accelerator runs: counted by operators, FLOPs and time, because a device can hold nearly all the FLOPs and little of the time.Explained in: SimEng 10 · interactive: operator coverage · SimEng 10 · operator coverage as a metric
Performance analysis
- The USE method
- For every resource (CPUs, memory, disks, network, locks), check utilisation, saturation (work queued because the resource is busy) and errors: a checklist that finds system bottlenecks without guessing.Explained in: SimEng 11 · USE for resources, profile first for code · SimEng 11 · USE in practice
- The profile-first loop
- Measure a baseline, profile, form one hypothesis, change one thing, re-measure, and check the answers did not change; repeat.Explained in: SimEng 11 · two methods · SimEng 11 · closing the loop: sort once · LLM Inference Simulators 08 · step zero: profile
- Linux perf and perf_event_paranoid
- The Linux profiler: hardware event counts (perf stat), sampled call stacks (perf record) and annotated code. Unprivileged access is governed by the kernel.perf_event_paranoid setting.Explained in: SimEng 12 · Linux perf · SimEng 11 · Linux perf: counters and call stacks
- Sampling profilers (py-spy)
- A profiler that records the call stack at a fixed rate; time spent shows up as samples. py-spy samples a Python process from outside, with no changes to it.Explained in: SimEng 12 · py-spy and sampling profilers · SimEng 11 · sampling profilers and flame graphs · SimEng 11 · reading a profile
- Flame graphs
- Sampled stacks merged by common prefix: a frame's width is its share of samples, the top edge of each tower is where time is spent, and the x-axis is not time.Explained in: SimEng 12 · flame graphs, on- and off-CPU · SimEng 11 · sampling profilers and flame graphs · SimEng 11 · interactive: flame graphs
- Self time and total time
- Self time counts samples where a function is running; total time counts samples where it is anywhere on the stack. Total locates the phase, self locates the code.Explained in: SimEng 11 · reading a profile: self and total time
- Cachegrind and instruction counts
- Valgrind's tool that runs a program on a simulated CPU, counting every instruction and modelling the caches: slow, but deterministic, so small differences can be compared without timing noise.Explained in: SimEng 12 · cachegrind and callgrind · SimEng 11 · counting instructions with cachegrind
- Benchmark distributions
- Repeated timings form a distribution, often with several clusters: report the median, the median absolute deviation and a bootstrap confidence interval rather than a mean and standard deviation.Explained in: SimEng 11 · benchmarks are distributions
- A/A tests and regression-gate design
- Running a timing gate on unchanged code measures its false-alarm rate; the number of runs per side and the margin then set what it catches. The margin must sit well below the regression to be caught.Explained in: SimEng 11 · interactive: designing a regression gate · SimEng 11 · regression detection in CI
Measurement tools and methods
- cProfile and snakeviz (instrumenting profilers)
- Python's deterministic profiler hooks every call and return: exact call counts and call trees, but the per-call cost inflates small, frequent functions. snakeviz draws the saved profile.Explained in: SimEng 12 · cProfile and snakeviz · LLM Inference Simulators 08 · step zero: profile
- Hardware performance counters and multiplexing
- The CPU's performance-monitoring unit counts events (cycles, instructions, cache and branch misses) at almost no cost; with more events than counters the kernel rotates them and scales the counts, which adds error.Explained in: SimEng 12 · hardware performance counters · SimEng 11 · perf stat
- GNU time (/usr/bin/time -v)
- Prints what the kernel accounted for a finished command: wall and CPU time, maximum resident set size, page faults and context switches. Free, and the first look in the USE method.Explained in: SimEng 12 · GNU time · SimEng 11 · USE in practice
- Load generators and coordinated omission
- A load generator issues timestamped requests to a system under test (MLPerf's LoadGen defines the standard scenarios). A closed-loop generator that waits for each response hides stalls from its own measurements: coordinated omission.Explained in: SimEng 12 · load generators
- tracemalloc (Python heap profiling)
- Records the allocating traceback of every live Python memory block, so snapshots show which lines hold memory; expensive, and blind to memory that native extensions allocate.Explained in: SimEng 12 · tracemalloc
- nvidia-smi and NVML
- NVIDIA's driver library for GPU monitoring and its command-line front end: utilisation, clocks, memory, power and an energy counter. Power readings are averaged over sensor windows that can miss most of a run.Explained in: SimEng 12 · nvidia-smi and NVML
- NVIDIA DCGM
- The Data Center GPU Manager samples numbered fields (power, energy, and profiling metrics such as SM and tensor-pipe activity) on every GPU, continuously; some profiling metrics cannot be collected together and are multiplexed.Explained in: SimEng 12 · DCGM
- Zeus (ML.ENERGY)
- A Python library measuring time and energy over named code windows from NVML's energy counter (and RAPL for CPUs); the measurement behind the ML.ENERGY leaderboard.Explained in: SimEng 12 · Zeus
- RAPL (CPU energy counters)
- Energy counters per CPU domain (package, cores, DRAM) in model-specific registers, read through Linux powercap or perf's power events. Package energy, not wall energy; root-only on current kernels.Explained in: SimEng 12 · RAPL
- External power analysers and board telemetry
- Meters at the wall or on a supply rail that measure voltage and current directly: the reference other power readings are calibrated against, including everything in the box.Explained in: SimEng 12 · power analysers
- PyTorch profiler
- Records an event per PyTorch operator (and GPU kernel), with shapes and memory on request; aggregates by operator or exports a Chrome trace.Explained in: SimEng 12 · PyTorch profiler
- ONNX Runtime profiler
- With profiling enabled, ONNX Runtime writes a Chrome-trace event per graph node per run, with operator type, execution provider and shapes, for the graph after its own optimisations.Explained in: SimEng 12 · ONNX Runtime profiler · LLM Inference Simulators 09 · ONNX and ONNX Runtime
- NVIDIA Nsight Systems
- A system-wide tracer of CPU threads, CUDA API calls, GPU kernels and copies on one timeline, for finding the idle gaps on a GPU.Explained in: SimEng 12 · Nsight Systems
- DRAMsim3
- A cycle-accurate DRAM controller and device simulator (DDR, LPDDR, GDDR, HBM) configured with JEDEC timing; used here as the reference that Memory_System_Sim was cross-checked against.Explained in: SimEng 12 · DRAMsim3 · SimEng 04 · cross-checked against DRAMsim3
- Ramulator
- A cycle-level DRAM simulator that describes each standard as a generic state machine, so new standards and mechanisms are easy to add.Explained in: SimEng 12 · Ramulator
- McPAT
- An analytical power, area and timing model of multicore processors built from their structure, driven by a performance simulator's activity counts; useful for relative comparisons, with documented sources of error.Explained in: SimEng 12 · McPAT
- CACTI
- An analytical model of SRAM and DRAM arrays and caches: from capacity, organisation and technology it gives access time, energy per access, leakage and area. Built and swept over capacity at 22 nm in SimEng 12.Explained in: SimEng 12 · CACTI · SimEng 12 · CACTI on this machine
- Accelergy and Timeloop
- Timeloop searches the mappings of a tensor workload onto an accelerator and counts the actions at every level; Accelergy prices those actions in energy and area from per-component models.Explained in: SimEng 12 · Accelergy and Timeloop
- Analytic lower bounds
- Bounds on run time from counts alone (the roofline, the busiest resource's busy time): optimistic by construction, so a simulated result below one is a bug.Explained in: SimEng 12 · analytic bounds
- PyTorch's FLOP counter (FlopCounterMode)
- A dispatch mode that adds up FLOPs from per-operator formulas as a model runs, even on the meta device; operators without a formula (fused CPU attention) are silently missed.Explained in: SimEng 12 · FlopCounterMode · SimEng 10 · operator coverage as a metric
Power, performance and area
- PPA: power, performance and area
- The three quantities a chip design is judged on, traded together at every level from architecture to layout: more SRAM saves traffic and costs area, a higher clock raises performance and power, a bigger die has longer wires.Explained in: SimEng 13 · what PPA is, and who trades it · SimEng 13 · the other trade-offs, named
- Dynamic power (αCV²f) and leakage
- Dynamic power is αCV²f: activity, switched capacitance, supply voltage squared, clock. Static power is V · Ileak, drawn whether busy or not. DVFS lowers V and f together, so energy per operation falls roughly as V².Explained in: SimEng 13 · dynamic, static, DVFS and dark silicon
- Dark silicon
- Since supply voltage stopped scaling with feature size, a chip has more transistors than its power budget can switch at once, so part of it must be idle or slowed at any moment; specialised accelerators are one response.Explained in: SimEng 13 · dark silicon
- What sets die area
- SRAM bitcells and their periphery, logic, I/O and PHYs (which shrink little with the node), and wiring with its repeaters and registers. In ARK's 7 nm breakdown the scratchpad is over half the die and most of the NTT and permutation units is wire.Explained in: SimEng 13 · area: what sets it · SimEng 13 · estimating area before layout
- Dies per wafer
- π(d/2)²/A − πd/√(2A) for wafer diameter d and die area A: the wafer area divided by the die area, minus the partial dies lost round the edge.Explained in: SimEng 13 · why area is cost
- Poisson yield
- Y = e−A·D0: the chance a die of area A has no defect when defects (density D0) land independently and uniformly. Pessimistic for large dies.Explained in: SimEng 13 · yield models, interactive
- Murphy yield
- Y = ((1 − e−A·D0)/(A·D0))²: yield when the defect density varies across the wafer. Agrees with Poisson for small dies and is higher for large ones.Explained in: SimEng 13 · yield models, interactive
- Reticle limit and chiplets
- A lithography field is about 26 × 33 mm (858 mm²), the largest die one exposure prints. Larger designs, or ones past the yield knee, are split into chiplets, which pay for die-to-die links, packaging and assembly.Explained in: SimEng 13 · the reticle · SimEng 13 · when to split the die
- perf/W
- Throughput per watt, which for a fixed task equals 1 / energy per task: it ranks designs by energy alone and ignores speed.Explained in: SimEng 13 · composite metrics
- perf/mm² and perf/$
- Throughput per square millimetre of silicon, or per dollar of manufacturing cost: what a design delivers for its die. Both ignore energy, and perf/mm² also ignores yield.Explained in: SimEng 13 · composite metrics · SimEng 13 · the scratchpad sweep
- EDP and ED²P
- Energy-delay product E·t weights energy and speed equally; E·t² weights speed more and is roughly independent of supply voltage under DVFS, so it compares designs rather than operating points.Explained in: SimEng 13 · composite metrics
- Total cost of ownership (TCO)
- Purchase cost amortised over the service life plus the energy bill (energy × price × datacentre overhead): the figure an operator optimises, and one that needs prices rarely known before silicon.Explained in: SimEng 13 · composite metrics
- Pareto front and dominance
- Design A dominates B if it is no worse on every axis and better on one; the Pareto front is the set nothing dominates. Adding an axis (area) can turn a single optimum into a front.Explained in: SimEng 13 · Pareto fronts · SimEng 13 · the live PPA explorer · FHE glossary · design sweeps and Pareto fronts
- An area model calibrated from papers and CACTI
- FHE_Accelerator_Sim's ppa.py: per-unit areas from ARK's published 7 nm breakdown, SRAM from a CACTI sweep scaled to ARK's scratchpad, HBM PHYs and uncore. Illustrative: good for ratios between designs, not absolute mm².Explained in: SimEng 13 · estimating area before layout · SimEng 13 · pitfalls
Accelerator models in SimPy
- Lowering a graph to tiles
- Turning each operator of a trace into tiles, the unit an accelerator loads, computes and stores: GEMMs are blocked so operands and result fit the on-chip buffer, other operators are cut into byte slices. Tiling sets both the extra memory traffic and how much load, compute and store overlap.Explained in: SimEng 14 · from a graph to tiles · SimEng 14 · a bigger buffer made it slower
- im2col: a convolution as a GEMM
- A convolution is a matrix product once each output position's input patch is laid out as a row: M = output positions, K = input channels per group × kernel size, N = output channels per group. It lets one matrix engine run both.Explained in: SimEng 14 · from a graph to tiles
- Systolic array (output-stationary)
- A grid of R × C multiply-accumulate cells that pass operands to their neighbours; in the output-stationary dataflow each cell accumulates one output. Its time per tile is about ⌈m/R⌉⌈n/C⌉k cycles plus a fill and drain, so PE efficiency drops when m or n is not a multiple of the array.Explained in: SimEng 14 · the timing model
- Cycle-approximate and cycle-accurate models
- A cycle-approximate model computes a block's cycle count from a formula (or a few state changes); a cycle-accurate one evaluates every register on every cycle and matches RTL exactly. The first is fast enough to explore designs; the second is worth building when the question lies inside the cycles, or to calibrate the first.Explained in: SimEng 14 · cycle-approximate against cycle-accurate · SimEng 14 · process-based against cycle-based
- Stall attribution
- Splitting a run's time into computing and the waits between tiles, each wait named by what the next load was waiting for (load bandwidth, buffer space, an earlier operator's results). The parts sum to the latency, and the largest names the hot-spot: a sharper tool than utilisation.Explained in: SimEng 14 · finding the stalls · SimEng 14 · where the time goes, interactive
- Read-after-write through memory
- When every operator writes its results to off-chip memory and the next reads them back, the next operator cannot start loading until the previous one is stored: the pipeline drains at every operator boundary. Fusion, keeping results on chip, removes the wait.Explained in: SimEng 14 · from a graph to tiles · SimEng 14 · dependency stalls
- Process-based against cycle-based modelling
- A process-based (event-driven) model jumps from one state change to the next; a cycle-based one evaluates every block on every clock edge, as RTL simulation does. On the same whole-cycle program they give identical answers; the cost scales with state changes for one and with cycles for the other.Explained in: SimEng 14 · process-based against cycle-based
- From an event loop to a recurrence
- When nothing in a model is left to arbitrate (no two engines ever contend), every event time follows from earlier ones by additions and maxima, so the event loop can be replaced by a straight-line recurrence: same answers, bit for bit, at a fraction of the cost.Explained in: SimEng 14 · the hot path in C++
- pybind11
- A header-only C++ library that exposes C++ functions and classes to Python, converting standard containers automatically; built here by setuptools with
Pybind11Extension. Bound objects are owned through a holder,std::unique_ptrby default.Explained in: SimEng 14 · the hot path in C++ with pybind11 · SimEng 14 · modern C++ in the port - Templates, concepts, move semantics and smart pointers
- Templates write code once for many types, and C++20 concepts state what a type must support; move semantics transfer a container's storage instead of copying it; smart pointers (
unique_ptr,shared_ptr) free memory when its owner goes, so nodeleteis written.Explained in: SimEng 14 · modern C++ in the port · SimEng 03 · modern C++ for models - How gem5 and SST are structured
- gem5: SimObjects configured from Python, connected by request and response ports that carry packets, with retries for flow control, and an event queue in ticks. SST: components connected by links with latencies, exchanging events, with clock handlers and conservative parallel simulation across ranks.Explained in: SimEng 14 · how gem5 and SST are structured
- Execution-provider partitioning (GetCapability)
- How ONNX Runtime hands part of a graph to a hardware backend: each execution provider says which nodes it can run, connected claimed nodes are fused into subgraphs, the rest falls back to the CPU provider, and tensors crossing between them cost transfers.Explained in: SimEng 14 · ONNX Runtime execution providers
- Polynomial products via the NTT
- Multiplying two polynomials modulo XN+1 by transforming both with the number-theoretic transform, multiplying pointwise and transforming back, after a twist by a 2N-th root of unity: O(N log N) instead of O(N2), and the core workload of lattice-based FHE.Explained in: SimEng 14 · an FHE workload
- Computing in the data path
- Doing work on data while it moves (in memory, in the network, in a link) instead of moving it to a processor first. It pays when the moved data needs little arithmetic per byte and the destination engine is the bottleneck; a stage slower than the stream throttles it.Explained in: SimEng 14 · computing in the data path (speculative)
Explained in the sister series
- Discrete-event simulation (DES)
- A simulator whose state changes only at instants (events). A clock, a state, a time-ordered event list and handlers; the clock jumps from one event to the next (next-event time advance), so idle time costs nothing.Explained in: the LLM Inference Simulators glossary
- SimPy
- A small, pure-Python DES library: Environment, processes, timeouts, events, Resource, Container, Store and interrupts. It has no built-in statistics or tracing; the probes are yours to design.Explained in: the LLM Inference Simulators glossary
- Process interaction and coroutines
- Writing each entity's life as one sequential function that suspends while it waits, instead of scattering it over event handlers. Python generators make this natural: yield an event, resume when it fires.Explained in: the LLM Inference Simulators glossary
- Determinism, tie-breaking and RNG streams
- Same configuration and seed must give bit-identical results: break ties between simultaneous events with a sequence number, prefer integer time for clocked models, and give each component its own random stream.Explained in: the LLM Inference Simulators glossary
- Cost model
- The part of a simulator that answers "how long does this take?" (and, with a power model, "how many joules?"). Kept separate from the event engine so a roofline can later be replaced by a measured table or an RTL-calibrated model.Explained in: the LLM Inference Simulators glossary
- Queueing theory checks (M/M/1, M/D/1)
- Closed-form results that a simulator's engine must reproduce: M/M/1 mean time in system 1/(μ−λ); M/D/1 mean wait λS²/(2(1−ρ)) (Pollaczek–Khinchine). The KV link is tested this way.Explained in: the LLM Inference Simulators glossary
- Little's law and the operational laws
- L = λW: mean population equals arrival rate times mean time in system, for any stable system. With the utilisation and forced-flow laws it gives free consistency checks on a simulator's probes.Explained in: the LLM Inference Simulators glossary
- The verification ladder
- Layers of evidence: unit tests against hand calculations, invariants, analytic checks, behavioural expectations, property-based tests, differential tests and, finally, correlation against measurement.Explained in: the LLM Inference Simulators glossary
- Property-based testing (Hypothesis)
- Generate many random configurations and check that invariants hold for all of them; Hypothesis shrinks any failure to a minimal example.Explained in: the LLM Inference Simulators glossary
- Differential testing
- Requiring two independent implementations to agree exactly on the same input: here the fast path against the baseline, and the browser JavaScript against the Python simulator.Explained in: the LLM Inference Simulators glossary
- Profiling a simulator
- Measure before optimising (cProfile, py-spy, perf). Time usually goes to model code that runs per entity per step, not to the event kernel.Explained in: the LLM Inference Simulators glossary
- Parallel runs and Amdahl's law
- Independent runs parallelise trivially (processes, job arrays), but each technique only speeds up the share of time it touches; serial overheads cap the gain.Explained in: the LLM Inference Simulators glossary
- CI, golden metrics and engineering metrics
- Run tests on every push and sweeps nightly (Jenkins or similar); fail the build if golden metrics drift; track correlation error, coverage, simulator speed and test health as engineering metrics.Explained in: the LLM Inference Simulators glossary
- Transaction-level modelling (SystemC TLM-2.0)
- Modelling bus reads and writes as function calls with annotated delays rather than signal wiggles. Loosely-timed TLM boots software; approximately-timed TLM models contention. The industry route to virtual platforms.Explained in: the LLM Inference Simulators glossary
- RTL simulation, co-simulation and golden models
- The software model is the executable specification and the golden reference (a scoreboard in the RTL testbench). RTL cycle counts calibrate the model, workload traces become directed tests, and co-simulation embeds RTL blocks in the system model.Explained in: the LLM Inference Simulators glossary
- RTL power and thermal analysis
- Switching activity from RTL simulation or emulation, run on real workload traces, gives per-block energies that calibrate the simulator; power maps feed thermal models that find thermal hot-spots.Explained in: the LLM Inference Simulators glossary
- Percentiles and tail-latency CCDFs
- Report p50, p90 and p99, not means, and plot the fraction of samples slower than x on log-log axes, where tails show. Say which population (all tokens or per request), and use streaming sketches when samples are many.Explained in: the LLM Inference Simulators glossary
- torch.export and FX graphs
- Capturing a PyTorch model ahead of time as one graph of ATen operators with fake-tensor shapes on every node; run_decompositions() lowers it to the smaller Core ATen set a cost model can cover.Explained in: the LLM Inference Simulators glossary
- torch.compile backends
- TorchDynamo captures graphs as Python runs and hands each to a backend callable, a clean hook for a simulator that estimates cost and then runs the graph eagerly.Explained in: the LLM Inference Simulators glossary
- Dispatch interception and the meta device
- A TorchDispatchMode sees every ATen call; on the meta device tensors have shapes but no storage, so a 70B model's exact operator trace can be recorded on a laptop.Explained in: the LLM Inference Simulators glossary
- ONNX and ONNX Runtime execution providers
- ONNX is a framework-neutral graph format; ONNX Runtime partitions graphs between execution providers, each claiming the nodes it can run, with fallback to the CPU.Explained in: the LLM Inference Simulators glossary
- FLOP and byte counting
- About 2 FLOPs per matmul parameter per token, plus attention that grows with context. A decode step must also read every weight and all of the KV cache. These two counts drive every timing and energy number in the series.Explained in: the LLM Inference Simulators glossary
- Roofline, arithmetic intensity and ridge point
- Attainable performance is the lower of peak compute and bandwidth × arithmetic intensity (FLOPs per byte). The ridge point is the intensity at which both limits meet; below it a kernel is memory-bound.Explained in: the LLM Inference Simulators glossary
- Warm-up, replications and confidence intervals
- A stochastic run is one experiment: drop the warm-up period, repeat with different seeds, and report mean ± a t-based confidence interval. Tail percentiles need many samples.Explained in: the LLM Inference Simulators glossary
- Hot-spot attribution
- Finding where speeding things up would help most. A stage breakdown that sums exactly to each request's latency, the largest waiting stage mapped to the resource that owns it, confirmed by what-if sensitivity runs.Explained in: the LLM Inference Simulators glossary
- Chrome trace events and Perfetto
- A JSON timeline format (complete slices, counters, metadata) that Perfetto and chrome://tracing display. The simulator, ONNX Runtime's profiler and RTL testbenches can all emit it, so timelines can be compared side by side.Explained in: the LLM Inference Simulators glossary
- Calibration and correlation
- Fitting a model's few free coefficients to measurements (or RTL), then tracking its error on workloads it was not fitted to. The correlation report is the deliverable that makes predictions trustworthy.Explained in: the LLM Inference Simulators glossary
- Prefill and decode
- Prefill processes the whole prompt in one pass (matrix-matrix, compute-bound) and produces the first token. Decode then generates one token per sequence per step (matrix-vector, memory-bound). They need different hardware balances.Explained in: the LLM Inference Simulators glossary
- KV cache
- The stored keys and values of every token of every live sequence, so attention need not recompute them. Its size (320 KiB per token for Llama-3-70B) limits batch size and is what disaggregation must move.Explained in: the LLM Inference Simulators glossary
- Utilisation, MFU and MBU
- Busy fraction says how often a resource works; model FLOPs and bandwidth utilisation say how much of its capability it uses. A continuous-batching decode engine is 100% busy at 4% MFU and about 77% MBU.Explained in: the LLM Inference Simulators glossary
- Common random numbers
- Feeding two designs the same random request stream, so noise cancels in their difference: a much narrower confidence interval on "which is better, and by how much" for the same number of runs.Explained in: the LLM Inference Simulators glossary
- Batch means
- One long run cut into batches long enough to be nearly independent, each giving one sample of the metric: warm-up is paid once instead of once per replication, at the risk of correlated batches.Explained in: the LLM Inference Simulators glossary
- Static and dynamic power
- Static power (leakage, clocks, always-on logic) is paid whenever a chip is on; dynamic power αCV²f is paid for switching. Energy per operation scales with V², which is why lowering voltage pays quadratically.Explained in: the LLM Inference Simulators glossary
- DVFS and power caps (the third roof)
- Lowering clock and voltage to cut power. A memory-bound step can slow its compute clock at no time cost; a power cap makes some steps power-bound, a third roof beside compute and memory.Explained in: the LLM Inference Simulators glossary
- TDP and peak power
- The board power limit a device throttles to. Steps that saturate compute and memory at once can exceed it, so the simulator enforces TDP by default and reports the peak power of each step.Explained in: the LLM Inference Simulators glossary
- Energy per token and energy–delay trade-off
- Joules per output token (or tokens per joule) is the efficiency headline. Capping power trades time for energy until static power over the longer run dominates, giving an energy-optimal operating point.Explained in: the LLM Inference Simulators glossary
- Continuous batching
- Re-forming the batch at every decode iteration: finished sequences leave and new ones join, so the GPU stays full. The scheduler loop at the heart of every serving engine and serving simulator.Explained in: the LLM Inference Simulators glossary
- TTFT, TPOT and inter-token latency
- Time to first token (prefill queueing and compute); time per output token after the first (DistServe's definition); and the gap between consecutive tokens, whose tail shows stalls the mean hides.Explained in: the LLM Inference Simulators glossary
- Disaggregated (prefill/decode) serving
- Running prefill and decode on separate machine pools (Splitwise, DistServe, Mooncake, NVIDIA Dynamo), so each phase is provisioned and power-managed for its own SLO. The cost is moving the KV cache.Explained in: the LLM Inference Simulators glossary
- The fidelity ladder
- The range of model detail from analytical formulas through discrete-event, transaction-level, cycle-level and RTL simulation to emulation, FPGA prototypes and silicon. Each rung is slower and more detailed; use the highest rung that answers the question.Explained in: the LLM Inference Simulators glossary
- Resources, stores, containers and back-pressure
- Resource: N identical servers held for a while (links, engines). Store: a buffer of items (FIFOs). Container: an amount (credits, bytes). A finite-capacity Store makes producers wait when it is full, which is back-pressure.Explained in: the LLM Inference Simulators glossary
- Parallel discrete-event simulation (PDES)
- Splitting one run across cores as logical processes: conservative (Chandy–Misra–Bryant, needs lookahead) or optimistic (Time Warp, rolls back). Worth it for large, loosely coupled models.Explained in: the LLM Inference Simulators glossary
- Number-theoretic transform (NTT)
- The FFT over a finite field Zq. It moves a limb between coefficient and evaluation form in (N/2) log N butterflies, so polynomial multiplication becomes element-wise. Every key switch and rescale runs NTTs.Explained in: the FHE Accelerator Simulators glossary
- Modular reduction (Barrett, Montgomery)
- Ways to compute a·b mod q without division, using precomputed constants. Every butterfly and multiply-add lane in an FHE accelerator contains one.Explained in: the FHE Accelerator Simulators glossary
- HBM and key streaming
- Off-chip high-bandwidth memory. Evaluation keys and DFT plaintexts are too big to keep on chip, so they stream from HBM, which is why bootstrapping is often memory-bound.Explained in: the FHE Accelerator Simulators glossary
- NTT-, MAC-, memory- and power-bound
- The simulator's verdict: the most-utilised resource sets the bound (HBM, NTT units, multiply-add lanes), or power when the power limit throttles a compute-bound design.Explained in: the FHE Accelerator Simulators glossary
- Kernel-level trace
- The contract between a scheme model or compiler and the hardware model: HE operations with their inputs, outputs, keys, plaintexts, levels and primitive kernels.Explained in: the FHE Accelerator Simulators glossary
- Scratchpad, LRU and spills
- The accelerator's on-chip SRAM, modelled as a least-recently-used store of ciphertexts, keys and plaintexts. A miss loads from HBM; evicting a live ciphertext writes it back (a spill).Explained in: the FHE Accelerator Simulators glossary
- Calibration and out-of-sample prediction
- Fit as few model coefficients as possible to measurements (here, one rate to an OpenFHE multiplication), then judge the model on quantities it was not fitted to.Explained in: the FHE Accelerator Simulators glossary
- Key switching and evaluation keys
- After a multiplication (relinearisation) or a rotation, part of the ciphertext is under the wrong key; key switching fixes it using an evaluation key (78–210 MiB for the N = 216 sets used here). It dominates the cost of FHE.Explained in: the FHE Accelerator Simulators glossary
- Design sweeps and Pareto fronts
- Simulate a grid of designs; a design is Pareto-optimal if no other is at least as good on every metric (latency, energy, and now area) and better on one. The front is the set of sensible choices; everything else is dominated.Explained in: the FHE Accelerator Simulators glossary
- Min-KS, seeded keys and on-the-fly plaintexts
- Algorithmic ways to cut key and plaintext traffic: rotate iteratively by one step so a BSGS loop reuses one key (Min-KS, from ARK), regenerate half of each key from a seed, and generate DFT plaintexts on chip.Explained in: the FHE Accelerator Simulators glossary
- ENOB and exact rounding
- Effective number of bits: the real precision of an analogue chain. Rounding to the exact integer needs the error below one half, which sets a minimum ENOB for each block size and digit width.Explained in: the FHE Accelerator Simulators glossary
- Dynamic power manager against worst-case clocking
- Worst-case clocking picks one clock at which everything at full rate fits the TDP; a dynamic manager gives each kernel the highest clock that fits the power actually free at that moment.Explained in: the FHE Accelerator Simulators glossary
- Silicon area and the area model
- Die area in mm² per component (NTT, MAC and permutation units, scratchpad, uncore, HBM PHYs). The simulator's model is at 7 nm, calibrated from ARK's published breakdown and a CACTI sweep, and illustrative; area never changes a simulated time or energy.Explained in: the FHE Accelerator Simulators glossary
- Die yield and silicon cost
- The fraction of dies with no fatal defect: Poisson e−A·D0 or Murphy's model, which is kinder to large dies. With dies per wafer it turns area into cost per good die, which grows faster than area.Explained in: the FHE Accelerator Simulators glossary
- RNS polynomials and limbs
- The ciphertext modulus Q is a product of word-sized primes, so each polynomial is stored as one row ('limb') of N residues per prime. A ciphertext at level ℓ has ℓ+1 limbs; all arithmetic stays in 64-bit words.Explained in: the FHE Accelerator Simulators glossary
- Bootstrapping
- Evaluating decryption homomorphically to turn a level-0 ciphertext into one at a high level: ModRaise, CoeffToSlot, EvalMod, SlotToCoeff. It costs hundreds of key switches and gigabytes of keys.Explained in: the FHE Accelerator Simulators glossary
- Fourier optics and the 4f system
- A lens Fourier-transforms the light field; two lenses with a mask between them multiply in the frequency domain, which is a convolution. The transform is free; the converters and lasers around it are not.Explained in: the FHE Accelerator Simulators glossary
- HEIR
- Google's MLIR-based FHE compiler: it chooses parameters, packing, rotations, levels and bootstrap placement, and emits code for libraries or hardware. Its ckks-dialect output drives this simulator.Explained in: the FHE Accelerator Simulators glossary