Simulation Engineering Toolkit — Presentation 12

Measurement Tools and Methods

A catalogue of the tools and methods the simulation series measure with, grouped by what they measure, each with its principle, overhead, accuracy and pitfalls, when to use it, and where this GitHub uses it: profilers and flame graphs, benchmarks, simulated caches and hardware counters, GPU and CPU energy telemetry, trace viewers and framework profilers, RTL coverage and power flows, DRAM simulators, the CACTI, McPAT and Accelergy estimators (CACTI built and swept here), the statistics of simulation output, and agreement checks. Overheads measured on this series' own simulators.

Profilers Hardware counters Energy telemetry Traces RTL coverage CACTI Output statistics
Question → Tool → Overhead → Accuracy → Measure → Trust
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

01

How to Read This Catalogue

The other decks in this series and its sisters use a lot of measurement tools: profilers, counters, coverage tools, trace viewers, DRAM simulators, power estimators. This deck is where each one is explained, so that a deck naming a tool can link here. Every tool and method is a card with the same five fields:

measuring software that runs here measuring hardware, and hardware that does not exist yet time (03-05)cProfile, py-spy, perf,criterion, LoadGen memory, cache (06)cachegrind, counters,tracemalloc timelines (09)Perfetto, PyTorch andORT profilers (Nsight: docs) trust in resultsreplications, CRN (14);golden, bounds (15) power, energy (07-08)NVML, DCGM, Zeus,RAPL, power analysers RTL (10)coverage, toggles,SAIF/VCD power flows memory systems (11)DRAMsim3, Ramulator unbuilt chips (12-13)McPAT, CACTI,Accelergy, Timeloop green: run on this machine (an Intel i7-3770 desktop, no NVIDIA GPU) and measured; amber: some tools need hardware or privileges this machine lacks. Those cards are written from the tools' documentation and papers, and say so. No reading is invented.

Every number from this machine comes from snippets/t12/ (run_t12.py, framework_profilers.py, rtl_coverage.py and cacti/run_cacti.py), recorded in RESULTS.md. Where deck 11 already measured something, the card links to it instead of measuring again. Tool versions and vendor products change quickly; each card links to its documentation, which is the authority.

02

Interactive: Which Tool Should I Use?

Answer two questions; the chooser names the tools to start with and links to their cards. It encodes the "when to use it" field of the cards, nothing more.

1. What do you want to know?
2.  
Pick a question.
03

Time: Instrumenting Against Sampling Profilers

There are two ways to find out where a program's time goes. An instrumenting profiler runs code at every function entry and exit and records exact counts and times. A sampling profiler looks at the call stack every few milliseconds and counts what it sees. The first is exact about calls and distorts time; the second is statistical about time and leaves the program almost untouched.

instrumenting (cProfile) run() a hook (red) at every call and return: exact counts, but each hook costs time, most for many small calls sampling (py-spy, perf) run() a stack sample (blue) at a fixed rate: cheap, but time is estimated from counts, and shortcalls can fall between samples The same simulation under each tool, wall clock (slowdown against no tool): the table below. Coverage and tracemalloc are there for comparison: they hook every line and every allocation.
ToolRunsMedianSlowdownRuns (s)
no tool51.23 s1.00x1.23, 1.26, 1.19, 1.21, 1.23
cProfile (instrumenting every call)32.51 s2.04x2.52, 2.51, 2.49
py-spy record, 100 Hz31.29 s1.05x1.30, 1.29, 1.26
py-spy record, 500 Hz31.78 s1.45x1.84, 1.78, 1.57
coverage run (line coverage)33.18 s2.59x3.28, 3.18, 3.06
coverage run --branch33.49 s2.85x3.53, 3.48, 3.49
python -X tracemalloc=1 (one frame per allocation)35.79 s4.72x5.84, 5.79, 5.78
python -X tracemalloc=25 (25 frames)331.45 s25.63x31.60, 31.42, 31.45

Source: snippets/RESULTS.md in _simeng_build

cProfile and snakeviztool

Principle
Python's built-in deterministic profiler hooks every function call and return (sys.setprofile-style events) and records call counts, own time (tottime) and time including callees (cumtime) per function. snakeviz draws the saved profile as an interactive icicle or sunburst of the call tree.
Overhead
2.04x on Disaggregated_Inference_Sim, which makes millions of small calls. The cost is per call, so code made of many tiny functions slows most.
Accuracy and pitfalls
Counts are exact; times are inflated, unevenly: the hook's cost lands on whichever function is small and frequent, so they look worse than they are. Time inside C functions is attributed to the built-in, with no detail below it. Profile a realistic workload, not the first second of one.
When to use it
When you need exact call counts or a call tree (who calls this a million times?), on a run short enough to repeat. For "where does the time go" on a long or live run, a sampling profiler distorts less.
Where this GitHub uses it
InfSim 08, step zero: profile found the per-step bookkeeping with it; its current report is in RESULTS.md (T12).

py-spy (sampling profilers)tool

Principle
Another process reads the target Python process's memory at a fixed rate and reconstructs its Python call stack; each sample adds one count to every frame on the stack. Time is estimated as share of samples. Works on a running process (py-spy record --pid) with no code changes.
Overhead
Measured on the same run: 1.05x at 100 Hz and 1.45x at 500 Hz. By default py-spy pauses the process for each sample so the stack is consistent; --nonblocking avoids the pause at the risk of torn stacks.
Accuracy and pitfalls
Statistical: a function with few samples has a wide error bar (the relative error of a share falls roughly as one over the square root of its sample count). By default it only counts threads that are running; --idle includes waiting ones. Native extensions need --native. At 500 Hz here, py-spy now and then lost a race with the child's exit and had to be rerun (RESULTS.md).
When to use it
The default first look at any Python program, and the only safe one in production. Use cProfile when you need counts; perf when the time is in native code.
Where this GitHub uses it
SimEng 11: sampling profilers and reading a profile, on Disaggregated_Inference_Sim and Memory_System_Sim.

Linux perftool

Principle
The kernel's interface to the CPU's performance-monitoring unit. perf stat counts events (cycles, instructions, misses) over a run; perf record samples: every N events or at a frequency it records the instruction pointer and, with --call-graph, the stack. Works for any language whose frames it can unwind.
Overhead
Counting mode costs next to nothing (the hardware counts). Sampling costs an interrupt per sample, so it scales with the rate (-F); DWARF unwinding copies a chunk of stack per sample and costs far more than frame-pointer unwinding.
Accuracy and pitfalls
Needs permission (kernel.perf_event_paranoid); stacks need frame pointers or DWARF info, which release builds and interpreters often lack; more events than counters means multiplexing and scaled estimates (counters card). An interpreter shows up as its own C functions unless it cooperates (Python 3.12's -X perf).
When to use it
Compiled code, interpreter internals, and "why" questions (IPC, misses) that a Python profiler cannot answer. Where perf is not allowed, as in many containers and CI runners, use py-spy or callgrind.
Where this GitHub uses it
SimEng 11: Linux perf, counters and call stacks (the Rust port and the Python simulator); slide 06 here (multiplexing).

GNU time (/usr/bin/time -v)tool

Principle
Runs a command and prints what the kernel accounted for it when it exited (via wait4): wall time, user and system CPU time, maximum resident set size, page faults, and voluntary and involuntary context switches. Not the shell's time keyword, which prints only times.
Overhead
None during the run: the kernel keeps these counts anyway.
Accuracy and pitfalls
One number per run, so it says nothing about where. Maximum RSS is a peak and counts shared libraries; CPU time above wall time means parallelism; many involuntary switches mean more runnable threads than CPUs.
When to use it
The first look in the USE method: is it CPU-bound, waiting, swapping, or oversubscribed? And the cheapest way to record peak memory in CI.
Where this GitHub uses it
SimEng 11: USE in practice (a rayon sweep on one and eight threads); the RSS figure on slide 06 here.
Docs: time(1)
04

Flame Graphs, On-CPU and Off-CPU

A profile of thousands of stacks needs a picture. SimEng 11 slide 06 has interactive flame graphs of four real profiles from this series; this card is the method behind them.

Flame graphsmethod

Principle
Sampled stacks are merged by common prefix and drawn as towers: the root at the bottom, callees above. A frame's width is the share of samples with it on the stack (total time); the exposed top edge is self time. Siblings are sorted alphabetically, so the x-axis is not time. An on-CPU flame graph samples running threads; an off-CPU one weighs stacks by the time threads spent blocked (sleeping, waiting for I/O or a lock), recorded from scheduler events.
Overhead
The picture is free; the cost is the profiler's that collected the stacks (slide 03). Off-CPU tracing records every context switch, which on a busy system costs much more than sampling.
Accuracy and pitfalls
Inherits the sampling error of thin frames. Missing frames (no frame pointers, inlining) make towers look shallower and move time to the wrong parent. An on-CPU graph of a program that mostly waits looks small and misleading: waiting is invisible to it. Comparing two graphs by eye is weak; a differential flame graph colours the change.
When to use it
Whenever a profile has more than a screenful of functions. Off-CPU when wall time is much larger than CPU time (GNU time shows that), which for simulators usually means I/O or a lock rather than compute.
Where this GitHub uses it
SimEng 11: sampling profilers and flame graphs; the folded stacks are in snippets/t11/out of the hub repository.
Docs: Flame graphs · Off-CPU flame graphs (Brendan Gregg)
05

Benchmarks and Load Generators

A profiler says where time goes; a benchmark says how long something takes, and whether a change moved it. SimEng 11 slide 10 measured why a benchmark result is a distribution, not a number (median, MAD and a bootstrap interval), and slide 11 designed a regression gate from that noise. Two tools that apply those statistics:

criterion (Rust microbenchmarks)tool

Principle
Runs a closure many times: a warm-up phase, then samples of increasing iteration counts, fits time against iterations by linear regression (so fixed per-sample overhead drops out), and reports the slope with a bootstrap confidence interval. It saves each run and reports whether the next one changed, with a significance test and a noise threshold.
Overhead
Wall-clock seconds per benchmark (by default 3 s of warm-up and 5 s of measurement); the code under test runs at full speed. black_box stops the compiler optimising the work away.
Accuracy and pitfalls
A microbenchmark measures a hot loop with warm caches and a trained branch predictor, which the real program may never see. "Change detected" at 5% on a noisy desktop can be noise; criterion's own noise threshold (2% by default) and SimEng 11's A/A test say how much. It compares runs on one machine only.
When to use it
For a kernel, a data structure or an event-queue operation whose cost you want to track. For whole-program speed, time the whole process repeatedly instead; for Python, pytest-benchmark plays the same role (not used in this series).
Where this GitHub uses it
SimEng 01: the toolchain (events per second of Rust_DES_Kernel); SimEng 06: a performance gate in CI.

Load generators (MLPerf LoadGen)tool

Principle
A load generator issues requests to a system under test and timestamps each one's issue and completion. MLPerf Inference's LoadGen defines scenarios: single-stream, multi-stream, offline, and server (Poisson arrivals with a latency bound), and checks that enough queries ran for the reported percentile to be meaningful.
Overhead
It shares the machine with the system under test; LoadGen is C++ to keep its own cost small, and must never be the bottleneck.
Accuracy and pitfalls
Coordinated omission: a closed-loop generator that waits for each response before sending the next stops sending when the system stalls, so the stall hides itself and the tail latency looks far better than users see. Measure from the intended send time, with open-loop (scheduled) arrivals.
When to use it
Benchmarking a serving system, and validating a serving simulator against one: a simulator that can emulate LoadGen's server scenario can be compared with published MLPerf results.
Where this GitHub uses it
InfSim 04: workloads and benchmarks; InfSim 06: open-loop arrivals. The simulators generate open-loop Poisson arrivals themselves.
06

Memory and Cache: Simulated Caches, Counters and Heaps

Valgrind: cachegrind and callgrindtool

Principle
Valgrind translates the program's machine code on the fly and runs it on a synthetic CPU. Cachegrind counts every instruction executed (and, with --cache-sim=yes, simulates a first-level and last-level cache); callgrind adds the call graph, so costs can be attributed to callers. Both annotate source lines.
Overhead
Measured: 93x on a short Rust run (Valgrind's start-up dominates) and 67x on the Python simulator (table below). The manual quotes 20–100×.
Accuracy and pitfalls
Instruction counts are exact and deterministic: the same run gives the same count, so a 1% change is visible where timing noise is several per cent (SimEng 11 slide 10). The cache model is simple (no prefetcher, no out-of-order overlap), so miss counts are indicative, and instructions are not time: a cache miss costs hundreds of cycles, an add one.
When to use it
Comparing two versions of code, in CI without timing noise, or where perf is not allowed. For real cache behaviour on real hardware, use counters.
Where this GitHub uses it
SimEng 11: counting instructions with cachegrind (Rust against Python; the sort that took half the instructions).

Hardware performance countersmethod

Principle
The CPU's performance-monitoring unit has a handful of programmable counters per core (Intel's recent cores have four to eight, plus a few fixed ones) that increment on events: cycles, instructions retired, cache references and misses, branch mispredictions, TLB misses. Instructions per cycle (IPC) and misses per thousand instructions come from their ratios.
Overhead
Essentially none in counting mode: the hardware counts while the program runs at full speed.
Accuracy and pitfalls
Ask for more events than counters and the kernel multiplexes: it rotates events onto the counters and scales each count by the time it was enabled. Here twelve events were each counted for 18–45% of the run (table below), and the scaled estimate moved: instructions: counted alone 2,620,114,505; multiplexed with eleven others 2,656,701,145 (+1.4%). Event definitions vary by CPU generation, some events do not exist (LLC-load-misses here), and raw counts mislead: normalise per instruction.
When to use it
When a profile says where but not why: low IPC with many misses points at memory, many mispredictions at control flow. Count a few events per run, and repeat (perf stat -r), rather than many at once.
Where this GitHub uses it
SimEng 11: perf stat on the Rust and Python simulators; the multiplexing table below.
Docs: perf tutorial (events, multiplexing and scaling)

tracemalloc (Python heap)tool

Principle
Hooks Python's memory allocators and records, for every live block, the traceback where it was allocated (one frame by default). Snapshots can be grouped by line and diffed between two points in a run. Valgrind's massif does the same for native heaps.
Overhead
Large and growing with the frames kept: 4.72x with one frame and 25.63x with 25, on the simulation of slide 03. Its own bookkeeping also uses memory.
Accuracy and pitfalls
Exact for Python allocations, blind to memory that C extensions allocate themselves, and not the process's footprint: Peak traced Python heap 49.0 MiB; still allocated at the end 35.2 MiB; the process's maximum resident set (/usr/bin/time -v, no tracing) 76.9 MiB.
When to use it
Finding which lines hold memory, or a leak (diff two snapshots). For the footprint, GNU time's maximum RSS; for NumPy and PyTorch buffers, their own tools.
Where this GitHub uses it
Here only: Disaggregated_Inference_Sim's largest holder is sim.py:132, the per-request inter-token latency lists that the metrics need at the end (table below).
Cachegrind's slowdown
ProgramNative (median of 5)Under cachegrindSlowdown
Rust disagg-rs, 500 requests9.6 ms0.89 s93x
Python disagg-sim, 500 requests276.6 ms18.58 s67x

Source: snippets/RESULTS.md in _simeng_build

tracemalloc: largest holders at the end of the run
Line (at the end of the run)SizeBlocks
src/disagg_sim/sim.py:13223.7 MiB766,474
src/disagg_sim/sim.py:4764.7 MiB46,254
src/disagg_sim/sim.py:4772.2 MiB61,672
src/disagg_sim/sim.py:4751.6 MiB30,836

Source: snippets/RESULTS.md in _simeng_build

perf stat with twelve events (multiplexed)
EventCount (scaled)Run-to-run variationCounted for
cycles1,266,070,3490.53%27%
instructions2,656,701,1450.63%36%
cache-references6,164,3102.00%36%
cache-misses2,985,4762.11%45%
branches275,974,2950.57%45%
branch-misses3,015,3372.65%45%
L1-dcache-loads497,366,3860.72%35%
L1-dcache-load-misses19,862,8111.65%35%
LLC-loads4,526,9508.05%18%
LLC-load-missesnot supported on this CPU--
dTLB-loads498,474,5230.59%18%
dTLB-load-misses1,301,4864.17%18%

Source: snippets/RESULTS.md in _simeng_build

07

GPU Telemetry: nvidia-smi, NVML, DCGM and Zeus

This machine has no NVIDIA GPU (nvidia-smi and dcgmi are not installed; RESULTS.md records the check), so nothing on this slide was measured here. The cards are written from NVIDIA's and the tools' documentation and from a published measurement study; the outputs shown are documentation excerpts, labelled as such. Drivers and field lists change between releases: the linked documentation is the authority.

nvidia-smi and NVMLtool

Principle
NVML is the C library the driver exposes for monitoring and management: utilisation, clocks, temperatures, memory, power draw, a cumulative energy counter (Volta and later), throttle reasons. nvidia-smi is the command-line front end; nvidia-smi --query-gpu=... --format=csv -lms 100 polls fields to CSV.
Overhead
A query is a driver call, cheap enough to poll at 10 Hz; the GPU's sensors update on their own schedule regardless of how often you ask.
Accuracy and pitfalls
"Utilisation" means at least one kernel was running during the sample period, not how busy the GPU's units were. Power is not instantaneous: Yang, Adamek and Armour found that on A100 and H100 the reading averages 25 ms out of every 100 ms, so 75% of the run is never sampled; newer drivers separate power.draw.instant from power.draw.average. Their recommended practice (many repetitions, randomised delays, discard the rise time) brought it to about 5% of an external meter.
When to use it
A quick look, scripts, and the energy counter around a long run. For many GPUs or profiling-level metrics, DCGM; for per-region energy in Python, Zeus.
Where this GitHub uses it
InfSim 07: calibrating the power model (read power while sweeping FLOP and byte rates, fit pJ/FLOP and pJ/byte).

NVIDIA DCGMtool

Principle
The Data Center GPU Manager runs a host engine that samples numbered fields for every GPU: device fields such as DCGM_FI_DEV_POWER_USAGE (W) and DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION (mJ since the driver loaded), and profiling fields read from the GPU's counters, such as SM active (1002), SM occupancy (1003), tensor pipe active (1004) and DRAM active (1005). dcgmi dmon streams them; dcgm-exporter feeds Prometheus.
Overhead
Designed to run continuously on production nodes. Profiling fields default to 1 Hz, with 100 ms as the minimum interval.
Accuracy and pitfalls
Profiling fields are fractions of time or cycles averaged over the interval, so a short burst is diluted. Not every profiling metric can be collected at once: some need several passes and are multiplexed, and which combinations work depends on the GPU (dcgmi profile -l lists them). Power fields inherit the sensor averaging of the NVML card.
When to use it
Fleet monitoring and long calibration runs: power against utilisation over hours, across many GPUs. For one program's timeline, Nsight Systems.
Where this GitHub uses it
InfSim 07: the power fidelity ladder (silicon measurement) and InfSim 11 (calibrate coefficients by regression on measured power).
Documentation excerpt, not a measurement: dcgmi dmon output format (NVIDIA DCGM user guide)
$ dcgmi dmon -e 1001,1004,1005
# Entity  GRACT  TENSO  DRAMA
    GPU 0  0.969  0.928  0.000

Zeus (ML.ENERGY)tool

Principle
A Python library that measures time and energy over named code windows: monitor.begin_window("step") … end_window("step"). On Volta and newer NVIDIA GPUs it reads NVML's energy counter at the window edges; on older ones it polls power in a separate process. It also reads CPU and DRAM energy through RAPL, which needs root. It underlies the ML.ENERGY leaderboard's energy per request.
Overhead
Its documentation gives under 10 ms per call, so windows should be long compared with that: a training step or a batch, not a kernel.
Accuracy and pitfalls
As good as the counters underneath: NVML's energy counter and RAPL, with their update periods and averaging. Energy of a window shared with other work on the same GPU is the GPU's, not your code's.
When to use it
Energy per request, per step or per epoch in Python code; baselines for validating a simulator's energy model.
Where this GitHub uses it
InfSim 07: measuring energy per request; InfSim 10's reading list.
08

CPU and Wall Energy: RAPL and Power Analysers

Both were tried here. What happened, as recorded:

So no CPU energy reading appears in this deck. That is the usual situation for an unprivileged user on a current kernel; the cards say what the reading would mean.

RAPL (CPU package energy)tool

Principle
Intel's Running Average Power Limit, also implemented by recent AMD CPUs, keeps energy counters per domain (package, cores, integrated GPU, and on servers DRAM) in model-specific registers. Linux exposes them as /sys/class/powercap/intel-rapl:*/energy_uj and as perf's power/energy-pkg/ events; energy over an interval is the difference of two reads.
Overhead
A register or file read, essentially free; the counters update about every millisecond.
Accuracy and pitfalls
The package domain is the CPU, not the wall: no power-supply losses, disks, fans or (on desktops) memory. On older parts such as this Ivy Bridge the energy is generally reported to be modelled from activity rather than measured. Khan et al. (RAPL in Action) report open issues: driver support, non-atomic register updates and unpredictable update timing. The counter wraps (here after 65.5 kJ, RESULTS.md), so long runs need several reads. Since 2020 Linux makes energy_uj root-only, because fine-grained energy readings leak secrets through a side channel (PLATYPUS).
When to use it
CPU energy of a long run, with privileges (root, CAP_PERFMON, or perf_event_paranoid 0 or lower). For the whole machine, an external meter.
Where this GitHub uses it
Here only, as the probe above; Zeus reads it for CPU energy.

External power analysers and board telemetrytool

Principle
A meter between the wall and the machine (or a shunt or current clamp on a supply rail) measures voltage and current and integrates power over time. Board telemetry is the same idea built in: current sensors on the accelerator card's rails, read through its management interface.
Overhead
None on the device: the measurement is outside it. The cost is equipment and a manual setup.
Accuracy and pitfalls
The most trustworthy number, and the least specific: a wall meter includes the power supply's losses and everything in the box, so it needs an idle baseline subtracted. A meter that samples slowly misses short transients; synchronising its readings with the program's phases takes care.
When to use it
As the reference that RAPL, NVML and simulator energy models are calibrated against, and for anything quoted as "wall power".
Where this GitHub uses it
InfSim 07: silicon-level power measurement; the nvidia-smi study above validated against one.
09

Traces and Framework Profilers

A profile aggregates; a trace keeps every event with its start and duration, so overlap, gaps and ordering are visible. The simulators in this series write traces in the same format the framework profilers do, which is the point: a simulated timeline and a measured one open side by side. Both framework profilers ran here on Torch_Sim_Frontend's two-layer tiny Llama:

RuntimePlainProfiledOverheadEvents recorded
PyTorch eager, torch.profiler (CPU activity, shapes)8.19 ms9.42 ms1.15x15,270 (30 passes)
ONNX Runtime CPU EP, enable_profiling5.69 ms9.64 ms1.70x6,512 (6,288 KB of JSON for 35 runs)

Source: snippets/RESULTS.md in _simeng_build

PyTorch operatorCallsSelf CPU shareONNX operator typeShare of node time
aten::mm45050.9%MatMul54.8%
aten::_scaled_dot_product_flash_attention_for_cpu608.9%Mul8.2%
aten::mul6607.8%Add4.9%
aten::index_select305.5%Transpose4.1%
aten::silu604.2%Where2.9%
aten::add4203.5%Pow2.5%

Source: snippets/RESULTS.md in _simeng_build

Perfetto and the Chrome trace formattool

Principle
The Chrome trace-event format is JSON: complete events ("ph":"X", start and duration in microseconds), counters ("C") and metadata ("M"), grouped by process and thread. Perfetto (and chrome://tracing) draws them as a timeline and can query them with SQL.
Overhead
The cost is in size: one event per kernel or operator per run grows fast (the ONNX Runtime row of the table above: megabytes of JSON for a few dozen runs of a tiny model). Trace a window, sample, or trace only events over a threshold.
Accuracy and pitfalls
As accurate as whoever wrote the timestamps; clocks of different sources must be aligned before a measured and a simulated trace are compared. Large JSON traces load slowly; Perfetto's binary format scales better.
When to use it
Whenever the question is "why", not "how much": idle gaps, serialisation, a queue that never drains.
Where this GitHub uses it
InfSim 06: traces in Perfetto; disagg-sim --trace (trace.py) and fhe-sim --trace (trace.py).

PyTorch profilertool

Principle
torch.profiler.profile records an event for every operator dispatched (aten::mm, …) with CPU time, and, on GPUs, kernel times through NVIDIA's CUPTI. Options record input shapes, memory and stacks; key_averages() aggregates, export_chrome_trace writes a trace for Perfetto, and a schedule skips warm-up steps.
Overhead
Measured here: 1.15x with CPU activity and shapes. Stacks and memory recording cost more.
Accuracy and pitfalls
Self time against total time matters as in any profiler: aten::linear contains aten::mm. The first iterations include one-off costs, so profile after warm-up. On a GPU, CPU-side operator time is launch time, not execution time; look at the kernel rows.
When to use it
Attributing a model's time to PyTorch operators, and finding the operators a cost model must cover.
Where this GitHub uses it
Here (table above: matrix multiplies take half the self CPU time); SimEng 10 traces the same models without running them.

ONNX Runtime profilertool

Principle
With SessionOptions.enable_profiling = True, ONNX Runtime writes one Chrome-trace event per graph node per run (operator type, execution provider, input and output shapes, memory deltas) plus session events; end_profiling() returns the file.
Overhead
Measured here: 1.70x, because this tiny graph has many cheap nodes, each of which writes an event. ONNX Runtime: session initialisation 45.8 ms; first run 11.0 ms against a median of 9.6 ms afterwards.
Accuracy and pitfalls
Node names come from the graph after ONNX Runtime's optimiser has fused and folded it, so they may not match the exported graph. The first run includes allocation and kernel selection. Times are per node; parallel nodes overlap.
When to use it
Attributing time to ONNX operators and execution providers, and comparing a deployed graph's timeline with a simulator's in the same viewer.
Where this GitHub uses it
InfSim 09: ONNX and ONNX Runtime; the table above, on the graph exported by Torch_Sim_Frontend's own exporter.

NVIDIA Nsight Systemstool

Principle
A system-wide tracer (nsys profile) that records CPU thread activity, CUDA API calls, GPU kernels and memory copies, and user annotations (NVTX ranges) on one timeline. Nsight Compute is its per-kernel counterpart. Not run here: no NVIDIA GPU.
Overhead
Low for API and kernel tracing, so whole training or serving steps can be captured; CPU sampling and many NVTX ranges add more. Captures are large.
Accuracy and pitfalls
Shows overlap and gaps, not why a kernel is slow (that is Nsight Compute's job). Annotate phases with NVTX, or the timeline is a wall of kernel names.
When to use it
When a GPU's utilisation is low and the question is where the idle gaps are: launch overhead, host-side work, synchronisation, copies.
Where this GitHub uses it
Not used; it is the GPU-side reference for the timelines the simulators produce.
One node event from the ONNX Runtime profile recorded here (abridged; RESULTS.md)
{"cat": "Node", "pid": 3327413, "tid": 3327413, "dur": 66, "ts": 61078, "ph": "X", "name": "node_embedding_kernel_time", "args": {"op_name": "Gather", "provider": "CPUExecutionProvider", "input_type_shape": [{"float": [1000, 256]}, {"int64": [1, 128]}], "output_type_shape": [{"float": [1, 128, 256]}]}}
10

Hardware and RTL: Coverage, Activity and Power

Verification progress is measured, too. RTL_CoSim_NTT's butterfly, rebuilt with Verilator's coverage on, run with uniform and with constrained-random stimulus, with its seeded Barrett bug and without:

BuildStimulusTest resultLineBranchToggleFunctional-coverage holesBuildRun
plainuniformpassed---a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows13 s1.73 s
plainconstrainedpassed---none13 s1.52 s
plain + buguniformpassed (bug missed)---a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows12 s1.66 s
plain + bugconstrainedFAILED (bug caught)---none12 s1.58 s
covuniformpassed7/74/41664/1836a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows30 s1.95 s
covconstrainedpassed7/74/41664/1836none30 s1.96 s
cov + buguniformpassed (bug missed)4/46/61664/1836a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows30 s1.97 s
cov + bugconstrainedFAILED (bug caught)4/46/61664/1836none30 s1.93 s

Source: snippets/RESULTS.md in _simeng_build

Uniform stimulus reaches every line and branch, and the same toggles as constrained stimulus, and still misses the bug. Only the functional-coverage holes (corr_2_x_sum_wraps, corr_2_x_diff_borrows) show that the situations that expose it never happened. The toggle points neither stimulus reaches are bits that cannot move for this prime and width.

Code coverage (coverage.py, llvm-cov, gcov)tool

Principle
Records which lines, branches or regions the tests executed: coverage.py (pytest-cov) traces Python lines; cargo llvm-cov and gcov/gcovr use counters the compiler inserts.
Overhead
Measured on the simulation of slide 03: 2.59x for line coverage and 2.85x with branches. Compiled instrumentation costs less.
Accuracy and pitfalls
Exact about what ran, silent about what was checked: executing a line is not asserting its result. A low number is a real warning; a high one is not proof (SimEng 06's medium suite had 100% line and branch coverage and a 67% mutation score).
When to use it
Always, cheaply, as a floor; then mutation testing to measure the checks.
Where this GitHub uses it
SimEng 06: coverage, and what it does not tell you; every Jenkins pipeline in SimEng 07 records it.

Verilator coverage and functional coveragetool

Principle
Built with --coverage, Verilator counts line (block) and branch execution and every toggle of every signal bit, and writes coverage.dat at the end; verilator_coverage merges and annotates. Functional coverage is different: named bins the testbench samples from each transaction (operand at its maximum, a sum that wraps) and crosses of them, counting the situations the stimulus created.
Overhead
Measured above: the coverage build took about twice as long to compile and the run was slightly slower (most of a run here is cocotb's Python). Functional coverage costs what the sampling code costs.
Accuracy and pitfalls
Code and toggle coverage measure what the design did, not whether the stimulus was interesting; as the table shows, they saturate early. Functional coverage is only as good as the bins someone thought to write; the crosses that make a bug observable are the ones that matter.
When to use it
Code and toggle coverage to find dead logic and untested pins; functional coverage to decide when verification is done.
Where this GitHub uses it
SimEng 05: functional coverage and crosses (coverage.py in RTL_CoSim_NTT), Verilator in CI; the table above.

Toggle counts and switching activitymethod

Principle
Dynamic power is roughly activity × capacitance × V² × f. Counting how often registers change during a workload estimates the activity term before a netlist exists, and shows how strongly power depends on the data.
Overhead
A comparison per sampled bit per cycle in the testbench, or Verilator's toggle coverage; either is a fraction of the simulation's cost.
Accuracy and pitfalls
Every register bit counts equally; combinational nets, glitches, clock tree and memories are missing. It ranks workloads; it does not give watts.
When to use it
Early, to compare data patterns or architectures; replaced by an RTL power flow once one exists.
Where this GitHub uses it
SimEng 05: switching activity as a power proxy (random 50-bit data is the worst case, and FHE data is random).

RTL power flows (SAIF/VCD into a power tool)method

Principle
Simulation writes switching activity, either per-signal toggle rates and time at each value (SAIF) or full waveforms (VCD, or the proprietary FSDB). A power-analysis tool annotates the activity onto the RTL or the gate-level netlist and applies the cell library's energy and leakage data. The tools are commercial (the large EDA vendors each sell one); this series does not use them.
Overhead
Waveform dumps are huge and slow the simulation many times over; SAIF is far smaller. The power run itself takes minutes to hours per window of activity.
Accuracy and pitfalls
Better the later the stage: RTL-level estimates guess the synthesised logic, gate-level after placement knows the wires. The vectors decide everything: a toy test bench gives toy activity, which is why the simulator's workload traces should drive it.
When to use it
Once RTL exists: to sign off a power budget, and to extract per-operation energies that replace a simulator's illustrative coefficients.
Where this GitHub uses it
InfSim 07: from the simulator to RTL power (described, not run).

Mutation testing (mutmut, cargo-mutants)method

Principle
Measures the tests rather than the code: a tool makes small deliberate bugs (mutants: >= to >, + to -) and runs the suite against each; the mutation score is the share it catches.
Overhead
One test-suite run per mutant, and a crate yields hundreds of mutants (Rust_DES_Kernel's table is in SimEng 06 slide 09), so a run takes the suite's time hundreds of times over: it runs in Jenkins, not on every commit.
Accuracy and pitfalls
Equivalent mutants change code but not behaviour and can never be caught, so 100% is not always reachable; boundary mutants on floats are mostly noise. A survivor is a question, not automatically a missing test.
When to use it
After coverage is high, to find code that runs under test but is never checked.
Where this GitHub uses it
SimEng 06: mutation testing and slide 09 (Rust_DES_Kernel, after three rounds of tests written for survivors).
11

Memory Systems: DRAMsim3 and Ramulator

Memory_System_Sim, this series' own command-level DRAM model, was validated four ways: closed forms, an independent protocol checker on Hypothesis-generated traces, behavioural checks, and a cross-check against DRAMsim3 on the same traces, timing and address mapping (SimEng 04 slide 11). The reference simulators:

DRAMsim3tool

Principle
A cycle-accurate model of a DRAM controller and devices (DDR3/4, LPDDR, GDDR, HBM), configured by an .ini file of JEDEC timing parameters and organisation; it takes a trace or a CPU simulator's requests, issues commands cycle by cycle under every timing constraint, and reports bandwidth, latency, power and, optionally, temperature.
Overhead
Run time grows with simulated cycles, not with requests, so idle periods still cost; it is usually run as a component of a larger simulator or on traces.
Accuracy and pitfalls
The devices follow the standard; the controller is one design among many (per-bank queues, scheduling and page policy). The SimEng 04 cross-check agreed within a few per cent where the device dominates and differed by tens of per cent on interleaved streams, purely because the two controllers queue differently. A cross-check that agrees everywhere usually means shared assumptions.
When to use it
As a reference model for a new DRAM model, or plugged into a CPU or accelerator simulator when memory timing matters more than a bandwidth figure.
Where this GitHub uses it
SimEng 04: cross-checked against DRAMsim3; the cross-check data in Memory_System_Sim.

Ramulatortool

Principle
A cycle-level DRAM simulator built around a generic state machine of each standard's hierarchy (channel, rank, bank group, bank, row), so new standards are added as tables rather than new code. Ramulator 2.0 is a modular rewrite.
Overhead
Comparable in kind to DRAMsim3: cycle-by-cycle, run on traces or inside a CPU simulator.
Accuracy and pitfalls
As with DRAMsim3, the controller policy you configure decides most results that the device's timing does not; and how closely each standard's model has been checked against real parts varies, so look up the one you use in its papers and repository before quoting absolute numbers. Not run here.
When to use it
When the standard or the research mechanism (refresh schemes, RowHammer mitigations) is one Ramulator models and DRAMsim3 does not.
Where this GitHub uses it
Named as an alternative reference in SimEng 04 and InfSim 04; not run.
12

Architecture Estimators: McPAT, CACTI, Accelergy and Timeloop

Before RTL exists, power and area come from estimators that build a component from its structure and technology parameters. They are fast, which is the point: they rank design options. They are not sign-off, and their error bars come from their own validation papers.

architecture description
→
component models (CACTI for arrays)
→
energy per action, area
→
× action counts from a simulator or mapper

McPATtool

Principle
An integrated power, area and timing model for multicore processors: cores, caches (through CACTI), networks on chip, memory controllers and clocking, described in an XML file with technology nodes from 90 to 22 nm. Activity statistics from a performance simulator (gem5 is the usual companion) turn per-access energies into runtime power.
Overhead
Analytical: seconds per configuration, so it can sit in a design-space sweep.
Accuracy and pitfalls
The original paper validated it against four published processors (Niagara, Niagara 2, Alpha 21364 and a Xeon); Xi et al. (HPCA 2015) found that its predictions can have significant error, because some models are incomplete, too high-level, or assume implementations that differ from the core being studied. Use it for relative comparisons inside one configuration family, not absolute watts.
When to use it
Early processor exploration with a CPU simulator. For accelerators, Accelergy fits better.
Where this GitHub uses it
Named in InfSim 07's power ladder and InfSim 10's reading list; not run.

CACTItool

Principle
An analytical model of SRAM and DRAM arrays, caches and scratchpads: given capacity, block size, associativity, ports, banks, technology node and cell type, it searches internal organisations (subarray splits, wire types) and reports access time, cycle time, dynamic energy per read and write, leakage and area for the best one under its optimisation target.
Overhead
CACTI took 0.27 to 3.51 s per configuration (26 s for all 14).
Accuracy and pitfalls
Its technology data are ITRS projections, not a foundry's, so absolute numbers at a node are estimates; the CACTI 6.0 report validated its new wire and bitline models within 12–13% of SPICE on a 65 nm predictive model, components rather than a whole array against silicon. The transistor type dominates leakage (slide 13: over a thousand-fold between itrs-hp and itrs-lstp). The smallest node with data is 22 nm; scaling beyond is your assumption.
When to use it
Sizing an on-chip memory: how area, energy and latency grow with capacity, for a scratchpad, cache or buffer. McPAT and Accelergy call it for their arrays.
Where this GitHub uses it
Slide 13 here (a 22 nm capacity sweep, recorded for the area model in FHE_Accelerator_Sim); named in InfSim 07.

Accelergy and Timelooptool

Principle
Timeloop describes an accelerator as a hierarchy of buffers, networks and arithmetic units, and a tensor workload as loop nests; its mapper searches the space of mappings (tiling, loop order, spatial unrolling) and, for each, counts accesses at every level and estimates cycles as the slowest component in a pipeline. Accelergy turns those action counts into energy and area: each component's per-action energy comes from plug-ins (CACTI for memories, tables for arithmetic) at a stated technology.
Overhead
Evaluating one mapping is analytical and fast; the cost is the search, because mapspaces are huge. In the Timeloop paper, 480,000 mappings of one convolution layer were all within 5% of peak performance yet varied nearly 19× in energy.
Accuracy and pitfalls
Timeloop was validated against an NVDLA-derived design's detailed simulator and against Eyeriss's published 65 nm results; Accelergy reports 95% accuracy against post-layout energy on Eyeriss. Its performance model assumes negligible pipeline stalls (reasonable with double buffering), and the energy is only as good as the per-action tables.
When to use it
Comparing dataflows and buffer sizes for tensor accelerators before RTL; always evaluate an architecture at its best mapping, or the comparison is unfair.
Where this GitHub uses it
Named in InfSim 04's landscape, InfSim 07 and InfSim 10; not run (CACTI, which it uses for arrays, is on slide 13).
13

CACTI on This Machine: SRAM Against Capacity

CACTI 7 (HewlettPackard/cacti, built here) swept over scratchpad capacity at 22 nm, with both of its SRAM transistor types and, for 64 MiB, more banks. Only the size, bank count and cell type differ from the upstream cache.cfg; every configuration file and output is in snippets/t12/cacti. These are model outputs from ITRS-based projections, not measurements of silicon.

CapacityCellsBanksArea (mm²)mm² per MiBRead energy (nJ)pJ per bit readWrite energy (nJ)Leakage (mW)Access time (ns)Cycle time (ns)Area efficiency
64 KiBitrs-hp10.0550.8830.01500.0290.045220.910.3780.34675.5%
256 KiBitrs-hp10.2931.1710.04730.0920.069977.250.6280.61356.9%
1 MiBitrs-hp11.0791.0790.11180.2180.1192296.231.1981.61961.8%
4 MiBitrs-hp14.2901.0720.24990.4880.25721,184.932.1401.61962.2%
16 MiBitrs-hp115.9770.9990.52401.0230.56944,916.853.7310.61366.8%
64 MiBitrs-hp157.2640.8951.03202.0161.122919,616.506.7030.61374.5%
64 KiBitrs-lstp10.0610.9750.02400.0470.03520.021.0131.27268.4%
256 KiBitrs-lstp10.2861.1440.06060.1180.06410.061.6222.55858.3%
1 MiBitrs-lstp11.2841.2840.13620.2660.13970.252.7162.55852.0%
4 MiBitrs-lstp13.8860.9710.27770.5420.25440.965.0687.45768.6%
16 MiBitrs-lstp115.1360.9460.62421.2190.63923.928.5612.96770.5%
64 MiBitrs-lstp157.0140.8911.21322.3701.228315.6915.4842.96774.9%
64 MiBitrs-hp866.4871.0391.07092.0921.085619,345.046.9961.61964.2%
64 MiBitrs-hp3266.6231.0411.04922.0491.064121,485.287.1581.84464.1%

Source: snippets/RESULTS.md in _simeng_build

14

Statistics of Simulation Output

A stochastic simulator's answer is a random variable. InfSim 06 slide 06 sets out the techniques and slide 07 runs replications and common random numbers live; this card collects them with their costs and pitfalls, and adds how many runs are enough.

Replications, confidence intervals, warm-up, batch means and common random numbersmethod

Principle
Replications: n independent runs (different seeds) give n samples of the metric; report the mean ± tn−1·s/√n. Warm-up deletion: discard the start, when the system is emptier than in steady state (choose the cut-off with Welch's graphical method or a rule such as MSER). Batch means: one long run cut into batches long enough to be nearly independent, which pays warm-up once. Common random numbers (CRN): give two designs the same random streams, so the noise cancels in their difference.
Overhead
Replications cost n runs (they parallelise perfectly; InfSim 08's parallel sweeps); warm-up costs the discarded simulated time in every replication; CRN costs nothing at run time but needs a separate random stream per source of randomness so the streams stay aligned.
Accuracy and pitfalls
The t-interval assumes roughly normal run means; with few runs and skewed metrics (tails) it is optimistic. Batches that are too short are correlated and the interval is too narrow. Not deleting warm-up biases every replication the same way, which more runs do not fix. CRN breaks when one design draws more random numbers than the other from a shared stream.
When to use it
Always for a stochastic simulator. How many runs: the half-width shrinks as 1/√n, so to halve it run four times as many; run a pilot of about ten, then n ≈ (t·s/target half-width)². Tail percentiles need many samples per run as well: a p99 from 800 requests rests on about eight.
Where this GitHub uses it
InfSim 06: statistics that survive review; every --compare in Disaggregated_Inference_Sim uses CRN; InfSim 08 on cutting the runs needed.
Further reading: InfSim 10's simulation textbooks (output analysis chapters)
15

Agreement as a Measurement

The last group measures the simulator itself: does it compute the model, and does the model match something independent? Each check produces a number (differences, error per metric, a bound), so it belongs in a catalogue of measurements.

Golden tests, differential tests and correlationmethod

Principle
A golden test compares today's output with a stored reference run and reports any difference. A differential test runs two independent implementations of the same model (a fast path and a reference, Python and Rust, Python and JavaScript) on the same inputs and requires agreement. A correlation report compares the model against measurement on a defined workload set, with an error per metric and a target.
Overhead
Golden tests are cheap; differential tests cost a second implementation and running both; correlation costs measurement campaigns on real hardware.
Accuracy and pitfalls
Golden files catch change, not error: a bug blessed into the reference is protected by it, so re-bless deliberately and review the diff. Two implementations written from the same misunderstanding agree perfectly. Decide the tolerance per quantity: exact for deterministic code, a tolerance only where maths libraries genuinely differ. A correlation fitted to the workloads it is quoted on is in-sample; hold some out.
When to use it
Golden tests on every simulator; differential tests for every port or fast path; correlation as the deliverable of validation.
Where this GitHub uses it
SimEng 06: golden tests and re-blessing; SimEng 02: golden and differential; InfSim 06: the verification ladder.

Analytic bounds (roofline, schedule bounds)method

Principle
A lower bound on run time from counts alone: the roofline bound max(FLOPs / peak FLOP rate, bytes / bandwidth), and the schedule bound, the busiest resource's total busy time, since nothing finishes before its busiest resource does. Queueing formulas (M/D/1) give exact answers for simple subsystems.
Overhead
Free: a few divisions over a trace.
Accuracy and pitfalls
Optimistic by construction: perfect overlap, peak efficiencies, no dependencies. A real schedule can sit far above it when dependencies serialise work. A simulated result below the bound is not a good result; it is a bug.
When to use it
As the first sanity check on any simulator output, and to say how much room an optimisation could possibly have.
Where this GitHub uses it
InfSim 03: the roofline; FHESim 03: lower bounds (search.py); InfSim 02: a queue checked against theory.

PyTorch's FLOP counter (FlopCounterMode)tool

Principle
A dispatch mode that intercepts every operator a model runs and adds its FLOPs from a per-operator formula (matmul, convolution, attention), by module. It works on the meta device, so even a 70-billion-parameter model is counted without weights.
Overhead
One trace of the model; on the meta device no arithmetic is done.
Accuracy and pitfalls
It counts only operators it has formulas for: on a CPU trace, fused attention (_scaled_dot_product_flash_attention_for_cpu) has none and its FLOPs silently vanish. Conventions differ (a multiply-add is two FLOPs; masked attention scores counted or not).
When to use it
As an independent count to check a front end or a closed-form cost model against.
Where this GitHub uses it
SimEng 10: operator coverage as a metric and the Torch_Sim_Frontend README (it agrees exactly on the meta trace, and misses fused CPU attention).

Two related methods are explained in their own decks: hot-spot attribution and causal profiling (InfSim 06), and the USE method (SimEng 11).

16

What to Take Away

Re-running this deck's measurements

On an idle machine, one at a time: run_t12.py (with valgrind and perf), framework_profilers.py (in Torch_Sim_Frontend's environment), rtl_coverage.py (in RTL_CoSim_NTT's, with Verilator) and cacti/run_cacti.py; then check_snippets.py --only=t12 regenerates RESULTS.md and the deck build inlines every number from it.