A catalogue of the tools and methods the simulation series measure with, grouped by what they measure, each with its principle, overhead, accuracy and pitfalls, when to use it, and where this GitHub uses it: profilers and flame graphs, benchmarks, simulated caches and hardware counters, GPU and CPU energy telemetry, trace viewers and framework profilers, RTL coverage and power flows, DRAM simulators, the CACTI, McPAT and Accelergy estimators (CACTI built and swept here), the statistics of simulation output, and agreement checks. Overheads measured on this series' own simulators.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.
The other decks in this series and its sisters use a lot of measurement tools: profilers, counters, coverage tools, trace viewers, DRAM simulators, power estimators. This deck is where each one is explained, so that a deck naming a tool can link here. Every tool and method is a card with the same five fields:
Every number from this machine comes from snippets/t12/ (run_t12.py, framework_profilers.py, rtl_coverage.py and cacti/run_cacti.py), recorded in RESULTS.md. Where deck 11 already measured something, the card links to it instead of measuring again. Tool versions and vendor products change quickly; each card links to its documentation, which is the authority.
Answer two questions; the chooser names the tools to start with and links to their cards. It encodes the "when to use it" field of the cards, nothing more.
There are two ways to find out where a program's time goes. An instrumenting profiler runs code at every function entry and exit and records exact counts and times. A sampling profiler looks at the call stack every few milliseconds and counts what it sees. The first is exact about calls and distorts time; the second is statistical about time and leaves the program almost untouched.
| Tool | Runs | Median | Slowdown | Runs (s) |
|---|---|---|---|---|
| no tool | 5 | 1.23 s | 1.00x | 1.23, 1.26, 1.19, 1.21, 1.23 |
| cProfile (instrumenting every call) | 3 | 2.51 s | 2.04x | 2.52, 2.51, 2.49 |
| py-spy record, 100 Hz | 3 | 1.29 s | 1.05x | 1.30, 1.29, 1.26 |
| py-spy record, 500 Hz | 3 | 1.78 s | 1.45x | 1.84, 1.78, 1.57 |
| coverage run (line coverage) | 3 | 3.18 s | 2.59x | 3.28, 3.18, 3.06 |
| coverage run --branch | 3 | 3.49 s | 2.85x | 3.53, 3.48, 3.49 |
| python -X tracemalloc=1 (one frame per allocation) | 3 | 5.79 s | 4.72x | 5.84, 5.79, 5.78 |
| python -X tracemalloc=25 (25 frames) | 3 | 31.45 s | 25.63x | 31.60, 31.42, 31.45 |
Source: snippets/RESULTS.md in _simeng_build
sys.setprofile-style events) and records call counts, own time (tottime) and time including callees (cumtime) per function. snakeviz draws the saved profile as an interactive icicle or sunburst of the call tree.py-spy record --pid) with no code changes.--nonblocking avoids the pause at the risk of torn stacks.--idle includes waiting ones. Native extensions need --native. At 500 Hz here, py-spy now and then lost a race with the child's exit and had to be rerun (RESULTS.md).perf stat counts events (cycles, instructions, misses) over a run; perf record samples: every N events or at a frequency it records the instruction pointer and, with --call-graph, the stack. Works for any language whose frames it can unwind.-F); DWARF unwinding copies a chunk of stack per sample and costs far more than frame-pointer unwinding.kernel.perf_event_paranoid); stacks need frame pointers or DWARF info, which release builds and interpreters often lack; more events than counters means multiplexing and scaled estimates (counters card). An interpreter shows up as its own C functions unless it cooperates (Python 3.12's -X perf)./usr/bin/time -v)toolwait4): wall time, user and system CPU time, maximum resident set size, page faults, and voluntary and involuntary context switches. Not the shell's time keyword, which prints only times.A profile of thousands of stacks needs a picture. SimEng 11 slide 06 has interactive flame graphs of four real profiles from this series; this card is the method behind them.
snippets/t11/out of the hub repository.A profiler says where time goes; a benchmark says how long something takes, and whether a change moved it. SimEng 11 slide 10 measured why a benchmark result is a distribution, not a number (median, MAD and a bootstrap interval), and slide 11 designed a regression gate from that noise. Two tools that apply those statistics:
black_box stops the compiler optimising the work away.pytest-benchmark plays the same role (not used in this series).--cache-sim=yes, simulates a first-level and last-level cache); callgrind adds the call graph, so costs can be attributed to callers. Both annotate source lines.LLC-load-misses here), and raw counts mislead: normalise per instruction.perf stat -r), rather than many at once.massif does the same for native heaps./usr/bin/time -v, no tracing) 76.9 MiB.sim.py:132, the per-request inter-token latency lists that the metrics need at the end (table below).| Program | Native (median of 5) | Under cachegrind | Slowdown |
|---|---|---|---|
| Rust disagg-rs, 500 requests | 9.6 ms | 0.89 s | 93x |
| Python disagg-sim, 500 requests | 276.6 ms | 18.58 s | 67x |
Source: snippets/RESULTS.md in _simeng_build
| Line (at the end of the run) | Size | Blocks |
|---|---|---|
src/disagg_sim/sim.py:132 | 23.7 MiB | 766,474 |
src/disagg_sim/sim.py:476 | 4.7 MiB | 46,254 |
src/disagg_sim/sim.py:477 | 2.2 MiB | 61,672 |
src/disagg_sim/sim.py:475 | 1.6 MiB | 30,836 |
Source: snippets/RESULTS.md in _simeng_build
| Event | Count (scaled) | Run-to-run variation | Counted for |
|---|---|---|---|
| cycles | 1,266,070,349 | 0.53% | 27% |
| instructions | 2,656,701,145 | 0.63% | 36% |
| cache-references | 6,164,310 | 2.00% | 36% |
| cache-misses | 2,985,476 | 2.11% | 45% |
| branches | 275,974,295 | 0.57% | 45% |
| branch-misses | 3,015,337 | 2.65% | 45% |
| L1-dcache-loads | 497,366,386 | 0.72% | 35% |
| L1-dcache-load-misses | 19,862,811 | 1.65% | 35% |
| LLC-loads | 4,526,950 | 8.05% | 18% |
| LLC-load-misses | not supported on this CPU | - | - |
| dTLB-loads | 498,474,523 | 0.59% | 18% |
| dTLB-load-misses | 1,301,486 | 4.17% | 18% |
Source: snippets/RESULTS.md in _simeng_build
This machine has no NVIDIA GPU (nvidia-smi and dcgmi are not installed; RESULTS.md records the check), so nothing on this slide was measured here. The cards are written from NVIDIA's and the tools' documentation and from a published measurement study; the outputs shown are documentation excerpts, labelled as such. Drivers and field lists change between releases: the linked documentation is the authority.
nvidia-smi is the command-line front end; nvidia-smi --query-gpu=... --format=csv -lms 100 polls fields to CSV.power.draw.instant from power.draw.average. Their recommended practice (many repetitions, randomised delays, discard the rise time) brought it to about 5% of an external meter.DCGM_FI_DEV_POWER_USAGE (W) and DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION (mJ since the driver loaded), and profiling fields read from the GPU's counters, such as SM active (1002), SM occupancy (1003), tensor pipe active (1004) and DRAM active (1005). dcgmi dmon streams them; dcgm-exporter feeds Prometheus.dcgmi profile -l lists them). Power fields inherit the sensor averaging of the NVML card.$ dcgmi dmon -e 1001,1004,1005
# Entity GRACT TENSO DRAMA
GPU 0 0.969 0.928 0.000monitor.begin_window("step") … end_window("step"). On Volta and newer NVIDIA GPUs it reads NVML's energy counter at the window edges; on older ones it polls power in a separate process. It also reads CPU and DRAM energy through RAPL, which needs root. It underlies the ML.ENERGY leaderboard's energy per request.Both were tried here. What happened, as recorded:
perf stat -e power/energy-pkg/ (the RAPL package-energy event perf lists: power/energy-cores/, power/energy-gpu/, power/energy-pkg/): exit status 1, "No supported events found."; with -a (system-wide): exit status 1. RAPL events are system-wide, so they need perf_event_paranoid 0 or lower, or CAP_PERFMON./sys/class/powercap/intel-rapl:0/energy_uj: mode 400, owned by uid 0; reading it as a user: PermissionError: Permission denied. The world-readable max_energy_range_uj next to it is 65,532,610,987 µJ: the counter wraps after 65.5 kJ, about 14 minutes at the package's 77 W long_term limit.So no CPU energy reading appears in this deck. That is the usual situation for an unprivileged user on a current kernel; the cards say what the reading would mean.
/sys/class/powercap/intel-rapl:*/energy_uj and as perf's power/energy-pkg/ events; energy over an interval is the difference of two reads.energy_uj root-only, because fine-grained energy readings leak secrets through a side channel (PLATYPUS).perf_event_paranoid 0 or lower). For the whole machine, an external meter.A profile aggregates; a trace keeps every event with its start and duration, so overlap, gaps and ordering are visible. The simulators in this series write traces in the same format the framework profilers do, which is the point: a simulated timeline and a measured one open side by side. Both framework profilers ran here on Torch_Sim_Frontend's two-layer tiny Llama:
| Runtime | Plain | Profiled | Overhead | Events recorded |
|---|---|---|---|---|
PyTorch eager, torch.profiler (CPU activity, shapes) | 8.19 ms | 9.42 ms | 1.15x | 15,270 (30 passes) |
ONNX Runtime CPU EP, enable_profiling | 5.69 ms | 9.64 ms | 1.70x | 6,512 (6,288 KB of JSON for 35 runs) |
Source: snippets/RESULTS.md in _simeng_build
| PyTorch operator | Calls | Self CPU share | ONNX operator type | Share of node time | |
|---|---|---|---|---|---|
aten::mm | 450 | 50.9% | MatMul | 54.8% | |
aten::_scaled_dot_product_flash_attention_for_cpu | 60 | 8.9% | Mul | 8.2% | |
aten::mul | 660 | 7.8% | Add | 4.9% | |
aten::index_select | 30 | 5.5% | Transpose | 4.1% | |
aten::silu | 60 | 4.2% | Where | 2.9% | |
aten::add | 420 | 3.5% | Pow | 2.5% |
Source: snippets/RESULTS.md in _simeng_build
"ph":"X", start and duration in microseconds), counters ("C") and metadata ("M"), grouped by process and thread. Perfetto (and chrome://tracing) draws them as a timeline and can query them with SQL.disagg-sim --trace (trace.py) and fhe-sim --trace (trace.py).torch.profiler.profile records an event for every operator dispatched (aten::mm, …) with CPU time, and, on GPUs, kernel times through NVIDIA's CUPTI. Options record input shapes, memory and stacks; key_averages() aggregates, export_chrome_trace writes a trace for Perfetto, and a schedule skips warm-up steps.aten::linear contains aten::mm. The first iterations include one-off costs, so profile after warm-up. On a GPU, CPU-side operator time is launch time, not execution time; look at the kernel rows.SessionOptions.enable_profiling = True, ONNX Runtime writes one Chrome-trace event per graph node per run (operator type, execution provider, input and output shapes, memory deltas) plus session events; end_profiling() returns the file.nsys profile) that records CPU thread activity, CUDA API calls, GPU kernels and memory copies, and user annotations (NVTX ranges) on one timeline. Nsight Compute is its per-kernel counterpart. Not run here: no NVIDIA GPU.{"cat": "Node", "pid": 3327413, "tid": 3327413, "dur": 66, "ts": 61078, "ph": "X", "name": "node_embedding_kernel_time", "args": {"op_name": "Gather", "provider": "CPUExecutionProvider", "input_type_shape": [{"float": [1000, 256]}, {"int64": [1, 128]}], "output_type_shape": [{"float": [1, 128, 256]}]}}Verification progress is measured, too. RTL_CoSim_NTT's butterfly, rebuilt with Verilator's coverage on, run with uniform and with constrained-random stimulus, with its seeded Barrett bug and without:
| Build | Stimulus | Test result | Line | Branch | Toggle | Functional-coverage holes | Build | Run |
|---|---|---|---|---|---|---|---|---|
| plain | uniform | passed | - | - | - | a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows | 13 s | 1.73 s |
| plain | constrained | passed | - | - | - | none | 13 s | 1.52 s |
| plain + bug | uniform | passed (bug missed) | - | - | - | a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows | 12 s | 1.66 s |
| plain + bug | constrained | FAILED (bug caught) | - | - | - | none | 12 s | 1.58 s |
| cov | uniform | passed | 7/7 | 4/4 | 1664/1836 | a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows | 30 s | 1.95 s |
| cov | constrained | passed | 7/7 | 4/4 | 1664/1836 | none | 30 s | 1.96 s |
| cov + bug | uniform | passed (bug missed) | 4/4 | 6/6 | 1664/1836 | a_zero, a_max, b_zero, b_max, w_one, w_max, corr_2_x_sum_wraps, corr_2_x_diff_borrows | 30 s | 1.97 s |
| cov + bug | constrained | FAILED (bug caught) | 4/4 | 6/6 | 1664/1836 | none | 30 s | 1.93 s |
Source: snippets/RESULTS.md in _simeng_build
Uniform stimulus reaches every line and branch, and the same toggles as constrained stimulus, and still misses the bug. Only the functional-coverage holes (corr_2_x_sum_wraps, corr_2_x_diff_borrows) show that the situations that expose it never happened. The toggle points neither stimulus reaches are bits that cannot move for this prime and width.
coverage.py (pytest-cov) traces Python lines; cargo llvm-cov and gcov/gcovr use counters the compiler inserts.--coverage, Verilator counts line (block) and branch execution and every toggle of every signal bit, and writes coverage.dat at the end; verilator_coverage merges and annotates. Functional coverage is different: named bins the testbench samples from each transaction (operand at its maximum, a sum that wraps) and crosses of them, counting the situations the stimulus created.>= to >, + to -) and runs the suite against each; the mutation score is the share it catches.Memory_System_Sim, this series' own command-level DRAM model, was validated four ways: closed forms, an independent protocol checker on Hypothesis-generated traces, behavioural checks, and a cross-check against DRAMsim3 on the same traces, timing and address mapping (SimEng 04 slide 11). The reference simulators:
.ini file of JEDEC timing parameters and organisation; it takes a trace or a CPU simulator's requests, issues commands cycle by cycle under every timing constraint, and reports bandwidth, latency, power and, optionally, temperature.Before RTL exists, power and area come from estimators that build a component from its structure and technology parameters. They are fast, which is the point: they rank design options. They are not sign-off, and their error bars come from their own validation papers.
itrs-hp and itrs-lstp). The smallest node with data is 22 nm; scaling beyond is your assumption.CACTI 7 (HewlettPackard/cacti, built here) swept over scratchpad capacity at 22 nm, with both of its SRAM transistor types and, for 64 MiB, more banks. Only the size, bank count and cell type differ from the upstream cache.cfg; every configuration file and output is in snippets/t12/cacti. These are model outputs from ITRS-based projections, not measurements of silicon.
| Capacity | Cells | Banks | Area (mm²) | mm² per MiB | Read energy (nJ) | pJ per bit read | Write energy (nJ) | Leakage (mW) | Access time (ns) | Cycle time (ns) | Area efficiency |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 64 KiB | itrs-hp | 1 | 0.055 | 0.883 | 0.0150 | 0.029 | 0.0452 | 20.91 | 0.378 | 0.346 | 75.5% |
| 256 KiB | itrs-hp | 1 | 0.293 | 1.171 | 0.0473 | 0.092 | 0.0699 | 77.25 | 0.628 | 0.613 | 56.9% |
| 1 MiB | itrs-hp | 1 | 1.079 | 1.079 | 0.1118 | 0.218 | 0.1192 | 296.23 | 1.198 | 1.619 | 61.8% |
| 4 MiB | itrs-hp | 1 | 4.290 | 1.072 | 0.2499 | 0.488 | 0.2572 | 1,184.93 | 2.140 | 1.619 | 62.2% |
| 16 MiB | itrs-hp | 1 | 15.977 | 0.999 | 0.5240 | 1.023 | 0.5694 | 4,916.85 | 3.731 | 0.613 | 66.8% |
| 64 MiB | itrs-hp | 1 | 57.264 | 0.895 | 1.0320 | 2.016 | 1.1229 | 19,616.50 | 6.703 | 0.613 | 74.5% |
| 64 KiB | itrs-lstp | 1 | 0.061 | 0.975 | 0.0240 | 0.047 | 0.0352 | 0.02 | 1.013 | 1.272 | 68.4% |
| 256 KiB | itrs-lstp | 1 | 0.286 | 1.144 | 0.0606 | 0.118 | 0.0641 | 0.06 | 1.622 | 2.558 | 58.3% |
| 1 MiB | itrs-lstp | 1 | 1.284 | 1.284 | 0.1362 | 0.266 | 0.1397 | 0.25 | 2.716 | 2.558 | 52.0% |
| 4 MiB | itrs-lstp | 1 | 3.886 | 0.971 | 0.2777 | 0.542 | 0.2544 | 0.96 | 5.068 | 7.457 | 68.6% |
| 16 MiB | itrs-lstp | 1 | 15.136 | 0.946 | 0.6242 | 1.219 | 0.6392 | 3.92 | 8.561 | 2.967 | 70.5% |
| 64 MiB | itrs-lstp | 1 | 57.014 | 0.891 | 1.2132 | 2.370 | 1.2283 | 15.69 | 15.484 | 2.967 | 74.9% |
| 64 MiB | itrs-hp | 8 | 66.487 | 1.039 | 1.0709 | 2.092 | 1.0856 | 19,345.04 | 6.996 | 1.619 | 64.2% |
| 64 MiB | itrs-hp | 32 | 66.623 | 1.041 | 1.0492 | 2.049 | 1.0641 | 21,485.28 | 7.158 | 1.844 | 64.1% |
Source: snippets/RESULTS.md in _simeng_build
itrs-hp is ITRS high-performance logic transistors, itrs-lstp low-standby-power ones; the difference in leakage (over three orders of magnitude) is the transistor model, not the array. Large on-chip memories are built from low-leakage cells; the high-performance numbers are a pessimistic bound.A stochastic simulator's answer is a random variable. InfSim 06 slide 06 sets out the techniques and slide 07 runs replications and common random numbers live; this card collects them with their costs and pitfalls, and adds how many runs are enough.
--compare in Disaggregated_Inference_Sim uses CRN; InfSim 08 on cutting the runs needed.The last group measures the simulator itself: does it compute the model, and does the model match something independent? Each check produces a number (differences, error per metric, a bound), so it belongs in a catalogue of measurements.
_scaled_dot_product_flash_attention_for_cpu) has none and its FLOPs silently vanish. Conventions differ (a multiply-add is two FLOPs; masked attention scores counted or not).Two related methods are explained in their own decks: hot-spot attribution and causal profiling (InfSim 06), and the USE method (SimEng 11).
On an idle machine, one at a time: run_t12.py (with valgrind and perf), framework_profilers.py (in Torch_Sim_Frontend's environment), rtl_coverage.py (in RTL_CoSim_NTT's, with Verilator) and cacti/run_cacti.py; then check_snippets.py --only=t12 regenerates RESULTS.md and the deck build inlines every number from it.