Simulation Engineering Toolkit — Presentation 09

Specifications, Requirements and Test Plans

Writing down what a simulator must do, and proving it does: requirement quality, EARS patterns with an interactive checker, non-functional requirements for simulators, requirements capture for mission-mode software, the V-model, a traceability matrix generated from the test run (which found a real gap), and worked templates for a specification, a test plan and a performance report.

EARS Requirement quality Mission mode V-model Traceability Test plans
Need → Requirement → Test → Trace → Report → Sign off
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

01

Why a Simulator Needs a Specification

A simulator's answers feed decisions: how many lanes, how much memory bandwidth, whether a design meets its target before silicon exists. Without a written statement of what it must do, and how well, nobody can say whether it is right, only whether it runs. A specification makes "right" checkable:

This deck uses one complete worked example throughout: the specification, test plan and traceability matrix of Torch_Sim_Frontend (deck 10), written for this purpose and kept in its repository. The test-level detail is in deck 06.

02

What Makes a Good Requirement

INCOSE's Guide to Writing Requirements and ISO/IEC/IEEE 29148 (requirements engineering) describe the same core characteristics. For an individual requirement these include being necessary, appropriate to its level, unambiguous, complete, singular (one thing), feasible, verifiable, correct and conforming to the agreed style. A set of requirements must also be complete and consistent as a whole.

Writing rule (paraphrased)Instead ofWrite
Structured statement, active voice, one "shall"Results shall be written and plotted.The simulator shall write the results to results.md.
No vague termsThe model shall be fast and accurate.The model shall report bootstrap time within 5% of the RTL reference.
No escape clauses..., where possible.(Decide. If it is optional, use the optional-feature pattern.)
No open-ended clauses... formats including JSON, CSV, etc.... in JSON and in CSV.
No superfluous infinitivesThe tool shall be able to trace 70B models.The tool shall trace models of up to 71 billion parameters.
Measurable... with low memory.... with a peak resident memory below 2 GB.

Verifiable is the one engineers break most: if no test, analysis, inspection or demonstration could fail it, it is a wish, not a requirement. Each requirement names its verification method (Test, Analysis, Inspection, Demonstration).

03

EARS: Five Patterns for Requirements

The Easy Approach to Requirements Syntax (Mavin, Wilkinson, Harwood and Novak, IEEE RE 2009, doi:10.1109/RE.2009.9) constrains natural language to a few templates. The keyword at the start says what kind of requirement it is, which removes a whole class of ambiguity (the patterns):

PatternTemplateFrom simfront's spec
UbiquitousThe <system> shall <response>SF-09: The roofline cost model shall use the same achievable FLOP and byte rates as Disaggregated_Inference_Sim's CostModel for the same device.
Event-drivenWhen <trigger>, the <system> shall <response>SF-02: When a model is traced on the meta device, the front end shall allocate no parameter storage.
State-drivenWhile <precondition>, the <system> shall <response>SF-08: While the torch.compile backend is active, the user's program shall produce outputs identical to eager execution.
Unwanted behaviourIf <trigger>, then the <system> shall <response>SF-06: If an operator has no cost rule, then the front end shall cost it at zero and name it in the coverage report.
Optional featureWhere <feature is included>, the <system> shall <response>SF-10: Where the offload cost model is selected, the cost model shall charge a link transfer the first time an operator reads a tensor resident on the other side, and not again.
ComplexWhile ..., when ..., the <system> shall ...While in mission mode, when the link reports an uncorrectable error, the controller shall ... (slide 06)

Unwanted-behaviour requirements are the ones teams forget: they come from asking "what if it goes wrong?" of every event. In simfront that question produced SF-06 (unknown operators) and SF-15 (a library upgrade that changes the arithmetic).

04

Interactive: Check a Requirement

Type or pick a requirement. The checker names its EARS pattern and flags the problems the writing rules warn about: vague and unmeasurable words, escape and open-ended clauses, superfluous infinitives, passive voice, non-binding verbs and more than one "shall". It is a set of simple patterns, not a judge: a clean result means only that nothing obvious is wrong.

05

Functional and Non-Functional Requirements for Simulators

Functional requirements say what the simulator does: which components and behaviours it models, which inputs it accepts, which outputs it produces. Non-functional requirements say how well, and for a simulator they carry most of the value:

QualityRequirement shapeEvidence in this series
AccuracyMetric X within tolerance T of reference R, on workloads WSystemC model against the SimPy model, op by op (deck 03); memsim against DRAMsim3 (deck 04)
DeterminismThe same inputs and seed shall give bit-identical outputsThe Rust port, bit-exact against Python (deck 02)
PerformanceAt least N simulated events per second on machine MRust_DES_Kernel's criterion gate (deck 07)
CapacityModels up to size S within memory BSF-14: every reference model traced within 2 GB
Fidelity boundariesWhat is not modelled, stated as scopeEvery repository's "Not modelled" section
MaintainabilityA change to X shall fail CI if it changes YSF-15; the cycle-count and behaviour gates

An accuracy requirement without a named reference and tolerance is unverifiable. Writing it forces the important conversation early: "accurate against what, and how close is close enough for the decision this supports?"

06

Requirements Capture for Mission-Mode Software

Usage varies between organisations, so define the term before using it. Here, mission-mode software means the software that runs on the product in its operational mode, doing the job the product exists for. That is as opposed to the test, bring-up, calibration, diagnostic and manufacturing software that runs on the same hardware at other times. Its requirements are captured differently because the stakes, the environment and the evidence expected are different:

ConcernBring-up and test softwareMission-mode software
Who depends on itEngineers in the labCustomers and their systems, unattended
ModesOften one, entered by handExplicit modes and transitions: state-driven requirements (While in mission mode, ...)
FaultsStop and reportDetect, contain, recover or degrade gracefully: unwanted-behaviour requirements for each failure
Timing and resourcesBest effortBudgets: latency, memory, power, with margins
EvidenceIt works on the benchTraceability from each requirement to verification, and in regulated domains a process standard
07

The V-Model

user needs requirements (spec.md) design implementation unit tests (rules) system tests (routes) validation validation: does it answer the right question? verification: traceability matrix unit tests
08

Traceability, Generated From the Test Run

simfront's tests name the requirements they verify (@pytest.mark.req("SF-04")). ci/trace_matrix.py runs the suite in-process, collects each test's markers and outcome, joins them to the requirement table in docs/spec.md and writes docs/traceability.md. It checks both directions: every requirement has a passing test, and every test traces to a requirement or is listed as untraced.

ci/trace_matrix.py: one row per requirement source
for r in reqs:
    tests = by_req.get(r["id"], [])
    if r["method"].startswith("T"):
        res = [col.outcome.get(t, "not run") for t in tests]
        ok = bool(tests) and all(x == "passed" for x in res)
        gaps += not ok
        names = "<br>".join(f"`{t.split('/')[-1]}`" for t in tests) or "**no test**"
        verdict = f"{sum(x == 'passed' for x in res)}/{len(tests)} passed" if tests else "**GAP**"
    else:
        names, verdict = demonstration_evidence(r), "demonstrated"
    lines.append(f"| **{r['id']}** {r['text']} | {r['pattern']} | {r['method']} | {names} | {verdict} |")
RequirementPatternMethodVerified byResult
SF-01 The front end shall record, for every captured operator, the shape, element size and identity of each input and output tensor and whether it is a model parameter.UbiquitousTtest_real_models.py::test_llama3_8b_trace_shapes<br>test_trace.py::test_json_round_trip2/2 passed
SF-02 When a model is traced on the meta device, the front end shall allocate no parameter storage.Event-drivenTtest_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-8b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-70b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[mistral-7b]3/3 passed
SF-03 For SwiGLU decoder configurations, the traced weight-matmul FLOPs of a prefill shall equal 2 × matmul parameters × tokens exactly.UbiquitousTtest_properties.py::test_traced_matmul_flops_equal_closed_form<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-8b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-70b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[mistral-7b]<br>test_real_models.py::test_qwen_biases_are_counted<br>test_real_models.py::test_matches_pytorch_flop_counter6/6 passed
SF-04 The four front ends shall report identical matmul FLOPs and weight bytes for the same model and input, apart from constant folding that is documented per case.UbiquitousTtest_real_models.py::test_fake_cpu_gives_fused_attention_with_the_same_arithmetic<br>test_routes.py::test_routes_agree_on_flops_and_weights[llama]<br>test_routes.py::test_routes_agree_on_flops_and_weights[gpt2]<br>test_routes.py::test_routes_agree_on_the_meta_device4/4 passed
SF-05 When a decode step is traced, the traced matmul FLOPs shall equal the closed form plus the rotary-frequency matmul, and the traced weight bytes shall equal the closed form's per-step weight traffic plus the RMSNorm weights.Event-drivenTtest_real_models.py::test_decode_step_matches_closed_form1/1 passed
SF-06 If an operator has no cost rule, then the front end shall cost it at zero and name it in the coverage report.Unwanted behaviourTtest_coverage.py::test_unknown_operator_is_reported<br>test_coverage.py::test_rule_coverage_of_real_traces_is_complete<br>test_rules.py::test_unknown_ops_are_zero_and_visible3/3 passed
SF-07 When a model is exported to ONNX without weights, the front end shall recover the shape and parameter identity of every weight.Event-drivenTtest_routes.py::test_onnx_export_runs_in_onnxruntime_and_matches_pytorch1/1 passed
SF-08 While the torch.compile backend is active, the user's program shall produce outputs identical to eager execution.State-drivenTtest_routes.py::test_compile_backend_returns_correct_outputs1/1 passed
SF-09 The roofline cost model shall use the same achievable FLOP and byte rates as Disaggregated_Inference_Sim's CostModel for the same device.UbiquitousTtest_cost.py::test_from_device_uses_the_inference_simulators_rates1/1 passed
SF-10 Where the offload cost model is selected, the cost model shall charge a link transfer the first time an operator reads a tensor resident on the other side, and not again.Optional featureTtest_cost.py::test_offload_with_everything_supported_is_the_roofline<br>test_cost.py::test_offload_moves_each_tensor_once_per_side2/2 passed
SF-11 When a model without data-dependent Python control flow is traced under fake tensors, the front end shall record the same operator sequence and shapes as a run with real tensors.Event-drivenTtest_routes.py::test_fake_tensors_trace_exactly_what_real_tensors_run1/1 passed
SF-12 The front end shall report operator coverage by operator count, FLOPs and time.UbiquitousTtest_coverage.py::test_device_coverage_three_ways1/1 passed
SF-13 The dispatch front end shall capture a Llama-3-8B 2,048-token prefill trace in less than 5 s on the CI machine.UbiquitousTtest_gate.py::test_gate_fails_when_capture_slows_down<br>test_real_models.py::test_capture_of_llama3_8b_prefill_is_fast2/2 passed
SF-14 The front end shall trace every reference model with a peak resident memory below 2 GB.UbiquitousD (results.md)examples/results.md: peak resident memory 642 MBdemonstrated
SF-15 If a library upgrade changes the traced FLOPs or weight bytes of a reference trace, then CI shall fail.Unwanted behaviourTtest_gate.py::test_gate_fails_on_an_arithmetic_change_and_tolerates_drift1/1 passed
SF-16 The front end shall run on a CPU-only machine with no GPU and no model weights.UbiquitousD (GitHub Actions)GitHub Actions runs the whole suite on ubuntu-24.04 runners with CPU-only PyTorchdemonstrated
SF-17 The cost rules shall count each operator's FLOPs and bytes by the conventions stated in rules.py and cost.py.UbiquitousTtest_cost.py::test_roofline_is_max_of_compute_and_memory_per_op<br>test_cost.py::test_fused_bound_is_below_unfused<br>test_cost.py::test_faster_hardware_is_never_slower<br>test_rules.py::test_mm_is_two_mkn<br>test_rules.py::test_bmm_addmm_linear_conv<br>test_rules.py::test_fused_attention_counts_qk_av_and_softmax<br>test_rules.py::test_views_are_free_and_copies_are_not<br>test_rules.py::test_elementwise_reduction_softmax<br>test_rules.py::test_embedding_reads_rows_not_the_table<br>test_rules.py::test_onnx_rules10/10 passed
SF-18 The accelerator model shall lower every operator that has a cost rule into tiles that fit the tile budget, and the tiles of a GEMM shall perform exactly its multiply-accumulates.UbiquitousTtest_accel_lower.py::test_gemm_dims_match_the_flop_rules<br>test_accel_lower.py::test_tiles_fit_and_cover_every_mac[4096]<br>test_accel_lower.py::test_tiles_fit_and_cover_every_mac[65536]<br>test_accel_lower.py::test_tiles_fit_and_cover_every_mac[1048576]4/4 passed
SF-19 While the on-chip buffer cannot hold the next tile, the load DMA shall not start its transfer, and buffer occupancy shall never exceed its capacity.State-drivenTtest_accel_sim.py::test_buffer_never_overflows[65536]<br>test_accel_sim.py::test_buffer_never_overflows[262144]<br>test_accel_sim.py::test_buffer_never_overflows[2097152]<br>test_accel_sim.py::test_backpressure_stalls_loads_when_compute_is_slow4/4 passed
SF-20 When an operator reads an activation, its first load shall not start before the operator that produced it has stored its results.Event-drivenTtest_accel_lower.py::test_dependencies_follow_activations_not_weights<br>test_accel_sim.py::test_loads_wait_for_producers_to_be_stored2/2 passed
SF-21 The accelerator model shall report latency, utilisation per component, a stall breakdown that sums to the latency, a hot-spot, a per-operator latency histogram, a timeline plot and a Chrome trace.UbiquitousTtest_accel_sim.py::test_stalls_sum_to_the_makespan[kw0-cnn]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw0-llama]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw0-gpt2]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw1-cnn]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw1-llama]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw1-gpt2]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw2-cnn]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw2-llama]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw2-gpt2]<br>test_accel_sim.py::test_hotspot_moves_with_the_bottleneck<br>test_accel_sim.py::test_outputs_report_histogram_timeline_and_chrome_trace11/11 passed
SF-22 While the two DMA engines cannot contend for a memory channel, the fast path (Python and C++) shall produce per-tile timings bit-identical to the SimPy model's.State-drivenTtest_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw0-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw0-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw0-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw1-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw1-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw1-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw2-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw2-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw2-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw3-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw3-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw3-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw4-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw4-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw4-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw0-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw0-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw0-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw1-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw1-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw1-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw2-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw2-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw2-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw3-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw3-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw3-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw4-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw4-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw4-gpt2]<br>test_accel_fastpath.py::test_random_programs_agree31/31 passed
SF-23 If the configuration lets the DMA engines contend for one memory channel, then the fast path shall refuse to run.Unwanted behaviourTtest_accel_fastpath.py::test_fast_path_refuses_a_shared_channel1/1 passed
SF-24 Where durations are quantised to whole cycles, the cycle-stepped twin shall produce per-tile timings identical to the event-driven model's.Optional featureTtest_accel_cycle.py::test_cycle_twin_equals_event_driven[kw0-cnn]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw0-llama]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw1-cnn]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw1-llama]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw2-cnn]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw2-llama]6/6 passed
SF-25 The NTT polynomial product shall equal schoolbook multiplication in Z_q[X]/(X^N + 1).UbiquitousTtest_accel_ntt.py::test_polymul_equals_schoolbook1/1 passed
SF-26 Where an in-transit stage is configured, the model shall run NTT operators in the read path, never faster than the stage's operations-per-byte budget allows.Optional featureTtest_accel_ntt.py::test_in_transit_stage_respects_its_budget[0.5]<br>test_accel_ntt.py::test_in_transit_stage_respects_its_budget[1.6]<br>test_accel_ntt.py::test_in_transit_stage_respects_its_budget[3.0]<br>test_accel_ntt.py::test_in_transit_stage_respects_its_budget[12.0]4/4 passed
SF-27 The FIFO model shall satisfy Little's law (time-averaged depth = throughput x mean time in the FIFO) on every run, with depth never above capacity.UbiquitousTtest_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-1]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-2]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-4]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-16]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-1]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-2]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-4]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-16]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-1]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-2]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-4]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-16]<br>test_accel_fifo.py::test_backpressure_stalls_a_fast_producer13/13 passed
SF-28 The execution-provider partitioner shall assign every ONNX node exactly once, and the ONNX Runtime profile shall name the provider of every node run.UbiquitousTtest_accel_ep.py::test_claim_assigns_every_node_once<br>test_accel_ep.py::test_ort_profile_names_the_provider_of_every_node2/2 passed

Source: docs/traceability.md in Torch_Sim_Frontend

Its first run found a real gap

SF-15 ("if a library upgrade changes the traced FLOPs or weight bytes of a reference trace, then CI shall fail") was claimed by the CI gate, with the method "T (CI gate)". The matrix reported GAP: nothing tested that the gate actually fails. tests/test_gate.py now changes one FLOP in a copy of the baseline and checks the gate's verdict. The spec's change history records the change of method (version 1.1).

09

Template 1: a Simulator Specification

SectionContentsIn simfront's spec
1. ScopeWhat the simulator is for, which decisions it informs, who uses it, what it is not"A design-exploration tool ... not mission-mode software"
2. DefinitionsEvery term a requirement relies on; the references accuracy is measured againstTrace, front end, weight bytes, closed form, reference models; tile, program, engines (added with the accelerator model)
3. RequirementsID, type, EARS pattern, text, verification method; one tableSF-01 to SF-17 (front end and cost models); SF-18 to SF-28 (the SimPy accelerator model, deck 14)
4. FidelityWhat is modelled, at what abstraction; what is notIn the README's "Not modelled"
5. Assumptions and constraintsVersions, platforms, input limitsPyTorch 2.6+, batch-1 static shapes
6. Change historyEvery change, with its reason1.1: SF-15's method changed after the gap; 1.2: SF-05 reworded when the closed form it checks was corrected; 1.3: SF-18 to SF-28 added for the accelerator model

The requirements table, as written (excerpt):

docs/spec.md source
| ID | Type | Pattern | Requirement | Verification |
|---|---|---|---|---|
| SF-01 | Functional | Ubiquitous | The front end shall record, for every captured operator, the shape, element size and identity of each input and output tensor and whether it is a model parameter. | T |
| SF-02 | Functional | Event-driven | When a model is traced on the meta device, the front end shall allocate no parameter storage. | T |
| SF-03 | Functional | Ubiquitous | For SwiGLU decoder configurations, the traced weight-matmul FLOPs of a prefill shall equal 2 × matmul parameters × tokens exactly. | T |
| SF-04 | Functional | Ubiquitous | The four front ends shall report identical matmul FLOPs and weight bytes for the same model and input, apart from constant folding that is documented per case. | T |
| SF-05 | Functional | Event-driven | When a decode step is traced, the traced matmul FLOPs shall equal the closed form plus the rotary-frequency matmul, and the traced weight bytes shall equal the closed form's per-step weight traffic plus the RMSNorm weights. | T |
| SF-06 | Functional | Unwanted behaviour | If an operator has no cost rule, then the front end shall cost it at zero and name it in the coverage report. | T |

A row per requirement, in a file under version control next to the code, reviewed in the same pull request as the change it describes.

10

Template 2: a Test Plan

A test plan says what will be tested, how a result is judged right, and when testing is finished. Its sections follow the usual contents (as in ISO/IEC/IEEE 29119-3), cut to what a small tool needs: test items and scope; approach; pass/fail criteria; entry and exit criteria; environment; deliverables; risks. The approach is the heart of it, because a test is only as good as its oracle, and each level here uses one that shares no code with what it checks:

LevelWhat is testedOracleRequirements
UnitEach cost ruleFLOP and byte counts worked by handSF-06, SF-17
Unit (property)Rules and cost models for any inputInvariants: 2MKN; faster hardware is never slowerSF-17
IntegrationEach front end on real configurationsDisaggregated_Inference_Sim's closed forms; PyTorch's FlopCounterModeSF-02, SF-03, SF-05
Integration (property)Random Llama-shaped configs, lengths, batchesThe closed form, exactly (Hypothesis)SF-03
DifferentialThe four front ends against each otherEach other: they must agreeSF-04, SF-07
FaithfulnessFake tensors and the compile backendA real run with data; eager outputsSF-08, SF-11
ExternalThe exported ONNX modelONNX Runtime's outputs against PyTorch'sSF-07
SystemCLI, gate, traceabilityExpected text; the gate's own failure casesSF-13, SF-15
UnitAccelerator loweringThe cost rules' FLOPs (2 x MACs); hand-worked im2col shapesSF-18
UnitAccelerator invariantsOccupancy within capacity; event order per tile; stalls sum to latencySF-19, SF-20, SF-21
DifferentialSimPy model against the fast path (Python, C++) and the cycle-stepped twinEach other, exactly, on real traces and on random programs (Hypothesis)SF-22, SF-23, SF-24
AnalyticFIFO modelLittle's law, exactly, on every runSF-27
Golden modelNTT polynomial productSchoolbook multiplication (Hypothesis)SF-25
BoundIn-transit stageThe stage's operations-per-byte budgetSF-26
ExternalExecution-provider partitioningONNX Runtime's own profileSF-28

Source: docs/test_plan.md in Torch_Sim_Frontend

Full plan: docs/test_plan.md.

11

Template 3: a Performance Report

Every code repository in this series has one: examples/results.md, written by examples/results.py, which is the only source of the numbers in its README and decks. The structure generalises to any simulation study:

SectionContentsExample
QuestionThe decision the study informs"Does a matmul-only engine need more operator coverage?"
EnvironmentVersions, machine, commit: enough to reproducesimfront results.md §1, with the commit hash
Method and configurationModel, workload, parameters; what is illustrative"Device numbers ... are datasheet-level or illustrative, as labelled"
ValidationWhy the model can be trusted for this question§4: the trace against the closed form; §7: against FlopCounterMode
ResultsTables and charts, with spread where there is noise§6: coverage counted by operators, FLOPs and time
LimitationsWhat could change the conclusion"Not modelled" plus the counting conventions

The environment section from simfront's report, generated, never typed:

ItemValue
Python3.12.12
PyTorch2.14.1+cpu
transformers5.18.0
onnx1.23.1
CPUx86_64
simfront commitcc64742

Source: examples/results.md in Torch_Sim_Frontend

A report generated by a script can be regenerated after any model change, which is what keeps the numbers in decks and READMEs true. Timing results also need repeat runs and a statement of spread (deck 11).

12

Reviews, Baselines and Sign-Off

13

What to Take Away