Writing down what a simulator must do, and proving it does: requirement quality, EARS patterns with an interactive checker, non-functional requirements for simulators, requirements capture for mission-mode software, the V-model, a traceability matrix generated from the test run (which found a real gap), and worked templates for a specification, a test plan and a performance report.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.
A simulator's answers feed decisions: how many lanes, how much memory bandwidth, whether a design meets its target before silicon exists. Without a written statement of what it must do, and how well, nobody can say whether it is right, only whether it runs. A specification makes "right" checkable:
This deck uses one complete worked example throughout: the specification, test plan and traceability matrix of Torch_Sim_Frontend (deck 10), written for this purpose and kept in its repository. The test-level detail is in deck 06.
INCOSE's Guide to Writing Requirements and ISO/IEC/IEEE 29148 (requirements engineering) describe the same core characteristics. For an individual requirement these include being necessary, appropriate to its level, unambiguous, complete, singular (one thing), feasible, verifiable, correct and conforming to the agreed style. A set of requirements must also be complete and consistent as a whole.
| Writing rule (paraphrased) | Instead of | Write |
|---|---|---|
| Structured statement, active voice, one "shall" | Results shall be written and plotted. | The simulator shall write the results to results.md. |
| No vague terms | The model shall be fast and accurate. | The model shall report bootstrap time within 5% of the RTL reference. |
| No escape clauses | ..., where possible. | (Decide. If it is optional, use the optional-feature pattern.) |
| No open-ended clauses | ... formats including JSON, CSV, etc. | ... in JSON and in CSV. |
| No superfluous infinitives | The tool shall be able to trace 70B models. | The tool shall trace models of up to 71 billion parameters. |
| Measurable | ... with low memory. | ... with a peak resident memory below 2 GB. |
Verifiable is the one engineers break most: if no test, analysis, inspection or demonstration could fail it, it is a wish, not a requirement. Each requirement names its verification method (Test, Analysis, Inspection, Demonstration).
The Easy Approach to Requirements Syntax (Mavin, Wilkinson, Harwood and Novak, IEEE RE 2009, doi:10.1109/RE.2009.9) constrains natural language to a few templates. The keyword at the start says what kind of requirement it is, which removes a whole class of ambiguity (the patterns):
| Pattern | Template | From simfront's spec |
|---|---|---|
| Ubiquitous | The <system> shall <response> | SF-09: The roofline cost model shall use the same achievable FLOP and byte rates as Disaggregated_Inference_Sim's CostModel for the same device. |
| Event-driven | When <trigger>, the <system> shall <response> | SF-02: When a model is traced on the meta device, the front end shall allocate no parameter storage. |
| State-driven | While <precondition>, the <system> shall <response> | SF-08: While the torch.compile backend is active, the user's program shall produce outputs identical to eager execution. |
| Unwanted behaviour | If <trigger>, then the <system> shall <response> | SF-06: If an operator has no cost rule, then the front end shall cost it at zero and name it in the coverage report. |
| Optional feature | Where <feature is included>, the <system> shall <response> | SF-10: Where the offload cost model is selected, the cost model shall charge a link transfer the first time an operator reads a tensor resident on the other side, and not again. |
| Complex | While ..., when ..., the <system> shall ... | While in mission mode, when the link reports an uncorrectable error, the controller shall ... (slide 06) |
Unwanted-behaviour requirements are the ones teams forget: they come from asking "what if it goes wrong?" of every event. In simfront that question produced SF-06 (unknown operators) and SF-15 (a library upgrade that changes the arithmetic).
Type or pick a requirement. The checker names its EARS pattern and flags the problems the writing rules warn about: vague and unmeasurable words, escape and open-ended clauses, superfluous infinitives, passive voice, non-binding verbs and more than one "shall". It is a set of simple patterns, not a judge: a clean result means only that nothing obvious is wrong.
Functional requirements say what the simulator does: which components and behaviours it models, which inputs it accepts, which outputs it produces. Non-functional requirements say how well, and for a simulator they carry most of the value:
| Quality | Requirement shape | Evidence in this series |
|---|---|---|
| Accuracy | Metric X within tolerance T of reference R, on workloads W | SystemC model against the SimPy model, op by op (deck 03); memsim against DRAMsim3 (deck 04) |
| Determinism | The same inputs and seed shall give bit-identical outputs | The Rust port, bit-exact against Python (deck 02) |
| Performance | At least N simulated events per second on machine M | Rust_DES_Kernel's criterion gate (deck 07) |
| Capacity | Models up to size S within memory B | SF-14: every reference model traced within 2 GB |
| Fidelity boundaries | What is not modelled, stated as scope | Every repository's "Not modelled" section |
| Maintainability | A change to X shall fail CI if it changes Y | SF-15; the cycle-count and behaviour gates |
An accuracy requirement without a named reference and tolerance is unverifiable. Writing it forces the important conversation early: "accurate against what, and how close is close enough for the decision this supports?"
Usage varies between organisations, so define the term before using it. Here, mission-mode software means the software that runs on the product in its operational mode, doing the job the product exists for. That is as opposed to the test, bring-up, calibration, diagnostic and manufacturing software that runs on the same hardware at other times. Its requirements are captured differently because the stakes, the environment and the evidence expected are different:
| Concern | Bring-up and test software | Mission-mode software |
|---|---|---|
| Who depends on it | Engineers in the lab | Customers and their systems, unattended |
| Modes | Often one, entered by hand | Explicit modes and transitions: state-driven requirements (While in mission mode, ...) |
| Faults | Stop and report | Detect, contain, recover or degrade gracefully: unwanted-behaviour requirements for each failure |
| Timing and resources | Best effort | Budgets: latency, memory, power, with margins |
| Evidence | It works on the bench | Traceability from each requirement to verification, and in regulated domains a process standard |
simfront's tests name the requirements they verify (@pytest.mark.req("SF-04")). ci/trace_matrix.py runs the suite in-process, collects each test's markers and outcome, joins them to the requirement table in docs/spec.md and writes docs/traceability.md. It checks both directions: every requirement has a passing test, and every test traces to a requirement or is listed as untraced.
for r in reqs:
tests = by_req.get(r["id"], [])
if r["method"].startswith("T"):
res = [col.outcome.get(t, "not run") for t in tests]
ok = bool(tests) and all(x == "passed" for x in res)
gaps += not ok
names = "<br>".join(f"`{t.split('/')[-1]}`" for t in tests) or "**no test**"
verdict = f"{sum(x == 'passed' for x in res)}/{len(tests)} passed" if tests else "**GAP**"
else:
names, verdict = demonstration_evidence(r), "demonstrated"
lines.append(f"| **{r['id']}** {r['text']} | {r['pattern']} | {r['method']} | {names} | {verdict} |")
| Requirement | Pattern | Method | Verified by | Result |
|---|---|---|---|---|
| SF-01 The front end shall record, for every captured operator, the shape, element size and identity of each input and output tensor and whether it is a model parameter. | Ubiquitous | T | test_real_models.py::test_llama3_8b_trace_shapes<br>test_trace.py::test_json_round_trip | 2/2 passed |
| SF-02 When a model is traced on the meta device, the front end shall allocate no parameter storage. | Event-driven | T | test_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-8b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-70b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[mistral-7b] | 3/3 passed |
| SF-03 For SwiGLU decoder configurations, the traced weight-matmul FLOPs of a prefill shall equal 2 × matmul parameters × tokens exactly. | Ubiquitous | T | test_properties.py::test_traced_matmul_flops_equal_closed_form<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-8b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[llama3-70b]<br>test_real_models.py::test_prefill_matmul_flops_equal_closed_form[mistral-7b]<br>test_real_models.py::test_qwen_biases_are_counted<br>test_real_models.py::test_matches_pytorch_flop_counter | 6/6 passed |
| SF-04 The four front ends shall report identical matmul FLOPs and weight bytes for the same model and input, apart from constant folding that is documented per case. | Ubiquitous | T | test_real_models.py::test_fake_cpu_gives_fused_attention_with_the_same_arithmetic<br>test_routes.py::test_routes_agree_on_flops_and_weights[llama]<br>test_routes.py::test_routes_agree_on_flops_and_weights[gpt2]<br>test_routes.py::test_routes_agree_on_the_meta_device | 4/4 passed |
| SF-05 When a decode step is traced, the traced matmul FLOPs shall equal the closed form plus the rotary-frequency matmul, and the traced weight bytes shall equal the closed form's per-step weight traffic plus the RMSNorm weights. | Event-driven | T | test_real_models.py::test_decode_step_matches_closed_form | 1/1 passed |
| SF-06 If an operator has no cost rule, then the front end shall cost it at zero and name it in the coverage report. | Unwanted behaviour | T | test_coverage.py::test_unknown_operator_is_reported<br>test_coverage.py::test_rule_coverage_of_real_traces_is_complete<br>test_rules.py::test_unknown_ops_are_zero_and_visible | 3/3 passed |
| SF-07 When a model is exported to ONNX without weights, the front end shall recover the shape and parameter identity of every weight. | Event-driven | T | test_routes.py::test_onnx_export_runs_in_onnxruntime_and_matches_pytorch | 1/1 passed |
SF-08 While the torch.compile backend is active, the user's program shall produce outputs identical to eager execution. | State-driven | T | test_routes.py::test_compile_backend_returns_correct_outputs | 1/1 passed |
SF-09 The roofline cost model shall use the same achievable FLOP and byte rates as Disaggregated_Inference_Sim's CostModel for the same device. | Ubiquitous | T | test_cost.py::test_from_device_uses_the_inference_simulators_rates | 1/1 passed |
| SF-10 Where the offload cost model is selected, the cost model shall charge a link transfer the first time an operator reads a tensor resident on the other side, and not again. | Optional feature | T | test_cost.py::test_offload_with_everything_supported_is_the_roofline<br>test_cost.py::test_offload_moves_each_tensor_once_per_side | 2/2 passed |
| SF-11 When a model without data-dependent Python control flow is traced under fake tensors, the front end shall record the same operator sequence and shapes as a run with real tensors. | Event-driven | T | test_routes.py::test_fake_tensors_trace_exactly_what_real_tensors_run | 1/1 passed |
| SF-12 The front end shall report operator coverage by operator count, FLOPs and time. | Ubiquitous | T | test_coverage.py::test_device_coverage_three_ways | 1/1 passed |
| SF-13 The dispatch front end shall capture a Llama-3-8B 2,048-token prefill trace in less than 5 s on the CI machine. | Ubiquitous | T | test_gate.py::test_gate_fails_when_capture_slows_down<br>test_real_models.py::test_capture_of_llama3_8b_prefill_is_fast | 2/2 passed |
| SF-14 The front end shall trace every reference model with a peak resident memory below 2 GB. | Ubiquitous | D (results.md) | examples/results.md: peak resident memory 642 MB | demonstrated |
| SF-15 If a library upgrade changes the traced FLOPs or weight bytes of a reference trace, then CI shall fail. | Unwanted behaviour | T | test_gate.py::test_gate_fails_on_an_arithmetic_change_and_tolerates_drift | 1/1 passed |
| SF-16 The front end shall run on a CPU-only machine with no GPU and no model weights. | Ubiquitous | D (GitHub Actions) | GitHub Actions runs the whole suite on ubuntu-24.04 runners with CPU-only PyTorch | demonstrated |
SF-17 The cost rules shall count each operator's FLOPs and bytes by the conventions stated in rules.py and cost.py. | Ubiquitous | T | test_cost.py::test_roofline_is_max_of_compute_and_memory_per_op<br>test_cost.py::test_fused_bound_is_below_unfused<br>test_cost.py::test_faster_hardware_is_never_slower<br>test_rules.py::test_mm_is_two_mkn<br>test_rules.py::test_bmm_addmm_linear_conv<br>test_rules.py::test_fused_attention_counts_qk_av_and_softmax<br>test_rules.py::test_views_are_free_and_copies_are_not<br>test_rules.py::test_elementwise_reduction_softmax<br>test_rules.py::test_embedding_reads_rows_not_the_table<br>test_rules.py::test_onnx_rules | 10/10 passed |
| SF-18 The accelerator model shall lower every operator that has a cost rule into tiles that fit the tile budget, and the tiles of a GEMM shall perform exactly its multiply-accumulates. | Ubiquitous | T | test_accel_lower.py::test_gemm_dims_match_the_flop_rules<br>test_accel_lower.py::test_tiles_fit_and_cover_every_mac[4096]<br>test_accel_lower.py::test_tiles_fit_and_cover_every_mac[65536]<br>test_accel_lower.py::test_tiles_fit_and_cover_every_mac[1048576] | 4/4 passed |
| SF-19 While the on-chip buffer cannot hold the next tile, the load DMA shall not start its transfer, and buffer occupancy shall never exceed its capacity. | State-driven | T | test_accel_sim.py::test_buffer_never_overflows[65536]<br>test_accel_sim.py::test_buffer_never_overflows[262144]<br>test_accel_sim.py::test_buffer_never_overflows[2097152]<br>test_accel_sim.py::test_backpressure_stalls_loads_when_compute_is_slow | 4/4 passed |
| SF-20 When an operator reads an activation, its first load shall not start before the operator that produced it has stored its results. | Event-driven | T | test_accel_lower.py::test_dependencies_follow_activations_not_weights<br>test_accel_sim.py::test_loads_wait_for_producers_to_be_stored | 2/2 passed |
| SF-21 The accelerator model shall report latency, utilisation per component, a stall breakdown that sums to the latency, a hot-spot, a per-operator latency histogram, a timeline plot and a Chrome trace. | Ubiquitous | T | test_accel_sim.py::test_stalls_sum_to_the_makespan[kw0-cnn]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw0-llama]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw0-gpt2]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw1-cnn]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw1-llama]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw1-gpt2]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw2-cnn]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw2-llama]<br>test_accel_sim.py::test_stalls_sum_to_the_makespan[kw2-gpt2]<br>test_accel_sim.py::test_hotspot_moves_with_the_bottleneck<br>test_accel_sim.py::test_outputs_report_histogram_timeline_and_chrome_trace | 11/11 passed |
| SF-22 While the two DMA engines cannot contend for a memory channel, the fast path (Python and C++) shall produce per-tile timings bit-identical to the SimPy model's. | State-driven | T | test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw0-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw0-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw0-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw1-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw1-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw1-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw2-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw2-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw2-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw3-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw3-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw3-gpt2]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw4-cnn]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw4-llama]<br>test_accel_fastpath.py::test_recurrence_equals_simpy_exactly[kw4-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw0-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw0-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw0-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw1-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw1-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw1-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw2-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw2-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw2-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw3-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw3-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw3-gpt2]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw4-cnn]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw4-llama]<br>test_accel_fastpath.py::test_cpp_equals_simpy_exactly[kw4-gpt2]<br>test_accel_fastpath.py::test_random_programs_agree | 31/31 passed |
| SF-23 If the configuration lets the DMA engines contend for one memory channel, then the fast path shall refuse to run. | Unwanted behaviour | T | test_accel_fastpath.py::test_fast_path_refuses_a_shared_channel | 1/1 passed |
| SF-24 Where durations are quantised to whole cycles, the cycle-stepped twin shall produce per-tile timings identical to the event-driven model's. | Optional feature | T | test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw0-cnn]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw0-llama]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw1-cnn]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw1-llama]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw2-cnn]<br>test_accel_cycle.py::test_cycle_twin_equals_event_driven[kw2-llama] | 6/6 passed |
| SF-25 The NTT polynomial product shall equal schoolbook multiplication in Z_q[X]/(X^N + 1). | Ubiquitous | T | test_accel_ntt.py::test_polymul_equals_schoolbook | 1/1 passed |
| SF-26 Where an in-transit stage is configured, the model shall run NTT operators in the read path, never faster than the stage's operations-per-byte budget allows. | Optional feature | T | test_accel_ntt.py::test_in_transit_stage_respects_its_budget[0.5]<br>test_accel_ntt.py::test_in_transit_stage_respects_its_budget[1.6]<br>test_accel_ntt.py::test_in_transit_stage_respects_its_budget[3.0]<br>test_accel_ntt.py::test_in_transit_stage_respects_its_budget[12.0] | 4/4 passed |
| SF-27 The FIFO model shall satisfy Little's law (time-averaged depth = throughput x mean time in the FIFO) on every run, with depth never above capacity. | Ubiquitous | T | test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-1]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-2]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-4]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.2-16]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-1]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-2]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-4]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.2-1.0-16]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-1]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-2]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-4]<br>test_accel_fifo.py::test_littles_law_and_capacity[1.0-1.0-16]<br>test_accel_fifo.py::test_backpressure_stalls_a_fast_producer | 13/13 passed |
| SF-28 The execution-provider partitioner shall assign every ONNX node exactly once, and the ONNX Runtime profile shall name the provider of every node run. | Ubiquitous | T | test_accel_ep.py::test_claim_assigns_every_node_once<br>test_accel_ep.py::test_ort_profile_names_the_provider_of_every_node | 2/2 passed |
Source: docs/traceability.md in Torch_Sim_Frontend
SF-15 ("if a library upgrade changes the traced FLOPs or weight bytes of a reference trace, then CI shall fail") was claimed by the CI gate, with the method "T (CI gate)". The matrix reported GAP: nothing tested that the gate actually fails. tests/test_gate.py now changes one FLOP in a copy of the baseline and checks the gate's verdict. The spec's change history records the change of method (version 1.1).
| Section | Contents | In simfront's spec |
|---|---|---|
| 1. Scope | What the simulator is for, which decisions it informs, who uses it, what it is not | "A design-exploration tool ... not mission-mode software" |
| 2. Definitions | Every term a requirement relies on; the references accuracy is measured against | Trace, front end, weight bytes, closed form, reference models; tile, program, engines (added with the accelerator model) |
| 3. Requirements | ID, type, EARS pattern, text, verification method; one table | SF-01 to SF-17 (front end and cost models); SF-18 to SF-28 (the SimPy accelerator model, deck 14) |
| 4. Fidelity | What is modelled, at what abstraction; what is not | In the README's "Not modelled" |
| 5. Assumptions and constraints | Versions, platforms, input limits | PyTorch 2.6+, batch-1 static shapes |
| 6. Change history | Every change, with its reason | 1.1: SF-15's method changed after the gap; 1.2: SF-05 reworded when the closed form it checks was corrected; 1.3: SF-18 to SF-28 added for the accelerator model |
The requirements table, as written (excerpt):
| ID | Type | Pattern | Requirement | Verification |
|---|---|---|---|---|
| SF-01 | Functional | Ubiquitous | The front end shall record, for every captured operator, the shape, element size and identity of each input and output tensor and whether it is a model parameter. | T |
| SF-02 | Functional | Event-driven | When a model is traced on the meta device, the front end shall allocate no parameter storage. | T |
| SF-03 | Functional | Ubiquitous | For SwiGLU decoder configurations, the traced weight-matmul FLOPs of a prefill shall equal 2 × matmul parameters × tokens exactly. | T |
| SF-04 | Functional | Ubiquitous | The four front ends shall report identical matmul FLOPs and weight bytes for the same model and input, apart from constant folding that is documented per case. | T |
| SF-05 | Functional | Event-driven | When a decode step is traced, the traced matmul FLOPs shall equal the closed form plus the rotary-frequency matmul, and the traced weight bytes shall equal the closed form's per-step weight traffic plus the RMSNorm weights. | T |
| SF-06 | Functional | Unwanted behaviour | If an operator has no cost rule, then the front end shall cost it at zero and name it in the coverage report. | T |
A row per requirement, in a file under version control next to the code, reviewed in the same pull request as the change it describes.
A test plan says what will be tested, how a result is judged right, and when testing is finished. Its sections follow the usual contents (as in ISO/IEC/IEEE 29119-3), cut to what a small tool needs: test items and scope; approach; pass/fail criteria; entry and exit criteria; environment; deliverables; risks. The approach is the heart of it, because a test is only as good as its oracle, and each level here uses one that shares no code with what it checks:
| Level | What is tested | Oracle | Requirements |
|---|---|---|---|
| Unit | Each cost rule | FLOP and byte counts worked by hand | SF-06, SF-17 |
| Unit (property) | Rules and cost models for any input | Invariants: 2MKN; faster hardware is never slower | SF-17 |
| Integration | Each front end on real configurations | Disaggregated_Inference_Sim's closed forms; PyTorch's FlopCounterMode | SF-02, SF-03, SF-05 |
| Integration (property) | Random Llama-shaped configs, lengths, batches | The closed form, exactly (Hypothesis) | SF-03 |
| Differential | The four front ends against each other | Each other: they must agree | SF-04, SF-07 |
| Faithfulness | Fake tensors and the compile backend | A real run with data; eager outputs | SF-08, SF-11 |
| External | The exported ONNX model | ONNX Runtime's outputs against PyTorch's | SF-07 |
| System | CLI, gate, traceability | Expected text; the gate's own failure cases | SF-13, SF-15 |
| Unit | Accelerator lowering | The cost rules' FLOPs (2 x MACs); hand-worked im2col shapes | SF-18 |
| Unit | Accelerator invariants | Occupancy within capacity; event order per tile; stalls sum to latency | SF-19, SF-20, SF-21 |
| Differential | SimPy model against the fast path (Python, C++) and the cycle-stepped twin | Each other, exactly, on real traces and on random programs (Hypothesis) | SF-22, SF-23, SF-24 |
| Analytic | FIFO model | Little's law, exactly, on every run | SF-27 |
| Golden model | NTT polynomial product | Schoolbook multiplication (Hypothesis) | SF-25 |
| Bound | In-transit stage | The stage's operations-per-byte budget | SF-26 |
| External | Execution-provider partitioning | ONNX Runtime's own profile | SF-28 |
Source: docs/test_plan.md in Torch_Sim_Frontend
Full plan: docs/test_plan.md.
Every code repository in this series has one: examples/results.md, written by examples/results.py, which is the only source of the numbers in its README and decks. The structure generalises to any simulation study:
| Section | Contents | Example |
|---|---|---|
| Question | The decision the study informs | "Does a matmul-only engine need more operator coverage?" |
| Environment | Versions, machine, commit: enough to reproduce | simfront results.md §1, with the commit hash |
| Method and configuration | Model, workload, parameters; what is illustrative | "Device numbers ... are datasheet-level or illustrative, as labelled" |
| Validation | Why the model can be trusted for this question | §4: the trace against the closed form; §7: against FlopCounterMode |
| Results | Tables and charts, with spread where there is noise | §6: coverage counted by operators, FLOPs and time |
| Limitations | What could change the conclusion | "Not modelled" plus the counting conventions |
The environment section from simfront's report, generated, never typed:
| Item | Value |
|---|---|
| Python | 3.12.12 |
| PyTorch | 2.14.1+cpu |
| transformers | 5.18.0 |
| onnx | 1.23.1 |
| CPU | x86_64 |
| simfront commit | cc64742 |
Source: examples/results.md in Torch_Sim_Frontend
A report generated by a script can be regenerated after any model change, which is what keeps the numbers in decks and READMEs true. Timing results also need repeat runs and a statement of spread (deck 11).