The fidelity ladder of pre-silicon models — analytical, discrete-event, transaction-level, cycle-accurate, RTL, emulation — what each answers, what each costs, and how software-level simulators and RTL simulation feed each other.
New to simulation? Start with Introduction to Simulation: the engineering-wide picture, from field solvers and SPICE to system models, of which this deck's fidelity ladder is the chip-design part.
A chip's architecture is fixed two or three years before anyone can run software on it. Every major decision is made before the thing exists: how many compute units, how much on-chip SRAM, how wide the memory interface, what the interconnect looks like, which datatypes to support. A mask set at an advanced node costs tens of millions of dollars, and a wrong answer is not patchable.
The only way to evaluate a design that does not exist is to model it. A simulator is the architecture group's laboratory: it lets you run tomorrow's workloads on tomorrow's hardware today, measure what happens, and change the design while change is still cheap.
Is it fast enough? Where is the bottleneck? What will it draw, and how many joules per token? What if the SRAM were twice the size? Is the extra link worth its area and power?
Can we compile and run real models before silicon? Does the driver work? How do we map an operator onto this machine?
What should the RTL produce for this input, and how fast? Does the RTL match the specification?
A change costs almost nothing in a spreadsheet, a day in a simulator, weeks in RTL, months after tape-out, and a respin after silicon. Simulation moves discovery to the left of that curve.
One word, four rather different products. Most in-house simulators end up doing all four, so it pays to know which job a given feature serves.
| Job | Question | What it needs | Typical form |
|---|---|---|---|
| Architecture exploration | Which design should we build? | Speed, quick re-parameterisation, sweeps | Analytical model, discrete-event simulator |
| Performance prediction | How fast will this workload run on that design? | Calibrated timing, realistic workloads, metrics | Discrete-event or cycle-approximate model |
| Software bring-up | Does the software stack run before silicon? | Functional correctness, register-level interface, speed to boot | Virtual platform (e.g. QEMU, SystemC TLM) |
| Verification reference | What should the RTL produce? | Bit-exact functional behaviour | Golden C/C++/Python model used as a scoreboard |
The job description of a "simulation & frameworks" engineer usually spans the middle two and touches the outer two: build the performance model, connect it to real frameworks so real applications run on it, and supply references and metrics for verification and sign-off.
Every model sits somewhere on a ladder that trades speed and agility for accuracy and detail. Speeds below are rough orders of magnitude for a large SoC; they vary hugely with design size and modelling style.
Use the highest rung that can answer the question. Detail you cannot calibrate is not accuracy; it is just slowness with extra parameters.
Abstraction is a choice about the unit of time and the unit of data. Each rung up the ladder deletes detail that is assumed not to matter for the question at hand.
Pick a workload and a fidelity level. The estimate uses the order-of-magnitude speeds from the ladder, so treat the answer as a factor-of-ten guide. The point is the spread: it covers about twelve orders of magnitude.
| Level | Assumed speed | Wall-clock time |
|---|
This is why LLM serving studies are done with discrete-event and analytical models: an hour of traffic, the minimum for stable p99 latencies, is centuries of RTL simulation.
Software-level simulators (analytical, discrete-event, TLM, cycle-approximate) and RTL simulation are often discussed as rivals. In a working chip project they are partners at different stages that check each other.
| Software-level simulator | RTL simulation | |
|---|---|---|
| Answers | "What should we build, and how fast will it be?" | "Did we build exactly what we specified?" |
| Exists | From day one, before any RTL | Once the design is written, block by block |
| Input | Workloads: models, traces, request streams | Stimulus: test vectors, constrained-random sequences |
| Truth | Approximate; only as good as its calibration | Exact for the logic; it is the design |
| Scope | Whole system, whole application | Usually one block or subsystem at a time |
| Speed | Seconds to hours for real workloads | Hours to weeks for microseconds of chip time |
| Languages | Python (SimPy), C++, Rust, SystemC | SystemVerilog, VHDL; testbenches in UVM, cocotb |
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.
Across a project the two streams run in parallel, with information flowing both ways. The model leads early; the RTL becomes the authority late; correlation keeps them honest in between.
"Writing code to help with pre-tape-out verification" usually means three things: golden models for scoreboards, workload-derived test streams, and a performance sign-off that shows the RTL meets the targets the model predicted. All three come from the software simulator.
The practical bridge between a Python system model and RTL is usually cocotb or a Verilator-compiled C++ model. Both let a high-level model drive the real design one transaction at a time.
import cocotb
from cocotb.triggers import RisingEdge
from golden import ntt_reference # the same function the simulator uses
@cocotb.test()
async def ntt_matches_model(dut):
for vec in workload_vectors("traces/fhe_bootstrap.json"):
await drive(dut, vec) # transactor: Python object to pin wiggles
got = await collect(dut)
assert got == ntt_reference(vec), f"mismatch on {vec.id}"
cycles.append(dut.cycle_count.value) # feed back to calibrate the cost model
The adaptor between abstraction levels: turn "read 2 KB from address A" into AXI handshakes and back. Write them once and reuse them for co-sim, testbenches and emulation.
The model advances in events; RTL advances in clock edges. A co-sim must agree on time, typically by letting the RTL run for a quantum and then synchronising, the same idea as TLM-2.0 temporal decoupling.
The system runs as fast as its slowest island. Keep RTL islands small and the interface coarse, or the whole run drops to RTL speed.
For AI accelerators the case for system-level simulation is stronger than for CPUs, and stronger again for unconventional compute such as photonics.
Deck 02 builds a discrete-event simulator from an empty file and then rebuilds it in SimPy: the hands-on tutorial for the rest of the series.