LLM Inference Simulators — Presentation 01

Why Simulate? From Spreadsheets to RTL

The fidelity ladder of pre-silicon models — analytical, discrete-event, transaction-level, cycle-accurate, RTL, emulation — what each answers, what each costs, and how software-level simulators and RTL simulation feed each other.

Architecture exploration Fidelity ladder TLM RTL simulation Co-simulation Correlation
Spreadsheet → DES → TLM → Cycle → RTL → Silicon
00

Topics We'll Cover

New to simulation? Start with Introduction to Simulation: the engineering-wide picture, from field solvers and SPICE to system models, of which this deck's fidelity ladder is the chip-design part.

01

The Pre-Silicon Problem

A chip's architecture is fixed two or three years before anyone can run software on it. Every major decision is made before the thing exists: how many compute units, how much on-chip SRAM, how wide the memory interface, what the interconnect looks like, which datatypes to support. A mask set at an advanced node costs tens of millions of dollars, and a wrong answer is not patchable.

The only way to evaluate a design that does not exist is to model it. A simulator is the architecture group's laboratory: it lets you run tomorrow's workloads on tomorrow's hardware today, measure what happens, and change the design while change is still cheap.

Architects ask

Is it fast enough? Where is the bottleneck? What will it draw, and how many joules per token? What if the SRAM were twice the size? Is the extra link worth its area and power?

Software asks

Can we compile and run real models before silicon? Does the driver work? How do we map an operator onto this machine?

Verification asks

What should the RTL produce for this input, and how fast? Does the RTL match the specification?

The cost curve

A change costs almost nothing in a spreadsheet, a day in a simulator, weeks in RTL, months after tape-out, and a respin after silicon. Simulation moves discovery to the left of that curve.

02

Four Jobs a Simulator Does

One word, four rather different products. Most in-house simulators end up doing all four, so it pays to know which job a given feature serves.

JobQuestionWhat it needsTypical form
Architecture explorationWhich design should we build?Speed, quick re-parameterisation, sweepsAnalytical model, discrete-event simulator
Performance predictionHow fast will this workload run on that design?Calibrated timing, realistic workloads, metricsDiscrete-event or cycle-approximate model
Software bring-upDoes the software stack run before silicon?Functional correctness, register-level interface, speed to bootVirtual platform (e.g. QEMU, SystemC TLM)
Verification referenceWhat should the RTL produce?Bit-exact functional behaviourGolden C/C++/Python model used as a scoreboard

The job description of a "simulation & frameworks" engineer usually spans the middle two and touches the outer two: build the performance model, connect it to real frameworks so real applications run on it, and supply references and metrics for verification and sign-off.

03

The Fidelity Ladder

Every model sits somewhere on a ladder that trades speed and agility for accuracy and detail. Speeds below are rough orders of magnitude for a large SoC; they vary hugely with design size and modelling style.

AnalyticalRoofline, queueing formulas, spreadsheets. Closed-form: a few equations per workload.µs per answer
Discrete-eventRequests, batches and transfers as events on a timeline (SimPy, custom C++/Rust). Contention and queueing emerge.105–107 events/s
Transaction-levelSystemC TLM-2.0: bus reads/writes as function calls with annotated delays. Fast enough to boot an OS.~10–100 MIPS (LT)
Cycle-approximatePipelines and queues modelled per clock, but not per wire (gem5 detailed CPU, Accel-Sim).~104–106 cycles/s
RTL simulationThe actual Verilog/VHDL, every signal, every edge (Verilator, Xcelium, VCS, Questa, xsim).~101–104 Hz of design clock
EmulationRTL mapped onto special-purpose hardware (Palladium, Veloce, ZeBu).~1 MHz
FPGA prototypeRTL on FPGAs (HAPS or custom boards), often partitioned across several devices.~10–100 MHz
SiliconThe real chip. Ground truth, and far too late to change the architecture.GHz
The rule of thumb

Use the highest rung that can answer the question. Detail you cannot calibrate is not accuracy; it is just slowness with extra parameters.

04

What Each Level Throws Away

Abstraction is a choice about the unit of time and the unit of data. Each rung up the ladder deletes detail that is assumed not to matter for the question at hand.

The same 2 KB DMA read, seen at four abstraction levels Analytical t = 2048 B / bandwidth + latency Discrete-event wait bus transfer (holds the bus) Transaction-level arb burst 0 (512 B) burst 1 burst 2 burst 3 + resp RTL time → (each rung adds events; RTL adds one per signal edge, ~10^3–10^5 per transfer)

What you keep going up

  • Throughput, latency and contention at the level of whole operations.
  • The ability to run whole applications (an LLM serving a million requests).
  • Hours, not months, to try a new architecture.

What you lose going up

  • Pipeline bubbles, arbitration corner cases, bank conflicts.
  • Exact bit-level behaviour (unless you add a functional model).
  • Anything that depends on timing within a transaction.
05

Interactive: How Long Would It Take to Simulate?

Pick a workload and a fidelity level. The estimate uses the order-of-magnitude speeds from the ladder, so treat the answer as a factor-of-ten guide. The point is the spread: it covers about twelve orders of magnitude.

1.5
10^4
LevelAssumed speedWall-clock time

This is why LLM serving studies are done with discrete-event and analytical models: an hour of traffic, the minimum for stable p99 latencies, is centuries of RTL simulation.

06

Software-Level Simulators and RTL Simulation

Software-level simulators (analytical, discrete-event, TLM, cycle-approximate) and RTL simulation are often discussed as rivals. In a working chip project they are partners at different stages that check each other.

Software-level simulatorRTL simulation
Answers"What should we build, and how fast will it be?""Did we build exactly what we specified?"
ExistsFrom day one, before any RTLOnce the design is written, block by block
InputWorkloads: models, traces, request streamsStimulus: test vectors, constrained-random sequences
TruthApproximate; only as good as its calibrationExact for the logic; it is the design
ScopeWhole system, whole applicationUsually one block or subsystem at a time
SpeedSeconds to hours for real workloadsHours to weeks for microseconds of chip time
LanguagesPython (SimPy), C++, Rust, SystemCSystemVerilog, VHDL; testbenches in UVM, cocotb

Five ways they connect

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.

07

The Co-Flow: Models and RTL Feeding Each Other

Across a project the two streams run in parallel, with information flowing both ways. The model leads early; the RTL becomes the authority late; correlation keeps them honest in between.

Software-level models RTL & verification analytical +DES model calibrated perfmodel + metrics co-sim: model +RTL islands sign-off perfreport block RTL +unit testbench subsystem / SoCsimulation emulation / FPGA+ tape-out spec + golden model cycle counts(calibrate) traces asdirected tests correlationerror < target time → (concept → RTL → integration → pre-tape-out sign-off)
Pre-tape-out verification

"Writing code to help with pre-tape-out verification" usually means three things: golden models for scoreboards, workload-derived test streams, and a performance sign-off that shows the RTL meets the targets the model predicted. All three come from the software simulator.

08

Co-Simulation in Practice

The practical bridge between a Python system model and RTL is usually cocotb or a Verilator-compiled C++ model. Both let a high-level model drive the real design one transaction at a time.

cocotb: the Python model is the scoreboard for the RTL
import cocotb
from cocotb.triggers import RisingEdge
from golden import ntt_reference          # the same function the simulator uses

@cocotb.test()
async def ntt_matches_model(dut):
    for vec in workload_vectors("traces/fhe_bootstrap.json"):
        await drive(dut, vec)                 # transactor: Python object to pin wiggles
        got = await collect(dut)
        assert got == ntt_reference(vec), f"mismatch on {vec.id}"
        cycles.append(dut.cycle_count.value)   # feed back to calibrate the cost model

Transactors

The adaptor between abstraction levels: turn "read 2 KB from address A" into AXI handshakes and back. Write them once and reuse them for co-sim, testbenches and emulation.

Time synchronisation

The model advances in events; RTL advances in clock edges. A co-sim must agree on time, typically by letting the RTL run for a quantum and then synchronising, the same idea as TLM-2.0 temporal decoupling.

Speed tax

The system runs as fast as its slowest island. Keep RTL islands small and the interface coarse, or the whole run drops to RTL speed.

09

Why AI and Novel Compute Need System Simulators

For AI accelerators the case for system-level simulation is stronger than for CPUs, and stronger again for unconventional compute such as photonics.

10

What to Take Away

Next

Deck 02 builds a discrete-event simulator from an empty file and then rebuilds it in SimPy: the hands-on tutorial for the rest of the series.