Six levels of simulation, the methods they share, and how to choose between them, with a link at every level into working simulators and decks on this GitHub
Each level stands on its own, then ends with an On this GitHub box that links the decks and code going deeper. Concepts the simulation series already explain link to their glossary entries in the LLM Inference Simulators, FHE Accelerator Simulators and Simulation Engineering Toolkit hubs; the rest are explained on the slide.
A simulation advances a model of a system through time, or across many parameter settings, to predict behaviour that you cannot measure, or would rather not. Every simulator has the same parts:
The next slide takes nine motivations in turn.
Questions 2 and 3 are the same at every level, which is why Part II exists.
| Motivation | What it buys | An example in this deck |
|---|---|---|
| Architectural and design-space exploration | Compare options while change is cheap | Sizing on-chip SRAM: an analytical sweep, then DES (How to Choose a Level) |
| Functional verification | Does the design do what the specification says? | RTL against a golden model, with coverage and regressions in CI (Level 3) |
| Performance and power prediction | Throughput, latency and energy before the hardware exists | Energy per FHE bootstrap under a power cap (FHESim 05: Power and Energy per Bootstrap) |
| Pre-silicon software, HW/SW co-design | Firmware and drivers start before the chip | Virtual platforms (Level 6) |
| Debug and visibility | Every signal and state observable, as real hardware rarely is | Emulators keep full visibility; FPGA prototypes keep little (Gate Level, Emulation and FPGA Prototypes) |
| What-if, capacity planning and sizing | Load, failures and scaling, before they happen | Hours of serving traffic in seconds (Level 5) |
| Experiments that are impossible, dangerous or too costly | Crashes, faults and extreme physics; evidence for safety and certification | Failures too rare to test for (Monte Carlo); credibility rules such as NASA-STD-7009 |
| Training and learning | Operators and learning agents practise where mistakes are free | FAA-qualified flight simulators; RL in Gymnasium; RC Flight Line here |
A mistake is cheapest to fix in a model. In a NIST study's example (software, marked "example only"), a design-stage defect that costs 1× to fix at once costs 10× at system test and 30× after release; published ratios vary widely. For a chip, a late mistake can mean a respin (InfSim 01: The Pre-Silicon Problem).
NIST Planning Report 02-3, The Economic Impacts of Inadequate Infrastructure for Software Testing (RTI, 2002), Table 5-1
Each motivation tends to live at one or two of the six levels, defined on The Levels at a Glance. Tools are typical examples, not recommendations. Interview practice: Interview_Simulation, why simulate.
| Motivation | Usual level(s) | Typical tools | Go deeper |
|---|---|---|---|
| Architectural exploration | 4 Architecture, 5 System; analytical first | Spreadsheets and rooflines, SimPy DES, gem5, SystemC TLM | InfSim 01: Four Jobs a Simulator Does · FHESim 05: Design-Space Map · SimEng 13: Pareto Fronts |
| Functional verification | 3 Digital logic (2 for analogue and mixed-signal) | Icarus, Verilator, commercial simulators; cocotb, UVM, assertions | SimEng 05: Two Golden Models and a Scoreboard · Free SystemVerilog Simulators · Interview_SystemVerilog |
| Performance and power prediction | 4, 5; 3 for RTL power | DES or cycle-level models with a power model; RTL switching activity | InfSim 07: The Simulator's Power Model · SimEng 05: Switching Activity as a Power Proxy |
| Pre-silicon software, HW/SW co-design | 6 Software; 3 emulation and FPGA prototypes | Instruction-set simulators, QEMU, virtual prototypes, emulators | Level 6: Software Virtual Platforms · RISC-V: Mini RISC-V Stepper (an instruction-set simulator) |
| Debug and visibility | 3 (waveforms); 4–5 (traces) | Waveform viewers; Perfetto traces; hot-spot attribution | InfSim 06: Hot-Spot Attribution · FHESim 03: Metrics and Hot-Spots |
| What-if, capacity and sizing | 5 System | DES, queueing models, network simulators (ns-3), Monte Carlo | Disaggregated_Inference_Sim · InfSim 05: Run the Simulator |
| Impossible or dangerous experiments; safety evidence | 1 Physics, 2 Circuit; Monte Carlo at any level | FEM, CFD and crash solvers; fault injection; importance sampling | CFD, TCAD and Multiphysics · Monte Carlo and Variance Reduction · Verification, Validation and Calibration |
| Training and learning | Real-time models, often with hardware in the loop | Flight and driving simulators; RL environments | Co-Simulation and Hardware-in-the-Loop · RC-Flight-Line (a blade-element flight model) |
| Cost, schedule and risk | Every level; cheapest at the highest one that answers | Whatever finds the mistake earliest | InfSim 01: The Pre-Silicon Problem · How to Choose a Level |
Name the decision the simulation will change, the accuracy it needs, and what you will validate against. If you cannot name all three, the simulator is not ready to be built (How to Choose a Level).
| Level | What is solved | How time advances | Unit of time · unit of data | Typical methods |
|---|---|---|---|---|
| 1 Physics | Partial differential equations over a mesh: Maxwell, Navier–Stokes, heat, drift–diffusion | Fixed steps (stability-limited), or one solve per frequency | fs to ms · mesh cells | FDTD, FEM, MoM, CFD, TCAD |
| 2 Circuit | Kirchhoff's laws plus device models: differential-algebraic equations | Adaptive implicit steps; breakpoints at source corners | ps to ms · node voltages | SPICE, switching (PWL) simulators, IBIS, IBIS-AMI |
| 3 Digital logic | Boolean or four-state logic, with or without delays | Events (signal changes) or one evaluation per clock | ps or cycles · signals | RTL and gate-level simulation, emulation, FPGA prototypes |
| 4 Architecture | Pipelines, caches, memories, interconnect, accelerators | Cycles, transactions or events; or no time at all | cycles to µs · transactions | Roofline, discrete-event, SystemC TLM, cycle-level |
| 5 System | Requests, packets, jobs, agents, failures | Events; random sampling | µs to hours · requests | DES, queueing networks, network simulators, Monte Carlo, agent-based |
| 6 Software | Instructions on a modelled CPU with its peripherals | Instructions, in blocks or time quanta | instructions · architectural state | Instruction-set simulators, QEMU, virtual prototypes |
Each level throws away detail that the level below keeps, and gains speed and reach in exchange: a field solver can afford nanoseconds of a connector; a system simulator can afford hours of a data centre. The hand-over between levels is a reduced model: S-parameters, a compact device model, a cycle count, a cost model.
Six levels, from fields on a mesh to software on a modelled CPU. For each: what is solved, how time advances, the tools, speed and accuracy, what it is validated against, and where this GitHub goes deeper.
Partial differential equations over space and time: Maxwell's equations (electromagnetics), Navier–Stokes (fluids), heat conduction, elasticity, Poisson plus drift–diffusion (semiconductors). Space is cut into a mesh, and the equations become a large system of algebraic ones:
A result is credible only after a refinement study: shrink the cells until the quantity you care about stops changing, and use the trend (Richardson extrapolation) to estimate the error that remains. Mesh error, model error (the physics left out) and data error (material properties) are different things and need different checks.
Open source: openEMS, Meep (FDTD), Elmer (FEM); commercial suites from the EDA and CAE vendors.
K. S. Yee, IEEE Trans. Antennas Propag., 1966 · R. Courant, K. Friedrichs, H. Lewy, Math. Annalen, 1928
An explicit scheme is stable only if a wave crosses no more than about one cell per step (Courant, Friedrichs and Lewy, 1928). Small cells force small steps: halving the cell size in 3-D multiplies the work by 16.
0.1 mm cells in air give Δt ≤ 0.19 ps. Assume 107 cells and 108 cell updates per second on a multi-core CPU: each step takes 0.1 s, so one wall-clock second simulates about 1.9 × 10−12 s. A nanosecond of a connector takes about 9 minutes. Redo the arithmetic with your own assumptions.
Navier–Stokes solved by finite volumes. Turbulence is usually modelled (Reynolds-averaged or large-eddy models) rather than resolved, which only direct numerical simulation does, at enormous cost. Engineers use it for airflow through a server, cooling of a heat sink, aerodynamics. Validated against wind-tunnel and thermal-chamber measurements. Open source: OpenFOAM.
Process simulation predicts the shapes and doping that implant, diffusion and etch steps leave behind. Device simulation solves Poisson's equation with drift–diffusion for electrons and holes on the device's mesh, giving I–V and C–V curves. Those curves are fitted by compact models, which are what SPICE uses. Commercial suites dominate; DEVSIM is an open-source device simulator.
Real problems couple fields: current heats a conductor and heat raises its resistance (electro-thermal); temperature bends a package (thermo-mechanical); air carries heat away (fluid-thermal). Solvers couple either monolithically (one large system) or by partitioning, where each solver runs its own physics and they exchange boundary values, which is co-simulation (Part II).
Each hand-over is a reduced-order model: the level above never sees the mesh, so the reduction must keep what it needs, such as causality and passivity for S-parameters.
Also: DC operating point (Newton, no time), AC small-signal (one complex solve per frequency), noise. The interactive demo is a one-node transient problem, solved four ways.
Open source: ngspice, Xyce. Commercial SPICE and fast-SPICE simulators from the EDA vendors, and free vendor-supplied SPICE tools for board-level power design.
L. W. Nagel, SPICE2, UC Berkeley ERL M520, 1975 · C.-W. Ho, A. Ruehli, P. Brennan, IEEE Trans. Circuits Syst., 1975
A buck converter switches every microsecond or so, a load-transient study needs milliseconds, and SPICE spends its steps on the edges. Two remedies:
The event-driven solver in the demo uses the same idea on an RC circuit.
IBIS Open Forum (IBIS and IBIS-AMI specifications) · Verilog-AMS (Accellera) · SIMetrix/SIMPLIS (commercial PWL simulator; see its documentation)
Every signal change is an event in a time wheel; only logic that reads a changed signal is re-evaluated. Four-state values (0, 1, X, Z), arbitrary # delays and any testbench construct. Updates within one time step happen in ordered zero-time iterations: delta cycles in VHDL and SystemC (glossary: the SystemC kernel and delta cycles), the active and non-blocking-assignment regions of a SystemVerilog time slot. Icarus Verilog, GHDL and the commercial simulators work this way.
Evaluate the whole design once per clock edge as compiled code. Assumes synchronous logic, usually two-state, no timing inside a cycle. Verilator compiles SystemVerilog into a C++ model this way; since version 5 it also accepts timing constructs (--timing). Fast regressions, weaker on X-propagation and testbench features.
| Simulator | Style | Compile | Clock cycles per second |
|---|---|---|---|
| Icarus Verilog version 12.0 (stable) | event-driven, interpreted | 0.03 s | 6,416 |
| Verilator 5.020 | cycle-based, compiled C++ | 4.38 s | 837,883 |
Verilator ran 131× more cycles per second and paid for it in compile time. Design: ntt_core from RTL_CoSim_NTT (50-bit words, N = 1024, 4 lanes), every output word checked against the golden NTT. Script and results: demo/rtl_speed (i7-3770).
A small core: a large SoC runs orders of magnitude slower in both simulators, and the ratio varies with the design and the testbench.
RTL simulation is exact for the logic, because it is the design. What it needs validating against is the specification: tests, assertions, coverage, and a golden model checking every output. Large SoCs simulate at tens to thousands of cycles per second (indicative).
After synthesis, the design is a netlist of library cells. SDF (Standard Delay Format) back-annotation gives each cell and wire the delays computed by static timing analysis or extraction, and setup and hold checks fire during simulation. Used for reset and X-propagation, power-up sequences and checking timing constraints; far slower than RTL, so run sparingly. Timing sign-off itself is static timing analysis, not simulation.
RTL compiled onto special-purpose hardware (processor-based or FPGA-based emulators). Around a megahertz, with full signal visibility: fast enough to boot an operating system and run real software before silicon, at a high price per seat.
The RTL on one or more FPGAs, partitioned when it does not fit. Tens of megahertz and real I/O, but little visibility and long compile times. Mostly for software teams and system validation.
| Platform | Design clock reached | An hour of simulation covers |
|---|---|---|
| RTL simulation | 10 Hz to 10 kHz | under a second of chip time |
| Gate level with SDF | slower than RTL | less than RTL does |
| Emulation | ~1 MHz | an OS boot |
| FPGA prototype | 10 to 100 MHz | real workloads, slowly |
| Silicon | GHz | everything, too late to change |
Rules of thumb from InfSim 01's fidelity ladder; they vary hugely with design size and modelling style.
Closed-form performance: the roofline (attainable throughput = the smaller of peak compute and bandwidth × arithmetic intensity), queueing formulas, a spreadsheet of FLOPs and bytes. Microseconds per answer, so whole design spaces can be swept; blind to contention, queueing and transients.
The machine as components exchanging timed events (a kernel starts, a DMA finishes, a memory request returns), and the clock jumps from one event to the next. Contention and queueing emerge rather than being assumed. Each event's duration comes from a cost model. (Glossary: discrete-event simulation.)
Hardware counters and timings on an existing chip, RTL cycle counts for new blocks, and published results; the gap is reported as a calibration and correlation figure, not hidden.
Analytical: microseconds per design point. DES: about 105 to 107 events per second in a compiled kernel, fewer in Python (indicative; InfSim 01). The number of events, set by the abstraction, matters more than the language: see InfSim 08: Fewer Events.
A bus transfer is a function call that carries a payload and an annotated delay, not a sequence of pin wiggles.
Pipelines, caches and queues updated every cycle, as in gem5's detailed CPU models or GPU simulators. About 104 to 106 cycles per second (indicative). More detail is not more accuracy unless it is calibrated against the real machine.
N. Binkert et al., The gem5 simulator, 2011 · J. Lowe-Power et al., The gem5 Simulator: Version 20.0+, 2020 · A. Akram, L. Sawalha, A Survey of Computer Architecture Simulation Techniques and Tools, 2019
The events are now requests, batches, transfers and failures, across thousands of components and hours of traffic. Measured on this PC: Disaggregated_Inference_Sim (SimPy) serves 3,000 LLM requests, 770.9 s of simulated time, with 90,766 events in 0.923 s: 835× faster than real time, or 1,566× with its exact fast path. Script: demo/des_speed.
Servers and queues with arrival and service distributions. Closed forms such as Little's law (L = λW) and the M/M/1 queue give instant answers for simple cases, and double as unit tests for a simulator: a correct DES must reproduce them.
Packet-level simulators (ns-3, OMNeT++) model protocols and topologies packet by packet; flow-level models trade packet detail for speed. Agent-based models give many autonomous agents simple local rules (people, vehicles, traders) and let the aggregate behaviour emerge.
G. F. Riley, T. R. Henderson, The ns-3 Network Simulator, 2010 · C. M. Macal, M. J. North, Tutorial on agent-based modelling and simulation, 2010
Estimate an expectation by averaging random samples of the model. The standard error falls only as the square root of the number of samples, so each extra decimal digit costs a hundred times more runs.
Engineering uses: manufacturing yield under process variation (Monte Carlo SPICE), reliability and failure rates, risk and option prices, and the randomness inside every stochastic DES.
N. Metropolis, S. Ulam, The Monte Carlo Method, JASA, 1949
A Monte Carlo answer is an estimate with an error bar: report the confidence interval and the number of samples, and check that it shrinks as expected when N grows.
Fetch, decode and execute one instruction at a time against modelled registers and memory. Simple, and exact at the level of the instruction set. Spike, the RISC-V reference ISS, is used as the golden model when verifying processor RTL.
QEMU's code generator (TCG) translates each block of guest instructions into host code once and caches it. In the original paper, user-mode emulation was about 4× slower than native on integer code and 10× on floating point, with a further factor of 2 for the software MMU in full-system mode (QEMU 0.4.2; today's figures depend on the guest, host and workload).
F. Bellard, QEMU, a Fast and Portable Dynamic Translator, USENIX 2005
A whole platform (CPUs, bus, memory map, peripherals) as fast functional models, often SystemC TLM loosely timed, so firmware, drivers and operating systems run before silicon. Register-accurate, not cycle-accurate: right for software bring-up, wrong for performance. Open source: Renode; CPU vendors and EDA vendors sell their own.
The FHE accelerator work on this GitHub models one design at five levels, and each level checks or feeds another. It is a compact example of how the levels work together on a real project.
| Level | Model | What it answers | Explained in |
|---|---|---|---|
| Analytical | Roofline and the NTT-, MAC-, memory- and power-bound tests | Which resource limits bootstrapping? | FHESim 05: NTT-Bound Against Memory-Bound |
| Discrete-event | FHE_Accelerator_Sim (SimPy, bit-exact JS port) | Latency, energy and hot-spots for a trace | FHESim 03: The SimPy Engine |
| Memory system | Memory_System_Sim (command-level DRAM/HBM) | Replaces "bandwidth × efficiency" | SimEng 04: Plugging It Into the FHE Simulator |
| Transaction-level | SystemC_Accelerator_Model (TLM-2.0, AT and LT) | Does a C++ model agree op by op? | SimEng 03: Checked Against the SimPy Model |
| RTL | RTL_CoSim_NTT (cocotb on Verilator) | Real NTT cycle counts | SimEng 05: Feeding the RTL Back Into the Simulator |
| Area and cost | The calibrated area, yield and cost model | Is a design worth its silicon? | SimEng 13: PPA Explorer (glossary: PPA) |
Field solver (Signal Integrity 01–04) → S-parameters (Matrix Methods) → pulse response and equalisers (SerDes Equalisation) → a compliance verdict (Signal Integrity 17).
The questions every level shares: how time advances, how the step is chosen, how randomness is handled, how simulators are joined together, how a model earns trust, and how to choose a level in the first place.
For continuous state (voltages, fields, temperatures): advance by a step h and evaluate the derivatives. Fixed steps are simple and suit real-time use; adaptive steps follow the dynamics and save work where little happens.
For state that changes only at instants (a packet arrives, a signal toggles): keep a priority queue of future events, pop the earliest, run its handler, which may schedule more. Simultaneous events need a deterministic tie-break. (Glossary: DES.)
Continuous dynamics with discrete events (a switch opens, a thermostat fires): integrate, detect the zero crossing, locate it, then restart from it. SPICE breakpoints, switching simulators and FMI's event mode all do this.
A system is stiff when it mixes very different time constants: picosecond parasitics with millisecond thermal drift, fast and slow chemistry. An explicit method's step is capped by the fastest mode even after that mode has died away; an implicit one can let the step follow the slow behaviour you care about. That is why SPICE integrates implicitly.
Embedded pairs such as Dormand–Prince 5(4) compute two solutions of different order from the same stages; their difference estimates the local error, and the step grows or shrinks to keep it under a tolerance. The estimate assumes the solution is smooth. At a discontinuity (a square-wave edge, a switch closing) the solver rejects step after step until it has pinned the edge down, unless it is told where the edges are: breakpoints.
On the RC circuit of the next slide, forward Euler at h = 2.2τ ends 25.8 V off a 1 V signal; backward Euler at the same step stays within 0.708 V. Adaptive Dormand–Prince at a tolerance of 10−6 needs 2,660 evaluations and rejects 218 steps; told where the edges are, it needs 560 and its error falls from 2.4 × 10−4 V to 2.5 × 10−7 V.
J. R. Dormand, P. J. Prince, A family of embedded Runge–Kutta formulae, 1980 · E. Hairer, G. Wanner, Solving Ordinary Differential Equations II: Stiff and Differential-Algebraic Problems
An RC low-pass filter (τ = RC = 1 ms) driven by a 0–1 V square wave of period 4τ, simulated for 20τ (10 edges). Error is the largest |v − vexact| at the solver's own time points. Every row is reproduced by demo/rc_solvers.py.
| Method | Setting | Accepted steps | Rejected | Right-hand-side evaluations | Max error (V) |
|---|
Try: forward Euler at h = 2.2τ (unstable) against backward Euler at the same step (stable, but smeared); Dormand–Prince with and without breakpoints at the same tolerance; the event-driven solver, exact to rounding in 10 events, because the input is constant between edges, as in a piecewise-linear switching simulator.
Same inputs, same outputs, on every run and every machine. That takes deliberate work: a fixed rule for simultaneous events, ordered floating-point reductions, no dependence on hash order or thread timing (glossary: determinism, tie-breaking and RNG streams). Ports to another language can even be bit-exact.
Two or more simulators, each with its own solver, exchange values at agreed communication points. The Functional Mock-up Interface (FMI) packages a model as an FMU: an XML description plus C code or binaries. In model exchange the host integrates the FMU's equations; in co-simulation the FMU brings its own solver, and a master algorithm steps every FMU and passes signals between them at each macro step. The step size trades accuracy and stability against speed.
T. Blochwitz et al., The Functional Mockup Interface for Tool independent Exchange of Simulation Models, Modelica 2011
The controller under test is real hardware (an engine control unit, a motor drive's controller board), wired to a real-time simulation of the plant it controls. The simulator must finish every step before its deadline, so plant models are simplified until they do. The cheaper stages come first: model-in-the-loop, software-in-the-loop (the controller code on a PC), processor-in-the-loop (on the target processor).
"A digital twin is a set of virtual information constructs that mimics the structure, context, and behavior of a natural, engineered, or social system (or system-of-systems), is dynamically updated with data from its physical twin, has a predictive capability, and informs decisions that realize value. The bidirectional interaction between the virtual and the physical is central to the digital twin."
US National Academies, Foundational Research Gaps and Future Directions for Digital Twins (2024), adapting a 2020 AIAA definition
| What you have | Live data | Predicts | Feeds decisions back | Call it |
|---|---|---|---|---|
| A CAD or simulation model | no | yes | no | a model |
| A dashboard of sensor data | yes | no | maybe | monitoring |
| A model kept in step with live data | yes | yes | no | sometimes a "digital shadow" |
| All of these, with decisions acted on | yes | yes | yes | a digital twin |
The term is used loosely in marketing: ask which of the four properties actually hold. Nothing on this GitHub is a twin in this strict sense; the nearest ingredient is calibration, such as calibrating a simulator against OpenFHE.
R. G. Sargent, Verification and validation of simulation models, J. Simulation, 2013
Rows run from most physical detail (top) to least. Bands are orders of magnitude and each spans decades, because model size matters as much as the level. Indicative bands follow InfSim 01's ladder (large SoC, 1 GHz); QEMU is from Bellard (2005); FDTD is the worked estimate on the field-solver slide.
Icarus and Verilator on a small NTT core sit at and beyond the fast end of the RTL band, because the core is small; the SimPy serving simulator runs 835× to 1,566× faster than real time. Scripts in demo/.
Detail costs model-building time and calibration data: a roofline takes an afternoon; a validated cycle-level or meshed 3-D model takes weeks; RTL arrives late. The trade-off is named in SimEng 13: The Other Trade-offs; feel the spread in InfSim 01's calculator.
Use the highest level that can answer the question, and drop to a lower level only for the part that needs it, through a hand-over model or co-simulation. Detail you cannot calibrate is not accuracy.
| Question | Level and method | Why |
|---|---|---|
| Impedance of a new board stack-up | 1: 2-D field solver | Geometry and materials set it |
| Eye opening at 28 GBd through a backplane | 1–2: S-parameters, then channel simulation with IBIS-AMI | Nanoseconds of waveform; statistical BER |
| Will the converter ring on a load step? | 2: averaged model, then PWL switching simulation | Milliseconds of switching |
| Does the RTL meet its specification? | 3: RTL simulation against a golden model | Exact logic; regressions in a compiled simulator |
| How big should the on-chip SRAM be? | 4: analytical sweep, then DES | Hundreds of design points |
| p99 time-to-first-token under bursty traffic | 5: system DES with replications | Hours of traffic, many seeds |
| Will the driver boot before silicon? | 6: virtual platform | Billions of instructions, no timing |
| Yield of a large die under process variation | Analytical yield model, then Monte Carlo | Closed forms, then the variation |
Worked versions: InfSim 04: Which Question, Which Level? · Build or Reuse? · SimEng 13: Dies per Wafer and Yield
Law, Simulation Modeling and Analysis (discrete-event) · Hairer & Wanner, Solving ODEs II (stiff problems) · Taflove & Hagness, Computational Electrodynamics: The FDTD Method · Sargent (2013) on V&V · the reading lists in InfSim 10
| Level | Start here | Then |
|---|---|---|
| 1 Physics | Signal Integrity series (17 decks; computed by a field solver) | Matrix Methods in Network Parameters · Matrix Articles models · Numerical Methods |
| 2 Circuit | DC-DC Control Techniques | SerDes Equalisation · Matrix Concepts in Digital Filters |
| 3 Digital logic | Free SystemVerilog Simulators | SimEng 05: verification bridge · RTL_CoSim_NTT · Hardware index (RTL repos) |
| 4 Architecture | LLM Inference Simulators (InfSim 01–11) | FHE Accelerator Simulators · SimEng 03: SystemC TLM · SimEng 04: memory systems · SimEng 13: PPA |
| 5 System | InfSim 05: Disaggregated Inference | InfSim 06: metrics and validation · LLM Inference Explained (the simulator live in a web page) · Wilmott 14: Monte Carlo · Markov chains |
| 6 Software | RISC-V deck (instruction stepper) | Interview_RISC_V · Cortex-M series |
| Engineering a simulator | Simulation Engineering Toolkit (13 decks) | Rust and PyO3 ports, testing, Jenkins, specifications, performance analysis, measurement tools |
LLM Inference Simulators (why and how to simulate inference hardware) · FHE Accelerator Simulators (one accelerator, end to end) · Simulation Engineering Toolkit (the engineering around a simulator). Each hub has a glossary linking every concept to the slides that explain it.