TECHNICAL PRESENTATION

Introduction to
Simulation

How engineers simulate, from field solvers to cloud systems
Numerical methods Circuits & logic Architecture Systems
Fields → Circuits → Logic → Architecture → Systems → Software

Six levels of simulation, the methods they share, and how to choose between them, with a link at every level into working simulators and decks on this GitHub

Model  ·  Solve  ·  Measure  ·  Validate
01

Topics

Part I: the levels

  • What a simulation is; why simulate, and when not to; the levels at a glance
  • Level 1: physics and numerical methods (FEM, FDTD, MoM, CFD, TCAD)
  • Level 2: circuits (SPICE, switching simulators, IBIS and IBIS-AMI)
  • Level 3: digital logic (event-driven, cycle-based, gate level, emulation, FPGA prototypes)
  • Level 4: architecture (analytical, discrete-event, transaction-level, cycle-level)
  • Level 5: system, network and cloud (DES at scale, queueing, Monte Carlo)
  • Level 6: software virtual platforms (ISS, QEMU, virtual prototypes)
  • One accelerator, every level

Part II: cross-cutting methods

  • Time-stepping and event-driven
  • Stiffness, stability and step control
  • Interactive: one circuit, four solvers
  • Deterministic and stochastic; seeds and replications
  • Co-simulation and hardware-in-the-loop
  • Digital twins, defined carefully
  • Verification, validation and calibration
  • Speed, accuracy and effort; how to choose a level

How to read it

Each level stands on its own, then ends with an On this GitHub box that links the decks and code going deeper. Concepts the simulation series already explain link to their glossary entries in the LLM Inference Simulators, FHE Accelerator Simulators and Simulation Engineering Toolkit hubs; the rest are explained on the slide.

02

What a Simulation Is

A model, run forward

A simulation advances a model of a system through time, or across many parameter settings, to predict behaviour that you cannot measure, or would rather not. Every simulator has the same parts:

  • State: node voltages, field values, register contents, queue lengths
  • Rules for how the state changes: differential equations, logic, or event handlers
  • Inputs: stimulus, a workload, random arrivals
  • A solver or kernel that advances time
  • Observers that turn the run into metrics

Why simulate

  • The thing does not exist yet: a chip is designed years before silicon (InfSim 01, the pre-silicon problem)
  • Testing it is too slow, costly or dangerous
  • You cannot see inside it: the current in a via, the occupancy of a queue
  • You want a thousand what-ifs, not one

The next slide takes nine motivations in turn.

Four words that get confused

  • Analysis: a closed-form or static answer with no time evolution: a DC operating point, a roofline, static timing analysis
  • Simulation: a model advanced through time or events in software
  • Emulation: the design, or a functional equivalent, executing on other hardware: a hardware emulator running RTL, QEMU running another CPU's binaries
  • Prototype: the real design in a different implementation, such as RTL on FPGAs

Three questions every simulator answers

  1. What is the state, and how much detail does it keep? (Part I)
  2. How does time advance: fixed steps, adaptive steps, or events?
  3. How do you know the answer is right: verification and validation?

Questions 2 and 3 are the same at every level, which is why Part II exists.

03

Why Simulate?

MotivationWhat it buysAn example in this deck
Architectural and design-space explorationCompare options while change is cheapSizing on-chip SRAM: an analytical sweep, then DES (How to Choose a Level)
Functional verificationDoes the design do what the specification says?RTL against a golden model, with coverage and regressions in CI (Level 3)
Performance and power predictionThroughput, latency and energy before the hardware existsEnergy per FHE bootstrap under a power cap (FHESim 05: Power and Energy per Bootstrap)
Pre-silicon software, HW/SW co-designFirmware and drivers start before the chipVirtual platforms (Level 6)
Debug and visibilityEvery signal and state observable, as real hardware rarely isEmulators keep full visibility; FPGA prototypes keep little (Gate Level, Emulation and FPGA Prototypes)
What-if, capacity planning and sizingLoad, failures and scaling, before they happenHours of serving traffic in seconds (Level 5)
Experiments that are impossible, dangerous or too costlyCrashes, faults and extreme physics; evidence for safety and certificationFailures too rare to test for (Monte Carlo); credibility rules such as NASA-STD-7009
Training and learningOperators and learning agents practise where mistakes are freeFAA-qualified flight simulators; RL in Gymnasium; RC Flight Line here

The payoff: cost, schedule, risk

A mistake is cheapest to fix in a model. In a NIST study's example (software, marked "example only"), a design-stage defect that costs 1× to fix at once costs 10× at system test and 30× after release; published ratios vary widely. For a chip, a late mistake can mean a respin (InfSim 01: The Pre-Silicon Problem).

NIST Planning Report 02-3, The Economic Impacts of Inadequate Infrastructure for Software Testing (RTI, 2002), Table 5-1

On this GitHub

04

Motivation, Level and Tool

Each motivation tends to live at one or two of the six levels, defined on The Levels at a Glance. Tools are typical examples, not recommendations. Interview practice: Interview_Simulation, why simulate.

MotivationUsual level(s)Typical toolsGo deeper
Architectural exploration4 Architecture, 5 System; analytical firstSpreadsheets and rooflines, SimPy DES, gem5, SystemC TLMInfSim 01: Four Jobs a Simulator Does · FHESim 05: Design-Space Map · SimEng 13: Pareto Fronts
Functional verification3 Digital logic (2 for analogue and mixed-signal)Icarus, Verilator, commercial simulators; cocotb, UVM, assertionsSimEng 05: Two Golden Models and a Scoreboard · Free SystemVerilog Simulators · Interview_SystemVerilog
Performance and power prediction4, 5; 3 for RTL powerDES or cycle-level models with a power model; RTL switching activityInfSim 07: The Simulator's Power Model · SimEng 05: Switching Activity as a Power Proxy
Pre-silicon software, HW/SW co-design6 Software; 3 emulation and FPGA prototypesInstruction-set simulators, QEMU, virtual prototypes, emulatorsLevel 6: Software Virtual Platforms · RISC-V: Mini RISC-V Stepper (an instruction-set simulator)
Debug and visibility3 (waveforms); 4–5 (traces)Waveform viewers; Perfetto traces; hot-spot attributionInfSim 06: Hot-Spot Attribution · FHESim 03: Metrics and Hot-Spots
What-if, capacity and sizing5 SystemDES, queueing models, network simulators (ns-3), Monte CarloDisaggregated_Inference_Sim · InfSim 05: Run the Simulator
Impossible or dangerous experiments; safety evidence1 Physics, 2 Circuit; Monte Carlo at any levelFEM, CFD and crash solvers; fault injection; importance samplingCFD, TCAD and Multiphysics · Monte Carlo and Variance Reduction · Verification, Validation and Calibration
Training and learningReal-time models, often with hardware in the loopFlight and driving simulators; RL environmentsCo-Simulation and Hardware-in-the-Loop · RC-Flight-Line (a blade-element flight model)
Cost, schedule and riskEvery level; cheapest at the highest one that answersWhatever finds the mistake earliestInfSim 01: The Pre-Silicon Problem · How to Choose a Level
05

When Simulation Is the Wrong Tool

Do not simulate when…

What to do instead

  • Write the closed form or the spreadsheet first, and simulate only the part with dynamics: queues, contention, feedback, distributions
  • Measure, prototype or put hardware in the loop when the real system, or part of it, exists (FPGA prototypes, hardware-in-the-loop)
  • Combine them: measure what you can, simulate the rest, and calibrate the model on the measurements

A test before you start

Name the decision the simulation will change, the accuracy it needs, and what you will validate against. If you cannot name all three, the simulator is not ready to be built (How to Choose a Level).

06

The Levels at a Glance

LevelWhat is solvedHow time advancesUnit of time · unit of dataTypical methods
1 PhysicsPartial differential equations over a mesh: Maxwell, Navier–Stokes, heat, drift–diffusionFixed steps (stability-limited), or one solve per frequencyfs to ms · mesh cellsFDTD, FEM, MoM, CFD, TCAD
2 CircuitKirchhoff's laws plus device models: differential-algebraic equationsAdaptive implicit steps; breakpoints at source cornersps to ms · node voltagesSPICE, switching (PWL) simulators, IBIS, IBIS-AMI
3 Digital logicBoolean or four-state logic, with or without delaysEvents (signal changes) or one evaluation per clockps or cycles · signalsRTL and gate-level simulation, emulation, FPGA prototypes
4 ArchitecturePipelines, caches, memories, interconnect, acceleratorsCycles, transactions or events; or no time at allcycles to µs · transactionsRoofline, discrete-event, SystemC TLM, cycle-level
5 SystemRequests, packets, jobs, agents, failuresEvents; random samplingµs to hours · requestsDES, queueing networks, network simulators, Monte Carlo, agent-based
6 SoftwareInstructions on a modelled CPU with its peripheralsInstructions, in blocks or time quantainstructions · architectural stateInstruction-set simulators, QEMU, virtual prototypes

Going up the table

Each level throws away detail that the level below keeps, and gains speed and reach in exchange: a field solver can afford nanoseconds of a connector; a system simulator can afford hours of a data centre. The hand-over between levels is a reduced model: S-parameters, a compact device model, a cycle count, a cost model.

On this GitHub

PART I

The Levels

Six levels, from fields on a mesh to software on a modelled CPU. For each: what is solved, how time advances, the tools, speed and accuracy, what it is validated against, and where this GitHub goes deeper.

1 Physics→ 2 Circuit→ 3 Logic→ 4 Architecture→ 5 System→ 6 Software
07

Level 1: Physics and Numerical Methods

What is solved

Partial differential equations over space and time: Maxwell's equations (electromagnetics), Navier–Stokes (fluids), heat conduction, elasticity, Poisson plus drift–diffusion (semiconductors). Space is cut into a mesh, and the equations become a large system of algebraic ones:

  • Finite differences (FDM): derivatives replaced by differences on a structured grid. Simple and fast; awkward on curved geometry
  • Finite volumes (FVM): fluxes balanced cell by cell, so mass and energy are conserved exactly. The workhorse of CFD
  • Finite elements (FEM): the solution approximated by basis functions on elements of any shape; the weak form becomes a large sparse linear system. Fits any geometry; meshes refine where the field changes fast

The mesh is part of the model

A result is credible only after a refinement study: shrink the cells until the quantity you care about stops changing, and use the trend (Richardson extrapolation) to estimate the error that remains. Mesh error, model error (the physics left out) and data error (material properties) are different things and need different checks.

Speed, accuracy, validation

  • 3-D problems take hours to days per run; 2-D cross-sections take seconds
  • Accuracy is limited by geometry, material data and the mesh, rarely by arithmetic
  • Validated against closed-form cases, published benchmark problems, and measurement: a vector network analyser, a wind tunnel, a thermocouple

On this GitHub

08

Field Solvers: FDTD, FEM and MoM

Three ways to solve Maxwell's equations

  • FDTD (time domain): electric and magnetic fields on two interleaved grids (Yee, 1966), updated in a leapfrog. One pulse-excited run gives the whole frequency response. Explicit, so the step has a hard limit
  • FEM (usually frequency domain): one sparse solve per frequency on a tetrahedral mesh that follows any 3-D shape: packages, connectors, antennas
  • MoM (method of moments, integral equations): mesh only the conductor surfaces; dense matrices; suited to open radiating structures and layered (2.5-D) boards
  • Quasi-static 2-D solvers: a cross-section gives per-unit-length L, C, R and G for a transmission line in seconds

Tools and sources

Open source: openEMS, Meep (FDTD), Elmer (FEM); commercial suites from the EDA and CAE vendors.

K. S. Yee, IEEE Trans. Antennas Propag., 1966 · R. Courant, K. Friedrichs, H. Lewy, Math. Annalen, 1928

The CFL condition

Δt ≤ Δx / (c·√3)   (3-D Yee grid, cubic cells)

An explicit scheme is stable only if a wave crosses no more than about one cell per step (Courant, Friedrichs and Lewy, 1928). Small cells force small steps: halving the cell size in 3-D multiplies the work by 16.

Worked estimate (illustrative)

0.1 mm cells in air give Δt ≤ 0.19 ps. Assume 107 cells and 108 cell updates per second on a multi-core CPU: each step takes 0.1 s, so one wall-clock second simulates about 1.9 × 10−12 s. A nanosecond of a connector takes about 9 minutes. Redo the arithmetic with your own assumptions.

On this GitHub

09

CFD, TCAD and Multiphysics

Computational fluid dynamics

Navier–Stokes solved by finite volumes. Turbulence is usually modelled (Reynolds-averaged or large-eddy models) rather than resolved, which only direct numerical simulation does, at enormous cost. Engineers use it for airflow through a server, cooling of a heat sink, aerodynamics. Validated against wind-tunnel and thermal-chamber measurements. Open source: OpenFOAM.

Technology CAD (TCAD)

Process simulation predicts the shapes and doping that implant, diffusion and etch steps leave behind. Device simulation solves Poisson's equation with drift–diffusion for electrons and holes on the device's mesh, giving I–V and C–V curves. Those curves are fitted by compact models, which are what SPICE uses. Commercial suites dominate; DEVSIM is an open-source device simulator.

On this GitHub

Multiphysics

Real problems couple fields: current heats a conductor and heat raises its resistance (electro-thermal); temperature bends a package (thermo-mechanical); air carries heat away (fluid-thermal). Solvers couple either monolithically (one large system) or by partitioning, where each solver runs its own physics and they exchange boundary values, which is co-simulation (Part II).

Where the levels hand over

  • TCAD → compact device model → SPICE
  • EM field solver → S-parameters → circuit and channel simulation
  • CFD → thermal boundary conditions → power and reliability models

Each hand-over is a reduced-order model: the level above never sees the mesh, so the reduction must keep what it needs, such as causality and passivity for S-parameters.

10

Level 2: Circuits, and What SPICE Does

Inside a SPICE transient run

  • Modified nodal analysis (MNA): the unknowns are node voltages plus the currents through voltage sources and inductors. Each element stamps a few entries into one sparse system (Ho, Ruehli and Brennan, 1975)
  • Newton–Raphson for nonlinear devices: each iteration replaces every diode and transistor by its linearisation (a conductance and a current source), solves the sparse system by LU factorisation, and repeats until the voltages settle
  • Implicit integration: backward Euler, trapezoidal or Gear (BDF) turns each capacitor and inductor into a similar companion model for the step
  • Timestep control: the step is chosen from an estimate of the local truncation error, and forced to land on the corners of pulse and piecewise-linear sources (breakpoints). If Newton fails to converge, the step is cut and retried
[ G  B ; C  D ] · [ v ; i ] = [ is ; vs ]   (the MNA system, re-solved every iteration)

Also: DC operating point (Newton, no time), AC small-signal (one complex solve per frequency), noise. The interactive demo is a one-node transient problem, solved four ways.

On this GitHub

Speed, accuracy, validation

  • Accuracy is set by the device models (compact models fitted to silicon by the foundry) and by parasitics extracted from layout
  • Thousands of transistors: nanoseconds to microseconds of circuit time per minute (indicative). Fast-SPICE tools partition and simplify to go larger
  • Validated against silicon measurements and the foundry's process corners

Tools

Open source: ngspice, Xyce. Commercial SPICE and fast-SPICE simulators from the EDA vendors, and free vendor-supplied SPICE tools for board-level power design.

L. W. Nagel, SPICE2, UC Berkeley ERL M520, 1975 · C.-W. Ho, A. Ruehli, P. Brennan, IEEE Trans. Circuits Syst., 1975

11

Switching and Behavioural Models

Power converters break SPICE's budget

A buck converter switches every microsecond or so, a load-transient study needs milliseconds, and SPICE spends its steps on the edges. Two remedies:

  • Piecewise-linear (PWL) switching simulators, in the style of SIMPLIS: each switch and device is a few linear segments, each linear topology is solved exactly between switching events, and the simulator only has to find when the next switching event happens. It can also search for the periodic steady state directly
  • Averaged models: replace the switching by its average over a cycle, for loop design and stability

The event-driven solver in the demo uses the same idea on an RC circuit.

Behavioural models

  • Verilog-A / Verilog-AMS: equations describing a block, simulated inside a circuit simulator
  • IBIS: an I/O buffer as I/V and V/t tables, so a board can be simulated without the vendor's transistor netlist
  • IBIS-AMI: SerDes transmit and receive equalisation as executable algorithmic models that a channel simulator calls, statistically or bit by bit, to predict eyes and bit error ratios

Speed, accuracy, validation

  • PWL: much faster than SPICE on converters (indicative), exact for its own PWL model; accuracy rests on how well the segments fit the devices
  • IBIS: validated against the vendor's transistor-level simulation and lab measurements

Sources

IBIS Open Forum (IBIS and IBIS-AMI specifications) · Verilog-AMS (Accellera) · SIMetrix/SIMPLIS (commercial PWL simulator; see its documentation)

12

Level 3: Digital Logic, Event-Driven and Cycle-Based

Event-driven

Every signal change is an event in a time wheel; only logic that reads a changed signal is re-evaluated. Four-state values (0, 1, X, Z), arbitrary # delays and any testbench construct. Updates within one time step happen in ordered zero-time iterations: delta cycles in VHDL and SystemC (glossary: the SystemC kernel and delta cycles), the active and non-blocking-assignment regions of a SystemVerilog time slot. Icarus Verilog, GHDL and the commercial simulators work this way.

Cycle-based

Evaluate the whole design once per clock edge as compiled code. Assumes synchronous logic, usually two-state, no timing inside a cycle. Verilator compiles SystemVerilog into a C++ model this way; since version 5 it also accepts timing constructs (--timing). Fast regressions, weaker on X-propagation and testbench features.

On this GitHub

Measured: the same RTL in both

SimulatorStyleCompileClock cycles per second
Icarus Verilog version 12.0 (stable)event-driven, interpreted0.03 s6,416
Verilator 5.020cycle-based, compiled C++4.38 s837,883

Verilator ran 131× more cycles per second and paid for it in compile time. Design: ntt_core from RTL_CoSim_NTT (50-bit words, N = 1024, 4 lanes), every output word checked against the golden NTT. Script and results: demo/rtl_speed (i7-3770).

A small core: a large SoC runs orders of magnitude slower in both simulators, and the ratio varies with the design and the testbench.

Speed, accuracy, validation

RTL simulation is exact for the logic, because it is the design. What it needs validating against is the specification: tests, assertions, coverage, and a golden model checking every output. Large SoCs simulate at tens to thousands of cycles per second (indicative).

Icarus Verilog · Verilator guide · GHDL

13

Gate Level, Emulation and FPGA Prototypes

Gate-level simulation

After synthesis, the design is a netlist of library cells. SDF (Standard Delay Format) back-annotation gives each cell and wire the delays computed by static timing analysis or extraction, and setup and hold checks fire during simulation. Used for reset and X-propagation, power-up sequences and checking timing constraints; far slower than RTL, so run sparingly. Timing sign-off itself is static timing analysis, not simulation.

Hardware emulation

RTL compiled onto special-purpose hardware (processor-based or FPGA-based emulators). Around a megahertz, with full signal visibility: fast enough to boot an operating system and run real software before silicon, at a high price per seat.

FPGA prototyping

The RTL on one or more FPGAs, partitioned when it does not fit. Tens of megahertz and real I/O, but little visibility and long compile times. Mostly for software teams and system validation.

Orders of magnitude (indicative, large SoC)

PlatformDesign clock reachedAn hour of simulation covers
RTL simulation10 Hz to 10 kHzunder a second of chip time
Gate level with SDFslower than RTLless than RTL does
Emulation~1 MHzan OS boot
FPGA prototype10 to 100 MHzreal workloads, slowly
SiliconGHzeverything, too late to change

Rules of thumb from InfSim 01's fidelity ladder; they vary hugely with design size and modelling style.

On this GitHub

14

Level 4: Architecture, Analytical and Discrete-Event

Analytical models

Closed-form performance: the roofline (attainable throughput = the smaller of peak compute and bandwidth × arithmetic intensity), queueing formulas, a spreadsheet of FLOPs and bytes. Microseconds per answer, so whole design spaces can be swept; blind to contention, queueing and transients.

Discrete-event simulation (DES)

The machine as components exchanging timed events (a kernel starts, a DMA finishes, a memory request returns), and the clock jumps from one event to the next. Contention and queueing emerge rather than being assumed. Each event's duration comes from a cost model. (Glossary: discrete-event simulation.)

Validated against

Hardware counters and timings on an existing chip, RTL cycle counts for new blocks, and published results; the gap is reported as a calibration and correlation figure, not hidden.

On this GitHub

Speed

Analytical: microseconds per design point. DES: about 105 to 107 events per second in a compiled kernel, fewer in Python (indicative; InfSim 01). The number of events, set by the abstraction, matters more than the language: see InfSim 08: Fewer Events.

15

Architecture: Transaction-Level and Cycle-Level

Transaction-level modelling (SystemC TLM-2.0)

A bus transfer is a function call that carries a payload and an annotated delay, not a sequence of pin wiggles.

  • Loosely timed (LT): each initiator runs ahead of global time by up to a quantum (temporal decoupling). Fast enough to boot an operating system
  • Approximately timed (AT): a four-phase protocol per transfer, so contention and pipelining are modelled; slower

Cycle-level simulation

Pipelines, caches and queues updated every cycle, as in gem5's detailed CPU models or GPU simulators. About 104 to 106 cycles per second (indicative). More detail is not more accuracy unless it is calibrated against the real machine.

N. Binkert et al., The gem5 simulator, 2011 · J. Lowe-Power et al., The gem5 Simulator: Version 20.0+, 2020 · A. Akram, L. Sawalha, A Survey of Computer Architecture Simulation Techniques and Tools, 2019

Trace-driven or execution-driven

  • Trace-driven: replay a recorded stream of instructions, memory accesses or kernels. Fast and repeatable, but the program cannot react when the timing changes
  • Execution-driven: run the program on the model, so timing-dependent behaviour (spinning on a lock, adaptive batching) is captured. Slower
  • Sampling: simulate representative intervals in detail and fast-forward the rest (glossary: sampling, checkpoints and mode switching)
16

Level 5: System, Network and Cloud

DES at system scale

The events are now requests, batches, transfers and failures, across thousands of components and hours of traffic. Measured on this PC: Disaggregated_Inference_Sim (SimPy) serves 3,000 LLM requests, 770.9 s of simulated time, with 90,766 events in 0.923 s: 835× faster than real time, or 1,566× with its exact fast path. Script: demo/des_speed.

Queueing networks

Servers and queues with arrival and service distributions. Closed forms such as Little's law (L = λW) and the M/M/1 queue give instant answers for simple cases, and double as unit tests for a simulator: a correct DES must reproduce them.

Network simulators and agent-based models

Packet-level simulators (ns-3, OMNeT++) model protocols and topologies packet by packet; flow-level models trade packet detail for speed. Agent-based models give many autonomous agents simple local rules (people, vehicles, traders) and let the aggregate behaviour emerge.

G. F. Riley, T. R. Henderson, The ns-3 Network Simulator, 2010 · C. M. Macal, M. J. North, Tutorial on agent-based modelling and simulation, 2010

On this GitHub

Speed, accuracy, validation

  • Speed depends on events per simulated second: macro-stepping and coarser events buy orders of magnitude
  • Accuracy rests on the workload model (arrival process, request mix) at least as much as on the component models
  • Validated against production traces and against the queueing closed forms above
17

Monte Carlo and Variance Reduction

The method

Estimate an expectation by averaging random samples of the model. The standard error falls only as the square root of the number of samples, so each extra decimal digit costs a hundred times more runs.

standard error = σ / √N

Engineering uses: manufacturing yield under process variation (Monte Carlo SPICE), reliability and failure rates, risk and option prices, and the randomness inside every stochastic DES.

N. Metropolis, S. Ulam, The Monte Carlo Method, JASA, 1949

Variance reduction: shrink σ, not just grow N

  • Antithetic variates: pair each random draw with its mirror image, so their errors partly cancel
  • Control variates: subtract a correlated quantity whose mean is known exactly
  • Common random numbers: compare two designs on the same random stream, so the difference is not swamped by noise
  • Importance sampling: draw more samples from the rare region that matters and reweight them. Essential for bit error ratios of 10−12 or for failure probabilities
  • Quasi-Monte Carlo: low-discrepancy points (Sobol, Halton) fill the space more evenly than random ones

Reporting it

A Monte Carlo answer is an estimate with an error bar: report the confidence interval and the number of samples, and check that it shrinks as expected when N grows.

18

Level 6: Software Virtual Platforms

Instruction-set simulators (ISS)

Fetch, decode and execute one instruction at a time against modelled registers and memory. Simple, and exact at the level of the instruction set. Spike, the RISC-V reference ISS, is used as the golden model when verifying processor RTL.

Dynamic binary translation: QEMU

QEMU's code generator (TCG) translates each block of guest instructions into host code once and caches it. In the original paper, user-mode emulation was about 4× slower than native on integer code and 10× on floating point, with a further factor of 2 for the software MMU in full-system mode (QEMU 0.4.2; today's figures depend on the guest, host and workload).

F. Bellard, QEMU, a Fast and Portable Dynamic Translator, USENIX 2005

Virtual prototypes

A whole platform (CPUs, bus, memory map, peripherals) as fast functional models, often SystemC TLM loosely timed, so firmware, drivers and operating systems run before silicon. Register-accurate, not cycle-accurate: right for software bring-up, wrong for performance. Open source: Renode; CPU vendors and EDA vendors sell their own.

Speed, accuracy, validation

  • Interpretive ISS: tens of millions of instructions per second at most; translated (QEMU) and LT platforms: hundreds of millions to billions (indicative)
  • Functionally exact if the models are right; no timing to speak of
  • Validated by running the same software on the platform and on silicon or an FPGA, and by architecture compliance test suites
19

One Accelerator, Every Level

The FHE accelerator work on this GitHub models one design at five levels, and each level checks or feeds another. It is a compact example of how the levels work together on a real project.

LevelModelWhat it answersExplained in
AnalyticalRoofline and the NTT-, MAC-, memory- and power-bound testsWhich resource limits bootstrapping?FHESim 05: NTT-Bound Against Memory-Bound
Discrete-eventFHE_Accelerator_Sim (SimPy, bit-exact JS port)Latency, energy and hot-spots for a traceFHESim 03: The SimPy Engine
Memory systemMemory_System_Sim (command-level DRAM/HBM)Replaces "bandwidth × efficiency"SimEng 04: Plugging It Into the FHE Simulator
Transaction-levelSystemC_Accelerator_Model (TLM-2.0, AT and LT)Does a C++ model agree op by op?SimEng 03: Checked Against the SimPy Model
RTLRTL_CoSim_NTT (cocotb on Verilator)Real NTT cycle countsSimEng 05: Feeding the RTL Back Into the Simulator
Area and costThe calibrated area, yield and cost modelIs a design worth its silicon?SimEng 13: PPA Explorer (glossary: PPA)

The flows between levels

  • RTL cycle counts calibrate the NTT throughput the DES assumes
  • The DES's traces drive the SystemC model, which must agree with it
  • The DRAM model replaces a fixed efficiency factor with scheduled commands

A second thread: one SerDes link

Field solver (Signal Integrity 01–04) → S-parameters (Matrix Methods) → pulse response and equalisers (SerDes Equalisation) → a compliance verdict (Signal Integrity 17).

PART II

Cross-Cutting Methods

The questions every level shares: how time advances, how the step is chosen, how randomness is handled, how simulators are joined together, how a model earns trust, and how to choose a level in the first place.

20

Time-Stepping and Event-Driven

Fixed steps same h everywhere; stability caps h Adaptive steps small steps at the input edges Event-driven nothing computed between events event queue: next (t=905, arrival), then (t=1210, done) …

Time-stepping

For continuous state (voltages, fields, temperatures): advance by a step h and evaluate the derivatives. Fixed steps are simple and suit real-time use; adaptive steps follow the dynamics and save work where little happens.

Event-driven

For state that changes only at instants (a packet arrives, a signal toggles): keep a priority queue of future events, pop the earliest, run its handler, which may schedule more. Simultaneous events need a deterministic tie-break. (Glossary: DES.)

Hybrid

Continuous dynamics with discrete events (a switch opens, a thermostat fires): integrate, detect the zero crossing, locate it, then restart from it. SPICE breakpoints, switching simulators and FMI's event mode all do this.

21

Stiffness, Stability and Step Control

Explicit and implicit

  • Forward Euler (explicit): vn+1 = vn + h·f(tn, vn). Cheap per step, but on dv/dt = −v/τ each step multiplies the error by (1 − h/τ)
  • Backward Euler, trapezoidal, BDF (implicit): f evaluated at the new point, so each step needs a solve (Newton, for a nonlinear circuit), but decaying modes stay stable for any h
|1 − h/τ| < 1  ⇔  0 < h < 2τ   (forward Euler on dv/dt = −v/τ)

Stiffness

A system is stiff when it mixes very different time constants: picosecond parasitics with millisecond thermal drift, fast and slow chemistry. An explicit method's step is capped by the fastest mode even after that mode has died away; an implicit one can let the step follow the slow behaviour you care about. That is why SPICE integrates implicitly.

On this GitHub

Error control

Embedded pairs such as Dormand–Prince 5(4) compute two solutions of different order from the same stages; their difference estimates the local error, and the step grows or shrinks to keep it under a tolerance. The estimate assumes the solution is smooth. At a discontinuity (a square-wave edge, a switch closing) the solver rejects step after step until it has pinned the edge down, unless it is told where the edges are: breakpoints.

Try it next

On the RC circuit of the next slide, forward Euler at h = 2.2τ ends 25.8 V off a 1 V signal; backward Euler at the same step stays within 0.708 V. Adaptive Dormand–Prince at a tolerance of 10−6 needs 2,660 evaluations and rejects 218 steps; told where the edges are, it needs 560 and its error falls from 2.4 × 10−4 V to 2.5 × 10−7 V.

J. R. Dormand, P. J. Prince, A family of embedded Runge–Kutta formulae, 1980 · E. Hairer, G. Wanner, Solving Ordinary Differential Equations II: Stiff and Differential-Algebraic Problems

22

Interactive: One Circuit, Four Solvers

An RC low-pass filter (τ = RC = 1 ms) driven by a 0–1 V square wave of period 4τ, simulated for 20τ (10 edges). Error is the largest |v − vexact| at the solver's own time points. Every row is reproduced by demo/rc_solvers.py.

MethodSettingAccepted stepsRejectedRight-hand-side evaluationsMax error (V)

Try: forward Euler at h = 2.2τ (unstable) against backward Euler at the same step (stable, but smeared); Dormand–Prince with and without breakpoints at the same tolerance; the event-driven solver, exact to rounding in 10 events, because the input is constant between edges, as in a piecewise-linear switching simulator.

23

Deterministic and Stochastic

Deterministic models still need reproducibility

Same inputs, same outputs, on every run and every machine. That takes deliberate work: a fixed rule for simultaneous events, ordered floating-point reductions, no dependence on hash order or thread timing (glossary: determinism, tie-breaking and RNG streams). Ports to another language can even be bit-exact.

Stochastic models

  • Random arrivals, request sizes, failures, process variation
  • Seed every random stream explicitly and record the seed; give each source its own stream, so changing one does not shift the others
  • One run is one sample. Run independent replications, discard the warm-up transient, and report a mean with a confidence interval; within one long run, use batch means

Pitfalls

  • Reporting a single run as the answer
  • Confidence intervals computed from correlated observations, such as consecutive request latencies
  • Comparing two designs on different random streams, so noise swamps the difference (use common random numbers)
  • Tail percentiles: a p99 needs far more samples than a mean
24

Co-Simulation and Hardware-in-the-Loop

Co-simulation and FMI

Two or more simulators, each with its own solver, exchange values at agreed communication points. The Functional Mock-up Interface (FMI) packages a model as an FMU: an XML description plus C code or binaries. In model exchange the host integrates the FMU's equations; in co-simulation the FMU brings its own solver, and a master algorithm steps every FMU and passes signals between them at each macro step. The step size trades accuracy and stability against speed.

T. Blochwitz et al., The Functional Mockup Interface for Tool independent Exchange of Simulation Models, Modelica 2011

Mixed-signal and HW/SW

  • Mixed-signal: a time-stepping analogue solver and an event-driven digital kernel, synchronised at the boundary signals; real-number models of the analogue blocks speed it up
  • Hardware and software: Python drives RTL through cocotb while a golden model checks outputs; or compiled RTL sits inside a C++ or SystemC platform (glossary: co-simulation)

Hardware-in-the-loop (HIL)

The controller under test is real hardware (an engine control unit, a motor drive's controller board), wired to a real-time simulation of the plant it controls. The simulator must finish every step before its deadline, so plant models are simplified until they do. The cheaper stages come first: model-in-the-loop, software-in-the-loop (the controller code on a PC), processor-in-the-loop (on the target processor).

On this GitHub

25

Digital Twins, Defined Carefully

A definition worth using

"A digital twin is a set of virtual information constructs that mimics the structure, context, and behavior of a natural, engineered, or social system (or system-of-systems), is dynamically updated with data from its physical twin, has a predictive capability, and informs decisions that realize value. The bidirectional interaction between the virtual and the physical is central to the digital twin."

US National Academies, Foundational Research Gaps and Future Directions for Digital Twins (2024), adapting a 2020 AIAA definition

Model, monitor or twin?

What you haveLive dataPredictsFeeds decisions backCall it
A CAD or simulation modelnoyesnoa model
A dashboard of sensor datayesnomaybemonitoring
A model kept in step with live datayesyesnosometimes a "digital shadow"
All of these, with decisions acted onyesyesyesa digital twin

What a real twin needs

  • A model fast enough to keep up with the asset: usually a reduced-order model from the levels in Part I
  • Continual calibration from data (data assimilation), with its uncertainty quantified
  • Verification and validation of the whole loop, not just the model

The term is used loosely in marketing: ask which of the four properties actually hold. Nothing on this GitHub is a twin in this strict sense; the nearest ingredient is calibration, such as calibrating a simulator against OpenFHE.

26

Verification, Validation and Calibration

Three different questions

  • Verification: did we build the model right? The code matches the intended model: unit tests, analytic cases, convergence studies, differential testing against an independent implementation
  • Validation: did we build the right model? Its output matches the real system to the accuracy the decision needs, over the range of conditions it will be used for
  • Calibration: tuning parameters until the model matches data. Calibrate on one data set and validate on another, or the validation proves nothing

R. G. Sargent, Verification and validation of simulation models, J. Simulation, 2013

Model credibility

  • State the intended use and the domain the model covers
  • Document assumptions and what each level leaves out
  • Report validation results with the conditions they cover, and their uncertainty
  • Keep golden runs in CI, so a change that moves an answer is noticed

A ladder of evidence

  1. Closed-form cases the model must reproduce
  2. Convergence as the mesh or step is refined
  3. Agreement with another simulator
  4. Agreement with measurement, out of sample

(Glossary: the verification ladder.)

27

Speed, Accuracy and Effort

10-14 10-12 10-10 10-8 10-6 10-4 10-2 100 102 104 106 108 1010 real time simulated seconds per wall-clock second (log scale; cycle counts converted at a 1 GHz clock) most physical detail at the top 3-D full-wave EM (FDTD) L1 SPICE, transistor level L2 Gate level + SDF timing L3 RTL simulation L3 Icarus, measured: 6.42e-06 s/s Verilator, measured: 0.000838 s/s Cycle-level architecture L4 Hardware emulation L3 FPGA prototype L3 TLM virtual platform (LT) L4/6 QEMU dynamic translation L6 System DES (LLM serving) L5 event per step: 835 s/s fast path: 1.57e+03 s/s Analytical models L4/5 indicative band cited (paper) worked estimate measured on this PC

Reading it

Rows run from most physical detail (top) to least. Bands are orders of magnitude and each spans decades, because model size matters as much as the level. Indicative bands follow InfSim 01's ladder (large SoC, 1 GHz); QEMU is from Bellard (2005); FDTD is the worked estimate on the field-solver slide.

Measured here

Icarus and Verilator on a small NTT core sit at and beyond the fast end of the RTL band, because the core is small; the SimPy serving simulator runs 835× to 1,566× faster than real time. Scripts in demo/.

Effort

Detail costs model-building time and calibration data: a roofline takes an afternoon; a validated cycle-level or meshed 3-D model takes weeks; RTL arrives late. The trade-off is named in SimEng 13: The Other Trade-offs; feel the spread in InfSim 01's calculator.

28

How to Choose a Level

Questions to ask first

  1. Which decision will the answer change, and how accurate must it be to change it?
  2. What is the smallest unit of time and data that affects the answer?
  3. How much simulated time, or how many samples, do you need? A p99 needs hours of traffic; a reflection needs nanoseconds
  4. What can you calibrate and validate against?
  5. When do you need it, and who will maintain it?

The rule of thumb

Use the highest level that can answer the question, and drop to a lower level only for the part that needs it, through a hand-over model or co-simulation. Detail you cannot calibrate is not accuracy.

QuestionLevel and methodWhy
Impedance of a new board stack-up1: 2-D field solverGeometry and materials set it
Eye opening at 28 GBd through a backplane1–2: S-parameters, then channel simulation with IBIS-AMINanoseconds of waveform; statistical BER
Will the converter ring on a load step?2: averaged model, then PWL switching simulationMilliseconds of switching
Does the RTL meet its specification?3: RTL simulation against a golden modelExact logic; regressions in a compiled simulator
How big should the on-chip SRAM be?4: analytical sweep, then DESHundreds of design points
p99 time-to-first-token under bursty traffic5: system DES with replicationsHours of traffic, many seeds
Will the driver boot before silicon?6: virtual platformBillions of instructions, no timing
Yield of a large die under process variationAnalytical yield model, then Monte CarloClosed forms, then the variation

Worked versions: InfSim 04: Which Question, Which Level? · Build or Reuse? · SimEng 13: Dies per Wafer and Yield

29

Takeaways

What to remember

  • Why simulate: to explore, verify, predict, start software early, see inside, ask what-if, do what reality will not allow, and train; not when an analytical answer or a measurement is cheaper
  • Six levels, from fields on a mesh to software on a modelled CPU; each trades detail for speed and reach, and hands a reduced model to the level above
  • Two ways to advance time: steps for continuous state, events for discrete changes; real simulators mix them
  • Explicit methods are limited by stability (CFL, h < 2τ); implicit methods by the cost of a solve per step. Stiff problems need implicit methods
  • Adaptive steps need help at discontinuities: breakpoints, or event-driven segments
  • Stochastic answers need seeds, replications and confidence intervals; deterministic ones need reproducibility
  • Verification, validation and calibration are different jobs, and calibration data must not double as validation data
  • Choose the highest level that answers the question

Next steps

  1. Run the four-solver demo until the stability limit and the breakpoint effect are obvious
  2. Build a discrete-event kernel in 30 lines: InfSim 02
  3. Time Icarus against Verilator on a design of your own with demo/rtl_speed
  4. Pick a level you rarely use and read its "On this GitHub" links

Further reading

Law, Simulation Modeling and Analysis (discrete-event) · Hairer & Wanner, Solving ODEs II (stiff problems) · Taflove & Hagness, Computational Electrodynamics: The FDTD Method · Sargent (2013) on V&V · the reading lists in InfSim 10

30

Where to Go Next

LevelStart hereThen
1 PhysicsSignal Integrity series (17 decks; computed by a field solver)Matrix Methods in Network Parameters · Matrix Articles models · Numerical Methods
2 CircuitDC-DC Control TechniquesSerDes Equalisation · Matrix Concepts in Digital Filters
3 Digital logicFree SystemVerilog SimulatorsSimEng 05: verification bridge · RTL_CoSim_NTT · Hardware index (RTL repos)
4 ArchitectureLLM Inference Simulators (InfSim 01–11)FHE Accelerator Simulators · SimEng 03: SystemC TLM · SimEng 04: memory systems · SimEng 13: PPA
5 SystemInfSim 05: Disaggregated InferenceInfSim 06: metrics and validation · LLM Inference Explained (the simulator live in a web page) · Wilmott 14: Monte Carlo · Markov chains
6 SoftwareRISC-V deck (instruction stepper)Interview_RISC_V · Cortex-M series
Engineering a simulatorSimulation Engineering Toolkit (13 decks)Rust and PyO3 ports, testing, Jenkins, specifications, performance analysis, measurement tools

The three simulation series

LLM Inference Simulators (why and how to simulate inference hardware) · FHE Accelerator Simulators (one accelerator, end to end) · Simulation Engineering Toolkit (the engineering around a simulator). Each hub has a glossary linking every concept to the slides that explain it.

Indexes

Hardware · Software · Mathematics