Presentations in This Series
- Why Simulate? From Spreadsheets to RTL →The fidelity ladder of pre-silicon models — analytical, discrete-event, transaction-level, cycle-accurate, RTL, emulation — what each answers, what each costs, and how software-level simulators and RTL simulation feed each other.
- Simulator Development: A Hands-On Tutorial →Build a discrete-event simulator from an empty file: the event queue, simulated time, processes, resources and probes; then SimPy idioms, how to structure a simulator so it survives contact with architects, modelling buses, memories and pipelines, and making the simulator itself fast.
- What You Are Simulating: LLM Inference Workloads →Prefill versus decode, the roofline, KV-cache arithmetic, batching and parallelism — the first-order physics every LLM inference simulator must get right, with the formulas to put in a cost model.
- The LLM Simulator Landscape →A guided tour of the tools: serving-level simulators (Vidur, LLMServingSim, SplitwiseSim), hardware-level models (LLMCompass, Timeloop, SCALE-Sim), system and network simulators (ASTRA-sim, gem5, SST, SystemC), and analytical calculators — sorted by the question each one answers.
- Disaggregated Inference, Simulated →Why splitting prefill and decode onto separate hardware raises goodput (Splitwise, DistServe, Mooncake, NVIDIA Dynamo), what the KV-cache transfer costs, and a live discrete-event simulator you can drive in the browser — a JavaScript port of the SimPy model in Disaggregated_Inference_Sim.
- Metrics, Hot-Spots & Validation →Turning a simulation into evidence: latency distributions, utilisation, MFU/MBU, stage breakdowns and hot-spot attribution; traces in Perfetto; statistics that survive review; and the verification ladder — unit, invariant, analytic, behavioural, correlation — wired into CI.
- Power & Energy in Inference Simulators →Speed is half of performance. Static and dynamic power, why data movement dominates energy, the simulator's calibrated power model, prefill running hot and decode cool, DVFS and power caps as a third roof, energy proportionality, the link to RTL power and thermal analysis, and what photonic systems must pay for — all measured on the companion simulator.
- Accelerating Inference Simulators →Making this kind of simulator fast without changing its answers: event abstraction, incremental state, lazy bookkeeping, exact macro-stepping, probe discipline, compiled kernels, parallel and multi-fidelity sweeps, surrogates, sampling, and parallel discrete-event simulation — each measured on the companion simulator.
- From PyTorch, ONNX & HEIR to a Simulator →How real applications reach a simulated accelerator: graph capture with torch.fx / torch.export and torch.compile backends, out-of-tree PyTorch devices, ONNX Runtime execution providers, MLIR, and Google's HEIR compiler for fully homomorphic encryption.
- Further Learning: Books, Courses, Papers, Tools →A curated reading list for simulator engineers: textbooks on discrete-event simulation and computer architecture, university courses, the key LLM-serving and simulator papers, and the open-source tools worth reading the source of.
- Role Primer: Simulation & Frameworks Engineer →The series applied to a real job: a photonic-computing architecture group's simulation and frameworks role, mapped requirement by requirement onto the decks, with the domain background (photonics, FHE) and a set of practice questions.
Companion Code
- </>Disaggregated_Inference_Sim →A SimPy discrete-event simulator of prefill/decode-disaggregated LLM serving: roofline cost model, continuous batching, KV-capacity admission, a shared KV link, colocated baseline; latency percentiles, goodput, utilisation, MFU/MBU, stage breakdown, hot-spot attribution, Perfetto traces; a power model (static + pJ/FLOP + pJ/byte + pJ/bit, DVFS, per-pool power caps, joules per token); an exact accelerated path, search utilities, heterogeneous pools, FFT-mixing models, an optical transform engine, KV hand-off compression, a Causal Encoder-Decoder (encoder-only prefill) option, 154 tests, and the JavaScript port used live in deck 05.
How to read this series. Decks 01–02 are about simulators in general and stand alone. Decks 03–05 apply them to LLM serving and culminate in the live simulator. Decks 06–09 cover what turns a simulator into an engineering tool: trustworthy metrics, power and energy, simulator speed, and integration with real frameworks. Deck 10 is a reading list; deck 11 applies the whole series to a real job specification. Every concept is explained in the glossary below.
Glossary: Concepts and Where They Are Explained
Every concept the decks rely on, with a short explanation and links to where it is explained in depth: in this series, in the related decks elsewhere on this GitHub (Local LLM Hosting, Key Publications), or in the sister series FHE Accelerator Simulators, whose glossary covers the FHE concepts in depth. Each deck's contents slide links to the entries it uses.
Simulation fundamentals
- The fidelity ladder
- The range of model detail from analytical formulas through discrete-event, transaction-level, cycle-level and RTL simulation to emulation, FPGA prototypes and silicon. Each rung is slower and more detailed; use the highest rung that answers the question.Explained in: InfSim 01 · the fidelity ladder · InfSim 01 · what each level throws away · InfSim 01 · how long it would take
- Discrete-event simulation (DES)
- A simulator whose state changes only at instants (events). A clock, a state, a time-ordered event list and handlers; the clock jumps from one event to the next (next-event time advance), so idle time costs nothing.Explained in: InfSim 02 · anatomy of a DES · InfSim 02 · a kernel in 30 lines · InfSim 02 · step through the event list
- Process interaction and coroutines
- Writing each entity's life as one sequential function that suspends while it waits, instead of scattering it over event handlers. Python generators make this natural: yield an event, resume when it fires.Explained in: InfSim 02 · generators as coroutines
- SimPy
- A small, pure-Python DES library: Environment, processes, timeouts, events, Resource, Container, Store and interrupts. It has no built-in statistics or tracing; the probes are yours to design.Explained in: InfSim 02 · the whole vocabulary · InfSim 02 · modelling hardware in SimPy
- Resources, stores, containers and back-pressure
- Resource: N identical servers held for a while (links, engines). Store: a buffer of items (FIFOs). Container: an amount (credits, bytes). A finite-capacity Store makes producers wait when it is full, which is back-pressure.Explained in: InfSim 02 · modelling hardware in SimPy
- FCFS hold versus processor sharing
- Two ways to model a shared link: one transfer holds it at full rate while others queue, or all active transfers share it equally. Throughput and utilisation match; latency distributions do not.Explained in: InfSim 02 · hold the link or share it?
- Cost model
- The part of a simulator that answers "how long does this take?" (and, with a power model, "how many joules?"). Kept separate from the event engine so a roofline can later be replaced by a measured table or an RTL-calibrated model.Explained in: InfSim 02 · structuring a simulator · InfSim 03 · a cost-model checklist · InfSim 04 · inside Vidur
- Determinism, tie-breaking and RNG streams
- Same configuration and seed must give bit-identical results: break ties between simultaneous events with a sequence number, prefer integer time for clocked models, and give each component its own random stream.Explained in: InfSim 02 · time, ordering and determinism pitfalls · InfSim 02 · the tie-breaking sequence number
- Transaction-level modelling (SystemC TLM-2.0)
- Modelling bus reads and writes as function calls with annotated delays rather than signal wiggles. Loosely-timed TLM boots software; approximately-timed TLM models contention. The industry route to virtual platforms.Explained in: InfSim 02 · beyond Python · InfSim 04 · system and network simulators
- RTL simulation, co-simulation and golden models
- The software model is the executable specification and the golden reference (a scoreboard in the RTL testbench). RTL cycle counts calibrate the model, workload traces become directed tests, and co-simulation embeds RTL blocks in the system model.Explained in: InfSim 01 · software-level simulators and RTL · InfSim 01 · the co-flow · InfSim 01 · co-simulation in practice
- Serving, accelerator and network simulators
- Serving simulators (Vidur, LLMServingSim, SplitwiseSim) model request streams and schedulers; accelerator models (LLMCompass, Timeloop, SCALE-Sim) model compute arrays; system simulators (ASTRA-sim, gem5, SST) model networks and memory.Explained in: InfSim 04 · which question, which level? · InfSim 04 · side by side · InfSim 04 · inside Vidur
LLM inference
- Prefill and decode
- Prefill processes the whole prompt in one pass (matrix-matrix, compute-bound) and produces the first token. Decode then generates one token per sequence per step (matrix-vector, memory-bound). They need different hardware balances.Explained in: InfSim 03 · two phases · InfSim 05 · the interference problem
- FLOP and byte counting
- About 2 FLOPs per matmul parameter per token, plus attention that grows with context. A decode step must also read every weight and all of the KV cache. These two counts drive every timing and energy number in the series.Explained in: InfSim 03 · counting FLOPs · InfSim 03 · counting bytes
- Roofline, arithmetic intensity and ridge point
- Attainable performance is the lower of peak compute and bandwidth × arithmetic intensity (FLOPs per byte). The ridge point is the intensity at which both limits meet; below it a kernel is memory-bound.Explained in: InfSim 03 · the roofline · InfSim 03 · interactive roofline · Key Publications 04 · FlashAttention as IO-awareness · SimEng 12 · the roofline as an analytic bound
- KV cache
- The stored keys and values of every token of every live sequence, so attention need not recompute them. Its size (320 KiB per token for Llama-3-70B) limits batch size and is what disaggregation must move.Explained in: InfSim 03 · KV-cache arithmetic · Local LLM 03 · why the KV cache dominates serving
- Continuous batching
- Re-forming the batch at every decode iteration: finished sequences leave and new ones join, so the GPU stays full. The scheduler loop at the heart of every serving engine and serving simulator.Explained in: InfSim 03 · batching · Local LLM 03 · iteration-level scheduling
- PagedAttention
- Storing the KV cache in fixed-size blocks with a per-sequence block table, like virtual memory, which removes fragmentation and allows prefix sharing.Explained in: Local LLM 03 · PagedAttention · Key Publications 04 · the vLLM paper · InfSim 03 · batching
- Chunked prefill
- Splitting long prompts into chunks and mixing them with decode tokens in each step, which bounds how long decodes stall. The alternative to disaggregation for the same interference problem.Explained in: InfSim 03 · batching · Local LLM 03 · prefix caching and chunked prefill
- Tensor, pipeline and expert parallelism
- Splitting a model across devices by weight matrix (TP: all-reduces per layer), by layer (PP: activations between stages) or by expert (EP: all-to-all). Each adds communication a simulator must cost.Explained in: InfSim 03 · parallelism and its bill · Local LLM 05 · tensor parallel · Local LLM 05 · expert parallel
- Quantisation
- Storing weights (and sometimes activations or KV) in fewer bits. It shrinks the dominant byte term of decode, so it helps bandwidth-bound decode far more than compute-bound prefill.Explained in: InfSim 03 · counting bytes · Local LLM 07 · why quantise · Local LLM 07 · FP8 and beyond
Serving metrics
- TTFT, TPOT and inter-token latency
- Time to first token (prefill queueing and compute); time per output token after the first (DistServe's definition); and the gap between consecutive tokens, whose tail shows stalls the mean hides.Explained in: InfSim 03 · the metrics users feel · InfSim 05 · the interference problem
- Goodput and SLO attainment
- The rate of requests that meet every SLO (for example TTFT ≤ 1 s and TPOT ≤ 25 ms), rather than raw throughput. The metric DistServe optimises and the simulator reports.Explained in: InfSim 03 · the metrics users feel · InfSim 05 · the idea: split the phases · InfSim 05 · results
- Percentiles and tail-latency CCDFs
- Report p50, p90 and p99, not means, and plot the fraction of samples slower than x on log-log axes, where tails show. Say which population (all tokens or per request), and use streaming sketches when samples are many.Explained in: InfSim 06 · distributions, not averages · InfSim 05 · the live CCDF
- Utilisation, MFU and MBU
- Busy fraction says how often a resource works; model FLOPs and bandwidth utilisation say how much of its capability it uses. A continuous-batching decode engine is 100% busy at 4% MFU and about 77% MBU.Explained in: InfSim 06 · utilisation and efficiency · InfSim 05 · the simulator's resource table
- Little's law and the operational laws
- L = λW: mean population equals arrival rate times mean time in system, for any stable system. With the utilisation and forced-flow laws it gives free consistency checks on a simulator's probes.Explained in: InfSim 06 · operational laws · InfSim 02 · exercise: test Little's law
- Hot-spot attribution
- Finding where speeding things up would help most. A stage breakdown that sums exactly to each request's latency, the largest waiting stage mapped to the resource that owns it, confirmed by what-if sensitivity runs.Explained in: InfSim 06 · hot-spot attribution · InfSim 05 · experiment: starve the link
- Chrome trace events and Perfetto
- A JSON timeline format (complete slices, counters, metadata) that Perfetto and chrome://tracing display. The simulator, ONNX Runtime's profiler and RTL testbenches can all emit it, so timelines can be compared side by side.Explained in: InfSim 06 · traces in Perfetto · InfSim 09 · ONNX Runtime's profiler · SimEng 12 · Perfetto, as a measurement tool
Disaggregated serving
- Prefill–decode interference
- On a colocated server a waiting prefill runs before the next decode step, so every sequence mid-generation stalls for the whole prefill: mean TPOT moves a little, p99 inter-token latency jumps about tenfold.Explained in: InfSim 05 · the interference problem · InfSim 03 · two phases
- Disaggregated (prefill/decode) serving
- Running prefill and decode on separate machine pools (Splitwise, DistServe, Mooncake, NVIDIA Dynamo), so each phase is provisioned and power-managed for its own SLO. The cost is moving the KV cache.Explained in: InfSim 05 · split the phases · InfSim 05 · in production · InfSim 05 · live simulator
- KV-cache transfer
- Bytes = prompt tokens × KV per token (671 MB for a 2k-token Llama-3-70B prompt): 1.5 ms on NVLink, 13 ms on InfiniBand NDR, 215 ms on 25 GbE. The link and its queue can become the hot-spot.Explained in: InfSim 05 · what the KV transfer costs · InfSim 05 · experiment: starve the link
- Encoder-only prefill (causal encoder-decoder)
- A model whose prefill runs only part of the network: in DeepSeek-V4.1-Flash's Causal Encoder-Decoder the bottom half of the layers is a causal encoder and every decoder layer projects its KV from the last encoder state, so a prompt token activates 8B parameters and a generated token 16B. Prefill and decode pools then run different amounts of the model, which changes how they are sized.Explained in: InfSim 05 · asymmetric models: encoder-only prefill · Modern Architectures 06 · where the decoder's KV comes from · Modern Architectures 06 · asymmetric prefill and decode pools
Statistics and validation
- Warm-up, replications and confidence intervals
- A stochastic run is one experiment: drop the warm-up period, repeat with different seeds, and report mean ± a t-based confidence interval. Tail percentiles need many samples.Explained in: InfSim 06 · statistics that survive review · InfSim 06 · interactive replications · SimEng 12 · costs, pitfalls and how many runs
- Batch means
- One long run cut into batches long enough to be nearly independent, each giving one sample of the metric: warm-up is paid once instead of once per replication, at the risk of correlated batches.Explained in: InfSim 06 · statistics that survive review · SimEng 12 · batch means among the output statistics
- Common random numbers
- Feeding two designs the same random request stream, so noise cancels in their difference: a much narrower confidence interval on "which is better, and by how much" for the same number of runs.Explained in: InfSim 06 · interactive CRN demo · InfSim 06 · statistics · SimEng 12 · CRN among the output statistics
- Queueing theory checks (M/M/1, M/D/1)
- Closed-form results that a simulator's engine must reproduce: M/M/1 mean time in system 1/(μ−λ); M/D/1 mean wait λS²/(2(1−ρ)) (Pollaczek–Khinchine). The KV link is tested this way.Explained in: InfSim 02 · a queue, checked against theory · InfSim 06 · the verification ladder
- The verification ladder
- Layers of evidence: unit tests against hand calculations, invariants, analytic checks, behavioural expectations, property-based tests, differential tests and, finally, correlation against measurement.Explained in: InfSim 06 · the verification ladder · InfSim 06 · testing frameworks
- Property-based testing (Hypothesis)
- Generate many random configurations and check that invariants hold for all of them; Hypothesis shrinks any failure to a minimal example.Explained in: InfSim 06 · testing frameworks
- Differential testing
- Requiring two independent implementations to agree exactly on the same input: here the fast path against the baseline, and the browser JavaScript against the Python simulator.Explained in: InfSim 06 · the verification ladder · InfSim 08 · exact macro-stepping · InfSim 08 · keep a reference · SimEng 12 · agreement as a measurement
- Calibration and correlation
- Fitting a model's few free coefficients to measurements (or RTL), then tracking its error on workloads it was not fitted to. The correlation report is the deliverable that makes predictions trustworthy.Explained in: InfSim 06 · validation is a number · InfSim 01 · the co-flow · InfSim 07 · calibrating power · SimEng 12 · agreement as a measurement
- CI, golden metrics and engineering metrics
- Run tests on every push and sweeps nightly (Jenkins or similar); fail the build if golden metrics drift; track correlation error, coverage, simulator speed and test health as engineering metrics.Explained in: InfSim 06 · CI/CD and engineering metrics · InfSim 06 · specs, test plans and reports
Power and energy
- Static and dynamic power
- Static power (leakage, clocks, always-on logic) is paid whenever a chip is on; dynamic power αCV²f is paid for switching. Energy per operation scales with V², which is why lowering voltage pays quadratically.Explained in: InfSim 07 · where the joules go · InfSim 07 · energy proportionality
- Energy of data movement
- Fetching a byte from DRAM or HBM costs tens to thousands of times the energy of an arithmetic operation on it (Horowitz, ISSCC 2014), so memory-bound decode is an energy problem too.Explained in: InfSim 07 · Horowitz's energy table · InfSim 03 · bytes cost energy
- The simulator's power model
- Static watts × time, plus pJ per FLOP, per HBM byte and per link bit, applied to the counts the roofline already has. Coefficients are illustrative and calibrated by regressing measured power on FLOP and byte rates.Explained in: InfSim 07 · the power model · InfSim 07 · prefill runs hot, decode cool
- DVFS and power caps (the third roof)
- Lowering clock and voltage to cut power. A memory-bound step can slow its compute clock at no time cost; a power cap makes some steps power-bound, a third roof beside compute and memory.Explained in: InfSim 07 · DVFS and power caps · InfSim 07 · interactive power cap
- TDP and peak power
- The board power limit a device throttles to. Steps that saturate compute and memory at once can exceed it, so the simulator enforces TDP by default and reports the peak power of each step.Explained in: InfSim 07 · the power roof · InfSim 07 · power metrics checklist
- Energy per token and energy–delay trade-off
- Joules per output token (or tokens per joule) is the efficiency headline. Capping power trades time for energy until static power over the longer run dominates, giving an energy-optimal operating point.Explained in: InfSim 07 · the energy-delay curve · InfSim 07 · latency against energy · InfSim 07 · metrics checklist
- Energy proportionality
- Because static power is paid busy or idle, energy per token falls steeply as load rises (fivefold from 0.5 to 6 req/s here). Static-heavy hardware must be kept busy or power-gated.Explained in: InfSim 07 · energy proportionality · InfSim 07 · the datacentre view
- RTL power and thermal analysis
- Switching activity from RTL simulation or emulation, run on real workload traces, gives per-block energies that calibrate the simulator; power maps feed thermal models that find thermal hot-spots.Explained in: InfSim 07 · simulator to RTL power · InfSim 01 · the co-flow · SimEng 12 · RTL power flows
- Power in photonic compute
- Optical operations are cheap, but lasers and thermal tuning burn static power, and every DAC/ADC conversion costs energy that grows with resolution. Gains need high reuse per conversion and a busy core.Explained in: InfSim 07 · power in photonic compute · InfSim 11 · photonic computing primer · FHE glossary · converter energy
Simulator acceleration
- Profiling a simulator
- Measure before optimising (cProfile, py-spy, perf). Time usually goes to model code that runs per entity per step, not to the event kernel.Explained in: InfSim 08 · step zero: profile
- Event abstraction
- Choosing what an event is: one per batch step instead of per token or per layer gave 14× and 80× fewer events, exactly, because nothing observable happens inside a step.Explained in: InfSim 08 · fewer events
- Incremental state and lazy bookkeeping
- Keeping running sums instead of recomputing them, and reconstructing per-entity details only when needed (from a per-instance step log), so each step costs O(1) instead of O(batch).Explained in: InfSim 08 · cheaper events
- Macro-stepping (time skipping)
- Between state changes the future is deterministic, so compute many steps in a tight loop and schedule one event; interrupt when new work arrives. Exact, and tested bit-identical to the baseline.Explained in: InfSim 08 · exact macro-stepping · InfSim 08 · measured
- Multi-fidelity search and bisection
- Use a cheap analytic bound to bracket the answer, then bisect with the simulator: 7 runs instead of a 40-point grid. Bayesian optimisation and surrogates go further.Explained in: InfSim 08 · smarter experiments
- Parallel runs and Amdahl's law
- Independent runs parallelise trivially (processes, job arrays), but each technique only speeds up the share of time it touches; serial overheads cap the gain.Explained in: InfSim 08 · parallel runs · InfSim 08 · Amdahl's law for simulators
- Sampling, checkpoints and mode switching
- Simulate representative regions in detail (SimPoint, SMARTS), fast-forward the rest in a cheaper mode, and fork experiments from checkpoints.Explained in: InfSim 08 · sampling and checkpoints
- Parallel discrete-event simulation (PDES)
- Splitting one run across cores as logical processes: conservative (Chandy–Misra–Bryant, needs lookahead) or optimistic (Time Warp, rolls back). Worth it for large, loosely coupled models.Explained in: InfSim 08 · PDES · InfSim 02 · making the simulator fast
Frameworks, compilers and FHE
- torch.export and FX graphs
- Capturing a PyTorch model ahead of time as one graph of ATen operators with fake-tensor shapes on every node; run_decompositions() lowers it to the smaller Core ATen set a cost model can cover.Explained in: InfSim 09 · graph capture with torch.export · InfSim 09 · three ways to drive a simulator
- torch.compile backends
- TorchDynamo captures graphs as Python runs and hands each to a backend callable, a clean hook for a simulator that estimates cost and then runs the graph eagerly.Explained in: InfSim 09 · a torch.compile backend
- Dispatch interception and the meta device
- A TorchDispatchMode sees every ATen call; on the meta device tensors have shapes but no storage, so a 70B model's exact operator trace can be recorded on a laptop.Explained in: InfSim 09 · dispatch interception
- PrivateUse1 out-of-tree devices
- PyTorch's dispatch key for vendor accelerators: rename it, register kernels, and users write model.to("device"). A functional simulator behind it shows accuracy as well as speed.Explained in: InfSim 09 · the simulator as a device
- ONNX and ONNX Runtime execution providers
- ONNX is a framework-neutral graph format; ONNX Runtime partitions graphs between execution providers, each claiming the nodes it can run, with fallback to the CPU.Explained in: InfSim 09 · ONNX and ONNX Runtime
- MLIR and dialects
- A compiler framework built from dialects (operation sets at one level of abstraction) and progressive lowering. A custom dialect for new hardware is the natural input to its simulator.Explained in: InfSim 09 · MLIR in one slide · FHE glossary · MLIR, dialects and SSA
- Fully homomorphic encryption (FHE)
- Computing on encrypted data. Ciphertexts are large RNS polynomials; the work is NTTs, modular arithmetic and key switching, and the data movement of ciphertexts and keys dominates.Explained in: InfSim 09 · FHE in one slide · InfSim 09 · FHE data-movement calculator · FHE glossary · CKKS · FHE glossary · bootstrapping
- CKKS bootstrapping
- Refreshing a CKKS ciphertext that has run out of levels by evaluating decryption homomorphically. Dominated by rotations and their key switches, so evaluation-key traffic often makes it memory-bound.Explained in: FHE glossary · bootstrapping · InfSim 11 · the FHE half: findings
- Number-theoretic transform (NTT)
- The FFT over a finite field, used for polynomial multiplication in FHE. The kernel optical transform engines target.Explained in: InfSim 09 · FHE in one slide · FHE glossary · the NTT · FHE glossary · optical transforms
- HEIR
- Google's MLIR-based FHE compiler: secret-annotated programs are lowered through scheme dialects to polynomial arithmetic and then to libraries such as OpenFHE, or to hardware.Explained in: InfSim 09 · HEIR · FHE glossary · HEIR as a simulator front end