LLM Inference Simulators — Presentation 04

The LLM Simulator Landscape

A guided tour of the tools: serving-level simulators (Vidur, LLMServingSim, SplitwiseSim), hardware-level models (LLMCompass, Timeloop, SCALE-Sim), system and network simulators (ASTRA-sim, gem5, SST, SystemC), and analytical calculators — sorted by the question each one answers.

Vidur LLMServingSim LLMCompass ASTRA-sim gem5 / SST SystemC TLM
Serving → Cluster → Accelerator → Micro-arch → Network
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.

01

A Map: Which Question, Which Level?

LLM-related simulators fall into a few bands. Each band takes a different input and answers a different question. The usual mistake is reaching for a tool from the wrong band.

Serving / cluster"What TTFT/TPOT/goodput for this request stream and deployment?" Vidur LLMServingSim SplitwiseSim DistServe sim / ours Analytical"Rough step time and memory for this model on that hardware?" LLM-Viewer GenZ Calculon roofline spreadsheets Accelerator architecture"How should the compute array, SRAM and dataflow be designed?" LLMCompass Timeloop+Accelergy SCALE-Sim Accel-Sim System & network"What do collectives, memory systems and interconnects cost?" ASTRA-sim gem5 SST SystemC TLM ns-3 / Garnet upper bands take request streams; lower bands take operators, traces or instructions
02

Serving-Level Simulators

These simulate a whole serving deployment (request arrivals, a scheduler, batching, KV-cache management, many replicas) and report the latency and throughput users experience. They are discrete-event simulators whose events are batch steps and requests, not cycles.

Vidur (Microsoft Research, MLSys 2024)

The reference serving simulator. Profiles real GPU operators once, trains per-operator runtime predictors, then simulates vLLM-, Orca- and Sarathi-style schedulers at scale. Vidur-Search explores deployment configurations (parallelism, batch size, scheduler) for a target workload at a fraction of the cost of trying them on GPUs.

LLMServingSim (IISWC 2024)

A hardware/software co-simulation infrastructure for LLM serving at scale. It combines an iteration-level serving model with accelerator and network simulators (it builds on ASTRA-sim) so that new NPU or memory designs can be judged on serving metrics rather than kernel speed alone.

SplitwiseSim (Microsoft, with Splitwise, ISCA 2024)

The discrete-event simulator released alongside the Splitwise paper on phase splitting. It models prefill and decode machine pools, KV transfer and cluster-level scheduling, driven by production-derived traces. It is the closest public relative of the simulator built in this series.

DistServe's simulator (OSDI 2024)

DistServe uses a simulator inside its placement algorithm: it searches over prefill and decode parallelism and instance counts, estimating goodput for each candidate by simulation, then deploys the best one.

Research in this band moves quickly; new serving simulators appear every few months, often as artifacts of systems papers. Check a paper's artifact link before writing your own.

03

Inside Vidur

Vidur's structure is the template for a calibrated serving simulator, and the same separation of concerns from deck 02 at production scale.

Operator profiling
GEMMs, attention, collectives on real GPUs
→
Runtime predictors
regressors per operator, interpolate unseen shapes
→
Event-driven engine
global and replica schedulers, batching, KV
→
Metrics & search
TTFT, TPOT, MFU; config search
04

Accelerator-Level Models

ToolFromWhat it modelsUse it to…
LLMCompassPrinceton, ISCA 2024Block-level hardware description (cores, systolic arrays, vector units, SRAM, HBM, links) with an automatic operator mapper, plus area and cost modelsCompare hypothetical LLM accelerator designs on performance and cost
Timeloop + AccelergyMIT / NVIDIALoop-nest mappings of tensor operations onto a memory hierarchy; energy per accessExplore dataflows and buffer sizes for one operator at a time
SCALE-SimGeorgia Tech / ArmCycle-level systolic-array GEMM and convolution, with memory trafficSize a systolic array, compare weight/output/input-stationary dataflows
Accel-Sim / GPGPU-SimUBC, Purdue and othersCycle-level NVIDIA-like GPU, trace- or execution-drivenStudy GPU micro-architecture changes on real kernels
LLM-Viewer, GenZAcademic, 2024Analytical roofline per layer and phaseQuick "what-if" on memory, bandwidth and FLOPs for many models
CalculonNVIDIA, SC 2023Analytical model of large-model training across parallelism strategiesCo-design training systems; a good example of a pure analytical model at scale
Novel compute

None of these knows about optical matrix units, ADC/DAC conversion or FHE polynomial arithmetic. That is why organisations building novel architectures write in-house simulators, often reusing the structure of these tools while replacing the hardware description.

05

System, Network and Micro-Architecture Simulators

ASTRA-sim

Distributed-ML system simulator from Georgia Tech with Meta and Intel. Three layers: a workload layer (now fed by MLCommons Chakra execution traces), a system layer (collective algorithms and scheduling), and a pluggable network layer (analytical, Garnet or ns-3). The tool for "what does this all-reduce cost on that topology?"

gem5

The long-running open-source computer-architecture simulator: CPU models from functional to out-of-order, the Ruby and Garnet memory and network models, full-system Linux boot. Slow but deep; the standard for memory-hierarchy research.

SST (Structural Simulation Toolkit)

Sandia's parallel, component-based framework. Components such as processors, memory (with the DRAMSim/Ramulator family of DRAM models) and networks are wired up from Python, and the core runs them in parallel with MPI. Built for HPC-scale co-design.

SystemC TLM and QEMU virtual platforms

The industry route to running real software pre-silicon: QEMU or instruction-set simulators for the CPU, SystemC TLM-2.0 models for the accelerators and peripherals. Used for driver and firmware bring-up rather than architectural studies.

06

Side by Side

ToolEngineMain languageInputKey outputs
VidurDiscrete-eventPythonRequest trace, model, deployment configTTFT, TPOT, E2E, MFU, KV use; best config
LLMServingSimIteration-level + HW simsPython / C++Requests, NPU and network configServing latency and throughput on new hardware
SplitwiseSimDiscrete-eventPythonProduction-style traces, cluster specLatency and throughput, cost and power by pool
LLMCompassAnalytical + mapperPythonHardware block description, operator shapesLatency per operator, area, cost
ASTRA-simDiscrete-eventC++Chakra traces, system and network configIteration time, communication breakdown
gem5Discrete-event (ticks)C++ with Python configBinaries or tracesCycles, cache and memory statistics
Disaggregated_Inference_SimDiscrete-event (SimPy)PythonPoisson or replayed requests, hardware presetsLatency percentiles, utilisation, MFU/MBU, stage breakdown, hot-spot, Perfetto trace

Note how often Python appears as the configuration and orchestration layer even where the core is C++. That is the shape of a modern in-house simulator: Python for workloads, configuration and analysis; a compiled core when speed demands it.

07

Workloads and Traces

A simulator is only as good as its inputs. For LLM serving the inputs are request streams; for hardware they are operator graphs and execution traces.

08

Build or Reuse?

Reuse when…

  • The hardware is a GPU or close to one (use Vidur-style profiling).
  • The question is about collectives on a known topology (use ASTRA-sim).
  • You need a full-system CPU and memory model (use gem5 or SST).
  • You need answers this week.

Build in-house when…

  • The compute primitive is new (optical, analogue, in-memory, FHE-specific).
  • The simulator must sit inside your own compiler and runtime toolchain.
  • It doubles as the pre-tape-out verification reference.
  • You must defend every assumption to customers and investors.
The pragmatic middle

Build the core model in-house, borrow the interfaces: accept Chakra or ONNX graphs as input, emit Chrome traces and standard serving metrics, and validate against Vidur or MLPerf numbers on a GPU baseline. Your simulator is then credible on known hardware before it predicts unknown hardware.

09

What to Take Away

Next

Deck 05 applies everything so far: a disaggregated prefill/decode server, simulated live in the browser.