A guided tour of the tools: serving-level simulators (Vidur, LLMServingSim, SplitwiseSim), hardware-level models (LLMCompass, Timeloop, SCALE-Sim), system and network simulators (ASTRA-sim, gem5, SST, SystemC), and analytical calculators — sorted by the question each one answers.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.
LLM-related simulators fall into a few bands. Each band takes a different input and answers a different question. The usual mistake is reaching for a tool from the wrong band.
These simulate a whole serving deployment (request arrivals, a scheduler, batching, KV-cache management, many replicas) and report the latency and throughput users experience. They are discrete-event simulators whose events are batch steps and requests, not cycles.
The reference serving simulator. Profiles real GPU operators once, trains per-operator runtime predictors, then simulates vLLM-, Orca- and Sarathi-style schedulers at scale. Vidur-Search explores deployment configurations (parallelism, batch size, scheduler) for a target workload at a fraction of the cost of trying them on GPUs.
A hardware/software co-simulation infrastructure for LLM serving at scale. It combines an iteration-level serving model with accelerator and network simulators (it builds on ASTRA-sim) so that new NPU or memory designs can be judged on serving metrics rather than kernel speed alone.
The discrete-event simulator released alongside the Splitwise paper on phase splitting. It models prefill and decode machine pools, KV transfer and cluster-level scheduling, driven by production-derived traces. It is the closest public relative of the simulator built in this series.
DistServe uses a simulator inside its placement algorithm: it searches over prefill and decode parallelism and instance counts, estimating goodput for each candidate by simulation, then deploys the best one.
Research in this band moves quickly; new serving simulators appear every few months, often as artifacts of systems papers. Check a paper's artifact link before writing your own.
Vidur's structure is the template for a calibrated serving simulator, and the same separation of concerns from deck 02 at production scale.
| Tool | From | What it models | Use it to… |
|---|---|---|---|
| LLMCompass | Princeton, ISCA 2024 | Block-level hardware description (cores, systolic arrays, vector units, SRAM, HBM, links) with an automatic operator mapper, plus area and cost models | Compare hypothetical LLM accelerator designs on performance and cost |
| Timeloop + Accelergy | MIT / NVIDIA | Loop-nest mappings of tensor operations onto a memory hierarchy; energy per access | Explore dataflows and buffer sizes for one operator at a time |
| SCALE-Sim | Georgia Tech / Arm | Cycle-level systolic-array GEMM and convolution, with memory traffic | Size a systolic array, compare weight/output/input-stationary dataflows |
| Accel-Sim / GPGPU-Sim | UBC, Purdue and others | Cycle-level NVIDIA-like GPU, trace- or execution-driven | Study GPU micro-architecture changes on real kernels |
| LLM-Viewer, GenZ | Academic, 2024 | Analytical roofline per layer and phase | Quick "what-if" on memory, bandwidth and FLOPs for many models |
| Calculon | NVIDIA, SC 2023 | Analytical model of large-model training across parallelism strategies | Co-design training systems; a good example of a pure analytical model at scale |
None of these knows about optical matrix units, ADC/DAC conversion or FHE polynomial arithmetic. That is why organisations building novel architectures write in-house simulators, often reusing the structure of these tools while replacing the hardware description.
Distributed-ML system simulator from Georgia Tech with Meta and Intel. Three layers: a workload layer (now fed by MLCommons Chakra execution traces), a system layer (collective algorithms and scheduling), and a pluggable network layer (analytical, Garnet or ns-3). The tool for "what does this all-reduce cost on that topology?"
The long-running open-source computer-architecture simulator: CPU models from functional to out-of-order, the Ruby and Garnet memory and network models, full-system Linux boot. Slow but deep; the standard for memory-hierarchy research.
Sandia's parallel, component-based framework. Components such as processors, memory (with the DRAMSim/Ramulator family of DRAM models) and networks are wired up from Python, and the core runs them in parallel with MPI. Built for HPC-scale co-design.
The industry route to running real software pre-silicon: QEMU or instruction-set simulators for the CPU, SystemC TLM-2.0 models for the accelerators and peripherals. Used for driver and firmware bring-up rather than architectural studies.
| Tool | Engine | Main language | Input | Key outputs |
|---|---|---|---|---|
| Vidur | Discrete-event | Python | Request trace, model, deployment config | TTFT, TPOT, E2E, MFU, KV use; best config |
| LLMServingSim | Iteration-level + HW sims | Python / C++ | Requests, NPU and network config | Serving latency and throughput on new hardware |
| SplitwiseSim | Discrete-event | Python | Production-style traces, cluster spec | Latency and throughput, cost and power by pool |
| LLMCompass | Analytical + mapper | Python | Hardware block description, operator shapes | Latency per operator, area, cost |
| ASTRA-sim | Discrete-event | C++ | Chakra traces, system and network config | Iteration time, communication breakdown |
| gem5 | Discrete-event (ticks) | C++ with Python config | Binaries or traces | Cycles, cache and memory statistics |
| Disaggregated_Inference_Sim | Discrete-event (SimPy) | Python | Poisson or replayed requests, hardware presets | Latency percentiles, utilisation, MFU/MBU, stage breakdown, hot-spot, Perfetto trace |
Note how often Python appears as the configuration and orchestration layer even where the core is C++. That is the shape of a modern in-house simulator: Python for workloads, configuration and analysis; a compiled core when speed demands it.
A simulator is only as good as its inputs. For LLM serving the inputs are request streams; for hardware they are operator graphs and execution traces.
Build the core model in-house, borrow the interfaces: accept Chakra or ONNX graphs as input, emit Chrome traces and standard serving metrics, and validate against Vidur or MLPerf numbers on a GPU baseline. Your simulator is then credible on known hardware before it predicts unknown hardware.
Deck 05 applies everything so far: a disaggregated prefill/decode server, simulated live in the browser.