LLM Inference Simulators — Presentation 11

Role Primer — Simulation & Frameworks Engineer

The series applied to a real job: a photonic-computing architecture group's simulation and frameworks role, mapped requirement by requirement onto the decks, with the domain background (photonics, FHE) and a set of practice questions.

Job mapping Photonic compute FHE Practice questions First 90 days
Role → Domain → Skills → Questions
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.

01

The Role at a Glance

This deck applies the series to a representative role: a simulation-and-frameworks software engineer in the architecture group of a photonic-computing company, as publicly advertised in 2026. Everything here comes from that public job description and general domain knowledge, not from inside knowledge of any company, and no company is named.

From the specification

  • Group: a newly formed Architecture group, architecting a novel photonic computing system.
  • Target markets: AI and Fully Homomorphic Encryption (FHE).
  • Level: Senior / Principal; permanent, full-time.
  • Location: UK remote, with travel to UK and US hubs.
  • Core: an in-house simulation system and its frameworks, across the software–hardware boundary, "enabling rapid bring-up, deep system insight of a novel compute architecture".

In this series' language

  • A discrete-event system simulator (SimPy is named) for a novel accelerator: decks 01, 02, 08.
  • That extracts latency, hot-spots and utilisation: deck 06.
  • That runs real applications through PyTorch, ONNX Runtime and HEIR: deck 09.
  • AI on photonic hardware: where an optical Fourier engine could and could not help LLM inference, with simulated results: the Fourier Optics for Inference series.
  • That supports pre-tape-out verification and bring-up: deck 01.
  • Delivered with specs, test plans, metrics and CI: deck 06.
  • And, for a photonic product above all, power and energy alongside speed: deck 07.
  • The FHE half of the role: the sister series FHE Accelerator Simulators (slide 07).
02

Responsibilities, Mapped to This Series

The last column points to the Simulation Engineering Toolkit, a companion series that fills the remaining gaps with decks and tested code.

Responsibility (from the specification)What it means in practiceDeckToolkit deck and code
Hands-on development of the in-house simulation environment of the photonics systemOwning DES engine code, component models, cost models, configuration02, 05, 0801, 02, 03; Rust_DES_Kernel, SystemC_Accelerator_Model
Build methods of extracting metrics: latency, hot-spots, utilisationProbes, stage breakdowns, attribution, traces, reports0611 (profiling the simulators themselves); 12 (every measurement tool and method: overhead, accuracy, when to use it)
Work with other tools developers so real applications run on the simulatorCompiler and runtime interfaces; operator coverage0910; Torch_Sim_Frontend
Integrate the simulator with PyTorch, ONNX Runtime and HEIRGraph capture, backends, execution providers, MLIR lowering; a HEIR front end for FHE programs09; FHE 0310: four working front ends; Torch_Sim_Frontend
Documenting and writing code to help pre-tape-out verificationGolden models, workload-derived tests, correlation with RTL01, 0605, 06; RTL_CoSim_NTT
Mentoring less experienced engineersCode review, design patterns, testing discipline02 (as teaching material)06 (as teaching material)
Helping in system bring-upThe simulator as the first target for the software stack01, 0903 (TLM models as virtual platforms)
Defining and requirements capture for mission-mode softwareTurning architecture behaviour into software requirements06 (specs)09: EARS, mission mode, traceability
Defining, tracking and reporting engineering metricsCorrelation error, coverage, simulator speed, test health06, 0808, 11: Jira, flow metrics, regression gates
On-time delivery against agreed metrics and sign-offMilestones tied to measurable criteria0608, 09: evidence for sign-off
03

Required Skills: What to Be Ready to Show

RequirementBe ready to demonstrateToolkit evidence
10+ years of Python or C/C++Idiomatic, tested, readable code; performance awareness (profiling, data structures); the C/C++–Python boundary02, 03; SystemC_Accelerator_Model, Rust_DES_Kernel
Python event simulators, e.g. SimPyProcesses, Resource/Store/Container, interrupts, determinism, validation against queueing theory; the companion simulator is a worked example01; Rust_DES_Kernel
Strong SoC and memory architecture knowledgeMemory hierarchies, HBM/DDR behaviour, DMA, interconnects and NoCs, arbitration, back-pressure, roofline reasoning04, 03; Memory_System_Sim
Various testing frameworkspytest (fixtures, parametrisation), property-based testing, golden tests; cocotb or UVM on the hardware side06, 05; RTL_CoSim_NTT
Specifications, test plans and reportsThe skeletons on deck 06, slide 11; examples of documents you have written09; the spec, test plan and matrix in Torch_Sim_Frontend
Communicating with expert and non-expert audiencesA one-paragraph answer first, evidence after; explaining a hot-spot to a manager and to an architectEvery deck's takeaways slide
Prepare evidence, not adjectives

For each row, have one concrete story (situation, what you did, measurable result) and, where possible, code you can show. A small, well-tested simulator with honest benchmarks is unusually strong evidence for this role.

04

Preferred Skills

Domain

  • AI/ML, HPC or cryptography: LLM inference physics (deck 03), FHE basics (deck 09); real models into an accelerator model (Toolkit 10).
  • Performance analysis: roofline, queueing, profiling, statistics (decks 03, 06, 08); flame graphs, cachegrind and regression gates on these simulators (Toolkit 11); what each measurement tool costs and when to use it, from profilers to CACTI (Toolkit 12).

Engineering

  • Rust: a faster simulator core behind a Python API (PyO3), see deck 08; built and measured, bit-exact, 71–74× faster than SimPy (re-measured on 2026-10-03, after the cost-model correction found by tracing the real model) (Toolkit 01, 02; Rust_DES_Kernel).
  • Jira: linking milestones to spec sections and metrics; JQL, workflows and flow metrics (Toolkit 08).
  • Jenkins and CI/CD: per-push tests, nightly sweeps, golden-metric gates (deck 06 Jenkinsfile); real pipelines for five repositories (Toolkit 07).
  • Git: branching, review, reproducible runs stamped with the commit SHA.
05

Photonic Computing: A Primer

Photonic computing uses light for part of the computation, typically the linear algebra or transforms, and electronics for the rest. The appeal is high throughput and potentially low energy for those operations; the challenges are at the boundaries.

What optics does well

  • Fourier transforms: a lens produces the Fourier transform of a field at its focal plane, in a single pass of light.
  • Matrix–vector products: interferometer meshes or wavelength-multiplexed weight banks.
  • Massive parallelism across space and wavelength, with propagation delay rather than clocked steps.

Where the difficulty lives

  • Conversion: DACs and modulators (electrical to optical), photodetectors and ADCs (optical to electrical) cost energy, area and time.
  • Precision and noise: analogue optics gives limited effective bits; exact arithmetic (as FHE needs) requires decomposition and correction schemes.
  • Feeding the core: memory bandwidth and data movement must keep pace, or the optical core idles (the same lesson as decks 03 and 09).
  • Calibration and stability over temperature and time.
Why the simulator is central

Whether a photonic system beats a GPU for a workload depends on all of the above together, which is why an architecture group needs a full-system simulator before committing to silicon. Read the company's own published material for its specific approach; this slide is general background.

06

What a Photonic-System Simulator Must Model

An illustrative hybrid electro-optical accelerator, drawn to show what a DES model would contain. It is not any company's actual architecture.

host / PCIeframework runtime HBM / DDRchannels, banks DMA + SRAMtiles, buffers DAC +modulators optical coretransform / MVMprecision, noise detectors+ ADC digital post-processingmod reduce, accumulate, NTT fix-up on-chip interconnect / NoC each box: a SimPy process or resource with a cost model; each arrow: a link with bandwidth, latency and back-pressure
Question the architects will askWhat the simulator needs
Is the optical core kept busy?Utilisation and stall attribution (memory, DMA, conversion)
How many converters, at what rate and precision?Converter models in the cost path; precision effects in a functional model
How large should on-chip SRAM be?Tiling and reuse models; traffic counts per level; and its area, since SRAM is bought in mm² (Toolkit 13)
Is the design worth its silicon? Which trade-off between speed, power and area should we make?Performance, power and area reported together, an area and yield model, and Pareto fronts over all three (Toolkit 13)
What latency and throughput for an LLM layer, or for FHE bootstrapping?Workloads lowered from PyTorch, ONNX and HEIR
Where is the hot-spot as the design changes?Stage breakdowns and what-if sensitivity sweeps
What does a token or a bootstrapping operation cost in joules, and how much of it is lasers, tuning and converters?Static and per-conversion energy models; energy breakdowns (deck 07)
07

The FHE Half: Findings and Questions

The job's second target market is fully homomorphic encryption. The sister series FHE Accelerator Simulators applies the same methods (SimPy, hot-spots, power under a TDP, a bit-exact JavaScript port) to an accelerator running CKKS bootstrapping, with recorded OpenFHE traces and HEIR-compiled programs. Its results give concrete, checkable material for an interview.

Finding (ARK-class design unless stated)NumberWhere
A baseline bootstrap is memory-bound: evaluation keys are the largest share of HBM traffic13.94 ms; keys 6.74 of 12.44 GBFHE 05
SlotToCoeff-first ordering (OpenFHE's sparse-slot path)13.94 → 11.51 ms; keys 6.74 → 5.73 GBFHE 02, 05
All algorithmic techniques together (Min-KS, seeded keys, on-the-fly plaintexts, SlotToCoeff first) turn it MAC-bound5.42 ms; HBM 12.44 → 0.72 GBFHE 05
Replaying a recorded OpenFHE bootstrap trace predicts the measured CPU time (scheme model alone: −30%)within 11%FHE 03
A HEIR-compiled CNN (LoLa) through the HEIR front end reproduces HEIR's own OpenFHE output55 rotations, 41 rotation keys, 2 relinearisations, exactlyFHE 03
Why is CKKS bootstrapping memory-bound on a compute-rich accelerator, and what fixes it?

Every rotation and relinearisation needs a key switch with its own evaluation key, tens to hundreds of MB at N = 216, so streaming keys from HBM dominates. Fixes: reuse keys across rotations (Min-KS, hoisting), generate half of each key on chip from a seed, generate plaintexts on the fly, reorder the bootstrap (SlotToCoeff first), and size the scratchpad so keys stay on chip. In the sister series these move the bottleneck from HBM to the MAC lanes. (FHE 01, 05.)

How would you validate an FHE accelerator model before any hardware exists?

Check operation counts against analytical formulas per bootstrap stage; record the real kernel stream from an instrumented library (OpenFHE) and compare structure; replay that trace on a CPU-like configuration and compare with measured time; run compiler-generated programs (HEIR) and require the same rotations and keys as the compiler's own library output; keep a bit-exact twin for differential tests. (FHE 03; deck 06.)

When does an optical NTT engine help, and when does it not?

Only when the design is NTT-bound: on a memory-bound configuration a faster transform moves nothing. Exact modular NTTs must be mapped onto analogue transforms by splitting coefficients into digits, which multiplies conversions, and every conversion costs energy that grows with resolution. Answer with the bound attribution first, then the precision and conversion budget. (FHE 04; deck 07.)

Where does HEIR meet the simulator?

HEIR lowers a secret-annotated program to a CKKS-dialect program with explicit rotations, relinearisations, rescales and bootstraps; a front end reads that and emits the simulator's operation trace. Cost feedback into HEIR's choices (for example where to place bootstraps) is the natural next step. (FHE 03; deck 09.)

08

Practice Questions: Simulation

Answer aloud before opening each outline. About 200 more, with worked answers and tested coding challenges: Interview_Simulation.

Model a DMA engine feeding a compute unit through a bounded FIFO in SimPy. How do you show back-pressure?

DMA and compute as processes; the FIFO as simpy.Store(capacity=k); memory channels as a Resource. yield fifo.put(tile) blocks when full, which is the back-pressure. Measure DMA stall time on put and compute stall on get; sweep k to find where throughput saturates. (Deck 02, slide 07.)

How do you validate a performance model of hardware that does not exist yet?
How does a software-level simulator relate to RTL simulation?

Different questions (what to build versus did we build it right) at different times. The model is the executable spec and golden reference (scoreboards via DPI-C or cocotb); RTL cycle counts calibrate the model; workload traces become directed tests; co-simulation mixes RTL islands into the system model. (Deck 01, slides 06–08.)

A SimPy simulation is too slow. What do you do, in order?

Profile. Reduce events (coarser granularity, macro-steps). Make events cheaper (incremental state, lazy bookkeeping, hoisting). Cut probe cost. Parallelise across runs. Do fewer runs (analytic bounds, bisection, Bayesian optimisation, common random numbers). Only then a compiled core or PDES. Prove each step exact with differential tests. (Deck 08.)

Two transfers share a link. FCFS or processor sharing, and does it matter?

Same throughput and utilisation; different latency distributions (10/20 ms against 20/20 ms for two 10 ms transfers). It matters for tail metrics. Choose from the hardware's actual arbitration, document it, and test it. (Deck 02, slide 08.)

How do you identify a hot-spot, and why is "highest utilisation" not enough?

Batched resources look 100% busy at low efficiency. Use a stage breakdown that sums to end-to-end latency, take the largest waiting stage, map it to its owning resource, and confirm with a what-if sensitivity run. (Deck 06, slide 04.)

09

Practice Questions: Workloads, Frameworks and Process

Why is LLM decode memory-bound, and what does that imply for a photonic accelerator?

Each step reads every weight and all the KV cache to produce one token per sequence: arithmetic intensity is about the batch size, far below the ridge point. Faster matmul alone barely helps decode; the design must address memory bandwidth, capacity and data movement. Prefill is compute-bound and benefits directly. (Deck 03.)

How would you get real PyTorch models running on the simulator?

Start with a TorchDispatchMode trace on the meta device (cheap, any model); then torch.export graphs for the compiler path; a torch.compile backend for one-line user adoption; a PrivateUse1 device for functional bring-up. Track operator coverage as a metric. (Deck 09.)

What is an NTT, and why does FHE need so many?

The number-theoretic transform is the FFT over a finite field. FHE ciphertexts are polynomials of degree 213–217 in RNS form, and polynomial multiplication, key switching and bootstrapping all run through NTTs, so NTT throughput, together with the memory traffic of ciphertexts and keys, dominates FHE cost. (Deck 09.)

Where does HEIR fit, and how would the simulator consume it?

HEIR is an MLIR-based FHE compiler: secret-annotated programs are lowered to scheme dialects (CKKS, BGV, CGGI) and then to polynomial and modular arithmetic, then to backends. A hardware team lowers into its own dialect and emits an operation stream (NTTs, modular arithmetic, key switches) for the simulator, using a CPU backend such as OpenFHE as the functional oracle. (Deck 09.)

How would you model power in a performance simulator, and what would you do with it?

Price the counts the simulator already has: static power × time, energy per FLOP or conversion, per byte of memory traffic, per bit on links. Add DVFS and power caps as a third roof on step time. Report average and peak power, joules per token and the static/compute/memory/link split. Calibrate coefficients by regression on measured power (DCGM) or from RTL power analysis driven by the simulator's workload traces. Use it to find free savings (memory-bound phases), the energy-optimal cap, and how load affects energy per token through static power. (Deck 07.)

Define engineering metrics that show progress towards sign-off.
How do you make simulation results reproducible and reviewable?

Configuration as data; seeds per component; outputs stamped with config hash and git SHA; deterministic tie-breaking; warm-up deletion; replications with confidence intervals; common random numbers for comparisons; CI that re-runs golden configurations. (Decks 02 and 06.)

10

A Plausible First 90 Days

A sketch to discuss, not a prescription: the right plan depends on where the in-house simulator is today.

PhaseGoalsVisible output
Days 1–30: learn and measureUnderstand the architecture and the simulator; run it on the reference workloads; profile it; read the test suiteA short written baseline: speed, coverage, test health, known gaps
Days 31–60: strengthen the coreAdd or harden metrics (stage breakdown, hot-spots, Chrome traces); add invariant and analytic tests; wire golden metrics into CIA metrics report on one real workload; green CI gate
Days 61–90: connect and correlateOne framework path end to end (for example a dispatch-mode trace from PyTorch); a correlation plan with the RTL team; an engineering-metrics dashboardA real model on the simulator; correlation targets agreed; metrics tracked weekly
11

Questions Worth Asking

About the simulator

  • What abstraction level is it today, and what is it expected to answer by tape-out?
  • SimPy throughout, or a compiled core planned (C++ or Rust)?
  • How is it validated now, and against what?
  • Is there a functional (bit-accurate) model as well as a timing model?

About the toolchain and the team

  • Which framework path comes first: PyTorch, ONNX Runtime or HEIR?
  • How do simulation, compiler, RTL and verification teams hand off to each other?
  • What does sign-off look like for the simulator's predictions?
  • How is the balance between AI and FHE workloads expected to evolve?
  • How are power and energy targets set, and how is the simulator's power model calibrated?
  • Which FHE schemes, parameter sets and workloads (bootstrapping, encrypted inference) matter most to the roadmap?
  • What CI and compute infrastructure exists for sweeps?
12

Checklist