The series applied to a real job: a photonic-computing architecture group's simulation and frameworks role, mapped requirement by requirement onto the decks, with the domain background (photonics, FHE) and a set of practice questions.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Local LLM Hosting and Key Publications decks, or the FHE Accelerator Simulators series.
This deck applies the series to a representative role: a simulation-and-frameworks software engineer in the architecture group of a photonic-computing company, as publicly advertised in 2026. Everything here comes from that public job description and general domain knowledge, not from inside knowledge of any company, and no company is named.
The last column points to the Simulation Engineering Toolkit, a companion series that fills the remaining gaps with decks and tested code.
| Responsibility (from the specification) | What it means in practice | Deck | Toolkit deck and code |
|---|---|---|---|
| Hands-on development of the in-house simulation environment of the photonics system | Owning DES engine code, component models, cost models, configuration | 02, 05, 08 | 01, 02, 03; Rust_DES_Kernel, SystemC_Accelerator_Model |
| Build methods of extracting metrics: latency, hot-spots, utilisation | Probes, stage breakdowns, attribution, traces, reports | 06 | 11 (profiling the simulators themselves); 12 (every measurement tool and method: overhead, accuracy, when to use it) |
| Work with other tools developers so real applications run on the simulator | Compiler and runtime interfaces; operator coverage | 09 | 10; Torch_Sim_Frontend |
| Integrate the simulator with PyTorch, ONNX Runtime and HEIR | Graph capture, backends, execution providers, MLIR lowering; a HEIR front end for FHE programs | 09; FHE 03 | 10: four working front ends; Torch_Sim_Frontend |
| Documenting and writing code to help pre-tape-out verification | Golden models, workload-derived tests, correlation with RTL | 01, 06 | 05, 06; RTL_CoSim_NTT |
| Mentoring less experienced engineers | Code review, design patterns, testing discipline | 02 (as teaching material) | 06 (as teaching material) |
| Helping in system bring-up | The simulator as the first target for the software stack | 01, 09 | 03 (TLM models as virtual platforms) |
| Defining and requirements capture for mission-mode software | Turning architecture behaviour into software requirements | 06 (specs) | 09: EARS, mission mode, traceability |
| Defining, tracking and reporting engineering metrics | Correlation error, coverage, simulator speed, test health | 06, 08 | 08, 11: Jira, flow metrics, regression gates |
| On-time delivery against agreed metrics and sign-off | Milestones tied to measurable criteria | 06 | 08, 09: evidence for sign-off |
| Requirement | Be ready to demonstrate | Toolkit evidence |
|---|---|---|
| 10+ years of Python or C/C++ | Idiomatic, tested, readable code; performance awareness (profiling, data structures); the C/C++–Python boundary | 02, 03; SystemC_Accelerator_Model, Rust_DES_Kernel |
| Python event simulators, e.g. SimPy | Processes, Resource/Store/Container, interrupts, determinism, validation against queueing theory; the companion simulator is a worked example | 01; Rust_DES_Kernel |
| Strong SoC and memory architecture knowledge | Memory hierarchies, HBM/DDR behaviour, DMA, interconnects and NoCs, arbitration, back-pressure, roofline reasoning | 04, 03; Memory_System_Sim |
| Various testing frameworks | pytest (fixtures, parametrisation), property-based testing, golden tests; cocotb or UVM on the hardware side | 06, 05; RTL_CoSim_NTT |
| Specifications, test plans and reports | The skeletons on deck 06, slide 11; examples of documents you have written | 09; the spec, test plan and matrix in Torch_Sim_Frontend |
| Communicating with expert and non-expert audiences | A one-paragraph answer first, evidence after; explaining a hot-spot to a manager and to an architect | Every deck's takeaways slide |
For each row, have one concrete story (situation, what you did, measurable result) and, where possible, code you can show. A small, well-tested simulator with honest benchmarks is unusually strong evidence for this role.
Photonic computing uses light for part of the computation, typically the linear algebra or transforms, and electronics for the rest. The appeal is high throughput and potentially low energy for those operations; the challenges are at the boundaries.
Whether a photonic system beats a GPU for a workload depends on all of the above together, which is why an architecture group needs a full-system simulator before committing to silicon. Read the company's own published material for its specific approach; this slide is general background.
An illustrative hybrid electro-optical accelerator, drawn to show what a DES model would contain. It is not any company's actual architecture.
| Question the architects will ask | What the simulator needs |
|---|---|
| Is the optical core kept busy? | Utilisation and stall attribution (memory, DMA, conversion) |
| How many converters, at what rate and precision? | Converter models in the cost path; precision effects in a functional model |
| How large should on-chip SRAM be? | Tiling and reuse models; traffic counts per level; and its area, since SRAM is bought in mm² (Toolkit 13) |
| Is the design worth its silicon? Which trade-off between speed, power and area should we make? | Performance, power and area reported together, an area and yield model, and Pareto fronts over all three (Toolkit 13) |
| What latency and throughput for an LLM layer, or for FHE bootstrapping? | Workloads lowered from PyTorch, ONNX and HEIR |
| Where is the hot-spot as the design changes? | Stage breakdowns and what-if sensitivity sweeps |
| What does a token or a bootstrapping operation cost in joules, and how much of it is lasers, tuning and converters? | Static and per-conversion energy models; energy breakdowns (deck 07) |
The job's second target market is fully homomorphic encryption. The sister series FHE Accelerator Simulators applies the same methods (SimPy, hot-spots, power under a TDP, a bit-exact JavaScript port) to an accelerator running CKKS bootstrapping, with recorded OpenFHE traces and HEIR-compiled programs. Its results give concrete, checkable material for an interview.
| Finding (ARK-class design unless stated) | Number | Where |
|---|---|---|
| A baseline bootstrap is memory-bound: evaluation keys are the largest share of HBM traffic | 13.94 ms; keys 6.74 of 12.44 GB | FHE 05 |
| SlotToCoeff-first ordering (OpenFHE's sparse-slot path) | 13.94 → 11.51 ms; keys 6.74 → 5.73 GB | FHE 02, 05 |
| All algorithmic techniques together (Min-KS, seeded keys, on-the-fly plaintexts, SlotToCoeff first) turn it MAC-bound | 5.42 ms; HBM 12.44 → 0.72 GB | FHE 05 |
| Replaying a recorded OpenFHE bootstrap trace predicts the measured CPU time (scheme model alone: −30%) | within 11% | FHE 03 |
| A HEIR-compiled CNN (LoLa) through the HEIR front end reproduces HEIR's own OpenFHE output | 55 rotations, 41 rotation keys, 2 relinearisations, exactly | FHE 03 |
Every rotation and relinearisation needs a key switch with its own evaluation key, tens to hundreds of MB at N = 216, so streaming keys from HBM dominates. Fixes: reuse keys across rotations (Min-KS, hoisting), generate half of each key on chip from a seed, generate plaintexts on the fly, reorder the bootstrap (SlotToCoeff first), and size the scratchpad so keys stay on chip. In the sister series these move the bottleneck from HBM to the MAC lanes. (FHE 01, 05.)
Check operation counts against analytical formulas per bootstrap stage; record the real kernel stream from an instrumented library (OpenFHE) and compare structure; replay that trace on a CPU-like configuration and compare with measured time; run compiler-generated programs (HEIR) and require the same rotations and keys as the compiler's own library output; keep a bit-exact twin for differential tests. (FHE 03; deck 06.)
Only when the design is NTT-bound: on a memory-bound configuration a faster transform moves nothing. Exact modular NTTs must be mapped onto analogue transforms by splitting coefficients into digits, which multiplies conversions, and every conversion costs energy that grows with resolution. Answer with the bound attribution first, then the precision and conversion budget. (FHE 04; deck 07.)
HEIR lowers a secret-annotated program to a CKKS-dialect program with explicit rotations, relinearisations, rescales and bootstraps; a front end reads that and emits the simulator's operation trace. Cost feedback into HEIR's choices (for example where to place bootstraps) is the natural next step. (FHE 03; deck 09.)
Answer aloud before opening each outline. About 200 more, with worked answers and tested coding challenges: Interview_Simulation.
DMA and compute as processes; the FIFO as simpy.Store(capacity=k); memory channels as a Resource. yield fifo.put(tile) blocks when full, which is the back-pressure. Measure DMA stall time on put and compute stall on get; sweep k to find where throughput saturates. (Deck 02, slide 07.)
Different questions (what to build versus did we build it right) at different times. The model is the executable spec and golden reference (scoreboards via DPI-C or cocotb); RTL cycle counts calibrate the model; workload traces become directed tests; co-simulation mixes RTL islands into the system model. (Deck 01, slides 06–08.)
Profile. Reduce events (coarser granularity, macro-steps). Make events cheaper (incremental state, lazy bookkeeping, hoisting). Cut probe cost. Parallelise across runs. Do fewer runs (analytic bounds, bisection, Bayesian optimisation, common random numbers). Only then a compiled core or PDES. Prove each step exact with differential tests. (Deck 08.)
Same throughput and utilisation; different latency distributions (10/20 ms against 20/20 ms for two 10 ms transfers). It matters for tail metrics. Choose from the hardware's actual arbitration, document it, and test it. (Deck 02, slide 08.)
Batched resources look 100% busy at low efficiency. Use a stage breakdown that sums to end-to-end latency, take the largest waiting stage, map it to its owning resource, and confirm with a what-if sensitivity run. (Deck 06, slide 04.)
Each step reads every weight and all the KV cache to produce one token per sequence: arithmetic intensity is about the batch size, far below the ridge point. Faster matmul alone barely helps decode; the design must address memory bandwidth, capacity and data movement. Prefill is compute-bound and benefits directly. (Deck 03.)
Start with a TorchDispatchMode trace on the meta device (cheap, any model); then torch.export graphs for the compiler path; a torch.compile backend for one-line user adoption; a PrivateUse1 device for functional bring-up. Track operator coverage as a metric. (Deck 09.)
The number-theoretic transform is the FFT over a finite field. FHE ciphertexts are polynomials of degree 213–217 in RNS form, and polynomial multiplication, key switching and bootstrapping all run through NTTs, so NTT throughput, together with the memory traffic of ciphertexts and keys, dominates FHE cost. (Deck 09.)
HEIR is an MLIR-based FHE compiler: secret-annotated programs are lowered to scheme dialects (CKKS, BGV, CGGI) and then to polynomial and modular arithmetic, then to backends. A hardware team lowers into its own dialect and emits an operation stream (NTTs, modular arithmetic, key switches) for the simulator, using a CPU backend such as OpenFHE as the functional oracle. (Deck 09.)
Price the counts the simulator already has: static power × time, energy per FLOP or conversion, per byte of memory traffic, per bit on links. Add DVFS and power caps as a third roof on step time. Report average and peak power, joules per token and the static/compute/memory/link split. Calibrate coefficients by regression on measured power (DCGM) or from RTL power analysis driven by the simulator's workload traces. Use it to find free savings (memory-bound phases), the energy-optimal cap, and how load affects energy per token through static power. (Deck 07.)
Configuration as data; seeds per component; outputs stamped with config hash and git SHA; deterministic tie-breaking; warm-up deletion; replications with confidence intervals; common random numbers for comparisons; CI that re-runs golden configurations. (Decks 02 and 06.)
A sketch to discuss, not a prescription: the right plan depends on where the in-house simulator is today.
| Phase | Goals | Visible output |
|---|---|---|
| Days 1–30: learn and measure | Understand the architecture and the simulator; run it on the reference workloads; profile it; read the test suite | A short written baseline: speed, coverage, test health, known gaps |
| Days 31–60: strengthen the core | Add or harden metrics (stage breakdown, hot-spots, Chrome traces); add invariant and analytic tests; wire golden metrics into CI | A metrics report on one real workload; green CI gate |
| Days 61–90: connect and correlate | One framework path end to end (for example a dispatch-mode trace from PyTorch); a correlation plan with the RTL team; an engineering-metrics dashboard | A real model on the simulator; correlation targets agreed; metrics tracked weekly |