Disaggregated_Inference_Sim extended with heterogeneous pools, FFT-mixing models and an optical transform device: TTFT, TPOT, goodput, joules per token and PPA against all-GPU baselines, the break-even point, the KV link options, and compressing the KV hand-off in transit, with the simulator live in the browser.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators, FHE Accelerator Simulators and the Simulation Engineering Toolkit for concepts those series already explain.
Deck 02 found by counting FLOPs that a Fourier-optical engine has little to do in LLM inference: 0.20% of a Hyena-2 model's prefill FLOPs at 2,048 tokens, and a large share only when the weights themselves are block-circulant, which is speculative at this scale. This deck runs the full serving system: a disaggregated cluster in Disaggregated_Inference_Sim, with an optical transform device in the prefill pool, against all-GPU baselines on the same request stream.
examples/results.py.Four extensions, each with every existing number unchanged (the 36 earlier tests pass, and sections 1–9 of results.md regenerate identically). The Python package and its JavaScript port (web/sim_engine.js, which this deck runs) agree bit for bit on every new configuration; the Rust port carries heterogeneous pools only.
| Extension | What it models | Python | JavaScript | Rust |
|---|---|---|---|---|
| Heterogeneous pools | A device and device count per pool: --prefill-device, --decode-device (the Splitwise idea, arXiv:2311.18677) | yes | bit-exact | bit-exact (Rust_DES_Kernel) |
| FFT-mixing models | Hyena-2, a 1:3 hybrid, block-circulant weights, direct or distilled decode; the op ledger equals deck 02's analysis to the FLOP | yes | bit-exact | rejected by name |
| Optical transform device | optical-fft: a transform engine co-packaged with an H100-class digital part (slide 03) | yes | bit-exact | rejected by name |
| KV hand-off compression | fp8, fp4 with block scales, frequency-domain keep-half; in the link or on the prefill GPU (slides 12–13) | yes | bit-exact | rejected by name |
The simulator's older optical device is an optical MAC: a “what if compute were nearly free” probe that speeds up every matmul. The new optical-fft is a transform engine: it takes only FFT and Fourier-plane work and leaves every matmul to its digital part. For a standard transformer it has nothing to do (deck 02, “A Standard Decoder Has No FFT”).
Per prefill step (a batch of prompts), following the phase-A design. All integer counts, so Python, JavaScript and any later port agree exactly:
Source: hardware.py (TransformEngine, CostModel.optical_terms). Every coefficient on this slide is illustrative.
The generally useful part of this work, and the part ported to Rust. Prefill is compute-bound and decode memory-bound (InfSim 03), so each pool can use the part that suits it:
| Pools | TTFT p99 | TPOT p99 | SLO met | J / token | tok/s per $1000 |
|---|---|---|---|---|---|
| H100 prefill + H100 decode (--device h100) | 353.6 ms | 8.8 ms | 100.0% | 0.326 | 2,850 |
| H100 prefill + A100 decode | 353.6 ms | 16.8 ms | 100.0% | 0.299 | 2,714 |
| A100 prefill + H100 decode | 47.1 s | 7.7 ms | 0.0% | 0.448 | 1,952 |
| A100 prefill + A100 decode | 47.1 s | 13.8 ms | 0.0% | 0.401 | 1,894 |
| 2x A100 prefill + H100 decode | 1,093.5 ms | 8.9 ms | 96.9% | 0.412 | 1,859 |
Source: results.md, section “10. Heterogeneous”
--prefill-device h100 --decode-device h100 explicitly reproduces the --device h100 run bit-identically (every request's timestamps and every summary number): True.Two models from results.md section 11 (the others are there too): Hyena-2, where transforms are 0.2% of prefill, and the most transform-heavy variant, block-circulant weights with the LM head on the last token only (84.5%).
Hyena-2
| Configuration | TTFT p99 | TPOT p99 | SLO met | J / token | Prefill optical-bound |
|---|---|---|---|---|---|
| Colocated x2, H100 | 252.2 ms | 26.7 ms | 92.1% | 0.389 | 0.0% |
| 1P1D H100 | 374.9 ms | 114.4 ms | 0.0% | 0.371 | 0.0% |
| 1P1D H100, GPU FFT at 1/16 | 406.1 ms | 114.4 ms | 0.0% | 0.377 | 0.0% |
| optical-fft prefill (defaults) + H100 decode | 172.0 s | 12.5 ms | 0.0% | 0.660 | 100.0% |
| optical-fft prefill (optimistic) + H100 decode | 372.8 ms | 114.4 ms | 0.0% | 0.383 | 0.0% |
Source: results.md, section “11. All-GPU”
Hyena-2 + circulant 256, last-token head
| Configuration | TTFT p99 | TPOT p99 | SLO met | J / token | Prefill optical-bound |
|---|---|---|---|---|---|
| Colocated x2, H100 | 6.7 ms | 3.4 ms | 100.0% | 0.208 | 0.0% |
| 1P1D H100 | 5.4 ms | 13.5 ms | 100.0% | 0.185 | 0.0% |
| 1P1D H100, GPU FFT at 1/16 | 40.0 ms | 13.5 ms | 100.0% | 0.210 | 0.0% |
| optical-fft prefill (defaults) + H100 decode | 322.1 s | 3.2 ms | 0.0% | 0.568 | 100.0% |
| optical-fft prefill (optimistic) + H100 decode | 90.5 ms | 13.5 ms | 100.0% | 0.194 | 100.0% |
Source: results.md, section “11. All-GPU”
The simulator itself, inlined from web/sim_engine.js (bit-exact with the Python package, tested), running the same 800 requests results.md used. The defaults reproduce section 11's circulant last-token row with the optimistic engine. Every control is a real simulator input; the coefficients are illustrative.
One 2,048-token prefill step, closed form. The transform share sets what the engine could take; the mask and the passes set what it costs:
| Variant | Optical share | H100 | H100 FFT at 1/16 | optical-fft defaults | optical-fft optimistic |
|---|---|---|---|---|---|
| Hyena-2 | 0.2% | 63.21 ms | 65.07 ms | 358.4 ms | 63.08 ms |
| Hybrid 1:3 | 0.2% | 62.17 ms | 63.56 ms | 283.8 ms | 62.07 ms |
| circulant 256, all-token head | 14.0% | 5.24 ms | 15.18 ms | 556.7 ms | 18.78 ms |
| circulant 1024, last-token head | 92.9% | 2.11 ms | 7.75 ms | 547.6 ms | 18.48 ms |
| circulant 256, last-token head | 84.5% | 2.12 ms | 11.23 ms | 553.5 ms | 18.78 ms |
| circulant 64, last-token head | 77.9% | 2.76 ms | 27.95 ms | 575.8 ms | 19.93 ms |
Source: results.md, section “12. Sweeps”
Mask rate (circulant 256, last-token head, ENOB 11, overlapped):
| Mask rate | Rewrites per step | Mask time | Step time | TTFT p99 (full run) | SLO met |
|---|---|---|---|---|---|
| 30 Hz | 277 | 9,233.33 ms | 9,238.26 ms | 6,721.4 s | 0.0% |
| 1,031 Hz | 277 | 268.67 ms | 273.60 ms | 100.6 s | 0.0% |
| 20,000 Hz | 277 | 13.85 ms | 18.78 ms | 90.5 ms | 100.0% |
| 100,000 Hz | 277 | 2.77 ms | 7.70 ms | 26.7 ms | 100.0% |
| 1,000,000 Hz | 277 | 0.28 ms | 5.21 ms | 16.5 ms | 100.0% |
Source: results.md, section “12. Sweeps”
ENOB (circulant 256, last-token head, 20 kHz mask, overlapped). Below 11 bits the averaging passes multiply the conversions; above it each conversion costs twice as much per extra bit:
| ENOB | Passes | pJ per pair | Step time | Conversion energy |
|---|---|---|---|---|
| 6 | 1,024 | 1.92 | 4,549.84 ms | 8.708 J |
| 7 | 256 | 3.84 | 1,148.22 ms | 4.354 J |
| 8 | 64 | 7.68 | 297.82 ms | 2.177 J |
| 9 | 16 | 15.36 | 85.22 ms | 1.089 J |
| 10 | 4 | 30.72 | 32.07 ms | 0.544 J |
| 11 | 1 | 61.44 | 18.78 ms | 0.272 J |
| 12 | 1 | 122.88 | 18.78 ms | 0.544 J |
| 13 | 1 | 245.76 | 18.78 ms | 1.089 J |
| 14 | 1 | 491.52 | 18.78 ms | 2.177 J |
Source: results.md, section “12. Sweeps”
Lasers and thermal tuning, charged for the whole run (optimistic engine, full runs):
| Lasers + tuning | J / token | Optical static share | Conversion share | TTFT p99 |
|---|---|---|---|---|
| 0 W | 0.184 | 0.0% | 0.6% | 90.5 ms |
| 10 W | 0.189 | 2.7% | 0.6% | 90.5 ms |
| 20 W | 0.194 | 5.3% | 0.5% | 90.5 ms |
| 50 W | 0.210 | 12.4% | 0.5% | 90.5 ms |
| 100 W | 0.236 | 22.0% | 0.4% | 90.5 ms |
| 200 W | 0.287 | 36.1% | 0.4% | 90.5 ms |
Source: results.md, section “12. Sweeps”
For comparison, all-GPU 1P1D: 0.185 J/token with GPU FFTs at the matmul rate, 0.210 at 1/16 (section 13).
Where does an optical prefill pool stop paying? Bisection on full simulator runs (as the sister series' search.py finds the maximum sustainable load), for the most transform-heavy variant and three assumptions about how fast GPUs run FFTs:
| GPU baseline | GPU TTFT p99 | Optical TTFT p99 | Mask rate at equal TTFT p99 |
|---|---|---|---|
| GPU FFT at 1x | 5.4 ms | 90.5 ms | never |
| GPU FFT at 1/4 | 10.4 ms | 90.5 ms | never |
| GPU FFT at 1/16 | 40.0 ms | 90.5 ms | 47,942 Hz |
Source: results.md, section “13. Break-even”
| GPU baseline | GPU J/token | Optical J/token | Static power at equal J/token |
|---|---|---|---|
| GPU FFT at 1x | 0.185 | 0.194 | 2.3 W |
| GPU FFT at 1/4 | 0.190 | 0.195 | 10.5 W |
| GPU FFT at 1/16 | 0.210 | 0.196 | 46.9 W |
Source: results.md, section “13. Break-even”
The PPA view reuses brief 03's method rather than inventing new coefficients: published GPU die areas, Murphy yield at D0 = 0.1/cm², $10,000 per 300 mm wafer, and the speculative optical-engine areas of FHE_Accelerator_Sim's area model (SimEng 13). Silicon only: no HBM, packaging or test.
| Device | Digital die mm² | Photonic die mm² | Total mm² | Silicon $ per good device |
|---|---|---|---|---|
| H100-SXM | 814 | 0 | 814 | $339 |
| A100-SXM | 826 | 0 | 826 | $348 |
| Optical-FFT + H100-class | 817 | 100 | 917 | $357 |
Source: results.md, section “10. Heterogeneous”
Throughput per dollar of silicon, from section 11 (circulant 256, last-token head):
| Configuration | SLO met | tok/J | tok/s per $1000 |
|---|---|---|---|
| Colocated x2, H100 | 100.0% | 4.80 | 2,928 |
| 1P1D H100 | 100.0% | 5.41 | 2,846 |
| 1P1D H100, GPU FFT at 1/16 | 100.0% | 4.76 | 2,846 |
| optical-fft prefill (defaults) + H100 decode | 0.0% | 1.76 | 689 |
| optical-fft prefill (optimistic) + H100 decode | 100.0% | 5.15 | 2,772 |
Source: results.md, section “11. All-GPU”
Photonic interconnect is not Fourier optics, but it moves the hand-off between pools. Two levers: the link, and what is handed off (deck 02: Hyena-2's direct-decode cache is 4× the GQA KV cache; a distilled state is a constant).
| Link | Name | GB/s | Latency (us) | pJ/bit |
|---|---|---|---|---|
| nvlink4 | NVLink 4 (one direction) | 450.0 | 5 | 5 |
| ib-ndr | InfiniBand NDR 400G | 50.0 | 10 | 15 |
| cpo-optical | Co-packaged optics (illustrative) | 200.0 | 5 | 3 |
| eth-100g | 100 GbE | 12.5 | 20 | 15 |
| eth-25g | 25 GbE | 3.1 | 20 | 15 |
Source: results.md, section “14. The KV hand-off”
Hyena direct cache (the biggest hand-off):
| Link | KV wait + transfer share of E2E | Link busy | TPOT p99 | SLO met | Link J / request |
|---|---|---|---|---|---|
| nvlink4 | 0.0% | 1.5% | 114.3 ms | 0.0% | 0.043 |
| ib-ndr | 0.2% | 13.8% | 114.4 ms | 0.0% | 0.129 |
| cpo-optical | 0.1% | 3.5% | 114.2 ms | 0.0% | 0.026 |
| eth-100g | 1.5% | 55.1% | 115.2 ms | 0.0% | 0.129 |
| eth-25g | 97.2% | 98.4% | 1,338 ms | 0.0% | 0.129 |
Source: results.md, section “14. The KV hand-off”
cpo-optical is an illustrative co-packaged-optics link (round numbers; optical I/O chiplets reach multi-Tb/s, Wade et al., IEEE Micro 2020).A different place for optics: computing in the link, on data that must move anyway. Published 2026 compute-in-transit prototypes put operations in an optical transceiver's data path. The number that decides what such a stage can do is its compute budget per byte at line rate: about 1–2 operations per byte. This simulator uses 1.6 (illustrative).
The mapping of compute-in-transit to disaggregated inference is this series' speculation, attributed to no one. Tolerable KV bit widths come from independent work: KIVI keeps quality at 2 bits (arXiv:2402.02750), KVQuant at 3 bits with under 0.1 perplexity loss (arXiv:2401.18079). The simulator does not model accuracy.
The same compression done in two places: in the link (no GPU time, +1 pJ/bit, +1 µs, limited by the budget) or on the prefill GPU (an elementwise pass: read the KV, write the compressed copy). TTFT does not include the hand-off (the first token leaves the prefill pool), so the effect shows in the hand-off latency (first token to KV on the decode pool), in TPOT and in the SLO. The transport-bound cases:
| Hand-off | Link | Load | Best preset in transit | Hand-off p99 uncompressed | in transit | at GPU |
|---|---|---|---|---|---|---|
| Hyena direct cache | eth-25g | 8 req/s | fp8 | 170.3 s | 34.4 s | 34.4 s |
| GQA KV cache | eth-25g | 10 req/s | fp8 | 750.0 ms | 160.5 ms | 159.1 ms |
| GQA KV cache | eth-25g | 12 req/s | fp8 | 2,959.0 ms | 163.7 ms | 163.7 ms |
| GQA KV cache | eth-25g | 14 req/s | fp8 | 9,657.8 ms | 166.6 ms | 166.6 ms |
Source: results.md, section “15. Compute in the transport”
| Hand-off | Load | SLO uncompressed | SLO in transit | SLO at GPU |
|---|---|---|---|---|
| Hyena direct cache | 8 req/s | 0.0% | 0.0% | 0.0% |
| GQA KV cache | 10 req/s | 100.0% | 100.0% | 100.0% |
| GQA KV cache | 12 req/s | 98.3% | 100.0% | 100.0% |
| GQA KV cache | 14 req/s | 41.0% | 100.0% | 100.0% |
Source: results.md, section “15. Compute in the transport”
Discussion only: not simulated, and all speculative. Other functions a compute-in-transit stage on the hand-off path could take, roughly from most to least promising:
| Function | Why it might fit | What would limit it |
|---|---|---|
| KV resharding in flight | Prefill and decode pools often use different tensor-parallel degrees, so the KV layout must be permuted or regathered: data movement, almost no arithmetic | The stage sees one link's stream; a regather across many links needs coordination |
| Confidential KV transport | KV crossing a network between pools or sites may need protection; post-quantum key exchange (ML-KEM) is built on NTTs, and FHE on NTTs too (FHE Accelerator Simulators) | Key management, and an FHE-encrypted KV cache multiplies its size |
| FEC and DSP offload | Forward error correction on the hand-off link; lower link latency and energy show up in TPOT tails | Already done in the link's DSP; the gain is in energy, not semantics |
| In-transit collectives | Tensor-parallel all-reduce partial sums with in-transit add or MAC inside the decode pool | Affects TPOT, not the hand-off; the simulator does not model all-reduce at all |
| A pooled KV tier over a photonic fabric | Decompress on fetch from a shared memory appliance (a photonic-CXL KV appliance, arXiv:2607.27187) | A different system design from point-to-point disaggregation |
Back to the series hub, or the LLM Inference Simulators, where the simulator is built.