Fourier Optics for Inference — Presentation 03

Optical Prefill Pools: Simulated Results

Disaggregated_Inference_Sim extended with heterogeneous pools, FFT-mixing models and an optical transform device: TTFT, TPOT, goodput, joules per token and PPA against all-GPU baselines, the break-even point, the KV link options, and compressing the KV hand-off in transit, with the simulator live in the browser.

Heterogeneous pools Transform device Break-even PPA Compute in transit Live simulator
Workload → Optical prefill → KV link → GPU decode → Metrics
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators, FHE Accelerator Simulators and the Simulation Engineering Toolkit for concepts those series already explain.

01

The Question, and the Short Answer

Deck 02 found by counting FLOPs that a Fourier-optical engine has little to do in LLM inference: 0.20% of a Hyena-2 model's prefill FLOPs at 2,048 tokens, and a large share only when the weights themselves are block-circulant, which is speculative at this scale. This deck runs the full serving system: a disaggregated cluster in Disaggregated_Inference_Sim, with an optical transform device in the prefill pool, against all-GPU baselines on the same request stream.

What the simulator says

  • Attention and Hyena models: no gain. At the default engine (ENOB 8, an 8-bit DMD mask) the optical prefill pool collapses: Hyena-2's TTFT p99 is 172.0 s against 374.9 ms on GPUs. With an optimistic engine it only ties, as Amdahl predicted.
  • The circulant variant: a narrow window. Even the optimistic engine (ENOB 11, a 20 kHz mask) wins TTFT only if GPUs run FFTs slowly. GPU FFT efficiency below which the optimistic optical pool wins TTFT p99: 0.032 (1/30.9). Against GPUs at 1/16 it needs a mask rewriting at 47,942 Hz (slide 09).
  • Static power and conversions decide the energy verdict, not the FLOPs.

How to read the numbers

  • Every number comes from examples/results.md, sections 10–15, written by examples/results.py.
  • Workload: a Llama-3-8B-shaped model, one device per instance, one prefill and one decode instance, Poisson arrivals at 8 req/s, 800 requests, 2,048-token prompts and 256-token outputs (cv 0.5), SLOs TTFT 1 s and TPOT 25 ms.
  • Optical coefficients are illustrative; the block-circulant LLM is speculative; the hardware mapping is this series' speculation, attributed to no company.
02

What Was Added to the Simulator

Four extensions, each with every existing number unchanged (the 36 earlier tests pass, and sections 1–9 of results.md regenerate identically). The Python package and its JavaScript port (web/sim_engine.js, which this deck runs) agree bit for bit on every new configuration; the Rust port carries heterogeneous pools only.

ExtensionWhat it modelsPythonJavaScriptRust
Heterogeneous poolsA device and device count per pool: --prefill-device, --decode-device (the Splitwise idea, arXiv:2311.18677)yesbit-exactbit-exact (Rust_DES_Kernel)
FFT-mixing modelsHyena-2, a 1:3 hybrid, block-circulant weights, direct or distilled decode; the op ledger equals deck 02's analysis to the FLOPyesbit-exactrejected by name
Optical transform deviceoptical-fft: a transform engine co-packaged with an H100-class digital part (slide 03)yesbit-exactrejected by name
KV hand-off compressionfp8, fp4 with block scales, frequency-domain keep-half; in the link or on the prefill GPU (slides 12–13)yesbit-exactrejected by name
Two different optical parts

The simulator's older optical device is an optical MAC: a “what if compute were nearly free” probe that speeds up every matmul. The new optical-fft is a transform engine: it takes only FFT and Fourier-plane work and leaves every matmul to its digital part. For a standard transformer it has nothing to do (deck 02, “A Standard Decoder Has No FFT”).

03

The Transform Engine Model

Per prefill step (a batch of prompts), following the phase-A design. All integer counts, so Python, JavaScript and any later port agree exactly:

Time

  • Passes k = 4max(0, ENOBreq − ENOB), with ENOBreq = 11 for BF16, 8 for INT8, 7 for FP8: averaging k passes buys ½ log2 k bits (Garg et al.). Intensity detection doubles k. Not FHE's digit planes: they do not help float workloads (deck 01).
  • Conversions = pairs per token × tokens × k, at 1012 samples/s.
  • Mask rewrites = ⌈filter spectra needed / 2 M⌉, at 1,031 Hz (an 8-bit DMD, Miscuglio et al., Optica 2020).
  • Step = digital part (matmuls, on its own roofline, DVFS and TDP) + optical time (or the max, if overlapped).

Energy and cost

  • Conversion energy per DAC+ADC pair = (10 + 20 fJ) × 2ENOB, the Walden rule with FHESim 04's figures of merit (FHESim 04).
  • Lasers and tuning: 20 W per device, charged for the whole run whether or not work arrives.
  • Decode on this device uses the digital part only: a direct cached dot product or a distilled recurrence needs no transform. Relaxed-tiling decode (FFTs at decode) is not simulated; deck 02 covers it analytically.
  • Area: 20 converter channels (0.15 mm² each) on the digital die plus a 100 mm² photonic die: brief 03's speculative figures (slide 10).

Source: hardware.py (TransformEngine, CostModel.optical_terms). Every coefficient on this slide is illustrative.

04

Heterogeneous Pools: H100 Prefill, A100 Decode

The generally useful part of this work, and the part ported to Rust. Prefill is compute-bound and decode memory-bound (InfSim 03), so each pool can use the part that suits it:

PoolsTTFT p99TPOT p99SLO metJ / tokentok/s per $1000
H100 prefill + H100 decode (--device h100)353.6 ms8.8 ms100.0%0.3262,850
H100 prefill + A100 decode353.6 ms16.8 ms100.0%0.2992,714
A100 prefill + H100 decode47.1 s7.7 ms0.0%0.4481,952
A100 prefill + A100 decode47.1 s13.8 ms0.0%0.4011,894
2x A100 prefill + H100 decode1,093.5 ms8.9 ms96.9%0.4121,859

Source: results.md, section “10. Heterogeneous”

05

Optical Prefill Against All-GPU Baselines

Two models from results.md section 11 (the others are there too): Hyena-2, where transforms are 0.2% of prefill, and the most transform-heavy variant, block-circulant weights with the LM head on the last token only (84.5%).

Hyena-2

ConfigurationTTFT p99TPOT p99SLO metJ / tokenPrefill optical-bound
Colocated x2, H100252.2 ms26.7 ms92.1%0.3890.0%
1P1D H100374.9 ms114.4 ms0.0%0.3710.0%
1P1D H100, GPU FFT at 1/16406.1 ms114.4 ms0.0%0.3770.0%
optical-fft prefill (defaults) + H100 decode172.0 s12.5 ms0.0%0.660100.0%
optical-fft prefill (optimistic) + H100 decode372.8 ms114.4 ms0.0%0.3830.0%

Source: results.md, section “11. All-GPU”

Hyena-2 + circulant 256, last-token head

ConfigurationTTFT p99TPOT p99SLO metJ / tokenPrefill optical-bound
Colocated x2, H1006.7 ms3.4 ms100.0%0.2080.0%
1P1D H1005.4 ms13.5 ms100.0%0.1850.0%
1P1D H100, GPU FFT at 1/1640.0 ms13.5 ms100.0%0.2100.0%
optical-fft prefill (defaults) + H100 decode322.1 s3.2 ms0.0%0.568100.0%
optical-fft prefill (optimistic) + H100 decode90.5 ms13.5 ms100.0%0.194100.0%

Source: results.md, section “11. All-GPU”

06

Interactive: The Simulator, Live

The simulator itself, inlined from web/sim_engine.js (bit-exact with the Python package, tested), running the same 800 requests results.md used. The defaults reproduce section 11's circulant last-token row with the optimistic engine. Every control is a real simulator input; the coefficients are illustrative.

11
20 W
TTFT p99
—
TPOT p99
—
SLO met / goodput
—
J / token
—
Prefill optical-bound
—
KV hand-off p99
—
Avg power
—
tok/s per $1000 of silicon
—
Where the average request's time goes
Where the energy goes

07

Where the Time Goes: Mask Rewrites and Passes

One 2,048-token prefill step, closed form. The transform share sets what the engine could take; the mask and the passes set what it costs:

VariantOptical shareH100H100 FFT at 1/16optical-fft defaultsoptical-fft optimistic
Hyena-20.2%63.21 ms65.07 ms358.4 ms63.08 ms
Hybrid 1:30.2%62.17 ms63.56 ms283.8 ms62.07 ms
circulant 256, all-token head14.0%5.24 ms15.18 ms556.7 ms18.78 ms
circulant 1024, last-token head92.9%2.11 ms7.75 ms547.6 ms18.48 ms
circulant 256, last-token head84.5%2.12 ms11.23 ms553.5 ms18.78 ms
circulant 64, last-token head77.9%2.76 ms27.95 ms575.8 ms19.93 ms

Source: results.md, section “12. Sweeps”

Mask rate (circulant 256, last-token head, ENOB 11, overlapped):

Mask rateRewrites per stepMask timeStep timeTTFT p99 (full run)SLO met
30 Hz2779,233.33 ms9,238.26 ms6,721.4 s0.0%
1,031 Hz277268.67 ms273.60 ms100.6 s0.0%
20,000 Hz27713.85 ms18.78 ms90.5 ms100.0%
100,000 Hz2772.77 ms7.70 ms26.7 ms100.0%
1,000,000 Hz2770.28 ms5.21 ms16.5 ms100.0%

Source: results.md, section “12. Sweeps”

08

Static Power and Energy per Token

ENOB (circulant 256, last-token head, 20 kHz mask, overlapped). Below 11 bits the averaging passes multiply the conversions; above it each conversion costs twice as much per extra bit:

ENOBPassespJ per pairStep timeConversion energy
61,0241.924,549.84 ms8.708 J
72563.841,148.22 ms4.354 J
8647.68297.82 ms2.177 J
91615.3685.22 ms1.089 J
10430.7232.07 ms0.544 J
11161.4418.78 ms0.272 J
121122.8818.78 ms0.544 J
131245.7618.78 ms1.089 J
141491.5218.78 ms2.177 J

Source: results.md, section “12. Sweeps”

Lasers and thermal tuning, charged for the whole run (optimistic engine, full runs):

Lasers + tuningJ / tokenOptical static shareConversion shareTTFT p99
0 W0.1840.0%0.6%90.5 ms
10 W0.1892.7%0.6%90.5 ms
20 W0.1945.3%0.5%90.5 ms
50 W0.21012.4%0.5%90.5 ms
100 W0.23622.0%0.4%90.5 ms
200 W0.28736.1%0.4%90.5 ms

Source: results.md, section “12. Sweeps”

For comparison, all-GPU 1P1D: 0.185 J/token with GPU FFTs at the matmul rate, 0.210 at 1/16 (section 13).

09

The Break-Even Point

Where does an optical prefill pool stop paying? Bisection on full simulator runs (as the sister series' search.py finds the maximum sustainable load), for the most transform-heavy variant and three assumptions about how fast GPUs run FFTs:

GPU baselineGPU TTFT p99Optical TTFT p99Mask rate at equal TTFT p99
GPU FFT at 1x5.4 ms90.5 msnever
GPU FFT at 1/410.4 ms90.5 msnever
GPU FFT at 1/1640.0 ms90.5 ms47,942 Hz

Source: results.md, section “13. Break-even”

GPU baselineGPU J/tokenOptical J/tokenStatic power at equal J/token
GPU FFT at 1x0.1850.1942.3 W
GPU FFT at 1/40.1900.19510.5 W
GPU FFT at 1/160.2100.19646.9 W

Source: results.md, section “13. Break-even”

10

Area, Cost and Performance per Dollar

The PPA view reuses brief 03's method rather than inventing new coefficients: published GPU die areas, Murphy yield at D0 = 0.1/cm², $10,000 per 300 mm wafer, and the speculative optical-engine areas of FHE_Accelerator_Sim's area model (SimEng 13). Silicon only: no HBM, packaging or test.

DeviceDigital die mm²Photonic die mm²Total mm²Silicon $ per good device
H100-SXM8140814$339
A100-SXM8260826$348
Optical-FFT + H100-class817100917$357

Source: results.md, section “10. Heterogeneous”

Throughput per dollar of silicon, from section 11 (circulant 256, last-token head):

ConfigurationSLO mettok/Jtok/s per $1000
Colocated x2, H100100.0%4.802,928
1P1D H100100.0%5.412,846
1P1D H100, GPU FFT at 1/16100.0%4.762,846
optical-fft prefill (defaults) + H100 decode0.0%1.76689
optical-fft prefill (optimistic) + H100 decode100.0%5.152,772

Source: results.md, section “11. All-GPU”

11

The KV Hand-off: Link and Payload

Photonic interconnect is not Fourier optics, but it moves the hand-off between pools. Two levers: the link, and what is handed off (deck 02: Hyena-2's direct-decode cache is 4× the GQA KV cache; a distilled state is a constant).

LinkNameGB/sLatency (us)pJ/bit
nvlink4NVLink 4 (one direction)450.055
ib-ndrInfiniBand NDR 400G50.01015
cpo-opticalCo-packaged optics (illustrative)200.053
eth-100g100 GbE12.52015
eth-25g25 GbE3.12015

Source: results.md, section “14. The KV hand-off”

Hyena direct cache (the biggest hand-off):

LinkKV wait + transfer share of E2ELink busyTPOT p99SLO metLink J / request
nvlink40.0%1.5%114.3 ms0.0%0.043
ib-ndr0.2%13.8%114.4 ms0.0%0.129
cpo-optical0.1%3.5%114.2 ms0.0%0.026
eth-100g1.5%55.1%115.2 ms0.0%0.129
eth-25g97.2%98.4%1,338 ms0.0%0.129

Source: results.md, section “14. The KV hand-off”

12

Compute in the Transport, Not the Model

A different place for optics: computing in the link, on data that must move anyway. Published 2026 compute-in-transit prototypes put operations in an optical transceiver's data path. The number that decides what such a stage can do is its compute budget per byte at line rate: about 1–2 operations per byte. This simulator uses 1.6 (illustrative).

What fits 1–2 ops per byte

  • Requantising the KV cache as it leaves prefill: BF16 to FP8 is about one operation per value (scale and round).
  • Frequency-domain compression along the token axis (FreqKV keeps half the DCT components, arXiv:2505.00570), if the transform itself is passive optics.
  • Moving data between layouts (resharding) and protecting it in flight (slide 14).

What does not

  • Matmuls. Prefill runs at thousands of FLOPs per byte of weights read (InfSim 03); a stage at 1.6 ops per byte is three orders of magnitude short.
  • FP4 with block scales, as modelled here: finding each block's maximum and scaling is 2 operations per value, and at a 3.8× compression ratio that is 3.8 operations per line byte. The stage becomes the bottleneck (slide 13).
  • A digital FFT at line rate: 2.5 log2 n operations per value.
Speculation, labelled

The mapping of compute-in-transit to disaggregated inference is this series' speculation, attributed to no one. Tolerable KV bit widths come from independent work: KIVI keeps quality at 2 bits (arXiv:2402.02750), KVQuant at 3 bits with under 0.1 perplexity loss (arXiv:2401.18079). The simulator does not model accuracy.

13

Compressing the Hand-off: In Transit or at the GPU

The same compression done in two places: in the link (no GPU time, +1 pJ/bit, +1 µs, limited by the budget) or on the prefill GPU (an elementwise pass: read the KV, write the compressed copy). TTFT does not include the hand-off (the first token leaves the prefill pool), so the effect shows in the hand-off latency (first token to KV on the decode pool), in TPOT and in the SLO. The transport-bound cases:

Hand-offLinkLoadBest preset in transitHand-off p99 uncompressedin transitat GPU
Hyena direct cacheeth-25g8 req/sfp8170.3 s34.4 s34.4 s
GQA KV cacheeth-25g10 req/sfp8750.0 ms160.5 ms159.1 ms
GQA KV cacheeth-25g12 req/sfp82,959.0 ms163.7 ms163.7 ms
GQA KV cacheeth-25g14 req/sfp89,657.8 ms166.6 ms166.6 ms

Source: results.md, section “15. Compute in the transport”

Hand-offLoadSLO uncompressedSLO in transitSLO at GPU
Hyena direct cache8 req/s0.0%0.0%0.0%
GQA KV cache10 req/s100.0%100.0%100.0%
GQA KV cache12 req/s98.3%100.0%100.0%
GQA KV cache14 req/s41.0%100.0%100.0%

Source: results.md, section “15. Compute in the transport”

14

Beyond Compression: Resharding, Confidential KV, FEC and Collectives

Discussion only: not simulated, and all speculative. Other functions a compute-in-transit stage on the hand-off path could take, roughly from most to least promising:

FunctionWhy it might fitWhat would limit it
KV resharding in flightPrefill and decode pools often use different tensor-parallel degrees, so the KV layout must be permuted or regathered: data movement, almost no arithmeticThe stage sees one link's stream; a regather across many links needs coordination
Confidential KV transportKV crossing a network between pools or sites may need protection; post-quantum key exchange (ML-KEM) is built on NTTs, and FHE on NTTs too (FHE Accelerator Simulators)Key management, and an FHE-encrypted KV cache multiplies its size
FEC and DSP offloadForward error correction on the hand-off link; lower link latency and energy show up in TPOT tailsAlready done in the link's DSP; the gain is in energy, not semantics
In-transit collectivesTensor-parallel all-reduce partial sums with in-transit add or MAC inside the decode poolAffects TPOT, not the hand-off; the simulator does not model all-reduce at all
A pooled KV tier over a photonic fabricDecompress on fetch from a shared memory appliance (a photonic-CXL KV appliance, arXiv:2607.27187)A different system design from point-to-point disaggregation
15

What This Does Not Show

16

What to Take Away

Back to the series hub, or the LLM Inference Simulators, where the simulator is built.