LLM Hub — Fourier Optics for Inference

Fourier Optics for Inference

Could a Fourier-optical engine, which computes a 2-D Fourier transform (and with a mask, a convolution) in one pass of light, accelerate disaggregated LLM inference? Three decks work it out: the physics and its precision and energy bill; where transforms appear in inference workloads, with an Amdahl analysis of prefill FLOP shares from a published script; and a simulator with heterogeneous pools and an optical transform device in the prefill pool, with the simulator live in the browser and an answer: where an optical prefill pool would and would not pay.

4f systemConvolution theoremENOBFNet / HyenaAmdahlDisaggregated serving3 decks live

Presentations in This Series

  1. Fourier Optics for Engineers →
    How a lens computes a Fourier transform and a 4f system a convolution, in one pass of light; coherent and incoherent light, spatial light modulators, detectors and the phase problem, integrated-photonic alternatives; and the precision and energy bill: ENOB, noise, passes, and the converters around the optics.
    live4f system · Convolution theorem · SLMs · Phase problem
  2. Transforms in Inference Workloads →
    Where a Fourier transform could appear in LLM inference, and how much it would matter: FNet, long convolutions (S4, H3, Hyena), causality at prefill and decode, an Amdahl analysis of prefill FLOP shares, structured weights, transform-domain KV compression, mask capacity, and an honest list of what does not map.
    liveFNet · Hyena · Causality · Amdahl
  3. Optical Prefill Pools: Simulated Results →
    Disaggregated_Inference_Sim extended with heterogeneous pools, FFT-mixing models and an optical transform device: TTFT, TPOT, goodput, joules per token and PPA against all-GPU baselines, the break-even point, the KV link options, and compressing the KV hand-off in transit, with the simulator live in the browser.
    liveHeterogeneous pools · Transform device · Break-even · PPA

Analysis and Code

  1. </>
    analysis/ →
    flop_share.py: the FLOP-share (Amdahl) analysis for a Llama-3-8B-shaped model in five variants, checked against Disaggregated_Inference_Sim's cost model to the FLOP; numerical checks of causal FFT convolution, relaxed tiling and the precision rule; conversion energy and mask capacity. It writes results.md, the source of every number in the decks. js/flop_model.js is its port for deck 02, tested to match exactly.
    codePython · NumPy · pytest · JavaScript
  2. </>
    Disaggregated_Inference_Sim →
    The SimPy disaggregated-serving simulator of the sister series, extended for deck 03 with heterogeneous pools, FFT-mixing models, an optical transform device and KV hand-off compression (Python and a bit-exact JavaScript port; heterogeneous pools also in the Rust port, Rust_DES_Kernel). examples/results.md is the source of every deck-03 number.
    codePython · SimPy

How to read this series. Deck 01 is the physics and the bill: how optics computes a Fourier transform, and what precision, conversions and static power cost. Deck 02 asks where LLM inference has transforms at all, and answers with an Amdahl analysis computed by a published script. Deck 03 puts an optical transform device into a disaggregated-serving simulator's prefill pool. Every number in decks 01-02 comes from analysis/results.md, written by analysis/flop_share.py; every number in deck 03 from Disaggregated_Inference_Sim's examples/results.md. Illustrative coefficients are labelled where they are used.

Optical-computing companies and this series. Several photonic-computing companies publicly describe systems for AI and for fully homomorphic encryption, and published optical computing for FHE centres on Fourier transforms (see FHESim 04). Everything here about applying a Fourier-optical engine to LLM inference is this series' own analysis and speculation, attributed to no company.

Glossary: Concepts and Where They Are Explained

Every concept the decks rely on, with a short explanation and links to the slides that explain it. Concepts the sister series already explain are linked, not repeated: LLM Inference Simulators (prefill and decode, the roofline, the KV cache, disaggregation, power), FHE Accelerator Simulators (the 4f system, digit planes, ENOB and exact rounding, the Walden figure of merit) and the Simulation Engineering Toolkit (Amdahl's law, PPA).

Optics

The lens as a Fourier transformer
In coherent light, the field in a lens's back focal plane is the 2-D Fourier transform of the field in its front focal plane; a point at distance x from the axis holds spatial frequency x/(λf). The transform is computed by propagation, in one pass, for every pixel at once.Explained in: FOptInf 01 · the lens as a Fourier transformer · FOptInf 01 · the 4f system
The convolution theorem
A convolution in one domain is a pointwise product in the other: F{h ∗ u} = F{h}·F{u}. A 4f system uses it optically (mask = filter spectrum); FFT convolution uses it digitally in O(n log n).Explained in: FOptInf 01 · the 4f system and the convolution theorem · FOptInf 01 · interactive FFT convolution · FOptInf 02 · long convolutions
Coherent and incoherent light
Coherent light (a laser) adds complex field amplitudes, so signed and complex values interfere; incoherent light adds non-negative intensities. Coherent systems compute Fourier transforms of fields; incoherent ones are limited to non-negative quantities unless values are offset or split.Explained in: FOptInf 01 · coherent and incoherent light · FOptInf 01 · detectors and the phase problem
Spatial light modulator (SLM)
A pixelated device that writes values into a light field. Liquid-crystal SLMs modulate phase or amplitude and settle in tens of milliseconds; a 2-megapixel digital micromirror device switches amplitude at about 1 kHz with 8-bit depth (20 kHz binary). Pixel count and frame rate bound how fast inputs and Fourier-plane masks can change.Explained in: FOptInf 01 · spatial light modulators · FOptInf 02 · mask capacity
Square-law detection and the phase problem
Photodetectors measure intensity |E|², losing the sign and phase of the field. Remedies: interfere with a reference beam (coherent detection), add a bias so the field stays positive, split positive and negative parts into two passes, or recover phase iteratively (phase retrieval).Explained in: FOptInf 01 · detectors and the phase problem
Mach–Zehnder interferometer (MZI) meshes
A triangular (Reck) or rectangular (Clements) grid of MZIs implements any N×N unitary, including the DFT, on a chip: N(N−1)/2 MZIs, each with a phase shifter that must be calibrated and held.Explained in: FOptInf 01 · integrated-photonic alternatives
Integrated optical Fourier transforms
On-chip alternatives to a free-space lens: diffractive or slab-lens structures, MZI meshes, and coupler and delay-line FFT circuits (used for all-optical OFDM). They are compact and fast but much smaller than a free-space 4f system.Explained in: FOptInf 01 · integrated-photonic alternatives

Precision and energy

The noise rule for floating-point workloads
For float workloads the analogue chain must match the number format's own rounding error, not be exact. Relative RMS error = crest × 2−ENOB × 2/√12; matching BF16 needs about 10.3 ENOB, INT8 8.0, FP8 E4M3 6.3. FHE's exact-rounding rule does not transfer to long convolutions.Explained in: FOptInf 01 · precision: ENOB, noise and crosstalk · FOptInf 02 · precision and energy per token
Crest factor
The ratio of a signal's peak to its RMS value. An ADC's full scale must cover the peak, so a high crest factor wastes resolution on the rare large values: the noise rule's error scales with it.Explained in: FOptInf 01 · precision: ENOB, noise and crosstalk
Dynamic precision by repeated passes
Repeating an analogue operation k times and averaging cuts random noise by √k, buying ½ log2 k bits: one extra bit costs four times the passes. Digit planes, which make FHE exact, do not help a float workload.Explained in: FOptInf 01 · buying precision: passes and planes · FOptInf 03 · static power and energy per token
Conversion-bound optics and the break-even ENOB
Every value entering or leaving the optics costs a DAC or ADC sample, at Walden-FoM × 2ENOB joules. With the series' illustrative FoMs a Hyena-style long convolution breaks even with a 1 pJ/FLOP digital FFT at about 12 ENOB.Explained in: FOptInf 01 · conversion energy and static power · FOptInf 02 · precision and energy per token · FOptInf 03 · static power and energy per token

Sequence models

FNet
An encoder that replaces self-attention with an unparameterised 2-D Fourier transform over sequence and hidden dimensions, keeping the real part. It mixes every token with every other, including later ones, so it is not causal and cannot serve a decoder as it stands.Explained in: FOptInf 02 · FNet: Fourier token mixing
Long convolution by FFT
A convolution whose filter is as long as the sequence, computed by FFT, pointwise multiply and inverse FFT in O(L log L) instead of O(L²). It is the token mixer of S4's convolution mode, H3 and Hyena.Explained in: FOptInf 02 · long convolutions · FOptInf 01 · interactive FFT convolution
Hyena
An attention-free operator: N+1 linear projections, a short depthwise convolution, then N rounds of a long implicit convolution (by FFT) and elementwise gating. Order 2 is the usual language-model setting.Explained in: FOptInf 02 · long convolutions · FOptInf 02 · the Amdahl analysis
State-space models (S4, H3, Mamba)
Sequence layers built on a linear recurrence. S4's time-invariant SSM can run as a long convolution (for training and prefill) or as a recurrence (for generation); Mamba's selective SSM makes its parameters input-dependent, which rules out the convolution form.Explained in: FOptInf 02 · long convolutions · FOptInf 02 · what does not map
Causal zero-padding
An FFT computes a circular convolution, which wraps late inputs onto early outputs. Padding input and filter to at least 2L before the FFT makes it the linear, causal convolution a decoder needs.Explained in: FOptInf 02 · causality at prefill and at decode · FOptInf 01 · interactive FFT convolution
Relaxed (tiled) online convolution
Exact token-by-token convolution in which each completed block of 2l inputs is convolved by FFT once and added to future outputs (Flash Inference): FFTs at decode, but small ones on the critical path.Explained in: FOptInf 02 · causality at prefill and at decode
Distilled recurrence (Laughing Hyena)
Fitting each long-convolution filter with a small state-space model after training, so decode runs as an O(1)-per-token recurrence with a constant-size state instead of a growing cache.Explained in: FOptInf 02 · causality at prefill and at decode · FOptInf 02 · what prefill hands to decode
Block-circulant weights
A weight matrix made of circulant blocks, each defined by one vector, so multiplying by it is an FFT, a pointwise product and an inverse FFT (CirCNN). It turns matmuls into transforms; at LLM scale it is speculative.Explained in: FOptInf 02 · structured weights
Monarch and butterfly matrices
Structured matrices built from products of block-diagonal and permutation factors, a family that includes the FFT. Monarch Mixer uses them along both sequence and hidden dimensions, with a causal variant.Explained in: FOptInf 02 · structured weights
Hybrid architectures
Models that interleave a few attention layers with cheaper mixers (Hyena, Mamba). They keep attention's recall where it matters; the transform share falls with the attention fraction.Explained in: FOptInf 02 · the Amdahl analysis

Inference mapping

Optical share and the Amdahl bound
The fraction f of prefill FLOPs a transform engine could take over (FFTs plus Fourier-plane multiplies). If it became free, prefill would speed up by at most 1/(1−f); for a Hyena-2 model of Llama-3-8B's shape f is 0.20% at 2,048 tokens.Explained in: FOptInf 02 · the Amdahl analysis · FOptInf 02 · interactive Amdahl explorer
Time-weighted share
FFTs run below a GPU's matmul rate, so their share of time exceeds their share of FLOPs. Weighting transform FLOPs by a relative efficiency r (an assumption) gives the share an engine could actually remove.Explained in: FOptInf 02 · the Amdahl analysis · FOptInf 02 · interactive Amdahl explorer
Fourier-plane mask capacity
A 4f pass multiplies by one mask, so every filter spectrum must be on the mask while its inputs pass. Mask values needed per forward pass divided by what the SLM holds gives rewrites per pass, each costing a frame.Explained in: FOptInf 02 · mask capacity · FOptInf 03 · mask rewrites in the simulator
Transform-domain KV compression
Compressing the KV cache or hidden states by keeping low-frequency DCT/Fourier coefficients along the sequence (FreqKV, FourierAttention, Fourier Transformer). It saves memory; its transforms are small.Explained in: FOptInf 02 · transform-domain compression
Prefill-to-decode hand-off
What a disaggregated server moves from the prefill pool to the decode pool: the KV cache for attention, the cached projection inputs for direct long-convolution decode (four times larger here), or a constant recurrent state after distillation.Explained in: FOptInf 02 · what prefill hands to decode · FOptInf 03 · the KV hand-off in the simulator
Transform engine against optical MAC
An optical MAC accelerates every matrix multiply (the simulator's hypothetical optical part); a Fourier transform engine accelerates only FFT and convolution work and leaves matmuls to a digital part.Explained in: FOptInf 02 · a standard decoder has no FFT · FOptInf 02 · why disaggregation fits · FOptInf 03 · the transform engine model
Optical links (co-packaged optics, optical I/O chiplets) that move bytes, such as the KV cache, between chips. They are photonics but not Fourier optics: they compute nothing.Explained in: FOptInf 02 · why disaggregation fits · FOptInf 03 · the KV hand-off in the simulator
Heterogeneous pools
Disaggregated serving with different hardware per phase: compute-heavy parts for prefill, bandwidth-rich parts for decode (Splitwise's proposal). An optical transform engine would join the prefill pool.Explained in: FOptInf 03 · heterogeneous pools in the simulator · FOptInf 02 · why disaggregation fits · LLM Inference Simulators 05 · limitations and extensions
GPU FFT efficiency
The fraction of its matmul rate a GPU reaches on FFT and Fourier-domain work. The simulator's default (1) is optimistic for GPUs; at 1/16 the GPU's own prefill of a transform-heavy model slows enough for an optical engine to have something to win. It moves the circulant verdict more than any optical parameter.Explained in: FOptInf 03 · where the time goes · FOptInf 03 · the break-even point
Break-even point
The value of a parameter (static power, ENOB, mask rate, GPU FFT efficiency) at which an optical prefill pool's J/token or TTFT p99 equals the all-GPU baseline's, found by bisection on full simulator runs.Explained in: FOptInf 03 · the break-even point
Compute in transit
Operations applied in a link's data path, between transmitter and receiver, on data that must move anyway. At about 1–2 operations per byte at line rate it suits requantising or frequency-compressing the KV hand-off, not matmuls; it helps latency only where the link is the bottleneck. Speculative here.Explained in: FOptInf 03 · compute in the transport · FOptInf 03 · in transit or at the GPU

Explained in the sister series

Prefill and decode · FLOP and byte counting · Roofline, arithmetic intensity and ridge point · KV cache · Disaggregated (prefill/decode) serving · KV-cache transfer · TTFT, TPOT and inter-token latency · Goodput and SLO attainment · Energy per token and energy–delay trade-off · Static and dynamic power · Power in photonic compute · Cost model · Prefill–decode interference · Quantisation · Fourier optics and the 4f system · Digit decomposition (bit slicing) · Bluestein's algorithm · ENOB and exact rounding · Walden figure of merit · Number-theoretic transform (NTT) · Amdahl's law for a port · PPA: power, performance and area · perf/W