Where a Fourier transform could appear in LLM inference, and how much it would matter: FNet, long convolutions (S4, H3, Hyena), causality at prefill and decode, an Amdahl analysis of prefill FLOP shares, structured weights, transform-domain KV compression, mask capacity, and an honest list of what does not map.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators, FHE Accelerator Simulators and the Simulation Engineering Toolkit for concepts those series already explain.
Deck 01 showed what a Fourier-optical engine does well: long transforms and convolutions, with the weights held on a mask, at a cost per converted value. The question for inference is whether the work contains such transforms, and how much of the work they are. Three places to look, and two things that are photonics but not Fourier optics:
| Where | What the transform does | Slides |
|---|---|---|
| Token mixing | FNet's Fourier mixing; long convolutions (S4's convolution mode, H3, Hyena) computed by FFT | 03–07 |
| Structured weights | Circulant, block-circulant or Monarch matrices, whose products are transforms | 08 |
| Transform-domain compression | DCT/Fourier coefficients of the KV cache or hidden states along the sequence | 09 |
| Optical MACs (not Fourier optics) | Matrix multiplies done optically: the simulator's hypothetical optical part | 02 |
| Photonic interconnect (not Fourier optics) | Moving the KV cache between pools over optical links | 13 |
Every number from here on comes from analysis/flop_share.py, which counts FLOPs for a model of Llama-3-8B's shape (Llama 3 report) in five variants and writes results.md.
A Llama-style decoder is matrix multiplies (projections, the SwiGLU MLP, the LM head) plus attention (QKT and AV), softmax, norms and activations. None of it is a Fourier transform. The analysis reproduces the companion simulator's FLOP count exactly (Disaggregated_Inference_Sim's cost model, checked to the FLOP), and splits it by class:
| Llama-3-8B prefill | GFLOP | Dense matmul | Attention |
|---|---|---|---|
| 512 tokens | 7,753.6 | 99.1% | 0.9% |
| 2,048 tokens | 31,839.1 | 96.5% | 3.5% |
| 32,768 tokens | 773,308.9 | 63.6% | 36.4% |
| 131,072 tokens | 6,470,935.2 | 30.4% | 69.6% |
Source: results.md, section “2. Prefill”. Lengths beyond the model's trained context are shape-only extrapolations.
The simulator's HYPOTHETICAL_OPTICAL device is a “what if compute were nearly free” probe: it speeds up every dense FLOP, and, as its comment in hardware.py says, decode barely moves because decode is bandwidth-bound.
A Fourier transform engine accelerates only FFT and convolution work and leaves matmuls to a digital part. For this model its share is zero. The rest of the deck asks which other models give it something to do.
FNet replaces each self-attention sublayer of a Transformer encoder with an unparameterised 2-D discrete Fourier transform over the sequence and hidden dimensions, keeping the real part. Lee-Thorp et al. report 92–97% of BERT's GLUE accuracy while training 80% faster on GPUs and 70% faster on TPUs at 512 tokens (arXiv:2105.03824).
The FFT is what makes FNet cheap. Making a cheap operation free buys little: the theme of this deck.
Causal sequence models that mix tokens with a long convolution (a filter as long as the sequence) can compute it by FFT in O(L log L):
| Model | Mixer | Convolution or recurrence? |
|---|---|---|
| S4 (Gu, Goel, Ré, arXiv:2111.00396) | A structured, time-invariant state-space model | Both: a long convolution for training, a recurrence for generation (reported 60× faster generation) |
| H3 (Fu, Dao et al., arXiv:2212.14052) | Two SSMs with multiplicative gating | FFT convolution at training and prefill |
| Hyena (Poli et al., arXiv:2302.10866) | N+1 projections, a short depthwise conv, N implicit long convolutions with gating (order 2 for language) | FFT convolution, with input and filter zero-padded to 2L − 1 for causality (the paper, section 3) |
| Mamba (Gu, Dao, arXiv:2312.00752) | A selective SSM: parameters depend on the input | Recurrence only: selectivity “prevents the use of efficient convolutions” (the abstract) |
| Mixer | Context | GFLOP / token | Mixer FLOPs / token | Transform share |
|---|---|---|---|---|
| transformer | 2,048 | 16.084 | 1,074.3 M | 0.00% |
| transformer | 32,768 | 32.190 | 17,180.4 M | 0.00% |
| hyena_direct | 2,048 | 17.694 | 1,074.3 M | 0.00% |
| hyena_direct | 32,768 | 33.800 | 17,180.4 M | 0.00% |
| hyena_distilled | 2,048 | 16.653 | 33.6 M | 0.00% |
| hyena_distilled | 32,768 | 16.653 | 33.6 M | 0.00% |
| hyena_tiled | 2,048 | 16.877 | 256.9 M | 1.40% |
| hyena_tiled | 32,768 | 17.046 | 425.7 M | 2.34% |
| hyena_recompute | 2,048 | 162.650 | 146,030.5 M | 85.82% |
| hyena_recompute | 32,768 | 3,040.278 | 3,023,658.5 M | 96.06% |
Source: results.md, section “4. Decode”
So the hypothesis needs care: decode can use FFTs, but only in small passes on the latency-critical path. Conversions per token per channel (DAC+ADC pairs): a prefill long convolution converts each input once and each output once, 1 pair per token. Relaxed tiling at decode converts, at each of its tile levels, a block of B inputs and B outputs every 2B steps, 0.5 pairs per token per level: 6 pairs at context 2,048 (12 levels); 8 pairs at context 32,768 (16 levels), in small passes on the latency-critical path.
Optical share f = FFT FLOPs plus the pointwise multiplies in the Fourier domain (what a 4f mask does), over all prefill FLOPs. If the engine made that work free, prefill would speed up by at most 1/(1−f) (Amdahl's law).
| Variant | 512 tokens | 2,048 tokens | 8,192 tokens | 32,768 tokens | 131,072 tokens |
|---|---|---|---|---|---|
| Transformer (Llama-3-8B) | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% |
| FNet-shaped (non-causal) | 0.08% | 0.09% | 0.10% | 0.11% | 0.12% |
| Hyena-2 | 0.17% | 0.20% | 0.23% | 0.26% | 0.29% |
| Hybrid 1:3 attention:Hyena | 0.13% | 0.15% | 0.17% | 0.18% | 0.15% |
| Hyena-2 + block-circulant weights | 13.63% | 13.99% | 14.35% | 14.70% | 15.05% |
Source: results.md, section “2. Prefill”
The FLOP ledger of flop_share.py, ported to JavaScript and tested to give identical numbers (js/flop_model.js). Choose a model, a prompt length, where the LM head runs, and how fast a GPU runs FFTs relative to matmuls (r, an assumption).
Things to try: switch the block-circulant model's LM head to the last token only, then set r = 1/16. That is the only corner of this space where a transform engine has most of the work.
If token mixing is too small a share, the transforms would have to come from the weights: matrices whose product with a vector is a transform.
The analysis makes every projection and MLP matrix block-circulant (block 256, illustrative) in the Hyena-2 model. The LM head then dominates, so where it runs decides the answer:
| Variant | Prompt | GFLOP | Optical share | Amdahl bound |
|---|---|---|---|---|
| Hyena-2 | 2,048 | 31,959.9 | 0.21% | 1.002x |
| Hyena-2 | 32,768 | 511,686.3 | 0.28% | 1.003x |
| Hyena-2 + block-circulant weights | 2,048 | 426.6 | 84.54% | 6.469x |
| Hyena-2 + block-circulant weights | 32,768 | 7,152.8 | 85.47% | 6.883x |
Source: results.md, section “2. Prefill”
This is the only variant where a transform engine has most of the work, and it is the least established: phase A of this work found no published LLM of this size with block-circulant weights throughout. Quality at that scale is unknown, and the weights live on the Fourier-plane mask (slide 11).
A third place transforms appear: compressing what the model stores, along the sequence, by keeping a few frequency coefficients. This literature exists and is recent:
| Work | What it transforms | Reported result |
|---|---|---|
| Fourier Transformer (He et al., arXiv:2305.15099) | Hidden states, with a DCT computed by FFT, progressively removing sequence redundancy | Inherits pretrained weights (BART); strong long-range benchmark results |
| FreqKV (Kai et al., arXiv:2505.00570) | The KV cache, iteratively compressed in the frequency domain, keeping low frequencies | Extends LLaMA-2-7B's context to 256K tokens with minimal training at 8K |
| FourierAttention (Liu et al., arXiv:2506.11886) | The long-context-insensitive head dimensions, projected onto fixed-length Fourier bases | Training-free; a fused Triton kernel |
Precision: float workloads need the analogue error to match the number format's own rounding (deck 01, “Precision: ENOB, Noise and Crosstalk”): about 8.0 ENOB for INT8-like error, 6.3 for FP8 E4M3 and 10.3 for BF16. Below that, averaging costs 4× the passes per bit.
Energy: DAC+ADC pairs per prompt token, and their energy at each ENOB, against the digital energy of the transform work they replace, in µJ per prompt token (Walden FoMs 10/20 fJ and 1 pJ/FLOP, both illustrative):
| Variant | Pairs / token | Digital uJ | ENOB 8 uJ | ENOB 12 uJ | Break-even ENOB |
|---|---|---|---|---|---|
| FNet-shaped (non-causal) | 131,072 | 11.1 | 1.0 | 16.1 | 11.5 |
| Hyena-2 | 262,144 | 33.0 | 2.0 | 32.2 | 12.0 |
| Hybrid 1:3 attention:Hyena | 196,608 | 24.8 | 1.5 | 24.2 | 12.0 |
| Hyena-2 + block-circulant weights | 2,162,688 | 176.1 | 16.6 | 265.8 | 11.4 |
Source: results.md, section “7. Conversion energy”
A 4f pass multiplies by one mask. Each Hyena channel's filter spectrum (or each circulant block) must be on the mask while its inputs pass, so the mask is the weight store, and rewriting it costs a frame of the spatial light modulator (deck 01, “Spatial Light Modulators”). Time to rewrite a 2-megapixel mask enough times for one forward pass, at three device rates: a micromirror device at 1-bit depth (20 kHz) and 8-bit depth (1.03 kHz), and a liquid-crystal SLM (30 Hz):
| Variant | Prompt | Rewrites | DMD 20 kHz | DMD 1.03 kHz | LC 30 Hz |
|---|---|---|---|---|---|
| Hyena-2 | 2,048 | 269 | 13.4 ms | 261 ms | 9.0 s |
| Hyena-2 | 32,768 | 4,296 | 214.8 ms | 4,167 ms | 143.2 s |
| Hyena-2 + block-circulant weights | 2,048 | 277 | 13.8 ms | 269 ms | 9.2 s |
| Hyena-2 + block-circulant weights | 32,768 | 4,303 | 215.2 ms | 4,174 ms | 143.4 s |
Source: results.md, section “8. Holding the filters”
Disaggregated serving moves state from the prefill pool to the decode pool once per request (LLM Inference Simulators 05, “What the KV Transfer Costs”). What that state is depends on the mixer:
| Decode style | Bytes per prompt token | Bytes for a 2,048-token prompt | Grows with the prompt |
|---|---|---|---|
| Transformer KV cache (GQA) | 131,072 | 268.4 MB | yes |
| Hyena direct: cached projection inputs | 524,288 | 1,073.7 MB | yes |
| Hyena distilled recurrence state | n/a | 16.8 MB | no |
Source: results.md, section “9. What prefill hands to decode”
Optical-computing companies publish little about inference mapping. Several publicly describe photonic computing systems for AI and for FHE, and published optical computing for FHE centres on Fourier transforms (FHESim 04). The mapping in this deck is the author's own analysis and speculation, attributed to no company.
| Work | Why a Fourier transform engine does not help |
|---|---|
| Attention (QKT, softmax, AV) | A data-dependent, causal, softmax-normalised product: not a convolution. At long context it is the biggest term (69.6% of Llama-3-8B prefill at 131,072 tokens), and none of it is a transform |
| Dense projections, MLP, LM head | General matmuls: work for an optical MAC, not a Fourier engine, unless the weights are structured (slide 08) |
| Mamba's selective scan | Input-dependent parameters rule out the convolution form; it is a recurrence |
| Decode, generally | One token per step: direct dot products or recurrences, memory-bound; relaxed tiling uses FFTs only in small latency-critical passes (slide 05) |
| Norms, activations, softmax, sampling, MoE routing | Non-linear or data-dependent; stay digital |
| FNet mixing in a decoder | Not causal (slide 03) |
| KV-cache reads | A bandwidth problem; optics could shrink it only through compression (slide 09) or move it over photonic links |
And the hidden costs that apply even where it maps: conversions per value, precision passes, mask rewrites, sign recovery at the detector, and static power whether or not work arrives.