How a lens computes a Fourier transform and a 4f system a convolution, in one pass of light; coherent and incoherent light, spatial light modulators, detectors and the phase problem, integrated-photonic alternatives; and the precision and energy bill: ENOB, noise, passes, and the converters around the optics.
4f systemConvolution theoremSLMsPhase problemENOBConversion energy
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators, FHE Accelerator Simulators and the Simulation Engineering Toolkit for concepts those series already explain.
A lens performs a two-dimensional Fourier transform on a light field as the light propagates through it. Every pixel of the transform is produced at once, by physics rather than by arithmetic, so the transform has no O(n log n) operation count, only the cost of getting data into and out of the light. That asymmetry is the whole story of this series.
What it offers
Massive parallelism: a megapixel field is transformed in one pass.
A free convolution: two lenses and a mask (the 4f system, slide 03) multiply in the frequency domain.
Low energy per operation inside the optics; the bill is paid at the edges (slide 11).
What it demands
A workload whose expensive part is a Fourier transform or a convolution. Deck 02 checks whether LLM inference has one.
Enough reuse per conversion to amortise the DACs, ADCs, lasers and modulators.
A caution worth keeping
McMahon's review of optical computing argues that the speed of light is not what gives optics an advantage; any advantage comes from combining several of optics' features (parallelism, passive linear operations, low-loss data movement) while avoiding the pitfalls of conversion and precision (arXiv:2308.00088). The sister series makes the same point about power (LLM Inference Simulators 07, “Power in Photonic and Novel Compute”).
02
The Lens as a Fourier Transformer
Place a transparency with field U0(x, y) in the front focal plane of a thin lens of focal length f, and illuminate it with a coherent plane wave of wavelength λ. In the back focal plane the field is, in the paraxial (Fresnel) approximation:
A Fourier transform, evaluated at spatial frequency (fx, fy) = (u, v) / (λf). A detector pixel at distance u from the axis reads the input's component at frequency u/(λf): longer focal lengths spread the spectrum out.
Front focal plane matters. With the input elsewhere, the transform picks up a quadratic phase factor; intensity measurements do not see it, but a following stage that uses the phase does.
Resolution is finite. The lens aperture and the pixel pitch of the devices at each plane bound the number of resolvable points (the space–bandwidth product), which plays the role of the FFT length.
It is 2-D by nature. A cylindrical lens transforms along one axis only, so one device can compute many independent 1-D transforms, one per row, in parallel. That is the form a sequence model's per-channel convolutions would take (deck 02).
The relation is the standard result of scalar diffraction theory, derived in any Fourier-optics text; the review by Wetzstein et al. puts it in the context of optical inference (Nature 588, 2020).
03
The 4f System and the Convolution Theorem
The convolution theorem: F{h ∗ u} = F{h} · F{u}. Lens 1 transforms the input, the mask multiplies it by the filter's spectrum H, and lens 2 transforms again. A second forward transform returns the result with coordinates mirrored, which is a relabelling, not a cost.
Correlation is the same machine with the conjugate spectrum on the mask: the classical optical correlator, used for pattern matching.
The mask holds the weights. Every different filter needs a different mask, so how fast and how large the mask can be rewritten matters as much as the transform (slide 06; deck 02 counts the rewrites for a sequence model).
Digital equivalent: FFT, pointwise multiply, inverse FFT, in O(n log n). An FFT computes a circular convolution; a causal, linear one needs zero-padding to at least twice the length, as the next slide shows.
Learned 4f layers have been demonstrated for image classification, with an optimised diffractive element as the mask (Chang et al., Sci. Rep. 8, 2018), and an amplitude-only Fourier-plane processor runs convolutions on about 1,000×1,000 matrices (Miscuglio et al., arXiv:2008.05853; Optica 2020).
04
Interactive: FFT Convolution in 1-D and 2-D
The digital twin of a 4f system: transform, multiply by the filter's spectrum, transform back. Compare it with a direct convolution, turn padding off to see the circular wrap, and add an ADC of a chosen ENOB at the output (noise of rms q/√12 with q = 2·peak/2ENOB, the definition of ENOB) to see the analogue error against the rule of slide 09.
ideal
FFT vs direct: max error / peak
—
Future leak (1-D)
—
ADC error (relative RMS)
—
Rule: crest × 2−ENOB × 2/√12
—
inputlog |input spectrum|output
The same checks, in NumPy, are in analysis/flop_share.py: padded FFT convolution against direct, 5.6e-16 of the peak; unpadded, 0.72.
05
Coherent and Incoherent Light
Coherent (a laser)
Fields add as complex amplitudes: two beams can cancel. Values can be signed or complex, encoded in amplitude and phase.
The lens transforms the field, so the 4f system computes a true complex convolution.
Costs: a stable laser, phase-accurate modulators, alignment, and speckle and coherent noise from every stray reflection.
Incoherent (LEDs, broadband light)
Intensities add, and intensities are never negative.
An incoherent system convolves intensity with the intensity point-spread function: a real, non-negative kernel only.
Signed values need a bias or a split into positive and negative parts, each costing dynamic range or a second pass.
Amplitude-only processing is a middle way: Miscuglio et al. modulate only amplitude in both planes and report robustness to coherence noise, with orders-of-magnitude lower delay than liquid-crystal phase devices (arXiv:2008.05853).
For inference, activations and filters are signed, so either the system is coherent and the detector recovers sign (slide 07), or every signed product is paid for twice. The simulator design for deck 03 makes this a parameter.
06
Spatial Light Modulators
Data enters a free-space system through a spatial light modulator (SLM): a pixel array that writes amplitude or phase into the beam. One sits at the input plane; another can hold the Fourier-plane mask.
Device
Writes
Update rate
Notes
Liquid-crystal SLM (e.g. LCoS)
Phase (or amplitude), many grey levels
Settles in tens of milliseconds (“tens of Hz” in Miscuglio et al.'s comparison)
High pixel counts; the usual choice for phase masks
Digital micromirror device (DMD)
Amplitude (on/off mirrors; grey levels by time-multiplexing)
2 million mirrors at 1,031 Hz with 8-bit depth, about 20 kHz binary, in Miscuglio et al.'s processor
At least two orders of magnitude faster than liquid crystal at the same resolution, per the same paper
Integrated modulators (on chip)
Amplitude or phase per waveguide
GHz: Feldmann et al. report operation above 14 GHz, limited by modulators and detectors
Demonstrations have few channels, not millions of pixels
The trade: free-space SLMs have millions of pixels but kHz-or-slower updates; integrated modulators run at GHz but have few channels. Throughput is pixels × rate either way, and the converters behind every pixel set the energy.
The mask is a weight store. A filter that changes every pass is limited by the mask's update rate; a filter held for many passes (a weight, reused across a batch) is nearly free. Deck 02 turns this into a count of mask rewrites per forward pass.
Rates and pixel counts move quickly with products; treat these as orders of magnitude and check the device datasheet. Sources: Miscuglio et al. (arXiv:2008.05853), Feldmann et al. (arXiv:2002.00281; Nature 589, 2021).
07
Detectors and the Phase Problem
Photodetectors are square-law: they measure intensity |E|2, not the field E. A coherent 4f system computes a signed or complex field, and the detector throws away its sign and phase. Four remedies:
Remedy
How
Cost
Coherent (homodyne) detection
Interfere the output with a reference beam; the cross term is linear in E
A phase-stable reference, extra optics or balanced detectors
Bias
Add a constant so the field stays positive, take √intensity, subtract
Dynamic range: the bias uses ADC codes
Split passes
Run positive and negative parts separately and subtract digitally
Twice the passes and conversions
Phase retrieval
Recover phase from intensity-only measurements iteratively
Many iterations; not exact in general
Direct phase determination: Macfaden, Gordon and Wilkinson decompose the input and use symmetries of the Fourier transform to determine phase directly from intensity measurements, giving an optical complex-to-complex DFT (Sci. Rep. 7, 2017).
Homodyne detection in optical neural networks: Hamerly et al. encode both inputs and weights optically and multiply them by interfering the two at the detectors (photoelectric multiplication, arXiv:1812.07614; Phys. Rev. X 9, 2019).
Phase retrieval algorithms are compared in Fienup's classic paper (Appl. Opt. 21, 1982).
The deck 03 design charges intensity detection as twice the passes, and coherent detection as one ADC sample per real output.
08
Integrated-Photonic Alternatives
A free-space 4f system is large and slow to reconfigure. On a chip there are three ways to compute a Fourier transform or a linear transform:
Approach
What it computes
Scale and limits
Example
MZI meshes
Any N×N unitary, the DFT included, as a grid of Mach–Zehnder interferometers
N(N−1)/2 MZIs, each with a calibrated phase shifter; loss and error grow with N
Photonic tensor cores perform matrix-vector products rather than transforms: Feldmann et al. (phase-change memory and frequency combs, Nature 589, 2021) and Xu et al. (11 TOPS, Nature 589, 2021). Those are optical MACs, the other branch deck 02 distinguishes.
An on-chip FFT has the converter bill of a free-space one without its millions of parallel pixels; its advantage is speed of reconfiguration and integration.
09
Precision: ENOB, Noise and Crosstalk
The effective number of bits (ENOB) of an analogue chain is defined from its signal-to-noise-and-distortion ratio, SINAD = 6.02·ENOB + 1.76 dB for a full-scale sine (Walden's ADC survey, IEEE JSAC 1999). It lumps every error source into one number: shot and thermal noise, laser intensity noise, crosstalk between pixels, aberrations, modulator non-linearity, and the converters' own quantisation.
For floating-point workloads the requirement is a noise level, not exactness. With the ADC's full scale set to the output's peak, the relative RMS error is crest × 2−ENOB × 2/√12, where the crest factor is the output's peak-to-RMS ratio. Measured on a causal FFT convolution:
The rule predicts the measurement (last column against the ENOB 8 column). Longer convolutions have a higher crest factor, so they need slightly more bits for the same error.
FHE needs exactness, which is much harder (FHESim 04, route 1). Exact rounding, by FHESim 04's rule 2^(ENOB-1) > d n (2^b - 1)^2 with 8-bit inputs and filters in 1-bit digits (d = 8, the most forgiving split; n = products summed per output), needs ENOB 9 for L = 16; ENOB 17 for L = 4,096; ENOB 21 for L = 65,536: feasible for FHE's 16-point blocks, not for sequence-length convolutions, so float workloads use the noise rule above instead.
10
Buying Precision: Passes and Planes
How many bits does inference need, and how do you buy more than the hardware has? The format's own rounding error sets the target (same metric, length 4,096):
Averaging works, expensively: k passes cut random noise by √k, so each extra bit costs 4× the passes (Garg et al., arXiv:2102.06365, who choose precision per layer). Reaching BF16's error from ENOB 8 takes 26 passes (first table, last column).
Digit planes do not help floats. They make FHE's integer arithmetic exact by shrinking each plane's values (FHESim 04, route 1), but for a float workload the top plane carries nearly all of the signal and is read at the same ENOB relative to its own peak.
FP8 and INT8 inference sit within reach of an 8-bit analogue chain; BF16 does not.
11
Conversion Energy and Static Power
Every value in and out costs a conversion
Energy per sample ≈ Walden FoM × 2ENOB: each extra bit doubles it. With the illustrative FoMs of FHESim 04 (DAC 10 fJ, ADC 20 fJ per conversion step):
The break-even question is how many digital operations each conversion replaces. A length-L FFT convolution does about 10 log2(2L) + 6 FLOPs per output sample digitally (a forward and an inverse real FFT of length 2L, 2.5·2L·log2(2L) each, and the pointwise product, all divided over L outputs; the filter's spectrum is precomputed), against one DAC and one ADC sample optically. Deck 02 works it through per token: the break-even is near ENOB 12 for the illustrative coefficients.
Optical energy per MAC can scale down with problem size, because one optical pass reuses each converted value many times: Anderson et al. find optical energy per MAC scaling as 1/d with Transformer width d in their simulations (arXiv:2302.10360). Large, reused operations are where optics competes.
12
What Optics Gives and What It Costs
Property
Free-space 4f engine
What it means for a workload
Transform cost
Time of flight; no operation count
Long transforms are as cheap as short ones, once the data is in
Parallelism
Millions of pixels; many 1-D transforms per frame
Wants batches of independent transforms (channels, sequences)
Precision
Analogue; this series assumes about 8 effective bits (illustrative)
Fine for FP8/INT8-like error, short of BF16 without averaging; exact integer arithmetic only for small blocks
Conversions
One DAC sample in and one ADC sample out per value, energy ∝ 2ENOB
Needs many operations per converted value to pay
Weights
On the Fourier-plane mask, rewritten at the SLM's frame rate
Wants weights held across a batch: prefill, not token-by-token decode
Signs and phase
Lost at a square-law detector unless detected coherently
Coherent detection, or twice the passes
Static power
Lasers, tuning, drivers
Must be kept busy
The sister series applied the same accounting to FHE, where exact modular arithmetic made the precision tax severe (FHESim 04, “When It Wins and When It Loses”). Deck 02 applies it to LLM inference.
13
What to Take Away
A lens is a Fourier transformer; two lenses and a mask are a convolver. The transform is free; the mask holds the weights.
FFT convolution is circular unless zero-padded to twice the length, which is what makes it causal (the interactive, slide 04).
Detectors see intensity: sign and phase need coherent detection, a bias, split passes or retrieval.
Float workloads need a noise level, not exactness: about 8.0 ENOB for INT8-like error and 10.3 for BF16-like; averaging buys bits at 4× passes per bit, and digit planes do not help.
Conversions and static power are the bill. Optics pays only when each converted value is reused by many operations and the engine is kept busy.
Next:deck 02 asks where LLM inference has transforms to offer, and how much of the work they are.