Fourier Optics for Inference — Presentation 01

Fourier Optics for Engineers

How a lens computes a Fourier transform and a 4f system a convolution, in one pass of light; coherent and incoherent light, spatial light modulators, detectors and the phase problem, integrated-photonic alternatives; and the precision and energy bill: ENOB, noise, passes, and the converters around the optics.

4f system Convolution theorem SLMs Phase problem ENOB Conversion energy
Input SLM → Lens → Fourier mask → Lens → Detector + ADC
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators, FHE Accelerator Simulators and the Simulation Engineering Toolkit for concepts those series already explain.

01

Why Fourier Optics, and Why Now

A lens performs a two-dimensional Fourier transform on a light field as the light propagates through it. Every pixel of the transform is produced at once, by physics rather than by arithmetic, so the transform has no O(n log n) operation count, only the cost of getting data into and out of the light. That asymmetry is the whole story of this series.

What it offers

  • Massive parallelism: a megapixel field is transformed in one pass.
  • A free convolution: two lenses and a mask (the 4f system, slide 03) multiply in the frequency domain.
  • Low energy per operation inside the optics; the bill is paid at the edges (slide 11).

What it demands

  • A workload whose expensive part is a Fourier transform or a convolution. Deck 02 checks whether LLM inference has one.
  • Tolerance of analogue precision (slide 09).
  • Enough reuse per conversion to amortise the DACs, ADCs, lasers and modulators.
A caution worth keeping

McMahon's review of optical computing argues that the speed of light is not what gives optics an advantage; any advantage comes from combining several of optics' features (parallelism, passive linear operations, low-loss data movement) while avoiding the pitfalls of conversion and precision (arXiv:2308.00088). The sister series makes the same point about power (LLM Inference Simulators 07, “Power in Photonic and Novel Compute”).

02

The Lens as a Fourier Transformer

Place a transparency with field U0(x, y) in the front focal plane of a thin lens of focal length f, and illuminate it with a coherent plane wave of wavelength λ. In the back focal plane the field is, in the paraxial (Fresnel) approximation:

Uf(u, v) = 1⁄iλf ∬ U0(x, y) · exp[−i 2π (xu + yv) / (λf)] dx dy

The relation is the standard result of scalar diffraction theory, derived in any Fourier-optics text; the review by Wetzstein et al. puts it in the context of optical inference (Nature 588, 2020).

03

The 4f System and the Convolution Theorem

f f f f input u(x)SLM or modulators lens 1F{u} Fourier planemask H = F{h} lens 2F{H·F{u}} outputdetector + ADC result at the output plane: (h ∗ u)(−x), a convolution, mirrored

Learned 4f layers have been demonstrated for image classification, with an optimised diffractive element as the mask (Chang et al., Sci. Rep. 8, 2018), and an amplitude-only Fourier-plane processor runs convolutions on about 1,000×1,000 matrices (Miscuglio et al., arXiv:2008.05853; Optica 2020).

04

Interactive: FFT Convolution in 1-D and 2-D

The digital twin of a 4f system: transform, multiply by the filter's spectrum, transform back. Compare it with a direct convolution, turn padding off to see the circular wrap, and add an ADC of a chosen ENOB at the output (noise of rms q/√12 with q = 2·peak/2ENOB, the definition of ENOB) to see the analogue error against the rule of slide 09.

ideal
FFT vs direct: max error / peak
—
Future leak (1-D)
—
ADC error (relative RMS)
—
Rule: crest × 2−ENOB × 2/√12
—

The same checks, in NumPy, are in analysis/flop_share.py: padded FFT convolution against direct, 5.6e-16 of the peak; unpadded, 0.72.

05

Coherent and Incoherent Light

Coherent (a laser)

  • Fields add as complex amplitudes: two beams can cancel. Values can be signed or complex, encoded in amplitude and phase.
  • The lens transforms the field, so the 4f system computes a true complex convolution.
  • Costs: a stable laser, phase-accurate modulators, alignment, and speckle and coherent noise from every stray reflection.

Incoherent (LEDs, broadband light)

  • Intensities add, and intensities are never negative.
  • An incoherent system convolves intensity with the intensity point-spread function: a real, non-negative kernel only.
  • Signed values need a bias or a split into positive and negative parts, each costing dynamic range or a second pass.
06

Spatial Light Modulators

Data enters a free-space system through a spatial light modulator (SLM): a pixel array that writes amplitude or phase into the beam. One sits at the input plane; another can hold the Fourier-plane mask.

DeviceWritesUpdate rateNotes
Liquid-crystal SLM (e.g. LCoS)Phase (or amplitude), many grey levelsSettles in tens of milliseconds (“tens of Hz” in Miscuglio et al.'s comparison)High pixel counts; the usual choice for phase masks
Digital micromirror device (DMD)Amplitude (on/off mirrors; grey levels by time-multiplexing)2 million mirrors at 1,031 Hz with 8-bit depth, about 20 kHz binary, in Miscuglio et al.'s processorAt least two orders of magnitude faster than liquid crystal at the same resolution, per the same paper
Integrated modulators (on chip)Amplitude or phase per waveguideGHz: Feldmann et al. report operation above 14 GHz, limited by modulators and detectorsDemonstrations have few channels, not millions of pixels
07

Detectors and the Phase Problem

Photodetectors are square-law: they measure intensity |E|2, not the field E. A coherent 4f system computes a signed or complex field, and the detector throws away its sign and phase. Four remedies:

RemedyHowCost
Coherent (homodyne) detectionInterfere the output with a reference beam; the cross term is linear in EA phase-stable reference, extra optics or balanced detectors
BiasAdd a constant so the field stays positive, take √intensity, subtractDynamic range: the bias uses ADC codes
Split passesRun positive and negative parts separately and subtract digitallyTwice the passes and conversions
Phase retrievalRecover phase from intensity-only measurements iterativelyMany iterations; not exact in general
08

Integrated-Photonic Alternatives

A free-space 4f system is large and slow to reconfigure. On a chip there are three ways to compute a Fourier transform or a linear transform:

ApproachWhat it computesScale and limitsExample
MZI meshesAny N×N unitary, the DFT included, as a grid of Mach–Zehnder interferometersN(N−1)/2 MZIs, each with a calibrated phase shifter; loss and error grow with NReck et al. (PRL 1994); Clements et al. (arXiv:1603.08788); Shen et al.'s coherent nanophotonic network (arXiv:1610.02365)
On-chip diffractionDiffractive layers in a silicon slab: a passive, learned linear transform (an on-chip lens is the Fourier special case)Passive and compact (0.15–0.3 mm2 footprints in the example); fixed once fabricatedFu et al. (Nat. Commun. 14, 2023)
Coupler and delay FFT circuitsThe FFT butterfly network built from couplers and delay linesFixed transform at very high line rate, used for optical OFDMHillerkuss et al., 26 Tbit/s with an all-optical FFT (Nat. Photon. 5, 2011)
09

Precision: ENOB, Noise and Crosstalk

The effective number of bits (ENOB) of an analogue chain is defined from its signal-to-noise-and-distortion ratio, SINAD = 6.02·ENOB + 1.76 dB for a full-scale sine (Walden's ADC survey, IEEE JSAC 1999). It lumps every error source into one number: shot and thermal noise, laser intensity noise, crosstalk between pixels, aberrations, modulator non-linearity, and the converters' own quantisation.

For floating-point workloads the requirement is a noise level, not exactness. With the ADC's full scale set to the output's peak, the relative RMS error is crest × 2−ENOB × 2/√12, where the crest factor is the output's peak-to-RMS ratio. Measured on a causal FFT convolution:

LengthCrestENOB 4ENOB 6ENOB 8ENOB 10ENOB 12ENOB 14Rule at ENOB 8
2562.971.1e-012.7e-026.7e-031.4e-034.1e-041.0e-046.7e-03
4,0963.641.3e-013.3e-028.3e-032.1e-035.1e-041.3e-048.2e-03
65,5364.431.6e-014.0e-021.0e-022.5e-036.2e-041.6e-041.0e-02

Source: results.md, section “6. Analogue precision”

10

Buying Precision: Passes and Planes

How many bits does inference need, and how do you buy more than the hardware has? The format's own rounding error sets the target (same metric, length 4,096):

FormatRounding error (relative RMS)ENOB to matchAveraged passes at ENOB 8
BF161.6e-0310.326
FP8 E4M32.6e-026.31
INT88.3e-038.01

Source: results.md, section “6. Analogue precision”

Two ways to buy precision at ENOB 8:

Length1 pass4 passes averaged16 passes averaged2 digit planes of 8 bits
2566.7e-033.3e-031.5e-036.7e-03
4,0968.3e-034.2e-032.1e-038.1e-03
65,5361.0e-025.0e-032.5e-039.9e-03

Source: results.md, section “6. Analogue precision”

11

Conversion Energy and Static Power

Every value in and out costs a conversion

Energy per sample ≈ Walden FoM × 2ENOB: each extra bit doubles it. With the illustrative FoMs of FHESim 04 (DAC 10 fJ, ADC 20 fJ per conversion step):

ENOBDAC pJADC pJPair pJ
40.160.320.48
60.641.281.92
82.565.127.68
1010.2420.4830.72
1240.9681.92122.88
14163.84327.68491.52
16655.361,310.721,966.08

Source: results.md, section “7. Conversion energy”

Static power never sleeps

  • Lasers burn power whether or not work arrives; wall-plug efficiency sets the electrical cost.
  • Thermal tuning holds resonant and interferometric devices on wavelength.
  • SLM drivers and their frame buffers.
  • Static power must be amortised over work: an idle optical engine is pure cost. This is energy proportionality, from the sister series (LLM Inference Simulators 07, “Energy Proportionality: Static Power and Load”).
12

What Optics Gives and What It Costs

PropertyFree-space 4f engineWhat it means for a workload
Transform costTime of flight; no operation countLong transforms are as cheap as short ones, once the data is in
ParallelismMillions of pixels; many 1-D transforms per frameWants batches of independent transforms (channels, sequences)
PrecisionAnalogue; this series assumes about 8 effective bits (illustrative)Fine for FP8/INT8-like error, short of BF16 without averaging; exact integer arithmetic only for small blocks
ConversionsOne DAC sample in and one ADC sample out per value, energy ∝ 2ENOBNeeds many operations per converted value to pay
WeightsOn the Fourier-plane mask, rewritten at the SLM's frame rateWants weights held across a batch: prefill, not token-by-token decode
Signs and phaseLost at a square-law detector unless detected coherentlyCoherent detection, or twice the passes
Static powerLasers, tuning, driversMust be kept busy

The sister series applied the same accounting to FHE, where exact modular arithmetic made the precision tax severe (FHESim 04, “When It Wins and When It Loses”). Deck 02 applies it to LLM inference.

13

What to Take Away