Optical NTT Engines: Precision, Conversion and Power
Fourier transforms in optics, and what it takes to map an exact modular NTT onto an analogue complex FFT: digit decomposition, small-prime RNS, rounding and error correction, ENOB, DAC/ADC energy from a Walden figure of merit, and laser and thermal-tuning power. The simulator then shows when an optical engine wins and when it loses.
Fourier opticsDigit planesBluesteinENOBWalden FoMLaser power
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Cryptography decks or LLM Inference Simulators.
A lens computes a two-dimensional Fourier transform physically: the optical field in its front focal plane appears, Fourier-transformed, in its back focal plane. Two lenses with a mask between them (a 4f system) multiply in the frequency domain, which is a convolution in space. The transform itself costs time-of-flight and no switching energy.
Published photonic FHE work includes PHAT (Yang et al., arXiv:2609.11613), a photonic accelerator for TFHE whose FFTs run on optically addressed phase-change memory, and OptoLink (Saiham et al., arXiv:2506.12962), which uses photonics for the memory bandwidth side.
02
The Mismatch: Exact Modular Arithmetic
What CKKS needs
An NTT over Zq with q a 28–60-bit prime: twiddles are integers mod q, results must be exact.
One wrong coefficient is a wrong ciphertext. CKKS tolerates approximation in the message, not in the ring arithmetic.
What optics provides
A complex DFT with unit-modulus twiddles, in analogue, with an effective precision (ENOB) set by noise, linearity and the converters: typically far fewer bits than a limb.
No modular reduction, no integers.
The NTT of a limb is not a complex DFT, so the optics cannot run it directly. What it can compute exactly, given enough precision, is an integer convolution: convolve small non-negative integers with a complex FFT and round. Every route below turns modular NTT work into integer convolutions small enough that rounding is exact. The same trick lets TFHE libraries multiply polynomials with 64-bit floating-point FFTs.
03
Route 1: Digit Decomposition
Split each operand into d digits of b bits, x = Σi xi·2bi. A convolution of the full operands becomes d2 convolutions of digit planes, recombined digitally with shifts and adds.
Output size of one plane-pair convolution of length n: at most n·(2b−1)2. Summing the d products of equal weight i+i' in the analogue domain before the ADC ("grouped") multiplies that by d, and needs 2d−1 output planes instead of d2.
Exactness: a signed ADC with 2ENOB codes over ±FS errs by at most FS/2ENOB; rounding is exact if that is below ½: 2ENOB−1 > d·n·(2b−1)2.
The trade: narrower digits need fewer bits of analogue precision but more planes, and every plane is another pass through the converters.
Precedent
PHAT reaches the 42-bit FFT precision its TFHE workload needs from 6-bit optical cells by exactly this kind of multi-word decomposition ("7 × 6 + 1 = 43 bits of precision (including 1 bit for sign)", arXiv:2609.11613).
04
Route 2: A Hybrid NTT via Bluestein
The mapping the companion simulator models, one illustrative route among several:
Digital NTT stages log2(N/n) butterfly stages split XN+1 into length-n pieces
→
Bluestein chirp each length-n modular DFT becomes a size-2n cyclic convolution
→
Optical FFT convolution on b-bit digit planes: d planes in, 2d−1 out
→
Digital correction round, recombine, chirp, reduce mod q
Bluestein's identity jk = (j2 + k2 − (k−j)2)/2 turns Σj xjωjk into a chirp multiply, a convolution with ψ−m2 (where ψ2 = ω, so q ≡ 1 mod 2n), and another chirp multiply. Both chirps and the convolution kernel are integers mod q, so the convolution is an integer convolution.
What is offloaded: only the last log2n of the log2N butterfly stages, which is (log2n)/2 butterflies per point. A 16-point block offloads 2 butterflies per point; a 4,096-point block, 6.
What it costs per point: (d + 2d−1) planes × 2n samples per n points, through a DAC or an ADC: 2(3d−1) conversions per point.
05
Route 3: Small Primes and Error Correction
Idea
What it buys
What it costs
In the model?
Small-prime RNS
Narrower limbs (CraterLake uses 28-bit words, SHARP 36-bit) mean fewer digits per operand
More limbs for the same modulus, so more NTTs; rescaling needs care (scale ≠ prime)
Yes, via limb bits (the optical precision depends on it)
Rounding with redundancy
Compute in a slightly bigger modulus or with a check residue; detect and correct rare analogue errors digitally
Extra digital work and some extra planes
No; the model requires worst-case exactness
Probabilistic exactness
Size the ENOB for typical rather than worst-case error, then rely on detection
A failure-rate argument a security reviewer must accept
No
Fewer conversions per point
Chain several optical stages without leaving the optical domain
Error accumulates across stages; needs higher ENOB at the end
No; an extension exercise
CraterLake's 28-bit words come from ARK's comparison table (arXiv:2205.00922, Table VII); SHARP's 36-bit words and 180 MB of on-chip memory are as reported in FHEmem (arXiv:2311.16293) and Osiris (arXiv:2408.09593).
06
The Precision Rule, Checked Functionally
precision.py implements route 2 end to end in pure Python: Bluestein chirps, digit planes, complex FFT convolutions, a signed ADC with 2ENOB codes, rounding and recombination. It is checked against a direct modular DFT. Here q = 12,289 (14-bit limbs) and the block is 16 points:
Digit bits b
Digits d
Predicted ENOB
ENOB used
Worst analogue error
Result
1
14
9
9
0.375
exact
1
14
9
8
0.750
wrong
2
7
11
11
0.484
exact
2
7
11
10
0.969
wrong
3
5
13
13
0.473
exact
3
5
13
12
0.953
wrong
The rule is tight: exact at the predicted ENOB, wrong one bit below, for every digit width. The hardware model's choice of b is tested against this functional model, the analogue counterpart of checking a cycle model against RTL.
07
The Precision Tax
What the rule implies for real limb widths (grouped planes; 2(3d−1) conversions per point; from examples/results.py):
Limb bits
Block
ENOB 8
ENOB 12
ENOB 16
ENOB 20
Butterflies per point offloaded
50
16
infeasible
b=1, d=50: 298
b=3, d=17: 100
b=5, d=10: 58
2
50
256
infeasible
infeasible
b=1, d=50: 298
b=3, d=17: 100
4
50
4,096
infeasible
infeasible
infeasible
b=1, d=50: 298
6
36
16
infeasible
b=1, d=36: 214
b=4, d=9: 52
b=6, d=6: 34
2
28
16
infeasible
b=2, d=14: 82
b=4, d=7: 40
b=6, d=5: 28
2
28
256
infeasible
infeasible
b=2, d=14: 82
b=4, d=7: 40
4
Each cell shows digit width b, digit count d and conversions per transform point.
Tens to hundreds of conversions buy two to six butterflies. Bigger blocks offload more stages but need more ENOB; at 8 ENOB nothing is feasible for these limb widths.
This is the general lesson of deck 07 of the sister series made specific: the optical win must be amortised over many operations per conversion, and exact FHE arithmetic makes conversions per point grow with limb width.
08
Conversion Energy and Static Power
Converters: the Walden figure of merit
Energy per conversion ≈ FoM × 2ENOB. With illustrative FoMs of 10 fJ/step (DAC) and 20 fJ/step (ADC):
ENOB
DAC + ADC pJ
8
7.7
12
122.9
16
1,966.1
20
31,457.3
Each extra bit doubles the energy, and the precision rule needs more bits for wider blocks.
Static power
Lasers run whether or not work arrives; wall-plug efficiency sets their electrical cost.
Thermal tuning holds resonant devices on wavelength.
The model charges 10 W + 10 W (illustrative) for the whole run, like digital static power.
Converter throughput is also a power budget: at ENOB 16, 5×1011 samples/s would need about 1 kW, so under a 300 W TDP the model runs the engine at 5×1010.
A simple comparison: a digital butterfly costs 10 pJ in the illustrative model, so one ENOB-8 conversion pair (7.7 pJ) is already comparable to the work it replaces, before precision multiplies the conversion count.
09
Interactive: Optical Engine Explorer
Set the engine and run a full ark-set bootstrap on the chosen digital design, with and without it. "Ideal" is a hypothetical engine that is exact at any precision (one plane in, one out), an upper bound on what optics could do here.
12
16
20
20
bootstrap latency (ms)
energy per bootstrap (mJ), stacked by consumer
10
When It Wins and When It Loses
Full ark-set bootstraps (baseline algorithm), from examples/results.py:
+ realistic engine: ENOB 16, 36-bit limbs, rate cut to fit the TDP
620.39 ms
89,183 mJ
bound by the optical engine
+ ideal engine: exact at any precision, 4,096-point blocks
12.59 ms
1,375 mJ
MAC-bound
+ ideal engine, 4× converter rate
11.63 ms
1,318 mJ
MAC-bound
ARK-class (memory-bound), no optics
13.94 ms
1,139 mJ
memory-bound
ARK-class + ideal engine
16.87 ms
1,632 mJ
memory-bound
Loses, with exact rounding at realistic ENOB: 23× slower and 37× more energy, because conversions per point dwarf the butterflies saved.
Wins on speed only in the ideal case, and only for an NTT-bound design (17.46 → 12.59 ms, 1.39×), where it moves the bound to the MAC lanes. Energy still rises, because lasers and tuning add static power.
Does nothing for a memory-bound design. On the ARK-class design it is worse, because at this converter rate it is slower than ARK-class digital NTT units and the design is waiting on keys anyway.
Energy break-even for the ideal engine (ENOB 8, energy question only): a conversion must cost under 30 pJ with 5 W of laser and tuning power, or 47 pJ with none; with 20 W it never breaks even.
11
What Would Change the Verdict
Precision per conversion. Only exactness forces digit planes. Workloads that tolerate approximation (CKKS encoding and decoding use complex FFTs; ML layers outside the encrypted domain) avoid the tax entirely. The sister series Fourier Optics for Inference works through that float case for LLM inference: a noise rule instead of exactness, averaging passes instead of digit planes, and a simulated optical prefill pool.
More work per conversion. Chaining optical stages, or doing a whole length-N transform per pass, raises butterflies offloaded per point, but needs more ENOB at the output.
Different mappings. Optical matrix-vector engines could take on base conversion, which is a matrix product, or polynomial products as convolutions; PHAT's TFHE design maps FFT-based polynomial multiplication. Each needs its own precision analysis.
Converter technology. The break-even energy above is a target: about 30 pJ per conversion pair at 8 ENOB, together with low static power.
The memory system first. As decks 03 and 05 show, a balanced FHE accelerator is limited by keys and plaintexts before NTTs; photonic interconnect (OptoLink's direction) addresses that bottleneck directly.
The general point
A simulator that charges for precision and conversion is what separates "optics computes FFTs for free" from a design decision. This deck's model is first-order and its coefficients are illustrative, but the structure (offload fraction × conversions per point × energy per conversion ± static power, set against a bound attribution) is the right shape for that argument.
12
What to Take Away
Optics computes complex FFTs and convolutions; FHE needs exact modular NTTs. The bridge is integer convolution on digit planes, here via Bluestein.
The exact-rounding rule 2ENOB−1 > d·n·(2b−1)2 is tight, as the functional model shows.
The precision tax: at 50-bit limbs and 12 ENOB, 298 conversions per point to offload 2 butterflies.
Converter energy grows as 2ENOB, and laser and tuning power is paid all the time.
Simulated: a realistic engine loses badly; an ideal one helps only NTT-bound designs and costs energy; neither helps a memory-bound design.
Next
Deck 05 collects the design-space results: SRAM, regimes, power and the techniques that matter most.