FHE Accelerator Simulators — Presentation 04

Optical NTT Engines: Precision, Conversion and Power

Fourier transforms in optics, and what it takes to map an exact modular NTT onto an analogue complex FFT: digit decomposition, small-prime RNS, rounding and error correction, ENOB, DAC/ADC energy from a Walden figure of merit, and laser and thermal-tuning power. The simulator then shows when an optical engine wins and when it loses.

Fourier optics Digit planes Bluestein ENOB Walden FoM Laser power
Digits → DAC → Optical FFT → ADC → Round → mod q
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Cryptography decks or LLM Inference Simulators.

01

Fourier Transforms in Optics

A lens computes a two-dimensional Fourier transform physically: the optical field in its front focal plane appears, Fourier-transformed, in its back focal plane. Two lenses with a mask between them (a 4f system) multiply in the frequency domain, which is a convolution in space. The transform itself costs time-of-flight and no switching energy.

input planeDAC + modulators lens 1 Fourier planemask: multiply lens 2 output planedetectors + ADC ffff
02

The Mismatch: Exact Modular Arithmetic

What CKKS needs

  • An NTT over Zq with q a 28–60-bit prime: twiddles are integers mod q, results must be exact.
  • One wrong coefficient is a wrong ciphertext. CKKS tolerates approximation in the message, not in the ring arithmetic.

What optics provides

  • A complex DFT with unit-modulus twiddles, in analogue, with an effective precision (ENOB) set by noise, linearity and the converters: typically far fewer bits than a limb.
  • No modular reduction, no integers.

The NTT of a limb is not a complex DFT, so the optics cannot run it directly. What it can compute exactly, given enough precision, is an integer convolution: convolve small non-negative integers with a complex FFT and round. Every route below turns modular NTT work into integer convolutions small enough that rounding is exact. The same trick lets TFHE libraries multiply polynomials with 64-bit floating-point FFTs.

03

Route 1: Digit Decomposition

Split each operand into d digits of b bits, x = Σi xi·2bi. A convolution of the full operands becomes d2 convolutions of digit planes, recombined digitally with shifts and adds.

Precedent

PHAT reaches the 42-bit FFT precision its TFHE workload needs from 6-bit optical cells by exactly this kind of multi-word decomposition ("7 × 6 + 1 = 43 bits of precision (including 1 bit for sign)", arXiv:2609.11613).

04

Route 2: A Hybrid NTT via Bluestein

The mapping the companion simulator models, one illustrative route among several:

Digital NTT stages
log2(N/n) butterfly stages split XN+1 into length-n pieces
→
Bluestein chirp
each length-n modular DFT becomes a size-2n cyclic convolution
→
Optical FFT convolution
on b-bit digit planes: d planes in, 2d−1 out
→
Digital correction
round, recombine, chirp, reduce mod q
05

Route 3: Small Primes and Error Correction

IdeaWhat it buysWhat it costsIn the model?
Small-prime RNSNarrower limbs (CraterLake uses 28-bit words, SHARP 36-bit) mean fewer digits per operandMore limbs for the same modulus, so more NTTs; rescaling needs care (scale ≠ prime)Yes, via limb bits (the optical precision depends on it)
Rounding with redundancyCompute in a slightly bigger modulus or with a check residue; detect and correct rare analogue errors digitallyExtra digital work and some extra planesNo; the model requires worst-case exactness
Probabilistic exactnessSize the ENOB for typical rather than worst-case error, then rely on detectionA failure-rate argument a security reviewer must acceptNo
Fewer conversions per pointChain several optical stages without leaving the optical domainError accumulates across stages; needs higher ENOB at the endNo; an extension exercise

CraterLake's 28-bit words come from ARK's comparison table (arXiv:2205.00922, Table VII); SHARP's 36-bit words and 180 MB of on-chip memory are as reported in FHEmem (arXiv:2311.16293) and Osiris (arXiv:2408.09593).

06

The Precision Rule, Checked Functionally

precision.py implements route 2 end to end in pure Python: Bluestein chirps, digit planes, complex FFT convolutions, a signed ADC with 2ENOB codes, rounding and recombination. It is checked against a direct modular DFT. Here q = 12,289 (14-bit limbs) and the block is 16 points:

Digit bits bDigits dPredicted ENOBENOB usedWorst analogue errorResult
114990.375exact
114980.750wrong
2711110.484exact
2711100.969wrong
3513130.473exact
3513120.953wrong

The rule is tight: exact at the predicted ENOB, wrong one bit below, for every digit width. The hardware model's choice of b is tested against this functional model, the analogue counterpart of checking a cycle model against RTL.

07

The Precision Tax

What the rule implies for real limb widths (grouped planes; 2(3d−1) conversions per point; from examples/results.py):

Limb bitsBlockENOB 8ENOB 12ENOB 16ENOB 20Butterflies per point offloaded
5016infeasibleb=1, d=50: 298b=3, d=17: 100b=5, d=10: 582
50256infeasibleinfeasibleb=1, d=50: 298b=3, d=17: 1004
504,096infeasibleinfeasibleinfeasibleb=1, d=50: 2986
3616infeasibleb=1, d=36: 214b=4, d=9: 52b=6, d=6: 342
2816infeasibleb=2, d=14: 82b=4, d=7: 40b=6, d=5: 282
28256infeasibleinfeasibleb=2, d=14: 82b=4, d=7: 404
08

Conversion Energy and Static Power

Converters: the Walden figure of merit

Energy per conversion ≈ FoM × 2ENOB. With illustrative FoMs of 10 fJ/step (DAC) and 20 fJ/step (ADC):

ENOBDAC + ADC pJ
87.7
12122.9
161,966.1
2031,457.3

Each extra bit doubles the energy, and the precision rule needs more bits for wider blocks.

Static power

  • Lasers run whether or not work arrives; wall-plug efficiency sets their electrical cost.
  • Thermal tuning holds resonant devices on wavelength.
  • The model charges 10 W + 10 W (illustrative) for the whole run, like digital static power.
  • Converter throughput is also a power budget: at ENOB 16, 5×1011 samples/s would need about 1 kW, so under a 300 W TDP the model runs the engine at 5×1010.

A simple comparison: a digital butterfly costs 10 pJ in the illustrative model, so one ENOB-8 conversion pair (7.7 pJ) is already comparable to the work it replaces, before precision multiplies the conversion count.

09

Interactive: Optical Engine Explorer

Set the engine and run a full ark-set bootstrap on the chosen digital design, with and without it. "Ideal" is a hypothetical engine that is exact at any precision (one plane in, one out), an upper bound on what optics could do here.

12
16
20
20
bootstrap latency (ms)
energy per bootstrap (mJ), stacked by consumer
10

When It Wins and When It Loses

Full ark-set bootstraps (baseline algorithm), from examples/results.py:

DesignBootstrapEnergyVerdict
Small digital (NTT-starved), no optics17.46 ms1,280 mJNTT-bound
+ realistic engine: 16-point blocks, ENOB 12, 50-bit limbs401.38 ms46,861 mJbound by the optical engine
+ realistic engine: ENOB 16, 36-bit limbs, rate cut to fit the TDP620.39 ms89,183 mJbound by the optical engine
+ ideal engine: exact at any precision, 4,096-point blocks12.59 ms1,375 mJMAC-bound
+ ideal engine, 4× converter rate11.63 ms1,318 mJMAC-bound
ARK-class (memory-bound), no optics13.94 ms1,139 mJmemory-bound
ARK-class + ideal engine16.87 ms1,632 mJmemory-bound
11

What Would Change the Verdict

The general point

A simulator that charges for precision and conversion is what separates "optics computes FFTs for free" from a design decision. This deck's model is first-order and its coefficients are illustrative, but the structure (offload fraction × conversions per point × energy per conversion ± static power, set against a bound attribution) is the right shape for that argument.

12

What to Take Away

Next

Deck 05 collects the design-space results: SRAM, regimes, power and the techniques that matter most.