CKKS seen from the datapath, not the proof: RNS polynomials in Zq[X]/(XN+1), how big ciphertexts and evaluation keys really are, the five primitive kernels (NTT, modular multiply-add, automorphism, base conversion, rescale), noise and levels, and why key switching dominates. Includes a live parameter calculator.
Concepts used here, and where they are explained. Each links to its series glossary entry: a short explanation, then links to the slides that explain it in depth, in this series, the Cryptography decks or LLM Inference Simulators.
Fully homomorphic encryption lets a server compute on data it cannot read. The price is speed: the F1 paper (Feldmann et al., MICRO 2021) puts software FHE at four to five orders of magnitude slower than computing on plaintext. Closing that gap is an architecture problem, so it is a simulation problem.
What makes it hard for hardware
Every value is a polynomial with tens of thousands of coefficients, stored as dozens of machine-word residues.
Multiplying two ciphertexts needs a key switch: hundreds of transforms and a read of an evaluation key over 100 MiB in size.
Noise accumulates; bootstrapping resets it, at a cost of hundreds of key switches (deck 02).
What this series does with it
Treats CKKS as a workload: shapes, sizes, kernel counts and bytes.
Simulates an accelerator running bootstrapping in SimPy (deck 03), live in the browser.
Asks whether an optical transform engine helps (deck 04) and what the design space looks like (deck 05).
Scope
This deck is about cost, not security. It uses the cryptography only as far as a datapath designer needs it, and points to the Cryptography series for the rest.
02
Fundamentals: Where to Learn the Mathematics
The mathematics is covered elsewhere on this GitHub, and this series stays consistent with its notation (one change: the ring degree is written N here, as in the bootstrapping literature, where Cryptography deck 08 writes n).
Cryptography 08: Fully Homomorphic EncryptionLWE and Ring-LWE, noise growth, BGV/BFV/CKKS/TFHE, relinearisation and key switching, RNS and NTT, libraries and compilers. Read this first if "Ring-LWE" is new.
A CKKS ciphertext is a pair of polynomials (c0, c1) in RQ = ZQ[X]/(XN+1). Q is a product of word-sized primes, Q = q0q1…qℓ, so each polynomial is stored as one row of N residues per prime: the residue number system (RNS). Arithmetic never needs multi-word integers.
Two forms. Coefficient form, or evaluation form after a number-theoretic transform (NTT) of each limb. Multiplication is element-wise in evaluation form, so ciphertexts normally live there.
Two levels of parallelism. N coefficients inside a limb, and limbs across primes. Both map naturally onto wide vector datapaths.
04
How Big Things Are
Sizes for the parameter sets used throughout the series, as computed by params.py. The ARK and Lattigo rows match Table III of ARK (Kim et al., MICRO 2022, arXiv:2205.00922) exactly.
Set
N
L
dnum
α
Ciphertext (top)
Evaluation key
log PQ (approx.)
ark
216
23
4
6
24.0 MiB
120.0 MiB
1,570
lattigo
216
24
5
5
25.0 MiB
150.0 MiB
1,560
gpu100x
216
34
5
7
35.0 MiB
210.0 MiB
2,180
openfhe-sparse
216
18
3
7
19.0 MiB
78.0 MiB
1,542
Ciphertext = 2 × N × (ℓ+1) × 8 bytes; it shrinks by one limb every level.
Evaluation key = dnum digits × 2 polynomials × (L+1+α) limbs × N × 8 bytes. The α extra limbs are the special primes P that key switching works in.
One key per rotation amount. A bootstrap uses dozens of distinct rotation keys, so gigabytes of key material per bootstrap (deck 02).
Security constrains log PQ for a given N. For 128-bit security at N = 216, total modulus sizes in the 1,500–1,800-bit range are typical; check the HomomorphicEncryption.org tables or your library's parameter checker rather than this approximation.
05
The Primitive Kernels
Every HE operation decomposes into five kernels. These are the units a hardware designer builds and a simulator models.
Kernel
What it does
Work per limb
Hardware
NTT / iNTT
Coefficient ↔ evaluation form; the finite-field FFT
(N/2) log2N butterflies: 524,288 at N = 216
Pipelined butterfly arrays (deck 10 of Cryptography)
Wide SIMD lanes with Barrett or Montgomery reduction
Automorphism
X → X5r: permutes coefficients to rotate slots
N words moved
A permutation network; no arithmetic
Base conversion
Re-expresses a polynomial from one prime set in another (ModUp, ModDown)
N × (input limbs) multiply-adds per output limb
Matrix-vector datapath; often the second-largest cost
Rescale
Divides by the last prime, dropping a limb
An iNTT, an NTT per remaining limb, a multiply-add
Built from the above
The companion simulator's trace format (deck 03) is exactly this list: every HE operation becomes a sequence of these kernels, each with an amount of work and a number of on-chip words touched.
06
Noise, Scale and Levels
CKKS in four lines
A real or complex vector is encoded with a scale Δ (for example 250), so the plaintext carries fixed-point numbers.
Encryption adds small noise; CKKS treats it as part of the approximation error.
A multiplication doubles the scale (Δ2). Rescale divides by a prime qℓ ≈ Δ to bring it back, and drops one limb.
So each multiplicative level costs one prime. A ciphertext at level 0 has one limb left and can no longer be multiplied.
What that means for hardware
Work shrinks with depth: an operation at level ℓ touches ℓ+1 limbs, so the same HMult is cheaper late in a computation.
L is a budget: more levels mean bigger ciphertexts and keys for every operation, not only the deep ones.
Bootstrapping raises a level-0 ciphertext back to a high level, but spends about 16 of the L = 23 levels in our model doing so (deck 02). The useful levels between bootstraps are what applications are left with.
The level is part of the workload
A cost model that ignores the level an operation runs at is wrong by up to a factor of L+1. The companion simulator tracks the level of every ciphertext object.
07
Key Switching, Step by Step
After a multiplication (relinearisation) or an automorphism (rotation), part of the ciphertext is "encrypted" under the wrong key. Key switching fixes it using an evaluation key. The hybrid method (Han & Ki, CT-RSA 2020, ePrint 2019/688) splits the limbs into dnum digits of α limbs each:
Transforms per key switch at level ℓ (with ℓ+1 limbs, k = α special primes and β = ⌈(ℓ+1)/α⌉ digits): (ℓ+1) + Σdigits(ℓ+1+k−αi) + 2k + 2(ℓ+1). For the ark set at the top level: 24 + 96 + 12 + 48 = 180.
08
Why Key Switching Dominates
One op, ark set, top level
NTT + iNTT limbs
Base-conversion multiply-adds
Other multiply-adds
Key loaded
Time on the ARK-class model
HMult (with relinearisation and rescale)
228
62.1 M
28.2 M
120 MiB
263.4 µs
HRotate
180
62.1 M
15.7 M
120 MiB
220.7 µs
Everything except the key switch is cheap. The tensor product of an HMult is 4(ℓ+1)N multiply-adds; the key switch adds 180 limb transforms and tens of millions of base-conversion operations.
The key is read once per use. 120 MiB per operation is 5× the ciphertext it operates on. At 1 TB/s that is 126 µs of HBM time alone, about half of the operation's time on this design.
Distinct keys defeat caches. Rotations by different amounts need different keys. On-chip SRAM helps only if a key is reused before it is evicted, which the algorithm has to arrange (deck 05).
Single-op times above are isolated operations on an otherwise idle machine; in a bootstrap, operations overlap and the binding resource is usually HBM.
The one-sentence version
FHE accelerators are NTT engines welded to a memory system that must stream evaluation keys, and the hard part is the memory system.
09
Interactive: Parameter Calculator
Choose N, L, dnum and the level an operation runs at. Sizes come from the same formulas as params.py; kernel counts come from the simulator's own trace generator (the JavaScript port, tested to match the Python exactly).
16
23
4
23
50
10
The dnum Trade-Off
dnum sets how many digits the key-switch input is split into. Fewer digits mean fewer, wider special primes: smaller keys and less work, but a larger total modulus PQ, which at fixed N reduces security or levels. From examples/results.py (N = 216, L = 23, top level):
dnum
α = k
log PQ
Evaluation key
Limb transforms per HRot
Base-conversion multiply-adds
HRot on ARK-class
1
24
2,650
48 MiB
144
121 M
143 µs
2
12
1,930
72 MiB
144
82 M
164 µs
4
6
1,570
120 MiB
180
62 M
221 µs
8
3
1,390
216 MiB
270
52 M
341 µs
24
1
1,270
600 MiB
650
46 M
832 µs
Small dnum is cheap but spends modulus bits on special primes; large dnum keeps PQ close to Q but multiplies key size and NTT count.
The 100x GPU paper (Jung et al., TCHES 2021, ePrint 2021/508) measured a 5.88× faster HMult on a V100 with dnum = 3 than with the maximum dnum = 45 (at a somewhat smaller modulus), because memory traffic, not multiply count, was the bottleneck. Parameter choice is an architectural decision, and a simulator has to sweep it.
11
What a Simulator Needs from All This
Must model
Because
In the companion simulator
Level of every ciphertext
Work and bytes scale with ℓ+1
Trace.levels; every kernel sized at its level
Key identity, not just key size
Reuse decides whether SRAM helps
Keys are named objects in the scratchpad
NTT, base conversion, multiply-add, automorphism separately
They map to different units with different throughputs
Four kernel kinds, three digital units plus an optional optical unit
Dependencies
Bootstrapping is a DAG with limited parallelism
Producer/consumer edges between HE ops
Parameter sets as data
N, L and dnum are design choices to sweep
CKKSParams presets and the calculator above
12
What to Take Away
A ciphertext is a matrix: 2 polynomials × (ℓ+1) RNS limbs × N words. At N = 216, L = 23 that is 24 MiB.
Five kernels (NTT, multiply-add, automorphism, base conversion, rescale) make up every HE operation.
Key switching dominates: 180–228 limb transforms plus a 120 MiB key read per HMult or HRot at the top level.
Levels are a budget and every operation's cost depends on its level.
dnum, L and N are architectural parameters, traded between work, key size and security.
Next
Deck 02 opens up bootstrapping, the operation that strings hundreds of these key switches together and turns FHE into a memory-bandwidth problem.