Why bandwidth times efficiency fails, and a command-level model that derives the efficiency instead: banks, rows and bank groups, the timing parameters, hits, misses and conflicts, FCFS and FR-FCFS, page policy, address mapping, refresh and tFAW, load and latency, validation four ways (including against DRAMsim3), and the model plugged into the FHE simulator.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.
Architecture simulators usually model memory as a pipe: bytes divided by peak bandwidth, sometimes times a derating factor of "about 70%". The FHE simulator's HBM works this way. The factor is not a property of the memory. It depends on the access pattern, the controller, the address mapping and the load. On a single HBM pseudo-channel with a single controller:
Achieved efficiency ranges from 0.10 to 0.93 of peak.
A stream reuses open rows; random traffic opens a row for every access and is capped by how fast rows can be opened.
The same HBM stream runs at 0.56 or 0.93 of peak depending on the scheduler; a DDR4 stream at 0.53 or 1.00 depending on where the bank-group bits sit.
Latency is flat until the queue fills, then climbs steeply. A fixed factor has no notion of queueing.
This deck builds a command-level model that derives the factor, checks it four ways, and plugs it into the FHE simulator. The code is Memory_System_Sim; every number comes from its examples/results.md.
Parallelism comes from banks: while one bank opens a row, others can transfer data. A DDR4 bank group shares some internal circuitry, so back-to-back accesses within a group must be further apart than accesses across groups. HBM2-class stacks have 8 channels, each split into two pseudo-channels with their own banks and a 64-bit data bus.
The controller issues ACT (open a row), RD/WR (one burst), PRE (close) and REF (refresh), at most one per cycle on the command bus. The JEDEC standards (JESD79-4 for DDR4, JESD235 for HBM) name the minimum gaps between commands. The values below are illustrative: in the range of a DDR4-3200 speed bin and an HBM2E-class pseudo-channel, not quoted from the standards. The DDR4 values coincide with DRAMsim3's DDR4_8Gb_x8_3200 configuration, which slide 11 uses.
| parameter | DDR4 | HBM |
|---|---|---|
| tck_ns | 0.625 | 0.625 |
| bus_bytes | 8 | 8 |
| bl | 8 | 4 |
| bankgroups | 4 | 4 |
| banks | 4 | 4 |
| row_bytes | 8192 | 1024 |
| CL | 22 | 23 |
| CWL | 16 | 7 |
| RCD | 22 | 23 |
| RP | 22 | 23 |
| RAS | 52 | 53 |
| RC | 74 | 76 |
| RRD_S | 4 | 4 |
| RRD_L | 8 | 7 |
| FAW | 34 | 26 |
| CCD_S | 4 | 2 |
| CCD_L | 8 | 4 |
| WR | 24 | 26 |
| WTR_S | 4 | 4 |
| WTR_L | 12 | 12 |
| RTP | 12 | 8 |
| RFC | 560 | 560 |
| REFI | 12480 | 6240 |
| peak per channel, GB/s | 25.6 | 25.6 |
Source: examples/results.md in Memory_System_Sim
The controller computes, for each candidate command, the earliest cycle at which every constraint it is subject to holds. Per bank group, not just per rank: that distinction was a bug (slide 10).
def t_act(self, ra: int, bg: int, b: _Bank) -> int:
t, rk = self.t, self.ranks[ra]
e = max(b.act_ok, rk.ref_until, rk.last_act + t.RRD_S, rk.act_bg[bg] + t.RRD_L)
if len(rk.acts) == 4:
e = max(e, rk.acts[0] + t.FAW)
return e
The row is already open: RD, then data after CL. CL + BL/2.
The bank is closed: ACT, wait RCD, RD. RCD + CL + BL/2.
Another row is open: PRE, wait RP, ACT, wait RCD, RD. RP + RCD + CL + BL/2.
Unloaded latencies, simulated against those closed forms, plus the bandwidth limits of the next three slides:
| device | check | simulated | closed form | formula |
|---|---|---|---|---|
| DDR4 | read, bank closed (cycles) | 48 | 48 | RCD + CL + BL/2 |
| DDR4 | write, bank closed (cycles) | 42 | 42 | RCD + CWL + BL/2 |
| DDR4 | read, row hit (cycles) | 26 | 26 | CL + BL/2 |
| DDR4 | read, row conflict (cycles) | 70 | 70 | RP + RCD + CL + BL/2 |
| DDR4 | streaming, no refresh (fraction of peak) | 0.999 | 1 | peak |
| DDR4 | refresh loss on a stream | 0.037 | 0.045 | RFC / REFI |
| DDR4 | same bank, next row each access | 0.054 | 0.054 | (BL/2) / RC |
| DDR4 | random reads, closed page | 0.457 | 0.471 | 4 (BL/2) / FAW (activate-limited bound) |
| HBM | read, bank closed (cycles) | 48 | 48 | RCD + CL + BL/2 |
| HBM | write, bank closed (cycles) | 32 | 32 | RCD + CWL + BL/2 |
| HBM | read, row hit (cycles) | 25 | 25 | CL + BL/2 |
| HBM | read, row conflict (cycles) | 71 | 71 | RP + RCD + CL + BL/2 |
| HBM | streaming, no refresh (fraction of peak) | 0.998 | 1 | peak |
| HBM | refresh loss on a stream | 0.094 | 0.090 | RFC / REFI |
| HBM | same bank, next row each access | 0.026 | 0.026 | (BL/2) / RC |
| HBM | random reads, closed page | 0.304 | 0.308 | 4 (BL/2) / FAW (activate-limited bound) |
Source: examples/results.md in Memory_System_Sim
Every latency matches exactly. A conflict costs almost three times a hit on this DDR4, so a controller's main job is to turn conflicts into hits by choosing the order of requests.
| pattern | FCFS open | FCFS closed | FR-FCFS open | FR-FCFS closed |
|---|---|---|---|---|
| streaming | 0.563 (0.97) | 0.679 (0.97) | 0.925 (0.97) | 0.926 (0.97) |
| streaming, 1 in 3 writes | 0.147 (0.96) | 0.149 (0.96) | 0.804 (0.97) | 0.803 (0.97) |
| random (1 GiB) | 0.039 (0.00) | 0.062 (0.00) | 0.270 (0.00) | 0.276 (0.00) |
| random, 1 in 3 writes | 0.039 (0.00) | 0.061 (0.00) | 0.230 (0.00) | 0.234 (0.00) |
| 4 interleaved streams | 0.041 (0.14) | 0.042 (0.14) | 0.484 (0.80) | 0.449 (0.73) |
| stride 4 KiB | 0.039 (0.00) | 0.075 (0.00) | 0.095 (0.00) | 0.095 (0.00) |
Source: examples/results.md in Memory_System_Sim
FR-FCFS wins everywhere, by up to 12× on interleaved streams, whose rows FCFS keeps closing. Writes cost a stream 13% even under FR-FCFS, because the bus has to turn around and a read must wait tWTR after a write. Page policy matters less than the scheduler here: with a full queue, the controller already knows whether another hit is coming.
The mapping decides which address bits select the channel, bank group, bank, row and column, and so which patterns spread across banks:
SCHEMES = {
"RoRaBaCoBgCh": ("ro", "ra", "ba", "co", "bg", "ch"), # default: a stream rotates over bank groups
"RoRaBgBaCoCh": ("ro", "ra", "bg", "ba", "co", "ch"), # a stream stays in one bank group (tCCD_L)
"RoCoRaBgBaCh": ("ro", "co", "ra", "bg", "ba", "ch"), # consecutive bursts rotate over banks
"ChRaBgBaRoCo": ("ch", "ra", "bg", "ba", "ro", "co"), # one bank per huge region (worst parallelism)
}
| mapping (MSB to LSB) | streaming | stride = one row (same bank) | stride 8 KiB | random |
|---|---|---|---|---|
| RoRaBaCoBgCh | 0.998 | 0.054 | 0.499 | 0.456 |
| RoRaBaCoBgCh + XOR | 0.996 | 0.456 | 0.841 | 0.457 |
| RoRaBgBaCoCh | 0.526 | 0.054 | 0.457 | 0.455 |
| RoRaBgBaCoCh + XOR | 0.526 | 0.456 | 0.457 | 0.456 |
| RoCoRaBgBaCh | 0.998 | 0.054 | 0.363 | 0.456 |
| RoCoRaBgBaCh + XOR | 0.998 | 0.456 | 0.841 | 0.456 |
Source: examples/results.md in Memory_System_Sim
Every tREFI the rank closes its banks and is busy for tRFC. A stream therefore loses about tRFC / tREFI: 4.5% on this DDR4, 9% on this HBM model, where the refresh interval is half as long. Measured on a stream: 3.7% and 9.4% (section 2 of the results).
At most four ACTs in any window of tFAW, a limit on peak current. Random traffic needs one ACT per burst, so its bandwidth is capped at 4 (BL/2) / tFAW of peak: 0.471 for DDR4, 0.308 for HBM. Closed-page random reads reach 0.457 and 0.304.
Two more bounds follow the same pattern. A stride that hits the same bank's next row each time runs at one access per tRC: simulated 0.054 and 0.026, against (BL/2) / tRC = 0.054 and 0.026. Without refresh, streaming reaches 0.999 of peak. These closed forms are tests in the repository (tests/test_controller.py), not just slide material.
Recorded results, not a live model. Pick a pattern and a policy to see the bandwidth it sustains. The chart shows mean and p99 latency against offered load for streaming and random traffic on one HBM pseudo-channel (FR-FCFS, open page). The dashed line marks the load you pick. Little's law (LLM Inference Simulators glossary) turns the latency into the number of requests in flight.
| offered load (fraction of peak) | streaming mean ns | streaming p99 ns | random mean ns | random p99 ns |
|---|---|---|---|---|
| 0.10 | 37 | 365 | 83 | 444 |
| 0.20 | 41 | 370 | 139 | 512 |
| 0.30 | 42 | 373 | 1250 | 2699 |
| 0.40 | 49 | 378 | ||
| 0.50 | 52 | 389 | ||
| 0.60 | 68 | 384 | ||
| 0.70 | 86 | 408 | ||
| 0.80 | 129 | 415 | ||
| 0.85 | 162 | 433 | ||
| 0.90 | 247 | 557 |
Source: examples/results.md in Memory_System_Sim
A memory model that ignores load would give the 37 ns number to a simulator running at 90% utilisation, and miss the tail completely.
The latencies and bounds of slides 04 and 07, asserted in tests/test_controller.py.
memsim.checker replays the command log and checks each rule as stated ("any two ACTs to one bank are at least tRC apart"), sharing no code with the controller.
Hypothesis generates random traces, schedulers, page policies, mappings, queue depths and refresh settings. Every command stream must pass the checker, and every request must complete.
DRAMsim3 on the same traces (next slide).
What it found. Hypothesis produced a trace with a write to bank group 0, then a write to group 1, then a read from group 0. The controller remembered only the last write, so it applied tWTR_S from the group-1 write. It should have applied tWTR_L from the group-0 write. The checker flagged "WR-RD same bank group (WTR_L): 28 < 32". The fix keeps per-group history; the shrunk example is now a regression test:
def test_regression_wtr_l_after_a_later_write_to_another_group():
"""Found by Hypothesis: WR (group 0), WR (group 1), RD (group 0). The read must wait WTR_L after
the group-0 write, not just WTR_S after the later group-1 write."""
t = PRESETS["ddr4"]
m = AddressMapping(t)
g0, g1 = 0, 1 << m.shift["bg"]
reqs = [Request(g0, write=True), Request(g1, write=True), Request(g0 + (1 << m.shift["co"]))]
checked(t, reqs, mapping=m, refresh=False)
wr0, rd = reqs[0], reqs[2]
assert rd.issue >= wr0.done + t.WTR_L
Who checks the checker? Its own tests plant violations in legal logs, and run a controller subclass with tFAW deliberately removed. The checker must flag every one.
DRAMsim3 (Li et al., IEEE CAL 2020) is a widely used cycle-accurate DRAM simulator. Both simulators ran the same traces, with the same timing (its DDR4_8Gb_x8_3200 config, one rank) and the same address mapping:
DRAMsim3 commit 2981759; 60,000 cycles per saturation run; throughput as a fraction of the data-bus peak.
| DRAMsim3 mapping | trace | metric | DRAMsim3 | memsim | memsim vs DRAMsim3 |
|---|---|---|---|---|---|
| rochrababgco | stream reads | throughput / peak | 0.670 | 0.619 | -7.6% |
| rochrababgco | stream, 1 in 3 writes | throughput / peak | 0.676 | 0.578 | -14.5% |
| rochrababgco | random reads (1 GiB) | throughput / peak | 0.450 | 0.440 | -2.2% |
| rochrababgco | random, 1 in 3 writes | throughput / peak | 0.445 | 0.442 | -0.7% |
| rochrababgco | same bank, next row each time | throughput / peak | 0.052 | 0.052 | +0.0% |
| rochrababgco | 4 interleaved streams | throughput / peak | 0.310 | 0.476 | +53.5% |
| rochrababgco | sparse random reads (latency) | mean read latency (cycles) | 73.0 | 72.1 | -1.2% |
| rochrabacobg | stream reads | throughput / peak | 0.955 | 0.958 | +0.3% |
| rochrabacobg | stream, 1 in 3 writes | throughput / peak | 0.903 | 0.911 | +1.0% |
| rochrabacobg | random reads (1 GiB) | throughput / peak | 0.450 | 0.438 | -2.8% |
| rochrabacobg | random, 1 in 3 writes | throughput / peak | 0.443 | 0.441 | -0.4% |
| rochrabacobg | same bank, next row each time | throughput / peak | 0.052 | 0.052 | +0.0% |
| rochrabacobg | 4 interleaved streams | throughput / peak | 0.814 | 0.684 | -15.9% |
| rochrabacobg | sparse random reads (latency) | mean read latency (cycles) | 70.7 | 69.9 | -1.2% |
Source: examples/results.md in Memory_System_Sim
A cross-check that agrees everywhere usually means the two models share assumptions. This one shows which results depend on the controller, and therefore need its design specified before they are quoted.
FHE_Accelerator_Sim moves HBM traffic as 4 MiB chunks through one first-come resource. It now accepts an optional memory model, Accelerator(memory=...), that times each chunk given its size, its direction and the direction of the chunk before it. Memory_System_Sim's adapter simulates one pseudo-channel's share of the chunk at command level, memoised:
def chunk_time(self, nbytes: int, write: bool, prev_write: bool | None, peak_gbps: float) -> float:
"""Seconds to move ``nbytes`` through HBM whose peak is ``peak_gbps`` (FHE_Accelerator_Sim protocol)."""
t = self.timing
channels = peak_gbps / t.channel_peak_gbps()
bursts = max(1, round(nbytes / t.access_bytes / channels))
return self.channel_cycles(bursts, write, prev_write) * t.tck_ns * 1e-9
| HBM model | baseline algorithm | Min-KS + seeded keys + OTF plaintexts |
|---|---|---|
| peak bandwidth (default) | 13.94 ms, 89%, memory-bound | 7.19 ms, 11%, MAC-bound |
| flat: 0.7 x bandwidth | 19.26 ms, 92%, memory-bound | 7.29 ms, 15%, MAC-bound |
| flat: 0.9 x bandwidth | 15.33 ms, 90%, memory-bound | 7.21 ms, 12%, MAC-bound |
| Memory_System_Sim: FR-FCFS, bank groups interleaved | 15.31 ms, 90%, memory-bound | 7.21 ms, 12%, MAC-bound |
| ... with 1 MiB chunks | 15.64 ms, 90%, memory-bound | 7.21 ms, 12%, MAC-bound |
| ... FCFS scheduler | 23.70 ms, 94%, memory-bound | 7.40 ms, 19%, MAC-bound |
| ... a row's bursts in one bank group | 24.43 ms, 94%, memory-bound | 7.42 ms, 19%, MAC-bound |
Source: examples/results.md in FHE_Accelerator_Sim
| HBM peak | peak bandwidth | flat 0.7 | flat 0.9 | Memory_System_Sim |
|---|---|---|---|---|
| 140 GB/s | 10.23 ms, memory-bound | 12.33 ms, memory-bound | 10.78 ms, memory-bound | 10.73 ms, memory-bound |
| 160 GB/s | 9.66 ms, MAC-bound | 11.46 ms, memory-bound | 10.10 ms, memory-bound | 10.06 ms, memory-bound |
| 200 GB/s | 8.93 ms, MAC-bound | 10.23 ms, memory-bound | 9.25 ms, MAC-bound | 9.23 ms, MAC-bound |
| 250 GB/s | 8.33 ms, MAC-bound | 9.34 ms, MAC-bound | 8.57 ms, MAC-bound | 8.56 ms, MAC-bound |
Source: examples/results.md in FHE_Accelerator_Sim
Reading it. FHE traffic is long sequential streams, so a good controller reaches about 0.90 of peak, mostly lost to refresh. A well-chosen flat 0.9 then agrees to 0.2%. A folklore 0.7 adds 38% to the bootstrap. An FCFS controller or a bank-group-unfriendly mapping adds 70–75%. Near the balance point the model changes the verdict: at 160 GB/s, peak bandwidth calls the optimised design MAC-bound and the HBM model calls it memory-bound. At 200 GB/s, flat 0.7 calls it memory-bound and the model agrees with peak that it is MAC-bound.
examples/results.md.