Simulation Engineering Toolkit — Presentation 04

Modelling Memory Systems: DRAM and HBM

Why bandwidth times efficiency fails, and a command-level model that derives the efficiency instead: banks, rows and bank groups, the timing parameters, hits, misses and conflicts, FCFS and FR-FCFS, page policy, address mapping, refresh and tFAW, load and latency, validation four ways (including against DRAMsim3), and the model plugged into the FHE simulator.

DRAM timing HBM FR-FCFS Address mapping Refresh DRAMsim3 Protocol checker
Map → Schedule → Constrain → Check → Cross-check → Plug in
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

01

Why "Bandwidth × Efficiency" Fails

Architecture simulators usually model memory as a pipe: bytes divided by peak bandwidth, sometimes times a derating factor of "about 70%". The FHE simulator's HBM works this way. The factor is not a property of the memory. It depends on the access pattern, the controller, the address mapping and the load. On a single HBM pseudo-channel with a single controller:

One device, one controller, six patterns (results.md §6)

Achieved efficiency ranges from 0.10 to 0.93 of peak.

Pattern

A stream reuses open rows; random traffic opens a row for every access and is capped by how fast rows can be opened.

Controller and mapping

The same HBM stream runs at 0.56 or 0.93 of peak depending on the scheduler; a DDR4 stream at 0.53 or 1.00 depending on where the bank-group bits sit.

Load

Latency is flat until the queue fills, then climbs steeply. A fixed factor has no notion of queueing.

This deck builds a command-level model that derives the factor, checks it four ways, and plugs it into the FHE simulator. The code is Memory_System_Sim; every number comes from its examples/results.md.

02

How DRAM Is Organised

channel: one command bus, one data bus rank (all chips answering together) bank group 0 bank group 1 bank group 2 bank group 3 4 banks each4 banks each one bank rows (65,536 x 8 KiB per rank) row buffer: the open row ACT copies a row into the buffer, RD/WR move one burst from it, PRE writes it back and closes it. HBM: each channel is two64-bit pseudo-channels.

Parallelism comes from banks: while one bank opens a row, others can transfer data. A DDR4 bank group shares some internal circuitry, so back-to-back accesses within a group must be further apart than accesses across groups. HBM2-class stacks have 8 channels, each split into two pseudo-channels with their own banks and a 64-bit data bus.

03

Commands and Timing Parameters

The controller issues ACT (open a row), RD/WR (one burst), PRE (close) and REF (refresh), at most one per cycle on the command bus. The JEDEC standards (JESD79-4 for DDR4, JESD235 for HBM) name the minimum gaps between commands. The values below are illustrative: in the range of a DDR4-3200 speed bin and an HBM2E-class pseudo-channel, not quoted from the standards. The DDR4 values coincide with DRAMsim3's DDR4_8Gb_x8_3200 configuration, which slide 11 uses.

parameterDDR4HBM
tck_ns0.6250.625
bus_bytes88
bl84
bankgroups44
banks44
row_bytes81921024
CL2223
CWL167
RCD2223
RP2223
RAS5253
RC7476
RRD_S44
RRD_L87
FAW3426
CCD_S42
CCD_L84
WR2426
WTR_S44
WTR_L1212
RTP128
RFC560560
REFI124806240
peak per channel, GB/s25.625.6

Source: examples/results.md in Memory_System_Sim

The controller computes, for each candidate command, the earliest cycle at which every constraint it is subject to holds. Per bank group, not just per rank: that distinction was a bug (slide 10).

src/memsim/controller.py: when may this ACT issue? source
def t_act(self, ra: int, bg: int, b: _Bank) -> int:
    t, rk = self.t, self.ranks[ra]
    e = max(b.act_ok, rk.ref_until, rk.last_act + t.RRD_S, rk.act_bg[bg] + t.RRD_L)
    if len(rk.acts) == 4:
        e = max(e, rk.acts[0] + t.FAW)
    return e
04

Row Hits, Misses and Conflicts

Hit

The row is already open: RD, then data after CL. CL + BL/2.

Miss

The bank is closed: ACT, wait RCD, RD. RCD + CL + BL/2.

Conflict

Another row is open: PRE, wait RP, ACT, wait RCD, RD. RP + RCD + CL + BL/2.

Unloaded latencies, simulated against those closed forms, plus the bandwidth limits of the next three slides:

devicechecksimulatedclosed formformula
DDR4read, bank closed (cycles)4848RCD + CL + BL/2
DDR4write, bank closed (cycles)4242RCD + CWL + BL/2
DDR4read, row hit (cycles)2626CL + BL/2
DDR4read, row conflict (cycles)7070RP + RCD + CL + BL/2
DDR4streaming, no refresh (fraction of peak)0.9991peak
DDR4refresh loss on a stream0.0370.045RFC / REFI
DDR4same bank, next row each access0.0540.054(BL/2) / RC
DDR4random reads, closed page0.4570.4714 (BL/2) / FAW (activate-limited bound)
HBMread, bank closed (cycles)4848RCD + CL + BL/2
HBMwrite, bank closed (cycles)3232RCD + CWL + BL/2
HBMread, row hit (cycles)2525CL + BL/2
HBMread, row conflict (cycles)7171RP + RCD + CL + BL/2
HBMstreaming, no refresh (fraction of peak)0.9981peak
HBMrefresh loss on a stream0.0940.090RFC / REFI
HBMsame bank, next row each access0.0260.026(BL/2) / RC
HBMrandom reads, closed page0.3040.3084 (BL/2) / FAW (activate-limited bound)

Source: examples/results.md in Memory_System_Sim

Every latency matches exactly. A conflict costs almost three times a hit on this DDR4, so a controller's main job is to turn conflicts into hits by choosing the order of requests.

05

Scheduling and Page Policy

patternFCFS openFCFS closedFR-FCFS openFR-FCFS closed
streaming0.563 (0.97)0.679 (0.97)0.925 (0.97)0.926 (0.97)
streaming, 1 in 3 writes0.147 (0.96)0.149 (0.96)0.804 (0.97)0.803 (0.97)
random (1 GiB)0.039 (0.00)0.062 (0.00)0.270 (0.00)0.276 (0.00)
random, 1 in 3 writes0.039 (0.00)0.061 (0.00)0.230 (0.00)0.234 (0.00)
4 interleaved streams0.041 (0.14)0.042 (0.14)0.484 (0.80)0.449 (0.73)
stride 4 KiB0.039 (0.00)0.075 (0.00)0.095 (0.00)0.095 (0.00)

Source: examples/results.md in Memory_System_Sim

FR-FCFS wins everywhere, by up to 12× on interleaved streams, whose rows FCFS keeps closing. Writes cost a stream 13% even under FR-FCFS, because the bus has to turn around and a read must wait tWTR after a write. Page policy matters less than the scheduler here: with a full queue, the controller already knows whether another hit is coming.

06

Address Mapping

The mapping decides which address bits select the channel, bank group, bank, row and column, and so which patterns spread across banks:

src/memsim/mapping.py: named schemes, most significant field first source
SCHEMES = {
    "RoRaBaCoBgCh": ("ro", "ra", "ba", "co", "bg", "ch"),    # default: a stream rotates over bank groups
    "RoRaBgBaCoCh": ("ro", "ra", "bg", "ba", "co", "ch"),    # a stream stays in one bank group (tCCD_L)
    "RoCoRaBgBaCh": ("ro", "co", "ra", "bg", "ba", "ch"),    # consecutive bursts rotate over banks
    "ChRaBgBaRoCo": ("ch", "ra", "bg", "ba", "ro", "co"),    # one bank per huge region (worst parallelism)
}
mapping (MSB to LSB)streamingstride = one row (same bank)stride 8 KiBrandom
RoRaBaCoBgCh0.9980.0540.4990.456
RoRaBaCoBgCh + XOR0.9960.4560.8410.457
RoRaBgBaCoCh0.5260.0540.4570.455
RoRaBgBaCoCh + XOR0.5260.4560.4570.456
RoCoRaBgBaCh0.9980.0540.3630.456
RoCoRaBgBaCh + XOR0.9980.4560.8410.456

Source: examples/results.md in Memory_System_Sim

07

Refresh and the Four-Activate Window

Refresh

Every tREFI the rank closes its banks and is busy for tRFC. A stream therefore loses about tRFC / tREFI: 4.5% on this DDR4, 9% on this HBM model, where the refresh interval is half as long. Measured on a stream: 3.7% and 9.4% (section 2 of the results).

tFAW

At most four ACTs in any window of tFAW, a limit on peak current. Random traffic needs one ACT per burst, so its bandwidth is capped at 4 (BL/2) / tFAW of peak: 0.471 for DDR4, 0.308 for HBM. Closed-page random reads reach 0.457 and 0.304.

Two more bounds follow the same pattern. A stride that hits the same bank's next row each time runs at one access per tRC: simulated 0.054 and 0.026, against (BL/2) / tRC = 0.054 and 0.026. Without refresh, streaming reaches 0.999 of peak. These closed forms are tests in the repository (tests/test_controller.py), not just slide material.

08

Interactive: Patterns, Policies and Load

Recorded results, not a live model. Pick a pattern and a policy to see the bandwidth it sustains. The chart shows mean and p99 latency against offered load for streaming and random traffic on one HBM pseudo-channel (FR-FCFS, open page). The dashed line marks the load you pick. Little's law (LLM Inference Simulators glossary) turns the latency into the number of requests in flight.

09

Load and Latency

offered load (fraction of peak)streaming mean nsstreaming p99 nsrandom mean nsrandom p99 ns
0.103736583444
0.2041370139512
0.304237312502699
0.4049378
0.5052389
0.6068384
0.7086408
0.80129415
0.85162433
0.90247557

Source: examples/results.md in Memory_System_Sim

A memory model that ignores load would give the 37 ns number to a simulator running at 90% utilisation, and miss the tail completely.

10

How We Know the Model Is Right

1. Closed forms

The latencies and bounds of slides 04 and 07, asserted in tests/test_controller.py.

2. An independent protocol checker

memsim.checker replays the command log and checks each rule as stated ("any two ACTs to one bank are at least tRC apart"), sharing no code with the controller.

3. Property-based traces

Hypothesis generates random traces, schedulers, page policies, mappings, queue depths and refresh settings. Every command stream must pass the checker, and every request must complete.

4. Another simulator

DRAMsim3 on the same traces (next slide).

What it found. Hypothesis produced a trace with a write to bank group 0, then a write to group 1, then a read from group 0. The controller remembered only the last write, so it applied tWTR_S from the group-1 write. It should have applied tWTR_L from the group-0 write. The checker flagged "WR-RD same bank group (WTR_L): 28 < 32". The fix keeps per-group history; the shrunk example is now a regression test:

tests/test_controller.py source
def test_regression_wtr_l_after_a_later_write_to_another_group():
    """Found by Hypothesis: WR (group 0), WR (group 1), RD (group 0). The read must wait WTR_L after
    the group-0 write, not just WTR_S after the later group-1 write."""
    t = PRESETS["ddr4"]
    m = AddressMapping(t)
    g0, g1 = 0, 1 << m.shift["bg"]
    reqs = [Request(g0, write=True), Request(g1, write=True), Request(g0 + (1 << m.shift["co"]))]
    checked(t, reqs, mapping=m, refresh=False)
    wr0, rd = reqs[0], reqs[2]
    assert rd.issue >= wr0.done + t.WTR_L

Who checks the checker? Its own tests plant violations in legal logs, and run a controller subclass with tFAW deliberately removed. The checker must flag every one.

11

Cross-Checked Against DRAMsim3

DRAMsim3 (Li et al., IEEE CAL 2020) is a widely used cycle-accurate DRAM simulator. Both simulators ran the same traces, with the same timing (its DDR4_8Gb_x8_3200 config, one rank) and the same address mapping:

DRAMsim3 commit 2981759; 60,000 cycles per saturation run; throughput as a fraction of the data-bus peak.

DRAMsim3 mappingtracemetricDRAMsim3memsimmemsim vs DRAMsim3
rochrababgcostream readsthroughput / peak0.6700.619-7.6%
rochrababgcostream, 1 in 3 writesthroughput / peak0.6760.578-14.5%
rochrababgcorandom reads (1 GiB)throughput / peak0.4500.440-2.2%
rochrababgcorandom, 1 in 3 writesthroughput / peak0.4450.442-0.7%
rochrababgcosame bank, next row each timethroughput / peak0.0520.052+0.0%
rochrababgco4 interleaved streamsthroughput / peak0.3100.476+53.5%
rochrababgcosparse random reads (latency)mean read latency (cycles)73.072.1-1.2%
rochrabacobgstream readsthroughput / peak0.9550.958+0.3%
rochrabacobgstream, 1 in 3 writesthroughput / peak0.9030.911+1.0%
rochrabacobgrandom reads (1 GiB)throughput / peak0.4500.438-2.8%
rochrabacobgrandom, 1 in 3 writesthroughput / peak0.4430.441-0.4%
rochrabacobgsame bank, next row each timethroughput / peak0.0520.052+0.0%
rochrabacobg4 interleaved streamsthroughput / peak0.8140.684-15.9%
rochrabacobgsparse random reads (latency)mean read latency (cycles)70.769.9-1.2%

Source: examples/results.md in Memory_System_Sim

A cross-check that agrees everywhere usually means the two models share assumptions. This one shows which results depend on the controller, and therefore need its design specified before they are quoted.

12

Plugging It Into the FHE Simulator

FHE_Accelerator_Sim moves HBM traffic as 4 MiB chunks through one first-come resource. It now accepts an optional memory model, Accelerator(memory=...), that times each chunk given its size, its direction and the direction of the chunk before it. Memory_System_Sim's adapter simulates one pseudo-channel's share of the chunk at command level, memoised:

src/memsim/fhe.py: the interface the FHE simulator calls source
def chunk_time(self, nbytes: int, write: bool, prev_write: bool | None, peak_gbps: float) -> float:
    """Seconds to move ``nbytes`` through HBM whose peak is ``peak_gbps`` (FHE_Accelerator_Sim protocol)."""
    t = self.timing
    channels = peak_gbps / t.channel_peak_gbps()
    bursts = max(1, round(nbytes / t.access_bytes / channels))
    return self.channel_cycles(bursts, write, prev_write) * t.tck_ns * 1e-9
HBM modelbaseline algorithmMin-KS + seeded keys + OTF plaintexts
peak bandwidth (default)13.94 ms, 89%, memory-bound7.19 ms, 11%, MAC-bound
flat: 0.7 x bandwidth19.26 ms, 92%, memory-bound7.29 ms, 15%, MAC-bound
flat: 0.9 x bandwidth15.33 ms, 90%, memory-bound7.21 ms, 12%, MAC-bound
Memory_System_Sim: FR-FCFS, bank groups interleaved15.31 ms, 90%, memory-bound7.21 ms, 12%, MAC-bound
... with 1 MiB chunks15.64 ms, 90%, memory-bound7.21 ms, 12%, MAC-bound
... FCFS scheduler23.70 ms, 94%, memory-bound7.40 ms, 19%, MAC-bound
... a row's bursts in one bank group24.43 ms, 94%, memory-bound7.42 ms, 19%, MAC-bound

Source: examples/results.md in FHE_Accelerator_Sim

HBM peakpeak bandwidthflat 0.7flat 0.9Memory_System_Sim
140 GB/s10.23 ms, memory-bound12.33 ms, memory-bound10.78 ms, memory-bound10.73 ms, memory-bound
160 GB/s9.66 ms, MAC-bound11.46 ms, memory-bound10.10 ms, memory-bound10.06 ms, memory-bound
200 GB/s8.93 ms, MAC-bound10.23 ms, memory-bound9.25 ms, MAC-bound9.23 ms, MAC-bound
250 GB/s8.33 ms, MAC-bound9.34 ms, MAC-bound8.57 ms, MAC-bound8.56 ms, MAC-bound

Source: examples/results.md in FHE_Accelerator_Sim

Reading it. FHE traffic is long sequential streams, so a good controller reaches about 0.90 of peak, mostly lost to refresh. A well-chosen flat 0.9 then agrees to 0.2%. A folklore 0.7 adds 38% to the bootstrap. An FCFS controller or a bank-group-unfriendly mapping adds 70–75%. Near the balance point the model changes the verdict: at 160 GB/s, peak bandwidth calls the optimised design MAC-bound and the HBM model calls it memory-bound. At 200 GB/s, flat 0.7 calls it memory-bound and the model agrees with peak that it is MAC-bound.

13

What to Take Away