An FHE accelerator tile in C++17 and SystemC, driven by the SimPy simulator's own traces: modern C++ for models, the SystemC kernel and delta cycles, TLM-2.0 sockets and payloads, approximately- and loosely-timed coding, temporal decoupling and the quantum, why tie-breaks are part of the model, op-by-op agreement with SimPy, GoogleTest and sanitizers.
Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.
A Python simulator is the fastest way to explore an architecture. Hardware teams standardise on C++ and SystemC (IEEE 1666-2023) when the model must live with the hardware: virtual platforms that boot firmware, models delivered to customers, and reuse of a vendor's IP models. The interfaces between them are standardised as TLM-2.0 (LLM Inference Simulators glossary; InfSim 02, slide 12).
| SimPy model | SystemC TLM model | |
|---|---|---|
| Strength | Fast to write and change; the scientific Python stack | Industry standard; composes with other vendors' models; runs software |
| Concurrency | Generators that yield events | Threads (coroutines) that wait() on events or time |
| Interfaces | Python calls | Sockets carrying a generic payload with a defined protocol |
| Speed knob | Fewer events (InfSim 08) | Loosely timed coding with temporal decoupling |
This deck builds a TLM-2.0 model of the FHE accelerator tile from FHE Accelerator Simulators 03. It is driven by that simulator's own trace format and checked against it on the same traces. The code is SystemC_Accelerator_Model.
A performance model is mostly plain data and arithmetic. C++17 lets that part read like the Python it ports, with no SystemC in sight, so it can be unit-tested on its own:
struct HeOp {
int id = 0;
std::string op, stage;
int level = 0;
std::vector<std::string> inputs;
std::string output;
std::optional<std::pair<std::string, std::int64_t>> key; // (key id, bytes)
std::vector<std::pair<std::string, std::int64_t>> pts; // (plaintext id, bytes)
std::vector<Kernel> kernels;
};
Segment segment(const Hardware& hw, int log_n, const Kernel& k) {
// hardware.py CostModel._seg and segments() (digital units only)
Unit u = unit_of(k.kind);
std::int64_t N = std::int64_t{1} << log_n;
std::int64_t work_i = u == Unit::Ntt ? k.amount * (N / 2) * log_n : k.amount;
double work = static_cast<double>(work_i);
double s = hw.clock;
double t_logic = work / (hw.rate(u) * s);
double t_sram = static_cast<double>(k.words * 8) / (hw.sram_gbps * 1e9 * s);
return {u, t_logic >= t_sram ? t_logic : t_sram, work};
}
std::vector, std::unique_ptr, std::shared_ptr for a gate shared by two processes): no manual delete, so the sanitizers have little to find.sc_spawn([this, pl]() { prefetch(pl); }).work / (rate * s), never "simplified", so every kernel's cost is bit-identical to the SimPy model's (deck 02, slide 06).sc_module) hold processes: SC_THREADs that can wait(), or methods that cannot. Processes are also spawned at run time (sc_spawn), one per HE operation here.sc_time, an integer count of a global resolution. This model sets 1 fs, so each delay is rounded to the nearest femtosecond: the only source of difference from the floating-point SimPy model.The C++ ports the parts of FHE_Accelerator_Sim that decide what happens: the cost model and the LRU scratchpad planner. SystemC schedules when. The issuer plans each operation in program order and spawns a process that waits for its inputs, loads what missed, runs its kernels on the units and writes its output:
void Tile::issuer() {
for (const auto& o : t_.ops) {
while (inflight_ >= hw_.window) wait(slot_ev_);
OpPlan pl = planner_.plan(o);
++inflight_;
sc_spawn([this, &o, pl]() { run_op(o, pl); });
}
}
Transaction-level modelling replaces pin-level signals with function calls that carry a generic payload: a command, an address, a data pointer and length, a response status, plus optional extensions. The tile's initiator socket binds to the HBM model's target socket, and either side can be swapped for another vendor's model.
| Coding style | Interface | Timing | Use it for |
|---|---|---|---|
| Loosely timed (LT) | b_transport(trans, delay): one blocking call | The target adds its latency to delay; processes may run ahead of simulated time | Software development, fast functional runs |
| Approximately timed (AT) | nb_transport_fw/bw with phases | Each phase happens at its simulated time; contention resolves in time order | Architecture exploration, performance |
Here the payload is timing-only: the data pointer refers to a dummy buffer the target never reads, and an extension carries the requesting operation's id for arbitration:
struct Requester : tlm::tlm_extension<Requester> {
int op = 0;
tlm::tlm_extension_base* clone() const override { return new Requester(*this); }
void copy_from(const tlm::tlm_extension_base& e) override { op = static_cast<const Requester&>(e).op; }
};
The base protocol has four phases: BEGIN_REQ and END_REQ on the forward path, BEGIN_RESP and END_RESP on the backward path. A target may shortcut a phase by returning an updated phase. This HBM accepts at once (END_REQ), waits for the channel, holds it for bytes / bandwidth, then answers with BEGIN_RESP. The initiator completes the transaction by returning TLM_COMPLETED, which implies END_RESP.
tlm::tlm_sync_enum Hbm::nb_transport_fw(tlm::tlm_generic_payload& trans, tlm::tlm_phase& phase, sc_time&) {
if (phase == tlm::BEGIN_REQ) {
// accept at once (END_REQ); a thread per request waits for the channel
sc_spawn([this, &trans]() { serve(&trans); });
phase = tlm::END_REQ;
return tlm::TLM_UPDATED;
}
if (phase == tlm::END_RESP) return tlm::TLM_COMPLETED;
SC_REPORT_ERROR("Hbm", "unexpected phase");
return tlm::TLM_COMPLETED;
}
void Hbm::serve(tlm::tlm_generic_payload* trans) {
Requester* who = nullptr;
trans->get_extension(who);
arb_->request(kHbmResource, who ? who->op : 0);
wait(seconds(static_cast<double>(trans->get_data_length()) / bw_));
arb_->release(kHbmResource);
trans->set_response_status(tlm::TLM_OK_RESPONSE);
++transactions_;
tlm::tlm_phase phase = tlm::BEGIN_RESP;
sc_time delay = SC_ZERO_TIME;
socket->nb_transport_bw(*trans, phase, delay); // the initiator completes it (END_RESP implied)
}
Every process stays in step with simulated time, so requests meet the channel in the order they happen. AT reproduces the SimPy model op by op (slide 09). The price is a context switch per phase.
In LT, a process keeps a local time offset and runs ahead of the kernel's time, synchronising only when its offset passes the global quantum or when it must wait for another process. The target books the channel from the caller's local time and returns the finish time as the annotated delay:
void Hbm::b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
// LT: book the channel from the initiator's local time (now + delay), in call order
double dt = static_cast<double>(trans.get_data_length()) / bw_;
sc_time start = std::max(busy_until_, sc_time_stamp() + delay);
busy_until_ = start + seconds(dt);
delay = busy_until_ - sc_time_stamp();
trans.set_response_status(tlm::TLM_OK_RESPONSE);
++transactions_;
}
} else {
sc_time delay = qk.get_local_time();
socket->b_transport(trans, delay);
qk.set(delay);
if (qk.need_sync()) qk.sync();
Fewer synchronisations make it faster. The cost: a process that runs ahead books the HBM channel or a unit before another process that would have asked earlier in simulated time. That is the accuracy LT trades away, and the quantum sets how much (slide 10).
Two operations finish waiting at the same instant and both want the NTT unit. Who goes first? A SimPy model answers by its event order. A SystemC kernel may answer either way: IEEE 1666 leaves the order of runnable processes unspecified. Left to the kernel, "first come" means whoever the implementation happens to run first. So the model makes the tie-break explicit:
void Arbiter::run() {
for (;;) {
wait(kick_);
// let every process that can still act at this instant do so, then decide
while (sc_pending_activity_at_current_time()) wait(SC_ZERO_TIME);
for (auto& r : res_) {
if (r.busy || r.pending.empty()) continue;
auto first = r.pending.begin();
r.busy = true;
first->second->notify();
r.pending.erase(first);
}
}
}
| case | kernel order: ops that differ from SimPy | horizon difference | explicit: ops that differ | horizon difference |
|---|---|---|---|---|
| bootstrap, ARK-class | 6 of 203 | +6.4e-15 | 0 | +6.4e-15 |
| bootstrap, small digital | 18 of 203 | +9.5e-15 | 0 | +9.5e-15 |
| bootstrap + Min-KS/seeded/OTF, small digital | 14 of 233 | +8.9e-16 | 0 | +8.9e-16 |
| 4 bootstraps, ARK-class | 24 of 812 | +1.2e-13 | 0 | +1.2e-13 |
Source: examples/results.md in SystemC_Accelerator_Model
With the kernel's order, 6–24 operations per run end at different times from SimPy's, although the totals agree. With "oldest operation first", none do. This is the same lesson as the Rust event list's sequence numbers (deck 01, slide 03): a deterministic tie-break is part of the model, so write it down.
tools/export_case.py writes three files from FHE_Accelerator_Sim: the trace (its own dump_trace JSON), the hardware, and the SimPy model's answers in worst-case clocking mode. The SystemC model reads the same trace and is compared with those answers:
| case | HE ops | SimPy horizon | SystemC horizon, relative difference | largest op-end difference | ops differing by more than 1 ps | HBM bytes identical | unit busy-time difference / horizon |
|---|---|---|---|---|---|---|---|
| one HMult (ark) | 1 | 0.2634 ms | -2.2e-16 | 0 fs | 0 | yes | 0e+00 |
| one HRot (ark) | 1 | 0.2207 ms | -6.7e-16 | 0 fs | 0 | yes | 0e+00 |
| bootstrap, ARK-class | 203 | 13.9354 ms | +6.4e-15 | 0 fs | 0 | yes | 0e+00 |
| bootstrap + Min-KS/seeded/OTF, ARK-class | 233 | 7.1914 ms | +4.4e-16 | 0 fs | 0 | yes | 0e+00 |
| bootstrap, small digital | 203 | 17.4644 ms | +9.5e-15 | 0 fs | 0 | yes | 0e+00 |
| bootstrap + Min-KS/seeded/OTF, small digital | 233 | 26.6480 ms | +8.9e-16 | 0 fs | 0 | yes | 0e+00 |
| bootstrap, ARK-class, 1 MiB chunks | 203 | 13.9734 ms | +9.8e-14 | 26.7 us | 19 | yes | 2e-17 |
| 4 bootstraps, ARK-class | 812 | 55.7134 ms | +1.2e-13 | 7 fs | 0 | yes | 0e+00 |
Source: examples/results.md in SystemC_Accelerator_Model
Four bootstraps (812 HE ops) on the ARK-class tile. Each point is a recorded run (median of five). Choose a quantum to see what LT buys and what it costs, against AT and SimPy.
Sixteen GoogleTest tests (deck 06, slide 10): the planner's bytes equal SimPy's exactly, unit busy times, AT agreement on four cases, LT invariants. One SystemC peculiarity shapes the suite: a process can elaborate and run a simulation only once, so each simulation runs in a forked child that sends its results back as JSON.
TEST_P(AtAgreement, MatchesSimPyToFemtoseconds) {
const std::string c = GetParam();
json r = simulate_in_child(c, Style::AT, 0), e = expected(c);
double H = e["horizon_s"];
EXPECT_NEAR(r["horizon"].get<double>(), H, 1e-12) << c; // within 1 ps on a ms-scale run
EXPECT_LT(max_abs_diff(r["op_end"], e["op_end"]), 1e-12) << c;
EXPECT_EQ(r["bytes"], e["bytes"]);
EXPECT_NEAR(r["busy"][0].get<double>(), e["busy"]["ntt"].get<double>(), 1e-12 * H);
EXPECT_NEAR(r["hbm_busy"].get<double>(), e["busy"]["hbm"].get<double>(), 1e-12 * H);
}
-DACCEL_SANITIZE=ON): GoogleTest, ASan + UBSan: 16 tests, 16 passed.sc_main replaces main (SystemC's library defines main), so the test binary calls RUN_ALL_TESTS() from sc_main.examples/results.md.