Simulation Engineering Toolkit — Presentation 03

C++ Performance Models with SystemC TLM-2.0

An FHE accelerator tile in C++17 and SystemC, driven by the SimPy simulator's own traces: modern C++ for models, the SystemC kernel and delta cycles, TLM-2.0 sockets and payloads, approximately- and loosely-timed coding, temporal decoupling and the quantum, why tie-breaks are part of the model, op-by-op agreement with SimPy, GoogleTest and sanitizers.

SystemC TLM-2.0 C++17 AT and LT Temporal decoupling GoogleTest ASan/UBSan
Trace → Plan → Schedule → Transact → Compare → Test
00

Topics We'll Cover

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

01

Why a C++ Model in SystemC?

A Python simulator is the fastest way to explore an architecture. Hardware teams standardise on C++ and SystemC (IEEE 1666-2023) when the model must live with the hardware: virtual platforms that boot firmware, models delivered to customers, and reuse of a vendor's IP models. The interfaces between them are standardised as TLM-2.0 (LLM Inference Simulators glossary; InfSim 02, slide 12).

SimPy modelSystemC TLM model
StrengthFast to write and change; the scientific Python stackIndustry standard; composes with other vendors' models; runs software
ConcurrencyGenerators that yield eventsThreads (coroutines) that wait() on events or time
InterfacesPython callsSockets carrying a generic payload with a defined protocol
Speed knobFewer events (InfSim 08)Loosely timed coding with temporal decoupling

This deck builds a TLM-2.0 model of the FHE accelerator tile from FHE Accelerator Simulators 03. It is driven by that simulator's own trace format and checked against it on the same traces. The code is SystemC_Accelerator_Model.

02

Modern C++ for Models

A performance model is mostly plain data and arithmetic. C++17 lets that part read like the Python it ports, with no SystemC in sight, so it can be unit-tested on its own:

include/accel/model.hpp: value types, std::optional for "no key" source
struct HeOp {
    int id = 0;
    std::string op, stage;
    int level = 0;
    std::vector<std::string> inputs;
    std::string output;
    std::optional<std::pair<std::string, std::int64_t>> key;    // (key id, bytes)
    std::vector<std::pair<std::string, std::int64_t>> pts;      // (plaintext id, bytes)
    std::vector<Kernel> kernels;
};
src/model.cpp: the cost model, in the same floating-point order as hardware.py source
Segment segment(const Hardware& hw, int log_n, const Kernel& k) {
    // hardware.py CostModel._seg and segments() (digital units only)
    Unit u = unit_of(k.kind);
    std::int64_t N = std::int64_t{1} << log_n;
    std::int64_t work_i = u == Unit::Ntt ? k.amount * (N / 2) * log_n : k.amount;
    double work = static_cast<double>(work_i);
    double s = hw.clock;
    double t_logic = work / (hw.rate(u) * s);
    double t_sram = static_cast<double>(k.words * 8) / (hw.sram_gbps * 1e9 * s);
    return {u, t_logic >= t_sram ? t_logic : t_sram, work};
}
03

The SystemC Kernel: Processes, Events, Delta Cycles

evaluaterun each runnable process updatesignals take new values delta notificationswake SC_ZERO_TIME waiters another delta cycle, same simulated time advance timeto the next timed event when nothing is runnable at this time
04

The Modelled Tile

Tile (sc_module, initiator socket) issuerwindow of 4 ops run_op x Nspawned processes prefetch, writebackspawned per op NTT MAC AUTO Arbiter: tie-break, grants units and HBM Planner: LRUscratchpad, decidesevery HBM byte Hbmtarget socket1000 GB/s,4 MiB transactions TLM-2.0

The C++ ports the parts of FHE_Accelerator_Sim that decide what happens: the cost model and the LRU scratchpad planner. SystemC schedules when. The issuer plans each operation in program order and spawns a process that waits for its inputs, loads what missed, runs its kernels on the units and writes its output:

src/tile.cpp: the issuer source
void Tile::issuer() {
    for (const auto& o : t_.ops) {
        while (inflight_ >= hw_.window) wait(slot_ev_);
        OpPlan pl = planner_.plan(o);
        ++inflight_;
        sc_spawn([this, &o, pl]() { run_op(o, pl); });
    }
}
05

TLM-2.0: Payloads, Sockets and Coding Styles

Transaction-level modelling replaces pin-level signals with function calls that carry a generic payload: a command, an address, a data pointer and length, a response status, plus optional extensions. The tile's initiator socket binds to the HBM model's target socket, and either side can be swapped for another vendor's model.

Coding styleInterfaceTimingUse it for
Loosely timed (LT)b_transport(trans, delay): one blocking callThe target adds its latency to delay; processes may run ahead of simulated timeSoftware development, fast functional runs
Approximately timed (AT)nb_transport_fw/bw with phasesEach phase happens at its simulated time; contention resolves in time orderArchitecture exploration, performance

Here the payload is timing-only: the data pointer refers to a dummy buffer the target never reads, and an extension carries the requesting operation's id for arbitration:

include/accel/tile.hpp: a payload extension source
struct Requester : tlm::tlm_extension<Requester> {
    int op = 0;
    tlm::tlm_extension_base* clone() const override { return new Requester(*this); }
    void copy_from(const tlm::tlm_extension_base& e) override { op = static_cast<const Requester&>(e).op; }
};
06

Approximately Timed: the Four-Phase Protocol

The base protocol has four phases: BEGIN_REQ and END_REQ on the forward path, BEGIN_RESP and END_RESP on the backward path. A target may shortcut a phase by returning an updated phase. This HBM accepts at once (END_REQ), waits for the channel, holds it for bytes / bandwidth, then answers with BEGIN_RESP. The initiator completes the transaction by returning TLM_COMPLETED, which implies END_RESP.

src/tile.cpp: the target's forward path source
tlm::tlm_sync_enum Hbm::nb_transport_fw(tlm::tlm_generic_payload& trans, tlm::tlm_phase& phase, sc_time&) {
    if (phase == tlm::BEGIN_REQ) {
        // accept at once (END_REQ); a thread per request waits for the channel
        sc_spawn([this, &trans]() { serve(&trans); });
        phase = tlm::END_REQ;
        return tlm::TLM_UPDATED;
    }
    if (phase == tlm::END_RESP) return tlm::TLM_COMPLETED;
    SC_REPORT_ERROR("Hbm", "unexpected phase");
    return tlm::TLM_COMPLETED;
}
src/tile.cpp: one thread per accepted request source
void Hbm::serve(tlm::tlm_generic_payload* trans) {
    Requester* who = nullptr;
    trans->get_extension(who);
    arb_->request(kHbmResource, who ? who->op : 0);
    wait(seconds(static_cast<double>(trans->get_data_length()) / bw_));
    arb_->release(kHbmResource);
    trans->set_response_status(tlm::TLM_OK_RESPONSE);
    ++transactions_;
    tlm::tlm_phase phase = tlm::BEGIN_RESP;
    sc_time delay = SC_ZERO_TIME;
    socket->nb_transport_bw(*trans, phase, delay);       // the initiator completes it (END_RESP implied)
}

Every process stays in step with simulated time, so requests meet the channel in the order they happen. AT reproduces the SimPy model op by op (slide 09). The price is a context switch per phase.

07

Loosely Timed: Temporal Decoupling and the Quantum

In LT, a process keeps a local time offset and runs ahead of the kernel's time, synchronising only when its offset passes the global quantum or when it must wait for another process. The target books the channel from the caller's local time and returns the finish time as the annotated delay:

src/tile.cpp: LT booking source
void Hbm::b_transport(tlm::tlm_generic_payload& trans, sc_time& delay) {
    // LT: book the channel from the initiator's local time (now + delay), in call order
    double dt = static_cast<double>(trans.get_data_length()) / bw_;
    sc_time start = std::max(busy_until_, sc_time_stamp() + delay);
    busy_until_ = start + seconds(dt);
    delay = busy_until_ - sc_time_stamp();
    trans.set_response_status(tlm::TLM_OK_RESPONSE);
    ++transactions_;
}
src/tile.cpp: the initiator side, with tlm_quantumkeeper source
} else {
    sc_time delay = qk.get_local_time();
    socket->b_transport(trans, delay);
    qk.set(delay);
    if (qk.need_sync()) qk.sync();

Fewer synchronisations make it faster. The cost: a process that runs ahead books the HBM channel or a unit before another process that would have asked earlier in simulated time. That is the accuracy LT trades away, and the quantum sets how much (slide 10).

08

Ties Are Part of the Model

Two operations finish waiting at the same instant and both want the NTT unit. Who goes first? A SimPy model answers by its event order. A SystemC kernel may answer either way: IEEE 1666 leaves the order of runnable processes unspecified. Left to the kernel, "first come" means whoever the implementation happens to run first. So the model makes the tie-break explicit:

src/tile.cpp: settle the instant, then grant the oldest request source
void Arbiter::run() {
    for (;;) {
        wait(kick_);
        // let every process that can still act at this instant do so, then decide
        while (sc_pending_activity_at_current_time()) wait(SC_ZERO_TIME);
        for (auto& r : res_) {
            if (r.busy || r.pending.empty()) continue;
            auto first = r.pending.begin();
            r.busy = true;
            first->second->notify();
            r.pending.erase(first);
        }
    }
}
casekernel order: ops that differ from SimPyhorizon differenceexplicit: ops that differhorizon difference
bootstrap, ARK-class6 of 203+6.4e-150+6.4e-15
bootstrap, small digital18 of 203+9.5e-150+9.5e-15
bootstrap + Min-KS/seeded/OTF, small digital14 of 233+8.9e-160+8.9e-16
4 bootstraps, ARK-class24 of 812+1.2e-130+1.2e-13

Source: examples/results.md in SystemC_Accelerator_Model

With the kernel's order, 6–24 operations per run end at different times from SimPy's, although the totals agree. With "oldest operation first", none do. This is the same lesson as the Rust event list's sequence numbers (deck 01, slide 03): a deterministic tie-break is part of the model, so write it down.

09

Checked Against the SimPy Model

tools/export_case.py writes three files from FHE_Accelerator_Sim: the trace (its own dump_trace JSON), the hardware, and the SimPy model's answers in worst-case clocking mode. The SystemC model reads the same trace and is compared with those answers:

caseHE opsSimPy horizonSystemC horizon, relative differencelargest op-end differenceops differing by more than 1 psHBM bytes identicalunit busy-time difference / horizon
one HMult (ark)10.2634 ms-2.2e-160 fs0yes0e+00
one HRot (ark)10.2207 ms-6.7e-160 fs0yes0e+00
bootstrap, ARK-class20313.9354 ms+6.4e-150 fs0yes0e+00
bootstrap + Min-KS/seeded/OTF, ARK-class2337.1914 ms+4.4e-160 fs0yes0e+00
bootstrap, small digital20317.4644 ms+9.5e-150 fs0yes0e+00
bootstrap + Min-KS/seeded/OTF, small digital23326.6480 ms+8.9e-160 fs0yes0e+00
bootstrap, ARK-class, 1 MiB chunks20313.9734 ms+9.8e-1426.7 us19yes2e-17
4 bootstraps, ARK-class81255.7134 ms+1.2e-137 fs0yes0e+00

Source: examples/results.md in SystemC_Accelerator_Model

10

Interactive: The Quantum Trade-Off

Four bootstraps (812 HE ops) on the ARK-class tile. Each point is a recorded run (median of five). Choose a quantum to see what LT buys and what it costs, against AT and SimPy.

11

GoogleTest and Sanitizers

Sixteen GoogleTest tests (deck 06, slide 10): the planner's bytes equal SimPy's exactly, unit busy times, AT agreement on four cases, LT invariants. One SystemC peculiarity shapes the suite: a process can elaborate and run a simulation only once, so each simulation runs in a forked child that sends its results back as JSON.

tests/test_tile.cpp: value-parameterised over four exported cases source
TEST_P(AtAgreement, MatchesSimPyToFemtoseconds) {
    const std::string c = GetParam();
    json r = simulate_in_child(c, Style::AT, 0), e = expected(c);
    double H = e["horizon_s"];
    EXPECT_NEAR(r["horizon"].get<double>(), H, 1e-12) << c;            // within 1 ps on a ms-scale run
    EXPECT_LT(max_abs_diff(r["op_end"], e["op_end"]), 1e-12) << c;
    EXPECT_EQ(r["bytes"], e["bytes"]);
    EXPECT_NEAR(r["busy"][0].get<double>(), e["busy"]["ntt"].get<double>(), 1e-12 * H);
    EXPECT_NEAR(r["hbm_busy"].get<double>(), e["busy"]["hbm"].get<double>(), 1e-12 * H);
}
12

What Did Not Work, and the Limits

13

What to Take Away