Simulation Engineering Toolkit — Presentation 07

Jenkins for Hardware and Simulation Teams

Continuous integration for simulators, from real pipelines that ran: declarative and scripted Jenkinsfiles, agents and labels for tools and licences, matrix builds as parameter sweeps, shared libraries, JUnit and coverage reports, nightly regressions, exact, speed and drift gates, credentials, the declarative linter, and how Jenkins and GitHub Actions divide the work.

Jenkinsfile Matrix builds Shared libraries JUnit & coverage Performance gates Credentials
Lint → Test → Cover → Gate → Results → Nightly
00

Topics We'll Cover

New to Jenkins? Start with Introduction to Jenkins: controller and agents, plugins, a first pipeline, credentials and test reports, from the same local Jenkins.

01

Why Simulators Need CI More Than Most Code

A simulator's product is its answers. A change can leave every unit test green and still move a bootstrap time by 3% or slow the nightly sweep to half speed. CI for a simulator therefore checks four things on every change: that it builds and passes its tests, that its answers agree with references, that its performance has not regressed, and that it produces the results the decks and reports quote.

Each code repository in this series has a Jenkinsfile, and every one of them was run on a local Jenkins (2.580.1, in user space). These are the pipelines:

RepositoryStagesGate
Rust_DES_Kernellint, Rust tests, Python differential tests, coverage, benchmarks, results, nightly sweepcriterion benchmarks against a baseline
RTL_CoSim_NTTlint, model tests, RTL tests on Verilator and on Icarus, results, nightly seed sweepexact RTL cycle counts
Memory_System_Simlint, tests, results, nightly load sweepexact efficiencies plus requests per second
SystemC_Accelerator_Modelbuild, GoogleTest (Release and ASan+UBSan), agreement with SimPy, results, nightly quantum sweepwall clock
Torch_Sim_Frontendlint, tests, traceability matrix, results, nightly sweepexact FLOPs and weight bytes, drift report, capture time

The Jenkins here runs on one machine, is reachable only from it, and has no separate agents. Slide 12 compares it with the GitHub Actions workflows that run the same tests on every push.

02

Declarative and Scripted Pipelines

A declarative pipeline is a fixed structure (pipeline, agent, parameters, stages, post) that Jenkins can validate before it runs and draw as a stage graph. A scripted pipeline is Groovy code inside node { }, with loops, maps and try/catch. Here is the same sweep as a scripted pipeline that also handles a failure on purpose:

snippets/t07/scripted.Jenkinsfile (ran as Demo_scripted #1) source
node('sim') {
    stage('Checkout') {
        git url: 'https://github.com/BrendanJamesLynskey/Memory_System_Sim.git', branch: 'main'
    }
    stage('Sweep') {
        def cells = [:]
        for (preset in ['ddr4', 'hbm']) {
            for (pattern in ['stream', 'random']) {
                def p = preset, q = pattern            // capture loop variables for the closure
                cells["${p}/${q}"] = {
                    sh "PYTHONPATH=src python3 -m memsim.cli --preset ${p} --pattern ${q} --n 3000"
                }
            }
        }
        parallel cells
    }
    stage('A failure that is handled') {
        try {
            sh 'PYTHONPATH=src python3 -m memsim.cli --preset no-such-preset'
        } catch (err) {
            echo "expected failure, handled: ${err}"
            currentBuild.description = 'sweep ok; bad preset rejected'
        }
    }
}
DeclarativeScripted
Checked before runningYes: the linter (slide 11) and the editorOnly when it runs
Logicwhen, matrix, post conditions; script { } blocks for the restAny Groovy (in the sandbox)
Use it forAlmost every repository pipelineGenerated stages, complex orchestration, and inside shared libraries

Its log shows the four parallel branches and the handled failure: expected failure, handled: hudson.AbortException: script returned exit code 2.

03

Agents, Labels and the Tools on Them

controller jobs and Jenkinsfilesbuild queuecredentials, librariesreports and history agent: labels sim, verilatorPython, Rust, Verilator, Icarus; 2 executors agent: labels vivado, licenseda commercial simulator and its licence server agent: label gpubenchmarks that need the device stages ask by label:agent { label 'sim' }

The controller schedules; agents run steps. A stage asks for an agent by label (agent { label 'sim' }), so a pipeline says what it needs, not which machine to use. For hardware and simulation teams, labels usually encode tools and licences:

Here there is one built-in node labelled sim, with one executor, run under systemd-run -p MemoryMax=6G: on an 8-core, 15 GB machine, two heavy builds at once have crashed it before.

A gotcha that cost three failed builds

RTL_CoSim_NTT's VERILATOR_ROOT parameter had the default "${env.HOME}/.local/opt/verilator-deb/…". The parameters { } directive is evaluated before any agent exists, so the default became the literal text null/.local/…, and through environment { } every sh step got PATH=null/…. The first two attempted fixes chased wrong theories: that parameters are null on a job's first build, and that an environment PATH doesn't reach sh. A later probe on Jenkins 2.580.1 (see Introduction to Jenkins) reproduced neither. The fix that worked resolves the tools in each step's own shell:

RTL_CoSim_NTT/Jenkinsfile source
environment {
    // Empty unless the VERILATOR_ROOT parameter is set. Never default a parameter to "${env.HOME}/...":
    // parameters {} is evaluated before the agent exists, so it becomes "null/..." (builds #1-#3 failed so).
    VR = "${params.VERILATOR_ROOT ?: ''}"
    MAKEFLAGS = '-j2'
    // HOME is resolved by each step's shell, so each step that needs Verilator sets it up itself
    VPATH = 'export VERILATOR_ROOT="${VR:-$HOME/.local/opt/verilator-deb/root/usr/share/verilator}"; [ -d "$VERILATOR_ROOT" ] && export PATH="$VERILATOR_ROOT/bin:$PATH" || unset VERILATOR_ROOT; '
}

And name such helpers carefully: Rust_DES_Kernel's first version of this fix called its helper CARGO, which cargo and maturin read as the path of the cargo binary, and its build failed until it was renamed WITH_CARGO.

04

Matrix Builds: a Parameter Sweep in CI

A declarative matrix runs the same stages for every combination of its axes, as parallel branches, minus the excludes. For a simulator, the axes are configurations: memory presets, access patterns, model sizes, simulator back ends.

snippets/t07/matrix.Jenkinsfile (ran as Demo_matrix #1) source
stage('Sweep') {
    matrix {
        agent { label 'sim' }
        axes {
            axis { name 'PRESET';  values 'ddr4', 'hbm' }
            axis { name 'PATTERN'; values 'stream', 'random', 'stride' }
        }
        excludes {
            exclude {                       // not a meaningful cell for this demo
                axis { name 'PRESET';  values 'hbm' }
                axis { name 'PATTERN'; values 'stride' }
            }
        }
        stages {
            stage('Simulate') {
                steps {
                    unstash 'src'
                    sh 'PYTHONPATH=src python3 -m memsim.cli --preset $PRESET --pattern $PATTERN --n 3000 | tee cell.txt'
                    sh 'mv cell.txt "cell-$PRESET-$PATTERN.txt"'
                    archiveArtifacts artifacts: "cell-${PRESET}-${PATTERN}.txt"
                }
            }
        }
    }
}
Jenkins pipeline graph of Demo_matrix build 1: a checkout stage, then five parallel matrix cells each running Simulate, all green

Demo_matrix #1 on the local Jenkins: two presets × three patterns, minus one excluded cell, so five cells.

05

Shared Libraries

When five repositories need the same stages, copy-paste drifts. A shared library is a Git repository of Groovy that pipelines load by name. A file in vars/ becomes a step. This one gives any simulator the house regression stages:

snippets/t07/shared-lib/vars/simRegression.groovy source
def call(Map cfg) {
    String py = cfg.get('python', 'python3')
    stage('Setup') {
        sh "${py} -m venv .venv && .venv/bin/pip install -q -e '.[test]'"
    }
    stage('Tests') {
        sh ".venv/bin/pytest -p no:logging --junitxml=junit.xml ${cfg.get('pytestArgs', '')}"
        junit 'junit.xml'
    }
    if (cfg.gate) {
        stage('Gate') {
            sh ".venv/bin/python ${cfg.gate}"
        }
    }
    if (cfg.results) {
        stage('Results') {
            sh ".venv/bin/python ${cfg.results} > /dev/null"
            archiveArtifacts artifacts: cfg.resultsFile
        }
    }
}
snippets/t07/library.Jenkinsfile (ran as Demo_library #1) source
@Library('simeng-ci') _

node('sim') {
    git url: 'https://github.com/BrendanJamesLynskey/Memory_System_Sim.git', branch: 'main'
    simRegression(gate: 'ci/perf_gate.py --margin 0.5', results: 'examples/results.py',
                  resultsFile: 'examples/results.md')
    stage('Publish (credential demo)') {
        withCredentials([string(credentialsId: 'results-upload-token', variable: 'TOKEN')]) {
            // The token is masked wherever it would be printed.
            sh 'echo "would upload results.md with token $TOKEN"'
        }
    }
}
Jenkins pipeline graph of Demo_library build 1: Setup, Tests, Gate, Results and Publish stages, all green

The library is configured once on the controller, with its repository and default version. The log opens with Loading library simeng-ci@main, and the stages it defined ran Memory_System_Sim's 31 tests and its gate. Pin a library to a tag in production pipelines, so a library change cannot break every build at once.

06

Test and Coverage Reports

Every test framework in this series writes JUnit XML (pytest, cargo-nextest, GoogleTest, cocotb's runner) and every coverage tool Cobertura XML (pytest-cov, cargo-llvm-cov). Jenkins reads both, keeps their history and draws trends (deck 06, slide 13):

Rust_DES_Kernel/Jenkinsfile source
stage('Coverage') {
    steps {
        sh(env.WITH_CARGO + 'cargo llvm-cov --release --cobertura --output-path coverage.xml')
        // cargo-llvm-cov lists each generic instantiation as its own method, and the
        // Coverage plugin's Cobertura parser rejects duplicate names: skip them.
        recordCoverage(tools: [[parser: 'COBERTURA', pattern: 'coverage.xml']],
                       sourceCodeRetention: 'LAST_BUILD', ignoreParsingErrors: true)
    }
}
Rust_DES_Kernel/Jenkinsfile: the first stage, added after build #4 source
// The workspace is reused between builds (it keeps the virtualenv and build caches), so
// delete the previous build's reports first. Without this a build that fails before its
// tests run publishes the last build's JUnit results as its own (Rust_DES_Kernel #4 did).
stage('Clean reports') {
    steps {
        sh 'rm -f target/nextest/ci/junit.xml pytest-junit.xml coverage.xml perf_report.md sweep.csv'
    }
}
Jenkins test report of Rust_DES_Kernel build 10: 47 tests passing, 30 Rust and 17 Python

Rust_DES_Kernel #10: 47 tests from two frameworks, one report.

07

Nightly Regressions and Parameters

Fast checks run on every change; the expensive ones (seed sweeps, load-latency curves, quantum sweeps, every model at every length) run nightly. One Jenkinsfile can serve both through a parameter:

Torch_Sim_Frontend/Jenkinsfile source
parameters {
    booleanParam(name: 'NIGHTLY', defaultValue: false, description: 'Also run the model x length sweep')
    string(name: 'PERF_MARGIN', defaultValue: '0.5',
           description: 'Allowed slow-down of trace capture against ci/perf_baseline.json')
}
Torch_Sim_Frontend/Jenkinsfile source
stage('Nightly sweep') {
    when { expression { params.NIGHTLY } }
    steps {
        sh '.venv/bin/python ci/sweep.py'
        archiveArtifacts artifacts: 'sweep.csv'
    }
}

Concepts used here, and where they are explained. Each links to a glossary entry: this series glossary, or the glossaries of LLM Inference Simulators and FHE Accelerator Simulators for concepts those series already explain.

08

Gating on Performance and Accuracy

A gate is a stage that compares this build with a stored baseline and fails the build if it got worse. For a simulator there are three kinds, and these repositories use all three:

KindComparedToleranceExample
Exact behaviourNumbers the model must reproduceNone: any change fails, and is re-blessed on purposeRTL cycle counts; memsim efficiencies; simfront FLOPs and weight bytes
SpeedWall time or events per secondA margin (25–50% here), best of several runscriterion benchmarks; capture time; requests per second
DriftNumbers that may legitimately changeReported, not failed (unless --strict)Operator counts after a PyTorch upgrade
Torch_Sim_Frontend/ci/perf_gate.py: one gate, three kinds of check source
def compare(base: dict, now: dict, margin: float, strict: bool = False) -> tuple[list[str], int]:
    """The gate's verdict: (report lines, number of failures)."""
    lines = ["# simfront gate", "", "| check | baseline | now | verdict |", "|---|---|---|---|"]
    bad = 0
    for k, b in base["traces"].items():
        n = now["traces"][k]
        for f in ("flops", "weight_bytes"):
            ok = n[f] == b[f]
            bad += not ok
            lines.append(f"| {k}: {f} | {b[f]:,} | {n[f]:,} | {'ok' if ok else 'CHANGED'} |")
        for f in ("ops", "bytes"):
            ok = n[f] == b[f]
            bad += strict and not ok
            lines.append(f"| {k}: {f} | {b[f]:,} | {n[f]:,} | {'ok' if ok else 'drift (framework)'} |")
    b, n = base["capture_s"], now["capture_s"]
    ok = n <= (1 + margin) * b
    bad += not ok
    lines.append(f"| capture llama3-8b prefill 2048 (s) | {b:.3f} | {n:.3f} | {'ok' if ok else 'SLOWER'} "
                 f"(margin {margin:.0%}) |")
    return lines, bad

The gate is itself tested: a test changes one FLOP in a copy of the baseline and checks the gate fails (deck 09). A speed gate on a shared machine needs a generous margin and best-of-N timing, or it will be flaky (deck 11).

09

Credentials

Pipelines need secrets: a token to publish results, a licence server, an SSH key for a lab machine. Jenkins stores them encrypted on the controller, and a pipeline refers to them by ID. withCredentials binds one to an environment variable for a block, and masks its value in the log:

snippets/t07/library.Jenkinsfile source
stage('Publish (credential demo)') {
    withCredentials([string(credentialsId: 'results-upload-token', variable: 'TOKEN')]) {
        // The token is masked wherever it would be printed.
        sh 'echo "would upload results.md with token $TOKEN"'
    }
}
Loading library simeng-ci@main ============================= 31 passed in 17.29s ============================== Masking supported pattern matches of $TOKEN + echo would upload results.md with token **** would upload results.md with token ****
10

Interactive: Every Build, and Why Some Failed

Every build on the local Jenkins, from its API. Choose a job to see its builds; the failures are where the lessons were.

Durations and results from the Jenkins API, recorded in snippets/RESULTS.md.

11

The Real Runs

Each job's page: the last artefacts, the coverage and test-result trends, and the stage view, one row per build. Choose a repository:

Jenkins job page with artefacts, coverage and test trends, and the stage view

Before any of these ran, each Jenkinsfile was checked by the declarative linter (POST /pipeline-model-converter/validate), which also catches typos that would otherwise surface only at run time:

Errors encountered validating Jenkinsfile: WorkflowScript: 6: Invalid condition "allways" - valid conditions are [always, changed, fixed, regression, aborted, success, unsuccessful, unstable, failure, notBuilt, cleanup] @ line 6, column 10. post { allways { echo "typo" } } ^
12

Jenkins and GitHub Actions

Every repository here uses both: GitHub Actions on every push, Jenkins for the full pipeline with its gates and nightly sweeps. They suit different jobs:

JenkinsGitHub Actions
Where it runsYour controller and agents: on site, next to licences, lab hardware and big machinesHosted runners (or self-hosted ones you register)
Pipeline definitionGroovy Jenkinsfile, shared librariesYAML workflows, reusable actions and workflows
Reports and trendsPlugins keep test, coverage and artefact history per jobLogs and artefacts per run; trends need extra tooling
Cost of ownershipYou run, upgrade and secure it, plugins includedLittle to run; minutes and runner sizes are the limits
Typical in hardware teamsRegressions that need EDA tools, licences or boardsOpen-source code, quick checks, public CI
13

What to Take Away