Automating the path from a developer's editor to production systems
main; CI across repositoriesBefore CI/CD, teams faced integration hell — developers worked in isolation for weeks, then merged everything at once.
The longer code sits unintegrated, the more expensive integration becomes. CI/CD makes integration a continuous, cheap activity rather than a periodic expensive one.
CI is a practice where developers integrate their changes into a shared branch frequently — ideally multiple times per day — with each integration verified by an automated build and test run.
The two CDs are often confused. The distinction is whether the final push to production is manual or automatic.
| Aspect | Continuous Delivery | Continuous Deployment |
|---|---|---|
| Definition | Code is always in a releasable state; deploy to prod requires a human click | Every commit that passes the pipeline is deployed to prod automatically |
| Human gate | Yes — explicit approval step | No — fully automated |
| Risk tolerance | Suitable when compliance, QA, or product sign-off is required | Requires very high test coverage and observability confidence |
| Typical users | Regulated industries, enterprise products | SaaS, high-cadence web services (e.g. Netflix, Etsy) |
| Prerequisite | Both require solid CI — you cannot have CD without CI | |
A pipeline is a sequence of automated stages. Each stage must pass before the next runs. A failure stops the pipeline and notifies the team.
Pipeline definitions live in the repo (e.g. .github/workflows/ci.yml). They are versioned, reviewed, and evolved alongside the application code.
The build stage converts source code into a runnable artefact — a compiled binary, a Docker image, a wheel package, or a bundled web app.
pip install, npm ci, go mod download)npm ci, not npm installin CI. ci is stricter: it uses the lockfile exactly and fails if it would need updating — preventing silent dependency drift.
# GitHub Actions build step
- name: Install deps
run: npm ci
- name: Lint
run: npm run lint
- name: Build
run: npm run build
- name: Upload artefact
uses: actions/upload-artifact@v4
with:
name: dist-${{ github.sha }}
path: dist/
Mike Cohn's test pyramid guides how to balance test types. More tests at the base (fast, cheap); fewer at the apex (slow, expensive).
Test a single function or class in isolation. No I/O, no network. Should run in < 1 s total for a module.
Test how components work together — e.g. service + database, or two microservices. Use real or containerised dependencies.
Drive the full stack through a browser or API client. Playwright, Cypress, Selenium. Run last; slowest.
A quality gate is a threshold that the pipeline enforces. If the code does not meet the standard, the pipeline fails and the change cannot proceed.
Require minimum test coverage — e.g. 80% line or branch coverage. Tools: pytest-cov, nyc, jacoco.
coverage: 82.4% ✔ (≥ 80%)
branch: 76.1% ✘ (< 75%) → FAIL
Run SAST tools: semgrep, bandit, eslint, SonarQube. Block merges that introduce new high-severity findings.
Scan for known CVEs in dependencies. npm audit, pip-audit, Snyk, Dependabot. Fail on HIGH or CRITICAL findings.
Start with loose gates and tighten over time. Enforcing 80% coverage on a legacy codebase from day one is demoralising — begin at your current level and ratchet upward.
More gates: lint, types, formatting and Lighthouse on slide 09; tests that pin results on slide 10.
Tests prove behaviour. These cheaper gates catch the rest in seconds, before a reviewer reads a line.
tsc --noEmit), so a wrong argument fails CI, not production.cargo fmt, ruff format) in check mode fails if any file is not formatted, so reviews never argue about layout."assertions": {
"categories:performance": ["error", { "minScore": 0.9 }],
"categories:accessibility": ["error", { "minScore": 0.9 }],
"categories:best-practices": ["error", { "minScore": 0.9 }]
}
From a Next.js site on this GitHub: its Lighthouse job fails the PR if the median score of any category drops below 0.9.
Simulators and numeric code need tests that pin numbers, not only behaviour. Every example here runs in CI on this GitHub.
.gitignore failed CI on 3 Oct 2026. An up-to-date check is CI regenerating such files and failing on any difference.next build + next start, not the dev server, so the bundle that ships is the one tested.Go deeper: Testing Frameworks for Simulators (fixtures, Hypothesis, golden tests, mutation testing).
How you branch determines when and what the pipeline runs.
Developers commit directly to main (or via very short-lived branches). CI runs on every push. Favoured by Google, Meta. Requires feature flags for incomplete work.
main is always deployable. Work happens in feature branches, merged via pull request. CI runs on each PR and on merge to main. Simple and effective for most teams.
Separate develop, release, and hotfix branches. More overhead — now considered heavyweight for most modern software.
| Trigger | Typical pipeline |
|---|---|
| Push to feature branch | Build + unit tests |
| Pull request opened | Full CI (build + all tests + analysis) |
Merge to main | Full CI + deploy to staging |
Tag push (v1.2.0) | Full CI + deploy to production |
| Scheduled (nightly) | Slow tests, security scans |
Don't run the full slow suite on every commit to a feature branch — developers lose patience and disable CI. Run a fast subset; run everything on merge.
main is what ships. Protection rules turn the pipeline's verdict from advice into a rule.
main. CI runs on it and reviewers comment before anything lands (GitLab calls it a merge request).CODEOWNERS file names who must review which paths.git push --force replaces the remote branch's history. Protection blocks it on main, and blocks deleting the branch.GH013 … Changes must be made through a pull request.main, one per change.main.On this GitHub: a Next.js site requires five checks (Lint & Typecheck, Unit Tests, Verify maths against PyTorch fixtures, E2E Tests, Lighthouse); the GitHub Actions deck runs a ruleset with an empty bypass list and shows a PR blocked (slides 19, 20). Plans differ for private repos: about rulesets, code owners.
A build artefact is the immutable, versioned output of the build stage. It is promoted through environments — never rebuilt.
Never rebuild the artefact for staging vs production. Rebuilding introduces the risk that the artefact that passed testing is not what gets deployed.
Containers solve the "works on my machine" problem. A Docker image bundles the application and its entire runtime — the same image runs in CI, staging, and production.
# Multi-stage Dockerfile (Python example)
FROM python:3.12-slim AS builder
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
FROM python:3.12-slim AS runtime
WORKDIR /app
COPY --from=builder /usr/local/lib/python3.12 \
/usr/local/lib/python3.12
COPY src/ .
RUN useradd -m appuser && chown -R appuser /app
USER appuser
CMD ["python", "main.py"]
Multi-stage builds keep the final image lean — the builder stage (with compilers and dev tools) is discarded. Runtime images should contain only what is needed to run.
docker build -t app:$SHA .trivy image app:$SHAghcr.io, ECR, Docker Hub:latest — avoid in production; ambiguous:abc1234 — git SHA, fully traceable:1.4.2 — semver for releases:main-20260306 — branch + date# .github/workflows/ci.yml
name: CI
on:
push:
branches: [main]
pull_request:
jobs:
build-and-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install dependencies
run: pip install -r requirements.txt
- name: Lint
run: ruff check .
- name: Unit tests
run: pytest tests/unit --cov=src \
--cov-fail-under=80
- name: Build Docker image
run: |
docker build \
-t ghcr.io/${{ github.repository }}:${{ github.sha }} .
- name: Push to registry
if: github.ref == 'refs/heads/main'
run: docker push \
ghcr.io/${{ github.repository }}:${{ github.sha }}
.github/workflows/Split slow test suites across multiple jobs that run concurrently. Use needs: to express dependencies between jobs.
Go deeper: Introduction to GitHub Actions, with every example run for real.
| Tool | Model | Config file | Strengths | Considerations |
|---|---|---|---|---|
| GitHub Actions | SaaS / self-hosted | .github/workflows/*.yml |
Tight GitHub integration, huge Marketplace, free for public repos | Costs scale with minutes on private repos |
| GitLab CI/CD | SaaS / self-hosted | .gitlab-ci.yml |
All-in-one DevOps platform, strong environments & review apps | GitLab hosting required (or self-host) |
| Jenkins | Self-hosted | Jenkinsfile (Groovy) |
Fully configurable, huge plugin ecosystem, runs anywhere | High operational burden; Groovy DSL has a learning curve |
| CircleCI | SaaS / self-hosted | .circleci/config.yml |
Fast, good caching, orbs for reusable config | Costs can surprise at scale |
| Tekton / ArgoCD | Kubernetes-native | CRD YAML manifests | Cloud-native, GitOps-friendly, very scalable | Steep learning curve; requires Kubernetes |
For a new project on GitHub, start with GitHub Actions. For an enterprise self-hosted requirement with Kubernetes, evaluate Tekton + ArgoCD.
When one repo installs another, a push to the library can break the app without a line of the app changing.
pkg @ git+URL installs whatever main is today; git+URL@<sha> pins one commit. Tracking catches breakage at once but can turn a green repo red overnight. Pinning is reproducible but needs deliberate bumps.GITHUB_SHA, within 30 days). An unpinned install resolves again, so a re-run tests the new upstream: gh run rerun <id>, or gh workflow run ci.yml where the workflow has a manual trigger.VENDORED.json) so a test can check the copy.pnpm-lock.yaml, Cargo.lock). pnpm install --frozen-lockfile fails rather than quietly changing it in CI.pkg @ git+… that is already installed, so a reused venv on a long-lived agent keeps testing the old upstream. The Jenkinsfiles force it:.venv/bin/pip install -q --force-reinstall --no-deps "disagg-sim @ git+https://github.com/BrendanJamesLynskey/Disaggregated_Inference_Sim"
repository_dispatch (GitHub Actions deck, slide 23); Dependabot can open the pin bumps (slide 14).Re-run rules: docs.github.com. pip VCS installs: pip docs.
Pipelines need credentials — API keys, registry passwords, cloud credentials. Mishandling them is one of the most common CI/CD security failures.
echo $MY_SECRET.env files with real values${{ secrets.MY_KEY }} — redacted from logsSLSA provenance and signed artefactsRunners should have the minimum IAM permissions needed. Use ephemeral runners (a fresh VM per job) rather than long-lived shared agents.
GitHub Actions supports OIDC federation with AWS, GCP, and Azure. The runner gets a short-lived token automatically — no stored secret needed.
How you deploy to production determines your blast radius and rollback speed.
Stop old version, start new version. Simple but causes downtime. Only for non-critical systems.
Replace instances one at a time. Zero downtime. Old and new versions briefly coexist — APIs must be backwards-compatible.
Two identical environments. Route traffic from blue (old) to green (new). Instant rollback by switching the load balancer back. Doubles infrastructure cost briefly.
Route a small percentage (e.g. 5%) of traffic to the new version. Monitor error rates and latency. Gradually shift 100% if healthy, or roll back quickly if not.
Deploy code but hide it behind a runtime flag. Decouple deployment from feature release. Allows trunk-based development with incomplete features safely in production.
Canary + feature flags is the combination used by most high-velocity teams (Netflix, Spotify, GitHub itself).
main) deploys to production and other branches and PRs get previews.vercel deploy --prod) ships the folder you run it in, so deploy a clean export (git archive HEAD), not a working tree with local files in it.main and the site updates within minutes.x-vercel-protection-bypass header lets automation through; keep it secret.DATABASE_URL), kept out of git and set separately for development, preview and production. Passwords and keys among them are secrets.vercel env pull returns it empty (Vercel now calls this type "Secret").Vercel environments · deployment protection · protection bypass · write-only variables · HTTP 308 · expand/contract (Fowler)
CI tested a copy. Production has its own database, runtime and settings, so check the live system too.
GET /api/health → 200
{"ok":true,"data":{"healthy":true,"db":"up","dbError":null,
"schema":{"status":"current","applied":1,"expected":1, …}}}
Health-check patterns: Kubernetes probes. The case study on slide 23 shows why each of these exists.
Deploying fast is only safe if you can detect failures quickly and recover faster than you broke things.
Configure alerts on key SLIs. If error rate exceeds threshold within a defined window after deploy, trigger automatic rollback — re-deploy the previous artefact SHA.
Deploy completes. Automated smoke tests run.
Monitor p50/p99 latency and error rate. Compare to pre-deploy baseline.
Check application logs for new error patterns.
Mark deploy as stable. Close the change window.
Both happened in 2026 to a small Next.js + Postgres site on Vercel, built from this GitHub. Both had green CI.
Lesson: every fallback needs a signal that it happened.
require() an ES module (a package written with import/export); some platform loaders cannot.require(). Node 20.19+ and 22.12+ allow that, so local runs and CI passed. The platform's function loader refused (ERR_REQUIRE_ESM), and every comment list with a comment in it returned 500: 14 errors in under 3 minutes.command:
"pnpm build && node --no-experimental-require-module node_modules/next/dist/bin/next start -p 3000",
A slow pipeline is a pipeline that developers work around. Target < 10 minutes for the core CI loop.
Cache dependency downloads between runs. Most platforms support keying the cache on the lockfile hash.
- uses: actions/cache@v4
with:
path: ~/.npm
key: npm-${{ hashFiles('**/package-lock.json') }}
Split test suites into shards. Run lint, unit tests, and security scans as parallel jobs. Use a matrix strategy for multi-version or multi-OS testing.
Only trigger expensive jobs when relevant files change. Skip the full test suite if only documentation changed.
on:
push:
paths:
- 'src/**'
- 'tests/**'
- 'requirements*.txt'
Order Dockerfile instructions from least to most frequently changing. Copy and install dependencies before copying application code.
The DORA (DevOps Research and Assessment) team identified four metrics that predict software delivery performance.
How often does the team successfully deploy to production? Elite teams deploy multiple times per day. Frequent, small deployments are safer than infrequent large ones.
Time from code committed to code running in production. Elite: < 1 hour. Measures the efficiency of the entire pipeline, including review and approvals.
Percentage of deployments that cause a production failure requiring rollback or hotfix. Elite: 0–5%. A high rate indicates inadequate testing or deployment strategy.
How long to restore service after a failure. Elite: < 1 hour. Requires fast detection (observability), fast rollback, and on-call processes.
These metrics are positively correlated with business outcomes — teams in the elite tier have 127× faster lead time than low performers (DORA State of DevOps Report).
Every CI/CD term used on this GitHub's projects, with the slide that explains it. Bare numbers are slides of this deck; GHA = Introduction to GitHub Actions, Jenkins = Introduction to Jenkins.
Pull request (PR) 12 a proposed merge, where CI and review happen
Branch protection 12 rules that guard a branch from direct pushes
Ruleset 12 GitHub's newer, layered, visible branch rules
Required status check 12 a named check that must pass to merge
Required reviews GHA 19 approvals needed before merging
CODEOWNERS 12 file naming who must review which paths
Force push 12 rewriting a branch's history on the remote
Bypass list / admin enforcement 12 who, if anyone, may skip the rules
Squash merge 12 a PR lands as one commit
Green at HEAD 12 the latest commit passed every check
CI status badge 12 README image of the latest run's result
Dependabot GHA 14 GitHub's bot that opens dependency-bump PRs
workflow_dispatch GHA 04 a manual Run workflow trigger
Re-running a workflow 17 running the same commit's checks again
Downstream (dependent) repo 17 a repo that installs yours
Pinning to a commit GHA 14 referencing one immutable SHA, not a branch
repository_dispatch GHA 23 one repo's CI starting another's
Webhook Jenkins 10 the forge calls CI when something happens
Scheduled (nightly) build Jenkins 10 a run on a timetable, not a push
Concurrency group GHA 18 runs that must not overlap
Lockfile / frozen install 17 exact versions; CI refuses to change them
Stale environment 17 a reused venv keeps an old dependency
Vendoring 17 a pinned, checked copy of another repo's file
Scheduled drift check 17 a timed run that flags upstream changes to pins
Pipeline as code 05 the pipeline lives in the repo as a file
Runner / agent 15 the machine a job runs on
Matrix build GHA 06 one job run per combination of values
Cache / cache key 24 saved downloads reused by later runs
Artefact 13 a build output stored and promoted
Service container GHA 07 a database container beside the job
Deployment environment GHA 13 a named target with its own secrets/rules
Shared library Jenkins 14 pipeline code shared by many repos
Multibranch pipeline Jenkins 15 one Jenkins job per branch and PR
Publishing to GitHub Pages 20 a branch served as a static site
Quality gate 08 a threshold the pipeline enforces
Lint 09 flag suspicious code without running it
Typecheck 09 the compiler checks every type
Format check 09 fail if any file isn't formatted
actionlint GHA 16 a linter for workflow files
Lighthouse 09 Google's page auditor, scores 0-100
Lighthouse CI (LHCI) 09 Lighthouse in CI, failing below a score
Performance budget 09 the score or size limit a gate enforces
Core Web Vitals 09 LCP, INP, CLS: real users' experience
Coverage target 08 minimum share of code the tests run
Unit / integration test 07 one unit alone / parts together
End-to-end (e2e) test 07 drive the whole app like a user
Playwright 10 a browser-automation test runner
e2e on a production build 10 test the bundle that ships
Test fixture 10 fixed inputs a test runs against
Golden / headline-number check 10 output must equal a reviewed answer
Parity test (bit-exact) 10 two implementations must agree exactly
Differential test 10 same random inputs through two versions
Up-to-date check 10 CI regenerates outputs and fails on a difference
Property-based test 10 generated inputs against a stated rule
Mutation testing 10 planted bugs the tests should catch
Flaky test 10 passes and fails on the same code
Performance regression gate 10 fail if a benchmark slows past noise
Silent fallback 23 an error swallowed and replaced by a default
Preview deployment 20 a throwaway copy built from a branch
Deployment protection / protection bypass 20 a login gate on deployment URLs; a token lets checks through
Production deployment 20 the one the real domain serves
Git-integrated vs CLI deploy 20 deploy on push / deploy a folder
Clean export (git archive) 20 deploy exactly the commit, nothing local
Environment variable 20 a setting the app reads at run time
Secret / masking GHA 13 encrypted setting, shown as *** in logs
Sensitive (write-only) variable 20 can't be read back after saving
OAuth callback URL 20 where sign-in returns the user
Custom domain 20 your own name for the production site
Redirect (308) 20 old URL permanently forwards to the new
OIDC 18 short-lived cloud credentials, no stored key
Least privilege 18 grant only the access a job needs
Schema migration 20 a versioned change to the database
Seeding 20 loading starter rows after migrating
Migrate before deploy 20 schema change first, then new code
Expand/contract 20 add, switch, then remove, across releases
Blue/green 19 two environments, switch traffic
Canary release 19 a small share of traffic first
Feature flag 19 code shipped dark, switched on at run time
Rollback 22 return to the previous good version
Post-deploy verification 21 checking the live system after a deploy
Health endpoint 21 URL reporting if dependencies are OK
Smoke check 21 fast, shallow test of the live system
Smoke test on real data 21 exercise code paths empty data skips
Runtime logs / error groups 21 the platform's record of live errors
Runbook 21 the written procedure for an operation
Render before write 21 do what can fail before storing
Works locally, fails on the platform 23 production's runtime differs from CI's
Module loader (ESM vs require) 23 how a runtime loads JS packages
DORA metrics 25 four delivery-performance measures
.github/workflows/ci.yml to your next projectmainHumble & Farley, Continuous Delivery (2010) · Kim et al., The DevOps Handbook (2nd ed. 2021) · DORA State of DevOps Reports · Martin Fowler's ContinuousIntegration article (martinfowler.com)