devopsdiary
Next chapter

diary / chapters / 05

chapter 05 · ~50 min · github lab

Continuous integration: the robot reviewer

CI is a machine that checks out every push, builds it, tests it and tells you within minutes whether it is safe. It is the cheapest quality improvement available to any team, and the foundation everything after this chapter stands on.

TOPIC 01The problem CI solves

Five developers work for two weeks on five branches. Each one's code works. On merge day, nothing works: two of them upgraded the same library to different versions, one renamed a function three others call, and a test that nobody ran in a fortnight has been failing since Tuesday. The team spends three days in "integration hell" — a phrase that existed because the pain was universal.

Continuous integration is the fix, and the name says it: integrate continuously. Merge small changes into the shared branch many times a day, and have a machine prove after every single push that the combined result still builds and passes its tests.

real life

Five cooks preparing one thali. Without CI, everyone cooks in a separate kitchen for two weeks and the dishes meet for the first time on the plate — three of them turn out to be the same sabzi and nothing is warm. With CI, everything goes onto a shared tasting spoon after every step: mistakes are caught while they are still one spoonful, not one banquet.

What you get for the price of one YAML file:

  • Fast feedback. A broken build is reported in minutes, while the change is still fresh in the author's head.
  • One consistent environment. The runner is clean every time, so "it works on my machine" stops being an argument.
  • A definition of "done" that nobody can skip. Tests, linters, format checks and scans run whether or not anyone remembered.
  • Reviewers who look at logic. Humans stop hunting for missing semicolons and start asking the questions only humans can.

TOPIC 02Anatomy of a workflow

Every CI system on earth uses the same four nested ideas, under slightly different names. Learn them once:

Trigger (event)
What starts it: a push, a pull request, a tag, a schedule, a manual click.
Job
A unit of work that runs on one machine. Jobs run in parallel unless you declare a dependency.
Runner
The machine (usually a fresh container or VM) the job runs on. Provided by the CI service, or self-hosted.
Step
One command, or one reusable action. Steps run in order and share the same filesystem.
the vocabulary map

GitHub Actions: workflow → jobs → steps. GitLab CI: pipeline → stages → jobs. Jenkins: pipeline → stages → steps. Same concepts. Learn one properly and you can read all three, which is why this chapter uses GitHub Actions and does not apologise for it.

TOPIC 03Your first pipeline

Create .github/workflows/ci.yml. On push, this checks out the code, installs dependencies, and runs lint plus tests.

.github/workflows/ci.yml
name: CI

on:
  push:
    branches: [main]
  pull_request:            # run on every PR — this is the important one

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4          # get the code

      - uses: actions/setup-node@v4        # install a toolchain
        with:
          node-version: '20'
          cache: 'npm'                     # cache deps between runs

      - run: npm ci                        # ci, not install: honours the lockfile
      - run: npm run lint
      - run: npm test -- --coverage

      - name: Upload coverage
        if: always()                       # run even when tests failed
        uses: actions/upload-artifact@v4
        with:
          name: coverage
          path: coverage/

Points that are easy to miss and matter a lot:

  • pull_request is the trigger that changes team behaviour. Checks appear on the PR, so nothing merges red.
  • npm ci (or pip install -r with hashes, or go mod download) installs exactly the lockfile. Using npm install in CI lets dependencies drift, and a build that is not reproducible is not evidence of anything.
  • Each step is a shell command. If it exits non-zero, the job fails. That is the whole contract — which means anything you can script, you can gate on.
  • if: always() is how you still collect logs, coverage and screenshots from a failed run, when you need them most.

TOPIC 04The testing pyramid

CI is only as useful as what it runs. The pyramid is about proportions: many cheap tests, few expensive ones.

shape of a healthy test suite
            ╱ E2E ╲            few · slow (minutes) · brittle · highest confidence
          ╱─────────╲          real browser, real stack: "can a user check out?"
        ╱ integration ╲        some · seconds · your code + a real DB or queue
      ╱─────────────────╲
    ╱       unit          ╲    many · milliseconds · one function, no I/O
  ╱─────────────────────────╲

  static checks (linter, formatter, type checker, secret scan) — instant, run first
real life

Building a house. Unit tests are checking each brick is not cracked. Integration tests are checking the wall stands. End-to-end tests are turning on a tap upstairs to see whether water arrives. You need all three — but if you only ever test by turning on taps, finding the one cracked brick takes a demolition crew.

Practical guidance: put static checks first, because they take seconds and fail loudest. Keep unit tests fast enough to run on every save. Reserve end-to-end tests for the two or three journeys that are the business (sign-up, checkout, search). And keep the whole pull-request pipeline under about ten minutes — beyond that, people start pushing without waiting, and CI quietly stops being a gate.

TOPIC 05Making it fast

Pipeline speed is a feature. Four levers, in order of payoff:

  • Cache dependencies. A cold npm ci can take two minutes; a warm cache takes ten seconds.
  • Run jobs in parallel. Lint, unit tests and build have no reason to wait for each other.
  • Fail fast. Order steps cheapest-first so a formatting error does not cost you a full test run.
  • Only run what changed in a monorepo, using path filters.
.github/workflows/ci.yml — parallel jobs with a gate
jobs:
  lint:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: '20', cache: 'npm' }
      - run: npm ci
      - run: npm run lint

  unit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: '20', cache: 'npm' }
      - run: npm ci
      - run: npm test

  build:
    needs: [lint, unit]        # only if both passed
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm ci && npm run build
      - uses: actions/upload-artifact@v4
        with: { name: dist, path: dist/ }

concurrency:                   # cancel superseded runs on the same branch
  group: ci-${{ github.ref }}
  cancel-in-progress: true
manual cache, when you need control
      - uses: actions/cache@v4
        with:
          path: ~/.npm
          key: npm-${{ hashFiles('**/package-lock.json') }}
          restore-keys: npm-        # fall back to the newest similar cache

The key is a hash of the lockfile, so the cache is reused while dependencies are unchanged and rebuilt the moment they change. Getting that key wrong is how teams end up with either a permanently cold cache or a stale one — both quietly expensive.

TOPIC 06Matrix builds

A matrix runs the same job across several versions or platforms in parallel — essential if you ship a library, and useful for any upgrade you are nervous about.

one job, six runs
  test:
    runs-on: ${{ matrix.os }}
    strategy:
      fail-fast: false          # let the others finish and report
      matrix:
        os: [ubuntu-latest, macos-latest]
        node: ['18', '20', '22']
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: ${{ matrix.node }} }
      - run: npm ci && npm test

TOPIC 07Artifacts: build once, deploy many

An artifact is the output of a build — a JAR, a dist/ folder, a container image. The rule that prevents a whole family of bugs: build once, then promote the same artifact through every environment.

real life

Printing exam papers. You print one master set and photocopy it for every room. If each room printed its own from the file, one printer would be low on toner and one question would come out unreadable — and you would only discover it during the exam. Rebuilding per environment is that separate printer.

If you rebuild for staging and again for production, the two builds can differ: a dependency published a new patch version, a base image moved, a timestamp changed. Then "it passed in staging" stops meaning anything. Chapter 09 turns this rule into a promotion pipeline; chapter 06 makes the artifact a container image, which is the most useful form of it.

TOPIC 08Secrets in CI

Pipelines need credentials — a registry password, a cloud role, a deploy key. Rules that are not negotiable:

  • Store them in the CI provider's secret store (GitHub → Settings → Secrets), never in the repository.
  • Reference them as ${{ secrets.NAME }}. They are masked in logs — but only if you never echo them yourself.
  • Scope them to the minimum: one secret per pipeline, with the least permission that works.
  • Prefer short-lived credentials via OIDC over long-lived keys. The pipeline exchanges its identity for a token that expires in an hour, so there is no static key to leak.
  • Be careful with untrusted pull requests: forks should never have access to production secrets. That is why GitHub withholds them from pull_request runs on forks by default — a default worth understanding rather than working around.
a secret used correctly
      - name: Log in to the registry
        run: echo "${{ secrets.REGISTRY_TOKEN }}" | docker login ghcr.io -u ${{ github.actor }} --password-stdin
        # --password-stdin keeps it out of the process list and shell history

      - name: Deploy
        env:
          DEPLOY_TOKEN: ${{ secrets.DEPLOY_TOKEN }}   # pass via env, not argv
        run: ./scripts/deploy.sh

TOPIC 09Testing against a real database

Mocks are fine for unit tests, but a query that works against a mock and fails against Postgres has taught you nothing. CI can start real services next to your job for the duration of the run.

integration tests with a real Postgres
  integration:
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:16
        env:
          POSTGRES_PASSWORD: testpass
        options: >-
          --health-cmd pg_isready
          --health-interval 5s
          --health-retries 10        # wait until it is actually ready
        ports: ['5432:5432']
    env:
      DATABASE_URL: postgres://postgres:testpass@localhost:5432/postgres
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: npm run migrate
      - run: npm run test:integration

The health check is the part beginners skip and then spend an afternoon on: without it, your tests start before Postgres is accepting connections and fail intermittently — a self-inflicted flaky test, which brings us to the next topic.

TOPIC 10Flaky tests and the red build

A flaky test passes and fails without the code changing. Flakiness is more dangerous than a plain failure, because it destroys the meaning of the signal: once people re-run red builds by reflex, a real failure will ship.

  • Common causes: a hard-coded sleep instead of waiting for a condition; tests sharing state or a database row; reliance on the current time or timezone; test order dependence; calling a real external API.
  • The rule: a flaky test is a bug with a ticket. Fix it, or quarantine it out of the gating suite — but never leave it as a coin-flip on the main pipeline.
  • Keep main green. If main goes red, fixing it is the team's top priority, ahead of features. A red main means nobody else can tell whether their change is broken.
the culture bit

CI only works if the team treats red as "stop". The technical setup takes an afternoon; the agreement that red blocks merging is the actual investment. Branch protection (chapter 04) is how you make that agreement structural instead of aspirational.

TOPIC 11Jenkins, GitLab, and the rest

ToolShapeWhere you'll meet it
GitHub ActionsYAML in the repo, hosted runners, huge action marketplaceMost new projects and open source
GitLab CI.gitlab-ci.yml, stages, built-in registry and environmentsCompanies self-hosting their whole DevOps stack
JenkinsSelf-hosted server, plugins, Jenkinsfile (Groovy)Established enterprises; enormous installed base
CircleCI / Buildkite / DroneHosted or hybrid, strong caching and parallelismTeams that outgrew a free tier
Argo Workflows / TektonPipelines as Kubernetes objectsKubernetes-native platforms (chapter 09)

Learn the concepts, not the syntax. Every one of these has a trigger, a runner, jobs, steps, caches, artifacts and secrets. When you change jobs and the tool changes with it, you are translating, not relearning.

remember this much
  • CI = every push is automatically built and tested on a clean machine, fast.
  • Trigger → job → runner → step. Same four ideas in every CI tool.
  • Static checks, then unit, then integration, then a couple of end-to-end journeys.
  • Install from the lockfile (npm ci), cache dependencies, run independent jobs in parallel, keep PR runs under ten minutes.
  • Build the artifact once and promote it; never rebuild per environment.
  • Secrets live in the CI secret store, scoped small and preferably short-lived.
  • Flaky tests are bugs. A red main branch stops the team.

LABA pipeline you would actually keep

50 minutes · a GitHub repo · free tier is plenty

From zero to a gate that blocks bad merges

  1. In any small app repo (Node, Python, Go — pick what you know), add .github/workflows/ci.yml with the first workflow above. Push and watch it run in the Actions tab.
  2. Prove it fails properly: break a test on purpose, push, and confirm the run goes red and the log shows exactly which assertion failed. A pipeline you have never seen fail is not yet trustworthy.
  3. Split into three jobs — lint, unit, build — with needs: [lint, unit] on build. Note in the Actions UI that the first two now run side by side.
  4. Add dependency caching. Compare the run duration before and after; write the two numbers down, because "I cut our CI from 4m to 90s" is an interview answer.
  5. Add the concurrency block. Push twice quickly and watch the first run get cancelled.
  6. Add an artifact upload for your build output, then download it from the run summary page and confirm it contains what you expect.
  7. Turn on branch protection for main: require these checks to pass. Now open a PR with a failing test and confirm the merge button is blocked. This step is the whole point of the chapter.
  8. Stretch: add an integration job with a services: Postgres and one test that really queries it. Then add a matrix over two language versions and see six runs report independently.

You now have the thing chapter 06 will package, chapter 09 will deploy, and chapter 11 will scan. Keep this repo — it becomes portfolio project one in chapter 12.

CHECKCheck yourself

Why use npm ci rather than npm install in a pipeline?

npm install may resolve a newer permitted version and rewrite the lockfile, which means today's green build and tomorrow's red build can come from identical source. npm ci installs the locked tree exactly, failing if the lockfile and manifest disagree. Every ecosystem has this pair — pip-sync, bundle install --deployment, go mod download — and the reproducibility argument is the same each time.

Your team rebuilds the application separately for staging and for production. What is the risk?

A transitive dependency publishing a patch, a base image moving, or a build timestamp is enough to make the production binary different from the tested one. Build once, tag the artifact with the commit SHA, and promote that exact artifact. The wasted CI minutes are real but trivial next to losing the meaning of your test results.

A test fails roughly one run in five with no code changes. The team re-runs the pipeline until it passes. What is the correct response?

Habitual re-running trains everyone to ignore red, which is how a genuine failure reaches production. Fix the cause — usually a missing wait-for-condition, shared state, or a time dependency. Blanket auto-retry hides the signal; deleting the test throws away coverage. Quarantine is an acceptable interim step precisely because it keeps the gate meaningful while the bug is worked.

saved in this browser only — no account needed