devopsdiary
Next chapter

diary / chapters / 09

chapter 09 · ~50 min · pipeline lab

Continuous delivery and releasing safely

CI proves a build is good. Delivery is everything after that: getting it in front of users in a way that is boring, reversible and observable. The goal is a deploy so unremarkable you would do one at 4 p.m. on a Friday.

TOPIC 01Delivery vs deployment

Three phrases get used interchangeably and should not be:

Continuous integration
Every push is built and tested automatically. (Chapter 05.)
Continuous delivery
Every build that passes is ready to go to production at the push of a button. The remaining step is a human decision, not a human process.
Continuous deployment
Every build that passes goes to production automatically, with no human step at all.

Most teams should aim for continuous delivery first. And there is a fourth distinction that matters more than any of them: deploy is not the same as release. Deploying puts code on the machines; releasing exposes the behaviour to users. Feature flags let you separate the two, which is what turns a risky launch into a dial you can turn.

real life

A new dish in a restaurant. Deploying is having it prepped and ready in the kitchen. Releasing is putting it on the menu. Doing both at the same instant, for every table, at the busiest hour, is how restaurants get bad reviews — offer it to one table first, watch their faces, then print the menus.

TOPIC 02Environments that mean something

EnvironmentPurposeDataWho deploys
localFast iterationFake / seeded (Compose, chapter 06)You, constantly
preview / PRReview a change in a live URLSeededAutomatic per pull request
stagingRehearsal for productionAnonymised copy of real shapesAutomatic on merge to main
productionReal usersRealAutomatic, or a button after approval

A staging environment only earns its keep if it resembles production: same artifact, same configuration shape, same infrastructure code with smaller numbers (chapter 08's module pattern). The classic failure is a staging environment with one instance, no TLS, and a tiny database — it will happily pass a change that falls over the moment it meets concurrency.

the one thing staging cannot give you

Real traffic. Real users click in strange orders, on old browsers, with slow networks and 20,000 items in a cart. That is why the strategies below exist: they use a small slice of production traffic as the final test, with a fast exit.

TOPIC 03Promoting one artifact

Chapter 05 established the rule; here is what it looks like end to end. Build the image once, tag it with the commit SHA, and move that exact digest through the environments.

one build, three environments
commit a91f3c
   │
   ├─ CI: test, build, scan ──▶ ghcr.io/acme/api:a91f3c   (immutable)
   │
   ├─ deploy to staging   → api:a91f3c   ✓ smoke tests pass
   ├─ deploy canary (5%)  → api:a91f3c   ✓ error rate flat for 10 min
   └─ deploy production   → api:a91f3c   ✓ same bytes, all the way through

# configuration differs per environment; the artifact never does
staging:    DATABASE_URL=…staging…   LOG_LEVEL=debug   REPLICAS=1
production: DATABASE_URL=…prod…      LOG_LEVEL=info    REPLICAS=6

The corollary: never build from a branch per environment. Environment branches guarantee that what you tested and what you shipped were compiled at different moments from different inputs, which quietly voids the test results.

TOPIC 04Release strategies

Recreate (stop the old, start the new)

Simple and honest: brief downtime. Fine for internal tools and batch jobs; unacceptable for anything customer-facing.

Rolling update

Replace instances a few at a time, waiting for each new one to become ready. This is the Kubernetes default (chapter 07, maxSurge/maxUnavailable). Zero downtime, cheap, and briefly runs two versions at once — which means your app and its database must tolerate that.

Blue-green

two identical environments, one switch
            ┌─────────────┐
  100% ────▶│ BLUE  v1.4  │  live
            └─────────────┘
            ┌─────────────┐
    0% ────▶│ GREEN v1.5  │  deployed, warmed, smoke-tested
            └─────────────┘
                   │
        flip the load balancer ──▶ green live, blue kept idle
        rollback = flip it back (seconds)

Fast, dramatic rollback; costs double the infrastructure during the switch, and long-lived connections need draining.

Canary

a small slice first, watching the numbers
  5%  ─▶ v1.5   watch error rate + p95 latency for 10 min
 25%  ─▶ v1.5   still flat?
 50%  ─▶ v1.5   still flat?
100%  ─▶ v1.5   promote, retire v1.4

 any step regresses ──▶ shift back to 0% automatically

The best risk/cost trade-off for anything with meaningful traffic, and the one worth learning to argue for. Its prerequisite is the next chapter: a canary without per-version metrics is just a slower way to break things.

real life

A canary release is tasting a spoonful before serving the pot; blue-green is cooking a second pot and swapping it in; rolling is replacing the dishes one table at a time. Which one you pick depends on how much a bad spoonful costs and how much a second pot costs.

TOPIC 05Feature flags

A flag is a runtime switch around new behaviour. Deploy the code dark, turn it on for yourself, then 1% of users, then everyone — and switch it off in seconds without a deploy.

a flag in code, and what it buys you
if (flags.enabled("new_checkout", { userId: user.id })) {
  return newCheckout(cart);
} else {
  return legacyCheckout(cart);
}

# the operational payoff:
#  deploy at 11:00 with the flag off        → zero user impact
#  enable for internal users at 11:10       → real production test
#  1% → 10% → 50% → 100% over two days      → gradual exposure
#  something looks wrong at 14:32           → off in 5 seconds, no rollback

Flags also decouple deployment from marketing: the launch can happen at 10 a.m. Monday without anyone deploying at 10 a.m. Monday. The discipline they demand is cleanup — every flag is a branch in your code, and a codebase with 300 stale flags is unreadable. Give each flag an owner and a removal date when you create it.

TOPIC 06Rollback: the real safety net

Confidence to deploy comes from the ability to undo, not from believing the change is perfect. Before any deploy you should be able to answer, out loud: how do I undo this, and how long does it take?

terminal · the undo commands, by platform
# kubernetes
kubectl rollout undo deployment/api
kubectl rollout status deployment/api

# plain docker host
docker compose up -d --no-deps api   # with the previous tag pinned

# terraform (infrastructure)
git revert <sha> && terraform plan   # revert the code, review, then apply

# gitops
git revert <sha> && git push          # the cluster reconciles itself back

# feature flag
# flip the switch — no deploy at all, seconds not minutes

Three habits that make rollback reliable: keep the previous artifact available (never delete the last few images), practise it deliberately in staging so the command is muscle memory, and note that forward-fixing under pressure has a much worse track record than reverting. Roll back first, understand second — the Incident Room game exists to drill exactly that instinct.

TOPIC 07Database migrations

Code rolls back easily. Data does not. This is the hardest part of safe releases, and it has a standard solution: never make a breaking schema change in one step.

expand → migrate → contract
Goal: rename column `name` to `full_name`.

✗ the one-step version
ALTER TABLE users RENAME COLUMN name TO full_name;
   During a rolling deploy, old pods still SELECT name → 500s.
   And a rollback now cannot work either. Both versions are broken.

✓ three deploys, always safe
1. EXPAND    add full_name (nullable). code writes BOTH, reads name.
2. MIGRATE   backfill full_name in batches. code reads full_name,
             still writes both. old pods keep working throughout.
3. CONTRACT  once no code touches name (and you are sure), drop it.

Rules to carry with you: migrations run before the new code and must be backwards compatible with the old code. Additive changes (new nullable column, new table, new index built concurrently) are safe. Destructive changes (drop, rename, narrow a type, add a NOT NULL without a default) are separate, later, deliberate deploys. And always know how to reverse a migration — or make it additive so you do not need to.

real life

Replacing a bridge while traffic still flows. You build the new span alongside, divert traffic gradually, and only then demolish the old one. Nobody closes the only bridge at rush hour and hopes the new one is finished by evening — but that is precisely what a one-step rename during a rolling deploy does.

TOPIC 08GitOps with Argo CD

In a push pipeline, CI holds cluster credentials and runs kubectl apply. In GitOps, that is inverted: a Git repository holds the desired state, and an agent inside the cluster continuously pulls it and reconciles reality to match — the chapter-07 thermostat, with Git as the dial.

the flow
  app repo ──CI──▶ image api:b72d10 ──▶ registry
                              │
                              ▼ (automated commit)
  config repo:  deployment.yaml  image: api:b72d10
                              │
                              ▼ Argo CD notices the commit (or is notified)
                        cluster reconciles → 6 pods on b72d10
                              │
  someone edits the cluster by hand ──▶ Argo marks it OUT OF SYNC
                                        and (optionally) reverts it

What you gain: every change to production is a reviewed commit; the current state of every environment is readable in a repo; rollback is git revert; drift is detected and can be self-healed; and no CI system needs standing cluster credentials. What it costs: a second repository to keep tidy, and a new failure mode when the agent's view and yours disagree.

terminal · argo cd, briefly
argocd app list
argocd app get notes-api            # Synced? Healthy? which revision?
argocd app diff notes-api           # repo vs cluster
argocd app sync notes-api
argocd app rollback notes-api 12

TOPIC 09A release checklist worth keeping

  • Before: CI green on the exact commit · the artifact is immutable and already scanned · migrations are additive and applied first · the rollback command is known and tested · someone other than you knows the deploy is happening.
  • During: deploy to a small slice first · watch error rate and p95 latency, not just "did the pods start" · give it long enough for slow failures (cache expiry, cron jobs) to appear.
  • After: run one real user journey yourself · check the dashboards you would check during an incident · leave the previous version available for a day · write down anything surprising while you still remember it.
  • Never: deploy something you cannot undo, at a time when nobody is around to notice, from an unpushed local branch.
remember this much
  • Delivery = always ready to ship; deployment = shipping automatically. Deploy ≠ release.
  • Build one artifact, promote the same digest; change configuration, never the bytes.
  • Rolling for cheap zero-downtime, blue-green for instant rollback, canary for the best risk/cost balance.
  • Feature flags separate exposure from deployment and turn a rollback into a switch.
  • Know your undo command before you deploy. Roll back first, diagnose second.
  • Schema changes: expand, migrate, contract — never a breaking change in one step.
  • GitOps makes Git the source of truth and rollback a revert.

LABA canary you can revert

50 minutes · your chapter 05 pipeline + chapter 07 cluster

Promote an artifact, ship a canary, then undo it under time pressure

  1. Extend your CI workflow: after tests pass, build the image and push it tagged ${{ github.sha }} to ghcr.io. No latest anywhere.
  2. Add a deploy-staging job that runs only on main, sets that exact tag on your Deployment, waits for kubectl rollout status, then curls /health as a smoke test. Make the job fail if the smoke test fails — a deploy that cannot fail is not a gate.
  3. Add a deploy-production job with needs: deploy-staging and a GitHub environment: that requires manual approval. Merge something and approve it. You now have continuous delivery.
  4. Hand-rolled canary: create a second Deployment api-canary with 1 replica on the new tag, sharing the Service's label so it receives roughly 1/(N+1) of traffic. Verify with a loop of curl that you see both versions respond.
  5. Make the canary bad on purpose: point it at an image whose /health returns 500. Watch the error rate in your loop. Now delete the canary Deployment and watch it clear. Time the whole detect-and-revert cycle and write the number down.
  6. Practise the real rollback: promote the bad version to the main Deployment, then recover with kubectl rollout undo. Do it twice, until you can do it without looking anything up.
  7. Migration drill: take any table and rename a column properly — deploy 1 adds the new column and writes both, deploy 2 reads the new one, deploy 3 drops the old. Confirm after each deploy that the previous version of the app would still work.
  8. Stretch: install Argo CD in your local cluster, point it at a config repo, and deploy by committing an image tag. Then git revert and watch the cluster roll itself back with nobody touching kubectl.

Steps 5 and 6 are the deliverable. Anyone can deploy; the interview answer that lands is "we detected a bad canary in 90 seconds and reverted in 30, and here is how I know".

CHECKCheck yourself

During a rolling update, a migration renames a column in one step. What happens?

A rolling update deliberately runs both versions at once, so the schema must satisfy both. That is why breaking changes are split into expand → migrate → contract. The nastiest part is the rollback trap: reverting the code returns you to a version that also cannot read the new schema, so you are stuck fixing forward during an incident.

Why is a canary release usually a better bet than deploying to 100% and watching?

A canary trades a little time for a much smaller blast radius: if the new version is bad, 5% of requests suffer briefly instead of all of them. It is strictly slower than a big-bang deploy, and it depends on having per-version metrics to compare — which is why chapter 10 comes next.

What is the main operational advantage of GitOps over a pipeline that runs kubectl apply?

The value is auditability and reconciliation: every production change is a commit someone approved, a hand-edit shows up as out-of-sync, and undo is git revert. As a bonus, no external CI system needs standing cluster credentials — the agent pulls rather than being pushed to.

saved in this browser only — no account needed