TOPIC 01Delivery vs deployment
Three phrases get used interchangeably and should not be:
- Continuous integration
- Every push is built and tested automatically. (Chapter 05.)
- Continuous delivery
- Every build that passes is ready to go to production at the push of a button. The remaining step is a human decision, not a human process.
- Continuous deployment
- Every build that passes goes to production automatically, with no human step at all.
Most teams should aim for continuous delivery first. And there is a fourth distinction that matters more than any of them: deploy is not the same as release. Deploying puts code on the machines; releasing exposes the behaviour to users. Feature flags let you separate the two, which is what turns a risky launch into a dial you can turn.
A new dish in a restaurant. Deploying is having it prepped and ready in the kitchen. Releasing is putting it on the menu. Doing both at the same instant, for every table, at the busiest hour, is how restaurants get bad reviews — offer it to one table first, watch their faces, then print the menus.
TOPIC 02Environments that mean something
| Environment | Purpose | Data | Who deploys |
|---|---|---|---|
| local | Fast iteration | Fake / seeded (Compose, chapter 06) | You, constantly |
| preview / PR | Review a change in a live URL | Seeded | Automatic per pull request |
| staging | Rehearsal for production | Anonymised copy of real shapes | Automatic on merge to main |
| production | Real users | Real | Automatic, or a button after approval |
A staging environment only earns its keep if it resembles production: same artifact, same configuration shape, same infrastructure code with smaller numbers (chapter 08's module pattern). The classic failure is a staging environment with one instance, no TLS, and a tiny database — it will happily pass a change that falls over the moment it meets concurrency.
Real traffic. Real users click in strange orders, on old browsers, with slow networks and 20,000 items in a cart. That is why the strategies below exist: they use a small slice of production traffic as the final test, with a fast exit.
TOPIC 03Promoting one artifact
Chapter 05 established the rule; here is what it looks like end to end. Build the image once, tag it with the commit SHA, and move that exact digest through the environments.
commit a91f3c │ ├─ CI: test, build, scan ──▶ ghcr.io/acme/api:a91f3c (immutable) │ ├─ deploy to staging → api:a91f3c ✓ smoke tests pass ├─ deploy canary (5%) → api:a91f3c ✓ error rate flat for 10 min └─ deploy production → api:a91f3c ✓ same bytes, all the way through # configuration differs per environment; the artifact never does staging: DATABASE_URL=…staging… LOG_LEVEL=debug REPLICAS=1 production: DATABASE_URL=…prod… LOG_LEVEL=info REPLICAS=6
The corollary: never build from a branch per environment. Environment branches guarantee that what you tested and what you shipped were compiled at different moments from different inputs, which quietly voids the test results.
TOPIC 04Release strategies
Recreate (stop the old, start the new)
Simple and honest: brief downtime. Fine for internal tools and batch jobs; unacceptable for anything customer-facing.
Rolling update
Replace instances a few at a time, waiting for each new one to become ready. This is the Kubernetes default (chapter 07, maxSurge/maxUnavailable). Zero downtime, cheap, and briefly runs two versions at once — which means your app and its database must tolerate that.
Blue-green
┌─────────────┐
100% ────▶│ BLUE v1.4 │ live
└─────────────┘
┌─────────────┐
0% ────▶│ GREEN v1.5 │ deployed, warmed, smoke-tested
└─────────────┘
│
flip the load balancer ──▶ green live, blue kept idle
rollback = flip it back (seconds)
Fast, dramatic rollback; costs double the infrastructure during the switch, and long-lived connections need draining.
Canary
5% ─▶ v1.5 watch error rate + p95 latency for 10 min 25% ─▶ v1.5 still flat? 50% ─▶ v1.5 still flat? 100% ─▶ v1.5 promote, retire v1.4 any step regresses ──▶ shift back to 0% automatically
The best risk/cost trade-off for anything with meaningful traffic, and the one worth learning to argue for. Its prerequisite is the next chapter: a canary without per-version metrics is just a slower way to break things.
A canary release is tasting a spoonful before serving the pot; blue-green is cooking a second pot and swapping it in; rolling is replacing the dishes one table at a time. Which one you pick depends on how much a bad spoonful costs and how much a second pot costs.
TOPIC 05Feature flags
A flag is a runtime switch around new behaviour. Deploy the code dark, turn it on for yourself, then 1% of users, then everyone — and switch it off in seconds without a deploy.
if (flags.enabled("new_checkout", { userId: user.id })) {
return newCheckout(cart);
} else {
return legacyCheckout(cart);
}
# the operational payoff:
# deploy at 11:00 with the flag off → zero user impact
# enable for internal users at 11:10 → real production test
# 1% → 10% → 50% → 100% over two days → gradual exposure
# something looks wrong at 14:32 → off in 5 seconds, no rollback
Flags also decouple deployment from marketing: the launch can happen at 10 a.m. Monday without anyone deploying at 10 a.m. Monday. The discipline they demand is cleanup — every flag is a branch in your code, and a codebase with 300 stale flags is unreadable. Give each flag an owner and a removal date when you create it.
TOPIC 06Rollback: the real safety net
Confidence to deploy comes from the ability to undo, not from believing the change is perfect. Before any deploy you should be able to answer, out loud: how do I undo this, and how long does it take?
# kubernetes kubectl rollout undo deployment/api kubectl rollout status deployment/api # plain docker host docker compose up -d --no-deps api # with the previous tag pinned # terraform (infrastructure) git revert <sha> && terraform plan # revert the code, review, then apply # gitops git revert <sha> && git push # the cluster reconciles itself back # feature flag # flip the switch — no deploy at all, seconds not minutes
Three habits that make rollback reliable: keep the previous artifact available (never delete the last few images), practise it deliberately in staging so the command is muscle memory, and note that forward-fixing under pressure has a much worse track record than reverting. Roll back first, understand second — the Incident Room game exists to drill exactly that instinct.
TOPIC 07Database migrations
Code rolls back easily. Data does not. This is the hardest part of safe releases, and it has a standard solution: never make a breaking schema change in one step.
Goal: rename column `name` to `full_name`.
✗ the one-step version
ALTER TABLE users RENAME COLUMN name TO full_name;
During a rolling deploy, old pods still SELECT name → 500s.
And a rollback now cannot work either. Both versions are broken.
✓ three deploys, always safe
1. EXPAND add full_name (nullable). code writes BOTH, reads name.
2. MIGRATE backfill full_name in batches. code reads full_name,
still writes both. old pods keep working throughout.
3. CONTRACT once no code touches name (and you are sure), drop it.
Rules to carry with you: migrations run before the new code and must be backwards compatible with the old code. Additive changes (new nullable column, new table, new index built concurrently) are safe. Destructive changes (drop, rename, narrow a type, add a NOT NULL without a default) are separate, later, deliberate deploys. And always know how to reverse a migration — or make it additive so you do not need to.
Replacing a bridge while traffic still flows. You build the new span alongside, divert traffic gradually, and only then demolish the old one. Nobody closes the only bridge at rush hour and hopes the new one is finished by evening — but that is precisely what a one-step rename during a rolling deploy does.
TOPIC 08GitOps with Argo CD
In a push pipeline, CI holds cluster credentials and runs kubectl apply. In GitOps, that is inverted: a Git repository holds the desired state, and an agent inside the cluster continuously pulls it and reconciles reality to match — the chapter-07 thermostat, with Git as the dial.
app repo ──CI──▶ image api:b72d10 ──▶ registry
│
▼ (automated commit)
config repo: deployment.yaml image: api:b72d10
│
▼ Argo CD notices the commit (or is notified)
cluster reconciles → 6 pods on b72d10
│
someone edits the cluster by hand ──▶ Argo marks it OUT OF SYNC
and (optionally) reverts it
What you gain: every change to production is a reviewed commit; the current state of every environment is readable in a repo; rollback is git revert; drift is detected and can be self-healed; and no CI system needs standing cluster credentials. What it costs: a second repository to keep tidy, and a new failure mode when the agent's view and yours disagree.
argocd app list argocd app get notes-api # Synced? Healthy? which revision? argocd app diff notes-api # repo vs cluster argocd app sync notes-api argocd app rollback notes-api 12
TOPIC 09A release checklist worth keeping
- Before: CI green on the exact commit · the artifact is immutable and already scanned · migrations are additive and applied first · the rollback command is known and tested · someone other than you knows the deploy is happening.
- During: deploy to a small slice first · watch error rate and p95 latency, not just "did the pods start" · give it long enough for slow failures (cache expiry, cron jobs) to appear.
- After: run one real user journey yourself · check the dashboards you would check during an incident · leave the previous version available for a day · write down anything surprising while you still remember it.
- Never: deploy something you cannot undo, at a time when nobody is around to notice, from an unpushed local branch.
- Delivery = always ready to ship; deployment = shipping automatically. Deploy ≠ release.
- Build one artifact, promote the same digest; change configuration, never the bytes.
- Rolling for cheap zero-downtime, blue-green for instant rollback, canary for the best risk/cost balance.
- Feature flags separate exposure from deployment and turn a rollback into a switch.
- Know your undo command before you deploy. Roll back first, diagnose second.
- Schema changes: expand, migrate, contract — never a breaking change in one step.
- GitOps makes Git the source of truth and rollback a revert.
LABA canary you can revert
Promote an artifact, ship a canary, then undo it under time pressure
- Extend your CI workflow: after tests pass, build the image and push it tagged
${{ github.sha }}toghcr.io. Nolatestanywhere. - Add a
deploy-stagingjob that runs only onmain, sets that exact tag on your Deployment, waits forkubectl rollout status, then curls/healthas a smoke test. Make the job fail if the smoke test fails — a deploy that cannot fail is not a gate. - Add a
deploy-productionjob withneeds: deploy-stagingand a GitHubenvironment:that requires manual approval. Merge something and approve it. You now have continuous delivery. - Hand-rolled canary: create a second Deployment
api-canarywith 1 replica on the new tag, sharing the Service's label so it receives roughly 1/(N+1) of traffic. Verify with a loop ofcurlthat you see both versions respond. - Make the canary bad on purpose: point it at an image whose
/healthreturns 500. Watch the error rate in your loop. Now delete the canary Deployment and watch it clear. Time the whole detect-and-revert cycle and write the number down. - Practise the real rollback: promote the bad version to the main Deployment, then recover with
kubectl rollout undo. Do it twice, until you can do it without looking anything up. - Migration drill: take any table and rename a column properly — deploy 1 adds the new column and writes both, deploy 2 reads the new one, deploy 3 drops the old. Confirm after each deploy that the previous version of the app would still work.
- Stretch: install Argo CD in your local cluster, point it at a config repo, and deploy by committing an image tag. Then
git revertand watch the cluster roll itself back with nobody touchingkubectl.
Steps 5 and 6 are the deliverable. Anyone can deploy; the interview answer that lands is "we detected a bad canary in 90 seconds and reverted in 30, and here is how I know".
CHECKCheck yourself
During a rolling update, a migration renames a column in one step. What happens?
A rolling update deliberately runs both versions at once, so the schema must satisfy both. That is why breaking changes are split into expand → migrate → contract. The nastiest part is the rollback trap: reverting the code returns you to a version that also cannot read the new schema, so you are stuck fixing forward during an incident.
Why is a canary release usually a better bet than deploying to 100% and watching?
A canary trades a little time for a much smaller blast radius: if the new version is bad, 5% of requests suffer briefly instead of all of them. It is strictly slower than a big-bang deploy, and it depends on having per-version metrics to compare — which is why chapter 10 comes next.
What is the main operational advantage of GitOps over a pipeline that runs kubectl apply?
The value is auditability and reconciliation: every production change is a commit someone approved, a hand-edit shows up as out-of-sync, and undo is git revert. As a bonus, no external CI system needs standing cluster credentials — the agent pulls rather than being pushed to.