TOPIC 01What DevOps is, in one paragraph
DevOps is a way of working where the same team builds a piece of software and is responsible for running it in production — supported by enough automation that releasing becomes boring. It is not a tool you install, not a person you hire, and not a job title you must have. It is a set of habits, plus the machinery that makes those habits cheap.
Notice the two halves, because most people only learn one of them. The cultural half is about ownership: you built it, you get the alert when it breaks. The technical half is about automation: if a release takes four humans and a checklist, nobody will do it often, so releases will be big, and big releases break things.
Think of a restaurant where the chefs never see the dining room and the waiters are not allowed in the kitchen. Chefs cook whatever they like; waiters get shouted at by customers about food they can't change. Both sides are working hard. The restaurant still fails — because nobody owns the whole meal. DevOps is knocking down that hatch and letting the kitchen hear the dining room.
Here is the definition you can give in an interview, and mean it: DevOps shortens the distance between an idea and a working feature in front of users, and shortens the distance between a failure and its fix. Those two distances are the whole game.
TOPIC 02The wall of confusion
For decades, most companies were organised like this: a development team wrote code and "threw it over the wall" to an operations team, whose job was to install it on servers and keep the servers alive. That wall had a logic to it — different skills, different tools — but it created a permanent conflict of interest.
| Developers were rewarded for | Operations were rewarded for |
|---|---|
| Shipping features fast | Nothing breaking |
| Change | Stability |
| Trying new libraries and versions | A boring, known, unchanging stack |
| Being finished at "code complete" | Being on call for someone else's code |
The predictable results — you will still meet all of these in real companies:
- "Works on my machine." The developer's laptop has different versions of everything than the server, so code that ran fine in testing dies in production.
- Release nights. Because deploying is risky and manual, changes get batched up for a monthly Saturday-night release. A hundred changes go out together, something breaks, and nobody knows which of the hundred did it.
- Blame instead of fixes. Ops says "bad code", dev says "bad servers", and the actual cause is never written down, so it happens again.
- Manual snowflakes. Someone SSHes into a server, edits a config to fix an outage at 2 a.m., and forgets. Six months later that server behaves differently from its twin and nobody knows why.
None of these problems are caused by lazy people. They are caused by a structure where the person who can fix the problem is not the person who feels it. Every practice in this diary is a way of putting feeling and fixing back into the same pair of hands.
TOPIC 03Waterfall → agile → DevOps
A quick history, because the vocabulary makes no sense without it.
- Waterfall (roughly 1970s–2000s)
- Plan everything, then design everything, then build everything, then test everything, then release once. Works for building a bridge, where you cannot pour half a foundation and ask users what they think. Terrible for software, where requirements change while you build.
- Agile (2001 onward)
- Build in small slices, show them to users, adjust. Two-week sprints, stand-ups, backlogs. Agile fixed how software was planned and built… and stopped exactly at the point where the code had to reach real servers.
- DevOps (roughly 2009 onward)
- Extends the same small-batch thinking all the way to production and beyond: automated builds, automated tests, automated deploys, monitoring, and the same team owning the result. Agile without DevOps means a team that can plan in two-week slices but still releases twice a year.
Waterfall is cooking a nine-course dinner for forty guests with no tasting until service. Agile is tasting each dish as you cook. DevOps is having a serving hatch wide enough that each dish reaches the table while it's hot — and a way to hear, immediately, that table six thinks the dal is over-salted.
TOPIC 04The pipeline, end to end
Here is the map for the rest of this diary. Every stage is a chapter, and the whole thing is what people mean when they say "the pipeline".
# you, on your laptop 1. plan ticket / issue: "checkout is slow on mobile" ch 01 2. code edit files, run it locally ch 02 3. commit git commit + push a branch ch 04 # the robots take over 4. build install deps, compile, produce one artifact ch 05 5. test unit + integration tests, linters ch 05 6. package build a container image, tag it with the sha ch 06 7. scan dependencies, secrets, image vulnerabilities ch 11 8. provision servers/cluster/db defined as code ch 08 9. deploy staging → smoke test → canary → production ch 07, 09 # and it never really ends 10. observe metrics, logs, traces, dashboards ch 10 11. alert page a human only when users are hurting ch 10 12. learn postmortem, then feed it back into "plan" ch 10, 12
That loop — plan, build, run, learn, plan again — is why the DevOps logo is an infinity symbol. Two things about it are worth noticing now:
- Every stage is automated except the first and the last. Thinking and learning are human work. Copying files to servers is not.
- Feedback gets more expensive as you move right. A typo caught by a linter costs five seconds. The same typo caught by a customer costs an incident, a rollback and a trust problem. This is why "shift left" is such a common phrase: move every check as early as it will go.
TOPIC 05CALMS: the five ingredients
CALMS is a checklist for whether a team is actually doing DevOps or just renamed their ops team. Learn it — it comes up in interviews, and it is genuinely useful for spotting what your own team is missing.
- C — Culture
- Shared ownership and blameless problem-solving. Test: when something breaks, does the conversation start with "what happened?" or with "who did it?"
- A — Automation
- Builds, tests, deploys and infrastructure are code, not clicks. Test: could you deploy today's change with one command, and would you be comfortable doing it on a Friday afternoon?
- L — Lean
- Small batches, less work-in-progress, no queues of half-finished work. Test: how long does an average change wait between "developer finished" and "users have it"?
- M — Measurement
- You have numbers for both speed and stability, and you look at them. Test: can anyone in the team say how often you deploy, and what percentage of deploys cause a problem?
- S — Sharing
- Knowledge is written down, not stored in one person's head. Test: if the one person who knows how the payments deploy works goes on holiday, does anything ship?
A hostel with one water tank: culture is everyone agreeing not to waste it, automation is a float valve instead of a person watching the tank, lean is filling little and often instead of a monthly flood, measurement is the level gauge on the side, and sharing is the note taped to the pump explaining how to restart it. Remove any one and someone eventually showers in cold water.
TOPIC 06Measuring it: four numbers
Years of research across thousands of teams (the DORA / Accelerate work) landed on four metrics that predict both software delivery performance and business outcomes. Two measure speed, two measure stability — and the striking finding is that they rise together. Teams that deploy more often are usually more stable, not less, because their changes are small and their recovery is practised.
| Metric | Question it answers | Weak | Strong |
|---|---|---|---|
| Deployment frequency | How often do we ship to production? | Once a month or less | Multiple times a day |
| Lead time for changes | Commit → running in production? | Weeks or months | Under a day |
| Change failure rate | What share of deploys cause a problem? | Around 40–60% | Under 15% |
| Time to restore service | How fast do we recover from failure? | Days | Under an hour |
Deploying 100 changes at once means 100 suspects when it breaks, and a rollback that removes 99 innocent changes. Deploying one change at a time means one suspect and a trivial rollback. Frequency is not recklessness — it is how you shrink the size of the thing that can go wrong.
TOPIC 07SRE, platform engineering, DevOps engineer
The job titles overlap and companies use them inconsistently. Here is the honest version:
- DevOps engineer
- Strictly, a contradiction — DevOps is a way of working, not a role. In practice, this title means: the person who builds and maintains the pipelines, the infrastructure code and the tooling other developers use. This is the job most of this diary trains you for.
- SRE (Site Reliability Engineer)
- Google's specific implementation of these ideas, with a strong emphasis on measurement: define reliability targets (SLOs), track an error budget, and use it to decide between shipping features and fixing reliability. An SRE is an engineer who treats operations as a software problem.
- Platform engineer
- The current evolution: instead of doing deploys for other teams, build an internal platform — golden pipelines, templates, a CLI — so product teams can self-serve safely. Success is measured by how little product teams need to ask you.
- Cloud engineer
- Weighted toward one provider's building blocks (AWS, Azure, GCP): networking, IAM, managed databases, cost. Heavy overlap with chapters 08 and 09.
Do not spend energy picking a title now. The underlying skill set — Linux, networking, Git, containers, pipelines, IaC, observability — is the same for all four, and it is exactly the list in the sidebar of this site.
TOPIC 08Five myths worth killing early
- "DevOps means no ops team." No — it means ops work is shared, automated and often turned into a platform. Specialists still exist; walls do not.
- "DevOps means Kubernetes." Plenty of excellent DevOps teams run on a couple of virtual machines and a shell script. Kubernetes is a tool for a specific problem (chapter 07 explains which).
- "You need certifications first." A working project you can explain beats a certificate you crammed. Chapter 12 gives you four such projects.
- "Automation means nobody is on call." Automation reduces toil, not responsibility. Someone still answers when the pager goes off — the point is that it goes off less, and the fix is faster.
- "You must learn every tool." Tools change every three years; the concepts do not. Learn why a category exists (config management, orchestration, secret storage) and picking up any specific tool takes a weekend.
- DevOps = shared ownership + automation, aimed at shortening two distances: idea → users, and failure → fix.
- The wall between dev and ops caused "works on my machine", risky release nights and blame; small batches and automation dissolve all three.
- CALMS is the ingredient list: culture, automation, lean, measurement, sharing.
- Four numbers tell you if it's working: deployment frequency, lead time, change failure rate, time to restore.
- Speed and stability are not a trade-off. Small, frequent, well-tested changes give you both.
LABMap your own pipeline
Draw the journey of one real change
This is the exercise consultants charge a lot of money for, and it works on a team of one.
- Pick any recent change to any software you have worked on — a college project counts, a website tweak counts.
- Write every step from "idea existed" to "a user could use it". Include waiting: "waited 2 days for review" is a step.
- Next to each step, write how long it took and how long it waited.
- Add up both columns. In almost every team, waiting massively exceeds working. That ratio is the thing DevOps attacks.
- Circle the single longest wait, then ask one question: could a machine do this, or could it be made smaller?
Why it matters: you now have a concrete story to tell in an interview — "our lead time was five days, and four of them were waiting for a manual test environment" — which is far more convincing than reciting the definition of CI.
CHECKCheck yourself
A team deploys once a month. Their manager wants fewer production incidents and proposes deploying once a quarter instead. What is wrong with that reasoning?
Risk lives in batch size, not in the calendar. One deploy with 300 changes has 300 suspects and an all-or-nothing rollback. The DORA research consistently finds high-frequency deployers have lower change failure rates, because each change is small and recovery is a well-practised routine.
Which of these is the clearest sign a team has adopted the culture half of DevOps, not just the tools?
Ownership of consequences is the load-bearing change. When the author of a change feels its failure, error handling, logging and rollback stop being someone else's problem — and every tool on the list gets used properly instead of ceremonially.
"Lead time for changes" measures the time from…
Commit → production. Incident start → resolved is time to restore service, the recovery metric. Keeping the two apart matters, because they are improved by completely different work: pipelines versus rollbacks and observability.