devopsdiary
Play the arcade

diary / chapters / 12

chapter 12 · ~35 min · the last one

Portfolio, interviews and what comes next

Eleven chapters of knowledge is worth very little until someone can see it working. This chapter turns what you have learned into four projects, a set of stories you can tell under pressure, and a habit that keeps you learning after the diary ends.

TOPIC 01What actually gets you hired

Hiring managers for DevOps roles are trying to answer one question: can this person be trusted with production? Everything on your CV is evidence for or against that. Ranked by how much it moves the needle:

  1. A project you built and can explain in depth, including what broke and how you found it. This is far and away the strongest signal.
  2. A story about a failure. "The cache invalidation was wrong, here is how I noticed and what I changed" tells them more than any success story.
  3. Fluency in the fundamentals. Linux, networking, Git. Interviewers probe these because tools change and fundamentals do not.
  4. Clear writing. Half of this job is a good pull-request description, a runbook, a postmortem. Your README is a work sample.
  5. Certifications. Useful for getting past a filter, rarely decisive in the conversation.
real life

Hiring a driver. A licence proves you passed a test once. What you actually want to know is how they handle a monsoon at night on a road with no lane markings. The projects below are your monsoon: they exist so you can say "here is what happened when it went wrong, and here is what I did".

PROJECT 01The full pipeline

chapters 04–07, 09 · a weekend

Commit to running container, with nothing done by hand

  • A small app with a /health endpoint and real tests (chapter 05).
  • CI on every pull request: lint, unit tests, integration tests against a real database service, all in parallel with caching. Under five minutes.
  • A multi-stage Dockerfile producing a small non-root image, tagged with the commit SHA and pushed to a registry (chapter 06).
  • Deploy to a local Kubernetes cluster with probes, resource limits and a ConfigMap (chapter 07).
  • Branch protection so nothing merges red (chapter 04).
  • The part that impresses: a short section in the README titled "how I roll this back", with the command and a screenshot of having done it.

The numbers to record: pipeline duration before and after caching; image size before and after multi-stage; time to roll back a bad deploy.

PROJECT 02Infrastructure as code

chapter 08 · one or two evenings

An environment that can be destroyed and rebuilt identically

  • Terraform for a small but real stack: network, one or two servers or a container service, a security group, DNS.
  • Modules called by a staging/ and a prod/ environment with different sizes — the same code, different numbers.
  • Remote state with locking, and the reasoning written down in the README.
  • Ansible (or cloud-init) to configure what Terraform created, demonstrably idempotent — show a second run reporting changed=0.
  • A CI job that runs fmt -check, validate, plan on pull requests and posts the plan.
  • The part that impresses: a documented drift exercise — you changed something by hand, the nightly plan detected it, and you reconciled it in code.

If you use a cloud free tier, put a destroy reminder in your calendar and say in the README that you did. Cost-awareness reads as maturity.

PROJECT 03The observability stack

chapter 10 · one evening

Prove you would notice before your users do

  • Your app exposing /metrics: a request counter with labels and a latency histogram.
  • Prometheus scraping it, Grafana showing exactly four panels — the golden signals.
  • Structured JSON logs with a trace ID, queryable in Loki.
  • One alert rule with a for: duration, a severity label and a runbook link.
  • A written SLO: one sentence defining the SLI, a target, and the monthly error budget in minutes.
  • The part that impresses: a screenshot of the alert firing because you broke the app on purpose, next to the graph showing the moment it recovered.

PROJECT 04The incident writeup

chapters 10–11 · two hours, and the highest value per hour of the four

A postmortem for something you broke deliberately

Almost nobody applying for a junior role brings one of these, and it demonstrates the exact judgement the job requires. Break your own project in a realistic way — deploy a version that exhausts the database connection pool, or set a memory limit that triggers OOM kills under load — and then write it up honestly.

  • Impact: what failed, for how long, measured — not "the site was slow" but "38% of requests to /checkout returned 500 for 11 minutes".
  • Timeline: timestamps from alert to mitigation to resolution, including the wrong turns.
  • Detection: how you found out, and how long it took. Then: what would have made that faster?
  • Contributing factors: plural, and systemic. No names.
  • Action items: three, each with an owner and a date, at least one of which you then actually implement — and link the commit.

Bring this to an interview and you will spend twenty minutes discussing it rather than answering trivia. That is a very good trade.

TOPIC 06Writing the README

Most portfolio projects fail not because the work is weak but because nobody can tell what was done. A reviewer gives you ninety seconds. Structure for that.

README.md — the shape that gets read
# Notes API — CI/CD and Kubernetes deployment

One paragraph: what this is, and what problem the setup solves.

## Architecture
A diagram or a plain ASCII sketch. Show the flow: commit → CI → registry
→ cluster → ingress → users.

## What I built
- CI on every PR: lint, unit, integration (real Postgres), 4m10s → 1m05s
  after dependency caching
- Multi-stage image: 1.1 GB → 94 MB, runs as non-root
- Rolling deploys with readiness probes; zero-downtime verified with a
  load generator during a rollout
- Rollback: `kubectl rollout undo` — measured at 34 seconds

## Decisions and trade-offs
Why Kustomize instead of Helm here. Why a managed database rather than
a StatefulSet. What I would do differently at 100x the traffic.

## How to run it locally
Three commands. Tested from a clean clone.

## What broke while building it
Two or three real problems, and how I diagnosed each one.

Those last two sections are the ones experienced engineers read first. "Decisions and trade-offs" shows you can reason about cost and complexity rather than following a tutorial; "what broke" shows you can debug. Numbers throughout — every "faster" and "smaller" should carry a figure.

TOPIC 07The free home lab

You do not need a cloud budget to practise any of this.

You wantFree option
Linux serversMultipass, VirtualBox, or a few containers with SSH
A Kubernetes clusterkind (multi-node from one YAML), minikube, or k3s on an old laptop
CI/CDGitHub Actions or GitLab CI free tier — generous for personal projects
Container registryGitHub Container Registry, free for public images
MonitoringPrometheus + Grafana + Loki in Compose
A public URL for a demoA small always-free cloud instance, or a tunnel like cloudflared
Cloud practiceAWS/GCP/Azure free tiers — set a budget alert on day one, before anything else
the one that catches everyone

Set a billing alert and a hard budget the moment you open a cloud account, then destroy resources when you finish for the day (terraform destroy exists for this). The classic story is a load balancer left running for a month. Being the person who mentions cost unprompted in an interview is a genuine advantage.

TOPIC 08The interview

A typical DevOps loop has four shapes. Prepare for each differently:

Fundamentals screen
Rapid Linux, networking, Git and container questions. This is what Interview Blitz and Terminal Trainer in the arcade are for — get to a point where the answers are automatic.
Troubleshooting / live debugging
"A service returns 502. Walk me through it." They are assessing method, not luck. Use the chapter-03 ladder out loud: name resolution, reachability, port, TLS, HTTP, logs. Say what you would check and what each result would rule out. Thinking aloud is the answer.
System design
"Design the deployment for a service with 10,000 requests per second." Start with questions — traffic shape, latency requirements, statefulness, budget. Then the pipeline, the release strategy, the observability, and the failure modes. Mentioning rollback and monitoring unprompted marks you out immediately.
Behavioural
"Tell me about a time something broke." Have three stories ready, each with situation, what you did, the outcome, and what you changed afterwards. Project 04 gives you one for free.

Two habits that improve every answer: say "I don't know, here is how I would find out" instead of guessing, and give a number wherever you can. And ask them questions — how often they deploy, how on-call works, what their last incident was. Their answers tell you whether the job is what the advert claimed.

TOPIC 09Questions, and what a good answer contains

  • "What is DevOps?" — Shared ownership plus automation, aimed at shortening idea → users and failure → fix. Not a tool, not a title. (Chapter 01.)
  • "Containers vs VMs?" — Shared kernel versus own kernel; seconds versus minutes; megabytes versus gigabytes; and name the security trade-off. (Chapter 06.)
  • "A pod is CrashLoopBackOff. What do you do?" — describe for Events, logs --previous for the dead instance, check config and env, check limits for OOM. Say the order. (Chapter 07.)
  • "How do you deploy without downtime?" — Rolling with readiness probes and maxUnavailable: 0; canary for risk; blue-green for instant rollback; backwards-compatible migrations throughout. (Chapters 07, 09.)
  • "How do you handle secrets?" — Never in Git or images; a secret store; injected at runtime; per environment; rotated; short-lived via OIDC where possible. (Chapter 11.)
  • "What would you alert on?" — User-visible symptoms with a duration and a runbook, not CPU. Mention SLOs and error budgets. (Chapter 10.)
  • "Terraform state — why does it matter?" — It maps code to real resources; remote, locked, encrypted, versioned, one per environment; without it Terraform duplicates everything. (Chapter 08.)
  • "Merge or rebase?" — Rebase your own branch for a clean review; merge shared branches; never rebase what others have pulled. (Chapter 04.)
  • "When would you not use Kubernetes?" — A static site, a single small service, a team without the capacity to operate it. Naming the limits of your favourite tool builds more trust than praising it. (Chapter 07.)
  • "How would you improve a legacy manual deploy?" — Write the steps down as a runbook, script the runbook, then automate the script, then add a gate. Nobody can automate what is not written down. (Chapters 02, 05.)

TOPIC 10Certifications, honestly

CertificationWorth it when
CKA (Certified Kubernetes Administrator)You want a hands-on credential — it is a practical, terminal-based exam, so preparing for it genuinely builds skill
AWS / Azure / GCP associate-levelYou are targeting a cloud-heavy role, or applying somewhere that filters on it
Terraform AssociateCheap, quick, and forces you to read the docs properly
Linux Foundation / RHCSAYou want to prove Linux depth without a work history to point at

The honest ranking: a project you can discuss beats a certificate; a certificate beats nothing. If you are choosing, do the practical exams (CKA, RHCSA) — the preparation itself is the value, because you cannot pass them by memorising.

TOPIC 11What to learn next

You now have the whole shape of the field. Directions worth taking, depending on what you enjoyed most:

  • Go deeper on one cloud. Networking (VPCs, peering, private endpoints), IAM in detail, managed databases, and cost. Depth in one provider transfers to the others.
  • Platform engineering. Golden paths, internal developer platforms, Backstage, self-service templates. Where the industry is heading.
  • SRE practice. Capacity planning, load testing, chaos experiments, error-budget policy. Read widely on incident response.
  • Learn a real programming language properly. Go or Python. The gap between "writes bash" and "writes tools" is the biggest single step in seniority.
  • Data and cost. Databases under load, backups you have actually restored, and FinOps. Being the person who cut the bill 30% is a career-making project.
  • Security depth. Threat modelling, supply chain, zero trust. Chapter 11 was the first ten percent.
the habit that matters more than any of them

Run something real, and keep it running. A single small service you own — with a pipeline, a dashboard, an alert that has actually woken you up, and a bill you pay attention to — teaches more in three months than any amount of reading. Everything in this diary was written by people who learned it that way.

remember this much
  • Projects you can explain in depth beat certificates and course lists.
  • Build four: the full pipeline, infrastructure as code, the observability stack, the incident writeup.
  • Break things on purpose and document it — a postmortem is the rarest and strongest artefact a junior candidate can bring.
  • Put numbers in your README: build times, image sizes, rollback duration.
  • In interviews, describe your method out loud and be comfortable saying "I don't know, here's how I'd find out".
  • Keep one real service running. That is the whole curriculum, repeated with your own hands.

CHECKCheck yourself

You have one weekend before an interview. What is the highest-value use of it?

Interviewers are hiring judgement under failure, and a postmortem is direct evidence of it — impact measured, timeline, systemic causes, action items. Breadth of tool names is easy to claim and easy to puncture with one follow-up question; a documented failure you fixed is hard to fake and immediately interesting to talk about.

Asked "how would you debug a 502 from our API?", what is the best-scoring answer?

They are grading method, because in production the cause is always new. A structured ladder shows you can eliminate whole classes of problem cheaply and will not thrash. Leading with one guess sounds decisive and falls apart when it is wrong; restarting first destroys the evidence, which is exactly the trap in the Incident Room game.

Which README section does an experienced reviewer usually value most?

Anyone can follow a tutorial and end up with the same stack. Explaining why you chose Kustomize over Helm, or a managed database over a StatefulSet, and what you would do differently at 100× the traffic, proves you understand cost and complexity — which is the actual job. Setup instructions matter too, but as hygiene rather than signal.

saved in this browser only — no account needed