devopsdiary
Next chapter

diary / chapters / 07

chapter 07 · ~70 min · local cluster lab

Kubernetes, explained slowly

Kubernetes has a reputation for being impossible. It is not — it is just twelve ideas introduced all at once by most tutorials. Here they arrive one at a time, each with the problem it solves, so by the end the YAML reads like sentences.

TOPIC 01What problem it solves

Chapter 06 gave you containers. Now put fifty of them across six machines and answer these questions:

  • A container crashes at 3 a.m. Who restarts it?
  • A whole machine dies. Who moves its containers elsewhere?
  • Traffic triples on Diwali. Who starts more copies, and who tells the load balancer they exist?
  • You are deploying v2. Who replaces containers gradually so the site never goes down, and who puts v1 back if v2 is broken?
  • Which machine has room for this new container?

You could answer all of that with shell scripts. Thousands of teams did, and then rewrote them, badly, one at a time. Kubernetes is the shared answer: you declare the desired state — "three copies of this image, reachable on this hostname, with this config" — and a control loop continuously makes reality match. That is the single most important sentence in this chapter.

real life

A thermostat, not a switch. A switch is imperative: "turn the heater on." A thermostat is declarative: "keep this room at 22°." It measures, acts, and keeps acting — if someone opens a window, it responds without being asked again. Kubernetes is a thermostat for infrastructure, and every object you write is a target setting.

TOPIC 02Do you even need it?

An honest section, because "we'll use Kubernetes" is the most expensive reflex in the industry.

SituationBetter answer
A static site or blogStatic hosting or a CDN
One app, modest traffic, small teamA VM with Compose, or a managed platform (App Runner, Cloud Run, Fly, Render)
Several services, several teams, need self-service deploys, autoscaling, zero-downtime releasesKubernetes earns its cost here
Regulated workloads across multiple clouds or on-premKubernetes, for the portable abstraction

Learn it anyway — it is on nearly every DevOps job description, and the concepts (declarative state, reconciliation, health probes, rolling updates) transfer to everything. Just be the person who can also say when it is overkill; that answer marks you out in an interview far more than reciting object types.

TOPIC 03The cluster, in parts

who does what
  CONTROL PLANE (the brain)               WORKER NODES (the muscle)
  ┌──────────────────────────┐            ┌──────────────────────────┐
  │ api-server               │◀── kubectl │ kubelet                  │
  │  the only way in; every  │            │  runs containers here,   │
  │  read and write goes here│───────────▶│  reports their health    │
  │                          │            │                          │
  │ etcd                     │            │ kube-proxy / CNI         │
  │  the database: desired   │            │  wires up pod networking │
  │  state of everything     │            │                          │
  │                          │            │ container runtime        │
  │ scheduler                │            │  containerd / CRI-O      │
  │  picks a node for each   │            │                          │
  │  new pod                 │            │  [pod] [pod] [pod]       │
  │                          │            │                          │
  │ controller-manager       │            └──────────────────────────┘
  │  the control loops that  │
  │  make reality == desired │            managed clusters (EKS/GKE/AKS)
  └──────────────────────────┘            run the control plane for you

You will spend 95% of your time talking to the api-server through kubectl, and the mental model that matters is the loop: you change desired state → a controller notices the difference → it acts → it checks again. Every strange Kubernetes behaviour you will ever debug is that loop doing exactly what you told it.

TOPIC 04Pods

A pod is the smallest thing Kubernetes runs: one or more containers that share a network address and can share volumes. Usually one container per pod. Containers in the same pod reach each other on localhost.

The crucial property: pods are disposable. They are never repaired — they are replaced, with a new name and a new IP. This is why you almost never create a pod directly, and why you never rely on a pod's address.

pod.yaml — useful for learning, rare in production
apiVersion: v1
kind: Pod
metadata:
  name: notes-api
  labels:
    app: notes-api
spec:
  containers:
    - name: api
      image: ghcr.io/you/notes-api:a91f3c
      ports:
        - containerPort: 3000
      env:
        - name: PORT
          value: "3000"        # quoted! env values are strings (see the arcade)
labels are the glue

labels are arbitrary key/value tags, and almost every connection in Kubernetes is made by matching them: a Service finds its pods by label, a Deployment tracks its pods by label, network policies select by label. Get a label wrong and two objects that look correct will simply never find each other — that is patient 2 in the YAML Doctor game.

TOPIC 05Deployments and ReplicaSets

A Deployment is what you actually write. It says: "keep N pods of this template running, and when I change the template, roll the change out gradually." It creates a ReplicaSet to hold the pods, and a new ReplicaSet for each version — which is exactly how rollback works.

deployment.yaml — the workhorse object
apiVersion: apps/v1              # NOT v1 — Deployments live in apps/v1
kind: Deployment
metadata:
  name: notes-api
spec:
  replicas: 3
  selector:
    matchLabels:
      app: notes-api             # must match template labels exactly
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1                # at most 1 extra pod during the roll
      maxUnavailable: 0          # never drop below 3 ready → zero downtime
  template:
    metadata:
      labels:
        app: notes-api           # ← the labels the selector hunts for
    spec:
      containers:
        - name: api
          image: ghcr.io/you/notes-api:a91f3c
          ports:
            - containerPort: 3000
          envFrom:
            - configMapRef: { name: notes-api-config }
            - secretRef:    { name: notes-api-secrets }
          resources:
            requests: { cpu: 100m, memory: 128Mi }
            limits:   { cpu: 500m, memory: 256Mi }
          readinessProbe:
            httpGet: { path: /health, port: 3000 }
            initialDelaySeconds: 5
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /health, port: 3000 }
            initialDelaySeconds: 20
            periodSeconds: 10
real life

A shift manager with a staffing rule: "three people on the floor at all times." Someone calls in sick — she calls in a replacement without asking you. New uniforms arrive — she swaps them one person at a time so the floor is never empty, and if the new uniform turns out to be unwearable, she has the old ones on the shelf ready. That is a Deployment: replicas, rolling update, rollback.

TOPIC 06Services

Pods come and go with new IP addresses, so you can never hard-code one. A Service is a stable name and address in front of a changing set of pods, with load balancing built in.

service.yaml
apiVersion: v1
kind: Service
metadata:
  name: notes-api                # other pods reach this as http://notes-api
spec:
  type: ClusterIP                # internal only (the default, and the right default)
  selector:
    app: notes-api               # any pod with this label is a backend
  ports:
    - name: http
      port: 80                   # the port ON THE SERVICE
      targetPort: 3000           # the port ON THE CONTAINER
TypeReachable fromUse for
ClusterIPInside the cluster onlyAlmost everything
NodePortA high port on every nodeLocal clusters, quick tests
LoadBalancerThe internet, via a cloud load balancerOne or two public entry points
Headless (clusterIP: None)DNS returns pod IPs directlyStatefulSets, client-side balancing

Inside the cluster, DNS gives you notes-api (same namespace) or notes-api.payments.svc.cluster.local (fully qualified). The port/targetPort confusion — service door versus container door — is patient 3 in the YAML Doctor, and it produces the maddening symptom of healthy pods behind a hanging service.

TOPIC 07Ingress

A LoadBalancer Service per app means a cloud load balancer per app, which is expensive and unmanageable. An Ingress is one entry point that routes by hostname and path — the chapter-03 reverse proxy, expressed as a Kubernetes object. It needs an ingress controller (nginx-ingress, Traefik, or your cloud's) running in the cluster to do the actual work.

ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: public
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt   # automatic TLS certificates
spec:
  ingressClassName: nginx
  tls:
    - hosts: [api.example.com]
      secretName: api-tls
  rules:
    - host: api.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: notes-api
                port: { number: 80 }
          - path: /admin
            pathType: Prefix
            backend:
              service:
                name: admin-ui
                port: { number: 80 }

TOPIC 08ConfigMaps and Secrets

One image, many environments — so configuration must come from outside the image. ConfigMaps hold non-sensitive settings; Secrets hold sensitive ones.

config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: notes-api-config
data:
  LOG_LEVEL: "info"
  FEATURE_NEW_CHECKOUT: "false"
---
apiVersion: v1
kind: Secret
metadata:
  name: notes-api-secrets
type: Opaque
stringData:                       # stringData: plain text in, base64 stored
  DATABASE_URL: postgres://app:hunter2@db:5432/app
Secrets are not encrypted by default

A Kubernetes Secret is base64-encoded, which is encoding, not encryption — anyone who can read Secrets in that namespace can read the value. Real protection needs: RBAC so almost nobody can read them, encryption at rest for etcd, and ideally an external store (External Secrets Operator, Vault, or your cloud's secret manager) so the value never sits in Git. And never commit a Secret manifest with a real value in it (chapter 11).

One practical detail that surprises everyone: changing a ConfigMap does not restart your pods. Either mount it as a file and reload on change, or roll the deployment: kubectl rollout restart deployment/notes-api.

TOPIC 09Probes

Kubernetes cannot know whether your process is actually working, so you tell it how to check. Three probes, three different jobs — and confusing them is the cause of a classic restart loop.

readinessProbe — "should I get traffic?"
Fails → the pod is removed from the Service's backends, but not restarted. This is what makes zero-downtime rollouts real: a new pod gets traffic only once it says it is ready.
livenessProbe — "am I still alive?"
Fails → the container is restarted. Use it for deadlocks and wedged event loops. Keep it forgiving; an aggressive liveness probe on a slow-starting app produces an endless restart cycle.
startupProbe — "have I finished booting?"
Holds the liveness probe off until the app has started. The correct fix for JVM-style slow starts, instead of inflating initialDelaySeconds.
real life

A shop's front door and its manager. Readiness is the OPEN sign: while the till is being counted, the sign stays off and nobody is sent in — the shop is fine, just not serving. Liveness is the manager checking whether the shopkeeper has fallen asleep; if so, wake them up (restart). Pointing traffic at a shop with the sign still off is a 502; sacking a shopkeeper who was merely counting stock is a restart loop.

TOPIC 10Requests and limits

  • requests = the capacity reserved for scheduling. The scheduler places a pod only on a node with this much free. Set too high, pods stay Pending; set too low, nodes get oversubscribed.
  • limits = the hard ceiling. Over the CPU limit and the container is throttled; over the memory limit and it is killed with OOMKilled (exit code 137).

Set both on every container. Without requests, the scheduler is guessing; without limits, one leaking pod can take a whole node down with it. Get real numbers from kubectl top pod under load rather than guessing, then leave headroom.

TOPIC 11Scaling and rollouts

terminal · day-to-day operations
kubectl apply -f k8s/                       # declarative: apply the whole folder
kubectl get pods -w                         # watch pods change state live
kubectl scale deployment notes-api --replicas=5

# deploy a new version and watch the rollout
kubectl set image deployment/notes-api api=ghcr.io/you/notes-api:b72d10
kubectl rollout status deployment/notes-api
kubectl rollout history deployment/notes-api
kubectl rollout undo deployment/notes-api     # ← the 3 a.m. command
kubectl rollout restart deployment/notes-api  # pick up new config
hpa.yaml — scale on load
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: notes-api
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: notes-api
  minReplicas: 3
  maxReplicas: 20
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70    # % of the CPU *request*

The HPA adds pods (horizontal). The Cluster Autoscaler adds nodes when pods cannot be scheduled. Both need sensible requests to work at all — utilisation is measured against the request, so a wrong request makes autoscaling nonsense.

TOPIC 12Storage and StatefulSets

Pods are ephemeral, so persistent data needs a PersistentVolumeClaim — a request for storage that the cluster fulfils from a StorageClass (usually a cloud disk). Stateful workloads that need stable names and their own disk per replica (databases, Kafka, Elasticsearch) use a StatefulSet instead of a Deployment: it gives pods predictable identities (db-0, db-1) and keeps each one's volume attached across restarts.

honest advice

Running a production database on Kubernetes is an advanced topic: backups, failover, upgrades and disk performance all become your problem. For most teams, a managed database (RDS, Cloud SQL) with your stateless apps on Kubernetes is the right split. Know StatefulSets exist and what they solve; do not make your first cluster a database cluster.

TOPIC 13Helm and Kustomize

You now have five YAML files, and you need them slightly different in staging and production. Two tools solve this:

Helm
Templated YAML with a values file — a package manager for Kubernetes. Best for installing third-party software (helm install ingress-nginx …) and for teams shipping the same chart to many environments.
Kustomize
Built into kubectl. A plain base plus overlay patches per environment, no templating language. Often the calmer choice for your own applications.
terminal · helm basics
helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx
helm install ingress ingress-nginx/ingress-nginx -n ingress --create-namespace
helm upgrade --install notes-api ./chart -f values.prod.yaml
helm template ./chart -f values.prod.yaml | less   # see the YAML before applying
helm rollback notes-api 3

TOPIC 14Debugging: the routine

terminal · in this order, every time
kubectl get pods                       # the STATUS column names your problem
kubectl describe pod notes-api-7c9    # scroll to Events at the bottom — the answer
kubectl logs notes-api-7c9            # what the app said
kubectl logs notes-api-7c9 --previous # what the CRASHED instance said ← key flag
kubectl get events --sort-by=.lastTimestamp | tail -20
kubectl exec -it notes-api-7c9 -- sh  # look around inside
kubectl port-forward svc/notes-api 8080:80  # test the service from your laptop
kubectl top pods                      # actual CPU/memory use
StatusMeaningWhere to look
PendingNot scheduled onto any nodedescribe Events: insufficient cpu/memory, no matching node, unbound PVC
ImagePullBackOffCannot fetch the imageTypo in the tag, private registry with no imagePullSecret
CrashLoopBackOffStarts, dies, restarts, repeatlogs --previous: bad config, missing env var, failed migration
OOMKilled / 137Exceeded its memory limitRaise the limit or fix the leak; check top
Running but 0/1 readyReadiness probe failingWrong path or port in the probe; app not actually up
Service returns nothingNo endpointskubectl get endpoints notes-api — usually a label or targetPort mismatch

kubectl get endpoints <service> deserves special mention: if it is empty, the Service is matching no pods, and you have a label or port problem rather than an application problem. That one command routinely saves an hour.

remember this much
  • You declare desired state; controllers continuously reconcile reality to it.
  • Pods are disposable. Deployments manage pods. Services give a stable address. Ingress routes from outside.
  • Labels and selectors connect everything — a mismatch means silence, not an error.
  • port is the Service's door, targetPort is the container's door.
  • Readiness controls traffic; liveness controls restarts; startup protects slow boots.
  • Always set requests and limits. OOMKilled = exit 137 = over the memory limit.
  • Debugging order: get pods → describe (read Events) → logs --previous → get endpoints.
  • kubectl rollout undo is the fastest fix in an incident.

LABA real deployment on a local cluster

70 minutes · kind, minikube, or Docker Desktop's Kubernetes

Deploy, expose, roll out, break, recover

  1. Create a cluster: kind create cluster (or minikube start). Verify with kubectl get nodes.
  2. Deploy something known-good first, so you separate "learning Kubernetes" from "debugging my app": kubectl create deployment web --image=nginx --replicas=3, then kubectl get pods -w.
  3. Expose it: kubectl expose deployment web --port=80, then kubectl port-forward svc/web 8080:80 and curl it.
  4. Now write real YAML. Create k8s/deployment.yaml and k8s/service.yaml for your chapter-06 image (load it into kind with kind load docker-image notes-api:v1). Include probes, resources, and a ConfigMap. kubectl apply -f k8s/.
  5. Watch a rolling update: in one terminal run kubectl get pods -w; in another, kubectl set image deployment/notes-api api=notes-api:v2. Observe old pods terminating only as new ones become ready.
  6. Break 1 — bad image: set the image to notes-api:nope. Find ImagePullBackOff, confirm it in describe, then kubectl rollout undo and watch it recover. Time yourself.
  7. Break 2 — label mismatch: change the Service selector to app: notes-apii, apply, and curl through the port-forward. It hangs. Run kubectl get endpoints notes-api and see the empty list — the fastest diagnosis in Kubernetes. Fix it.
  8. Break 3 — wrong targetPort: point targetPort at 8080 while the container listens on 3000. Same hang, different cause. Fix it.
  9. Break 4 — crash loop: remove a required environment variable from the ConfigMap and rollout restart. Find the reason with kubectl logs <pod> --previous.
  10. Break 5 — OOM: set limits.memory: 16Mi and watch the pod get OOMKilled with exit code 137.
  11. Stretch: install ingress-nginx with Helm, add an Ingress for notes.localtest.me, and reach your app through it. Then add an HPA and generate load with a busy loop.

Breaks 2 and 3 are the point of this lab. They produce identical symptoms from different causes, and having seen both you will never again stare at healthy pods behind a dead service.

CHECKCheck yourself

A pod is Running and shows 0/1 ready, so no traffic reaches it. Which probe is failing, and what happens next?

Readiness answers "should I get traffic?" and its failure only withdraws the pod from load balancing. Liveness failure restarts the container. Check the probe's path and port first — pointing readiness at a path your app does not serve is the most common cause, and the pod will sit at 0/1 forever without a single error in the app log.

Your Deployment's pods are all Running and healthy, but requests to the Service time out. What do you check first?

A Service is just a label selector plus a port mapping. If its endpoint list is empty it is pointing at nothing, so the fault is a label mismatch; if endpoints exist but requests still fail, suspect targetPort versus the port the container actually listens on. Both were labs in this chapter for exactly this reason.

A container is repeatedly killed with exit code 137. What does that mean and what is the fix?

137 = 128 + 9, meaning SIGKILL, and in Kubernetes that is almost always the memory limit being enforced. Confirm it in kubectl describe pod (Last State: Terminated, Reason: OOMKilled). Then decide honestly whether the limit was too tight or the application is leaking — raising the limit on a real leak just delays the same page.

saved in this browser only — no account needed