devopsdiary
Next chapter

diary / chapters / 10

chapter 10 · ~60 min · prometheus lab

Observability, alerting and on-call

Everything before this chapter puts software into production. This chapter is how you know whether it is actually working — and how you find out from a graph rather than from a customer on social media.

TOPIC 01Monitoring vs observability

Monitoring is watching things you already knew to watch: is CPU high, is the disk full, is the process up. It answers questions you thought of in advance.

Observability is being able to answer questions you did not think of in advance — "why are only Android users in Pune seeing checkout failures since 14:20?" — from the data your system already emits. Monitoring is the dashboard; observability is the ability to investigate.

real life

A car's dashboard versus a mechanic's diagnostic port. The dashboard has the gauges someone decided you would need — speed, fuel, temperature. The diagnostic port lets you ask a question nobody anticipated: which sensor reported what, in which order, in the ninety seconds before the warning light. You want the gauges for the daily drive and the port for the day it stalls.

The practical distinction: monitoring tells you something is wrong; observability tells you what. You need both, and the second one is mostly a consequence of instrumenting your code well.

TOPIC 02The three signals

SignalWhat it isBest atCost
MetricsNumbers over time, aggregated"How many, how fast, how often" — alerts and trendsCheap; keep months
LogsIndividual events with detail"What exactly happened to this one request"Expensive; keep days to weeks
TracesOne request's path across services, with timings"Which of our nine services is the slow one"Moderate; usually sampled

The usual investigation runs left to right: a metric alerts you (error rate up), a trace localises it (the payments call is taking 8 seconds), and a log explains it (connection pool exhausted, here's the stack trace). Teams that only have logs end up grepping for trends that a metric would have shown at a glance; teams that only have metrics know something is wrong and nothing else.

TOPIC 03Logs that are searchable

The single highest-value change you can make to logging: emit structured logs — one JSON object per line — instead of prose.

the same event, two ways
# ✗ unsearchable: you can grep it, and that is all
2026-07-30 14:22:07 ERROR Payment failed for user 4821 after 3 retries

# ✓ structured: every field is now a filter and a chart
{"ts":"2026-07-30T14:22:07Z","level":"error","msg":"payment failed",
 "service":"checkout","version":"b72d10","trace_id":"a1f9c2e4",
 "user_id":4821,"order_id":"ord_7781","attempt":3,
 "provider":"razorpay","duration_ms":8140,"error":"upstream_timeout"}

Now you can ask: how many upstream_timeout errors per minute, grouped by provider, only on version b72d10? That question is unanswerable with the first format and trivial with the second.

  • Log levels, used honestly. ERROR = something needs a human. WARN = suspicious but handled. INFO = notable state changes. DEBUG = off in production unless you are hunting something.
  • A correlation ID on every line. Generate a trace_id per incoming request and pass it to every downstream call. Without it, reconstructing one user's journey across services is guesswork.
  • Never log secrets or personal data. Tokens, passwords, card numbers, full addresses. Logs get shipped, indexed, and read by more people than you expect.
  • Write to stdout and let the platform collect it (chapter 06). Log files inside containers vanish with the container.
  • Sample the noisy ones. A million identical INFO lines per hour costs real money and hides the one line that mattered.

Typical stacks: Loki + Grafana (cheap, label-based), the ELK/OpenSearch family (powerful full-text search), or a managed service. The pattern is always the same: agent on the node ships stdout to a store, and you query it by labels plus text.

TOPIC 04Metrics and metric types

A metric is a name, a set of labels, and a number sampled over time: http_requests_total{service="api",status="500"}. Four types cover everything:

Counter
Only goes up (requests, errors, bytes). You almost always graph its rate, not its value.
Gauge
Goes up and down (queue depth, memory in use, active connections, temperature).
Histogram
Counts observations into buckets, so you can compute percentiles later. This is how you get p95 latency.
Summary
Pre-computed quantiles from the client. Cheaper, but cannot be aggregated across instances — which is usually the deal-breaker.
averages lie; percentiles do not

If 95 requests take 50 ms and 5 take 8 seconds, the average is 447 ms — a number no user experienced. The p95 is what your unhappiest reasonable user feels, and the p99 is what your loudest one does. Alert on percentiles. An average latency graph has hidden more outages than it has revealed.

The four golden signals to instrument for any request-serving service: latency (how long), traffic (how much), errors (how often it fails), saturation (how full the resource is). If you only ever build one dashboard, build that one.

TOPIC 05Prometheus and PromQL

Prometheus scrapes a /metrics endpoint on each target every 15 seconds and stores the numbers. Pull, not push — which means Prometheus also knows when a target stops answering.

what /metrics looks like
# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",route="/checkout",status="200"} 148021
http_requests_total{method="GET",route="/checkout",status="500"} 137

# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{route="/checkout",le="0.1"} 141002
http_request_duration_seconds_bucket{route="/checkout",le="0.5"} 147800
http_request_duration_seconds_bucket{route="/checkout",le="+Inf"} 148158
prometheus.yml
global:
  scrape_interval: 15s

scrape_configs:
  - job_name: notes-api
    static_configs:
      - targets: ['api:3000']       # or kubernetes_sd_configs in a cluster

rule_files:
  - alert.rules.yml

alerting:
  alertmanagers:
    - static_configs:
        - targets: ['alertmanager:9093']
PromQL — the six queries you will actually write
# requests per second, per route (rate over 5 minutes)
sum(rate(http_requests_total[5m])) by (route)

# error ratio as a percentage — the number that belongs in an alert
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
    / sum(rate(http_requests_total[5m]))

# p95 latency from a histogram
histogram_quantile(0.95,
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route))

# memory used vs its limit, per pod (saturation)
container_memory_working_set_bytes / container_spec_memory_limit_bytes

# is anything down? (up is 1 or 0)
min(up{job="notes-api"})

# restarts in the last hour — a leading indicator of trouble
increase(kube_pod_container_status_restarts_total[1h]) > 0

Two rules that fix most beginner PromQL: always wrap a counter in rate() (the raw value is meaningless), and make the range at least four scrape intervals ([5m] with 15-second scrapes) so a single missed scrape does not produce a hole in the graph.

TOPIC 06Dashboards people actually use

Grafana draws the graphs. The hard part is restraint. A dashboard with forty panels is a screensaver; a dashboard someone opens at 3 a.m. has one job — is it broken, and where?

  • Top row: the four golden signals for the service, at a glance. Request rate, error rate, p95 latency, saturation.
  • Second row: dependencies. Database, cache, queue depth, upstream APIs — because the fault is often not in your code.
  • Mark deploys on the time axis. The most common cause of a new problem is the most recent change, and seeing the two lined up saves ten minutes of arguing.
  • One dashboard per service, plus one overview. Not one giant dashboard for everything.
  • Label your axes and units. A panel titled "latency" with no unit is a trap for the next person.

TOPIC 07Traces

In a system with several services, "checkout is slow" needs a follow-up question: which part? A trace follows one request end to end, recording each hop as a span with a start, a duration and a parent.

one slow request, made obvious
trace a1f9c2e4  ·  POST /checkout  ·  total 8.31s
├─ api: validate cart                      12ms
├─ api: db SELECT cart items               31ms
├─ payments: POST /charge               8,140ms  ◀── there it is
│   └─ razorpay: HTTP POST              8,102ms      (no timeout set)
└─ api: db INSERT order                    24ms

Use OpenTelemetry — the vendor-neutral standard for emitting traces, metrics and logs — with a backend like Jaeger, Tempo or a managed service. Two details matter in practice: context propagation (the trace ID must be passed in headers to every downstream call, or the trace breaks in half) and sampling (keep a small percentage of normal traces, and ideally all the slow or failed ones).

TOPIC 08Alerts worth waking up for

Bad alerting is worse than none: it trains people to ignore their phone. The test for any alert is a single question — if this fires at 3 a.m., is there something a human must do right now? If not, it is a dashboard panel or a ticket, not a page.

Don't page onPage on
CPU above 80%Error rate above 2% for 5 minutes
A pod restarted onceCheckout success rate below 98%
Disk 61% fullDisk will be full within 4 hours (predicted)
A single slow requestp95 latency above 1s for 10 minutes
A cron job took longer than usualThe nightly job has not succeeded in 26 hours

The pattern: alert on symptoms users feel, not on causes you happen to be able to measure. High CPU with happy users is not an incident. Errors with low CPU very much is.

alert.rules.yml — a good alert has four parts
groups:
  - name: notes-api
    rules:
      - alert: HighErrorRate
        expr: |
          100 * sum(rate(http_requests_total{job="notes-api",status=~"5.."}[5m]))
              / sum(rate(http_requests_total{job="notes-api"}[5m])) > 2
        for: 5m                        # 1. duration: no flapping
        labels:
          severity: page               # 2. routing: page vs ticket
        annotations:
          summary: "5xx rate {{ $value | printf \"%.1f\" }}% on notes-api"
          impact: "Users see failed requests on checkout."   # 3. why it matters
          runbook: "https://wiki/runbooks/notes-api-5xx"     # 4. what to do

The for: clause and the runbook link are what separate a professional alert from a nuisance. A page at 3 a.m. with a link to the exact commands to run is a completely different experience from a page that says KubePodCrashLooping and nothing else.

TOPIC 09SLIs, SLOs and error budgets

SLI — Service Level Indicator
The measurement. "Percentage of checkout requests that return 2xx in under 500 ms."
SLO — Service Level Objective
Your internal target for that measurement. "99.9% over 30 days."
SLA — Service Level Agreement
A contract with customers, with financial penalties. Always looser than your SLO.

The error budget is the arithmetic that makes this operationally useful. A 99.9% SLO over 30 days allows 0.1% failure — about 43 minutes a month.

what the budget buys you
  99%     ≈ 7h 12m unhealthy per month
  99.9%   ≈ 43m           ← a sensible target for most services
  99.99%  ≈ 4m 19s        ← expensive: needs multi-region, heavy automation
  99.999% ≈ 26s           ← very few systems genuinely need this

  budget remaining ──▶ ship features, take risks, deploy on Friday
  budget spent     ──▶ freeze risky changes, spend the sprint on reliability
real life

A monthly data pack. Plenty left mid-month? Stream freely. Nearly out? Be careful until it resets. The point is not to hit zero downtime — that costs infinitely and slows you to a halt — but to have an agreed, numerical answer to "can we take this risk this week?" that does not depend on who argues loudest.

Pick one or two SLOs per service, on things users actually feel (availability and latency of the key journey), and review them monthly. Three well-chosen SLOs beat thirty metrics with opinions attached.

TOPIC 10Running an incident

Under stress, people either freeze or all dive at the same theory. Roles and a routine fix both.

  • Incident commander. Coordinates, decides, communicates. Does not debug — the moment the commander is head-down in logs, nobody is running the incident.
  • Operations / responders. The people actually investigating and applying fixes.
  • Scribe. Timestamps every action and finding in the channel. This is the postmortem, written as it happens, and it takes one person almost no effort.
  • Comms. Updates the status page and stakeholders on a fixed cadence, so nobody has to interrupt the responders to ask.
the routine that shortens every outage
1. ACKNOWLEDGE   take the page so nobody else is woken
2. ASSESS        what is the user impact, and how big? (dashboard, not code)
3. DECLARE       open a channel, name a commander, say what you know
4. MITIGATE      stop the bleeding — roll back, disable a flag, shed load
                 mitigation before diagnosis. always.
5. VERIFY        prove recovery with a real request and the graphs — every region
6. DIAGNOSE      now, with users safe, find out why
7. DOCUMENT      timeline, impact, cause, action items — within 48 hours
the habit worth drilling

Step 4 before step 6. The instinct to understand first is strong and expensive — every minute spent reading code is a minute of failing user requests. The Incident Room game in the arcade is built entirely around this ordering; play it before you are ever on call.

TOPIC 11Blameless postmortems

Blameless does not mean nobody is accountable. It means the investigation targets the conditions that allowed a mistake to become an outage, not the person whose keystroke was last. The reason is practical rather than sentimental: in a blaming culture people hide detail, and hidden detail is exactly what you need to prevent a repeat.

BlamingBlameless
"Ravi deployed without testing.""A deploy could reach production with no smoke test. Why was that possible, and what now makes it impossible?"
"Someone deleted the wrong table.""A single command could drop a production table with no confirmation and no recent backup verification."

A good postmortem is short and contains: a timeline with timestamps, the user impact in numbers (how many users, how long, what failed), the contributing factors, what made detection or recovery slow, and action items with an owner and a date. Action items without an owner are decoration — and a postmortem whose actions are never done teaches the team that writing them is theatre.

Two metrics to watch over time: MTTD (mean time to detect) and MTTR (mean time to restore). Improving detection is usually about alerting on the right symptom; improving restoration is usually about practised rollbacks. Both are learnable skills, and both are what "time to restore service" from chapter 01 was measuring all along.

remember this much
  • Monitoring answers questions you prepared; observability lets you ask new ones.
  • Metrics for trends and alerts, logs for detail, traces for "which service is slow".
  • Structured JSON logs with a trace ID. Never log secrets. Write to stdout.
  • Wrap counters in rate(). Alert on percentiles, never averages.
  • Four golden signals: latency, traffic, errors, saturation.
  • Page only on user-visible symptoms, with a duration and a runbook link.
  • 99.9% = 43 minutes a month. The error budget decides how much risk you can take.
  • In an incident: acknowledge, assess, declare, mitigate, verify, then diagnose.

LABWatch your own app

60 minutes · Docker Compose is enough

Instrument, graph, alert, then cause the alert on purpose

  1. Add a /metrics endpoint to your chapter-06 app using the Prometheus client for your language. Expose a counter for requests (labelled by route and status) and a histogram for duration.
  2. Add Prometheus and Grafana to your compose.yaml with the scrape config from topic 05. Bring it up and confirm your target is UP in the Prometheus UI (localhost:9090/targets).
  3. In the Prometheus expression browser, run the six PromQL queries from topic 05 against your own data. Generate traffic with a loop: while true; do curl -s localhost:3000/ > /dev/null; done.
  4. Build one Grafana dashboard with exactly four panels: request rate, error rate %, p95 latency, memory. Nothing else — practise restraint.
  5. Make errors on purpose: add a route that returns 500 for one in five requests. Hit it in a loop and watch your error-rate panel climb. Confirm the percentage matches what you expect — this is how you learn to trust a graph.
  6. Add the HighErrorRate rule from topic 08 with for: 1m (shorter, so the lab is not tedious). Watch it move through Inactive → Pending → Firing in the Prometheus Alerts tab. Notice that for: is what prevents a single blip from paging anyone.
  7. Structured logging: switch your app's logs to one JSON object per line with a trace_id. Add Loki to Compose and query {container="api"} | json | level="error" in Grafana.
  8. Latency, not just errors: add an endpoint that sleeps for a random 0–3 seconds. Watch your p95 panel rise while the average barely moves. This is the percentile lesson, made visible.
  9. Write an SLO: pick one journey, define the SLI precisely in one sentence, choose a target, and compute the monthly error budget in minutes. Write it in your README.
  10. Stretch: play Incident Room and try to close it in under 12 minutes. Then write a one-page postmortem for the outage in the game, with three action items that each have an owner.

CHECKCheck yourself

Average response time looks fine at 180 ms, but users complain the site is slow. What are you probably missing?

An average hides distribution. Ninety-five fast requests plus five eight-second ones average out to something reassuring while a real slice of users has a terrible time — and those users are often the ones with the largest carts or the most data. Graph and alert on p95 and p99, from a histogram.

Which alert is best designed?

It has all four properties of a good alert: it measures a symptom users feel, it has a duration so a blip does not page anyone, it is labelled for routing, and it links to what to do. High CPU with happy users is not an incident; one pod restart is what Kubernetes is for; 61% disk is a ticket at most.

You are paged at 3 a.m. Errors began four minutes after a deploy. What is the correct first substantive action?

Mitigate before you diagnose. The recent change is the prime suspect, and rolling back is fast, reversible and stops user pain immediately — after which you can read the diff calmly. Scaling only multiplies a bug that is not caused by saturation, and waiting simply spends more error budget.

saved in this browser only — no account needed