TOPIC 01Monitoring vs observability
Monitoring is watching things you already knew to watch: is CPU high, is the disk full, is the process up. It answers questions you thought of in advance.
Observability is being able to answer questions you did not think of in advance — "why are only Android users in Pune seeing checkout failures since 14:20?" — from the data your system already emits. Monitoring is the dashboard; observability is the ability to investigate.
A car's dashboard versus a mechanic's diagnostic port. The dashboard has the gauges someone decided you would need — speed, fuel, temperature. The diagnostic port lets you ask a question nobody anticipated: which sensor reported what, in which order, in the ninety seconds before the warning light. You want the gauges for the daily drive and the port for the day it stalls.
The practical distinction: monitoring tells you something is wrong; observability tells you what. You need both, and the second one is mostly a consequence of instrumenting your code well.
TOPIC 02The three signals
| Signal | What it is | Best at | Cost |
|---|---|---|---|
| Metrics | Numbers over time, aggregated | "How many, how fast, how often" — alerts and trends | Cheap; keep months |
| Logs | Individual events with detail | "What exactly happened to this one request" | Expensive; keep days to weeks |
| Traces | One request's path across services, with timings | "Which of our nine services is the slow one" | Moderate; usually sampled |
The usual investigation runs left to right: a metric alerts you (error rate up), a trace localises it (the payments call is taking 8 seconds), and a log explains it (connection pool exhausted, here's the stack trace). Teams that only have logs end up grepping for trends that a metric would have shown at a glance; teams that only have metrics know something is wrong and nothing else.
TOPIC 03Logs that are searchable
The single highest-value change you can make to logging: emit structured logs — one JSON object per line — instead of prose.
# ✗ unsearchable: you can grep it, and that is all
2026-07-30 14:22:07 ERROR Payment failed for user 4821 after 3 retries
# ✓ structured: every field is now a filter and a chart
{"ts":"2026-07-30T14:22:07Z","level":"error","msg":"payment failed",
"service":"checkout","version":"b72d10","trace_id":"a1f9c2e4",
"user_id":4821,"order_id":"ord_7781","attempt":3,
"provider":"razorpay","duration_ms":8140,"error":"upstream_timeout"}
Now you can ask: how many upstream_timeout errors per minute, grouped by provider, only on version b72d10? That question is unanswerable with the first format and trivial with the second.
- Log levels, used honestly. ERROR = something needs a human. WARN = suspicious but handled. INFO = notable state changes. DEBUG = off in production unless you are hunting something.
- A correlation ID on every line. Generate a
trace_idper incoming request and pass it to every downstream call. Without it, reconstructing one user's journey across services is guesswork. - Never log secrets or personal data. Tokens, passwords, card numbers, full addresses. Logs get shipped, indexed, and read by more people than you expect.
- Write to stdout and let the platform collect it (chapter 06). Log files inside containers vanish with the container.
- Sample the noisy ones. A million identical INFO lines per hour costs real money and hides the one line that mattered.
Typical stacks: Loki + Grafana (cheap, label-based), the ELK/OpenSearch family (powerful full-text search), or a managed service. The pattern is always the same: agent on the node ships stdout to a store, and you query it by labels plus text.
TOPIC 04Metrics and metric types
A metric is a name, a set of labels, and a number sampled over time: http_requests_total{service="api",status="500"}. Four types cover everything:
- Counter
- Only goes up (requests, errors, bytes). You almost always graph its rate, not its value.
- Gauge
- Goes up and down (queue depth, memory in use, active connections, temperature).
- Histogram
- Counts observations into buckets, so you can compute percentiles later. This is how you get p95 latency.
- Summary
- Pre-computed quantiles from the client. Cheaper, but cannot be aggregated across instances — which is usually the deal-breaker.
If 95 requests take 50 ms and 5 take 8 seconds, the average is 447 ms — a number no user experienced. The p95 is what your unhappiest reasonable user feels, and the p99 is what your loudest one does. Alert on percentiles. An average latency graph has hidden more outages than it has revealed.
The four golden signals to instrument for any request-serving service: latency (how long), traffic (how much), errors (how often it fails), saturation (how full the resource is). If you only ever build one dashboard, build that one.
TOPIC 05Prometheus and PromQL
Prometheus scrapes a /metrics endpoint on each target every 15 seconds and stores the numbers. Pull, not push — which means Prometheus also knows when a target stops answering.
# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",route="/checkout",status="200"} 148021
http_requests_total{method="GET",route="/checkout",status="500"} 137
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{route="/checkout",le="0.1"} 141002
http_request_duration_seconds_bucket{route="/checkout",le="0.5"} 147800
http_request_duration_seconds_bucket{route="/checkout",le="+Inf"} 148158
global:
scrape_interval: 15s
scrape_configs:
- job_name: notes-api
static_configs:
- targets: ['api:3000'] # or kubernetes_sd_configs in a cluster
rule_files:
- alert.rules.yml
alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093']
# requests per second, per route (rate over 5 minutes)
sum(rate(http_requests_total[5m])) by (route)
# error ratio as a percentage — the number that belongs in an alert
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# p95 latency from a histogram
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route))
# memory used vs its limit, per pod (saturation)
container_memory_working_set_bytes / container_spec_memory_limit_bytes
# is anything down? (up is 1 or 0)
min(up{job="notes-api"})
# restarts in the last hour — a leading indicator of trouble
increase(kube_pod_container_status_restarts_total[1h]) > 0
Two rules that fix most beginner PromQL: always wrap a counter in rate() (the raw value is meaningless), and make the range at least four scrape intervals ([5m] with 15-second scrapes) so a single missed scrape does not produce a hole in the graph.
TOPIC 06Dashboards people actually use
Grafana draws the graphs. The hard part is restraint. A dashboard with forty panels is a screensaver; a dashboard someone opens at 3 a.m. has one job — is it broken, and where?
- Top row: the four golden signals for the service, at a glance. Request rate, error rate, p95 latency, saturation.
- Second row: dependencies. Database, cache, queue depth, upstream APIs — because the fault is often not in your code.
- Mark deploys on the time axis. The most common cause of a new problem is the most recent change, and seeing the two lined up saves ten minutes of arguing.
- One dashboard per service, plus one overview. Not one giant dashboard for everything.
- Label your axes and units. A panel titled "latency" with no unit is a trap for the next person.
TOPIC 07Traces
In a system with several services, "checkout is slow" needs a follow-up question: which part? A trace follows one request end to end, recording each hop as a span with a start, a duration and a parent.
trace a1f9c2e4 · POST /checkout · total 8.31s ├─ api: validate cart 12ms ├─ api: db SELECT cart items 31ms ├─ payments: POST /charge 8,140ms ◀── there it is │ └─ razorpay: HTTP POST 8,102ms (no timeout set) └─ api: db INSERT order 24ms
Use OpenTelemetry — the vendor-neutral standard for emitting traces, metrics and logs — with a backend like Jaeger, Tempo or a managed service. Two details matter in practice: context propagation (the trace ID must be passed in headers to every downstream call, or the trace breaks in half) and sampling (keep a small percentage of normal traces, and ideally all the slow or failed ones).
TOPIC 08Alerts worth waking up for
Bad alerting is worse than none: it trains people to ignore their phone. The test for any alert is a single question — if this fires at 3 a.m., is there something a human must do right now? If not, it is a dashboard panel or a ticket, not a page.
| Don't page on | Page on |
|---|---|
| CPU above 80% | Error rate above 2% for 5 minutes |
| A pod restarted once | Checkout success rate below 98% |
| Disk 61% full | Disk will be full within 4 hours (predicted) |
| A single slow request | p95 latency above 1s for 10 minutes |
| A cron job took longer than usual | The nightly job has not succeeded in 26 hours |
The pattern: alert on symptoms users feel, not on causes you happen to be able to measure. High CPU with happy users is not an incident. Errors with low CPU very much is.
groups:
- name: notes-api
rules:
- alert: HighErrorRate
expr: |
100 * sum(rate(http_requests_total{job="notes-api",status=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="notes-api"}[5m])) > 2
for: 5m # 1. duration: no flapping
labels:
severity: page # 2. routing: page vs ticket
annotations:
summary: "5xx rate {{ $value | printf \"%.1f\" }}% on notes-api"
impact: "Users see failed requests on checkout." # 3. why it matters
runbook: "https://wiki/runbooks/notes-api-5xx" # 4. what to do
The for: clause and the runbook link are what separate a professional alert from a nuisance. A page at 3 a.m. with a link to the exact commands to run is a completely different experience from a page that says KubePodCrashLooping and nothing else.
TOPIC 09SLIs, SLOs and error budgets
- SLI — Service Level Indicator
- The measurement. "Percentage of checkout requests that return 2xx in under 500 ms."
- SLO — Service Level Objective
- Your internal target for that measurement. "99.9% over 30 days."
- SLA — Service Level Agreement
- A contract with customers, with financial penalties. Always looser than your SLO.
The error budget is the arithmetic that makes this operationally useful. A 99.9% SLO over 30 days allows 0.1% failure — about 43 minutes a month.
99% ≈ 7h 12m unhealthy per month 99.9% ≈ 43m ← a sensible target for most services 99.99% ≈ 4m 19s ← expensive: needs multi-region, heavy automation 99.999% ≈ 26s ← very few systems genuinely need this budget remaining ──▶ ship features, take risks, deploy on Friday budget spent ──▶ freeze risky changes, spend the sprint on reliability
A monthly data pack. Plenty left mid-month? Stream freely. Nearly out? Be careful until it resets. The point is not to hit zero downtime — that costs infinitely and slows you to a halt — but to have an agreed, numerical answer to "can we take this risk this week?" that does not depend on who argues loudest.
Pick one or two SLOs per service, on things users actually feel (availability and latency of the key journey), and review them monthly. Three well-chosen SLOs beat thirty metrics with opinions attached.
TOPIC 10Running an incident
Under stress, people either freeze or all dive at the same theory. Roles and a routine fix both.
- Incident commander. Coordinates, decides, communicates. Does not debug — the moment the commander is head-down in logs, nobody is running the incident.
- Operations / responders. The people actually investigating and applying fixes.
- Scribe. Timestamps every action and finding in the channel. This is the postmortem, written as it happens, and it takes one person almost no effort.
- Comms. Updates the status page and stakeholders on a fixed cadence, so nobody has to interrupt the responders to ask.
1. ACKNOWLEDGE take the page so nobody else is woken
2. ASSESS what is the user impact, and how big? (dashboard, not code)
3. DECLARE open a channel, name a commander, say what you know
4. MITIGATE stop the bleeding — roll back, disable a flag, shed load
mitigation before diagnosis. always.
5. VERIFY prove recovery with a real request and the graphs — every region
6. DIAGNOSE now, with users safe, find out why
7. DOCUMENT timeline, impact, cause, action items — within 48 hours
Step 4 before step 6. The instinct to understand first is strong and expensive — every minute spent reading code is a minute of failing user requests. The Incident Room game in the arcade is built entirely around this ordering; play it before you are ever on call.
TOPIC 11Blameless postmortems
Blameless does not mean nobody is accountable. It means the investigation targets the conditions that allowed a mistake to become an outage, not the person whose keystroke was last. The reason is practical rather than sentimental: in a blaming culture people hide detail, and hidden detail is exactly what you need to prevent a repeat.
| Blaming | Blameless |
|---|---|
| "Ravi deployed without testing." | "A deploy could reach production with no smoke test. Why was that possible, and what now makes it impossible?" |
| "Someone deleted the wrong table." | "A single command could drop a production table with no confirmation and no recent backup verification." |
A good postmortem is short and contains: a timeline with timestamps, the user impact in numbers (how many users, how long, what failed), the contributing factors, what made detection or recovery slow, and action items with an owner and a date. Action items without an owner are decoration — and a postmortem whose actions are never done teaches the team that writing them is theatre.
Two metrics to watch over time: MTTD (mean time to detect) and MTTR (mean time to restore). Improving detection is usually about alerting on the right symptom; improving restoration is usually about practised rollbacks. Both are learnable skills, and both are what "time to restore service" from chapter 01 was measuring all along.
- Monitoring answers questions you prepared; observability lets you ask new ones.
- Metrics for trends and alerts, logs for detail, traces for "which service is slow".
- Structured JSON logs with a trace ID. Never log secrets. Write to stdout.
- Wrap counters in
rate(). Alert on percentiles, never averages. - Four golden signals: latency, traffic, errors, saturation.
- Page only on user-visible symptoms, with a duration and a runbook link.
- 99.9% = 43 minutes a month. The error budget decides how much risk you can take.
- In an incident: acknowledge, assess, declare, mitigate, verify, then diagnose.
LABWatch your own app
Instrument, graph, alert, then cause the alert on purpose
- Add a
/metricsendpoint to your chapter-06 app using the Prometheus client for your language. Expose a counter for requests (labelled by route and status) and a histogram for duration. - Add Prometheus and Grafana to your
compose.yamlwith the scrape config from topic 05. Bring it up and confirm your target is UP in the Prometheus UI (localhost:9090/targets). - In the Prometheus expression browser, run the six PromQL queries from topic 05 against your own data. Generate traffic with a loop:
while true; do curl -s localhost:3000/ > /dev/null; done. - Build one Grafana dashboard with exactly four panels: request rate, error rate %, p95 latency, memory. Nothing else — practise restraint.
- Make errors on purpose: add a route that returns 500 for one in five requests. Hit it in a loop and watch your error-rate panel climb. Confirm the percentage matches what you expect — this is how you learn to trust a graph.
- Add the
HighErrorRaterule from topic 08 withfor: 1m(shorter, so the lab is not tedious). Watch it move through Inactive → Pending → Firing in the Prometheus Alerts tab. Notice thatfor:is what prevents a single blip from paging anyone. - Structured logging: switch your app's logs to one JSON object per line with a
trace_id. Add Loki to Compose and query{container="api"} | json | level="error"in Grafana. - Latency, not just errors: add an endpoint that sleeps for a random 0–3 seconds. Watch your p95 panel rise while the average barely moves. This is the percentile lesson, made visible.
- Write an SLO: pick one journey, define the SLI precisely in one sentence, choose a target, and compute the monthly error budget in minutes. Write it in your README.
- Stretch: play Incident Room and try to close it in under 12 minutes. Then write a one-page postmortem for the outage in the game, with three action items that each have an owner.
CHECKCheck yourself
Average response time looks fine at 180 ms, but users complain the site is slow. What are you probably missing?
An average hides distribution. Ninety-five fast requests plus five eight-second ones average out to something reassuring while a real slice of users has a terrible time — and those users are often the ones with the largest carts or the most data. Graph and alert on p95 and p99, from a histogram.
Which alert is best designed?
It has all four properties of a good alert: it measures a symptom users feel, it has a duration so a blip does not page anyone, it is labelled for routing, and it links to what to do. High CPU with happy users is not an incident; one pod restart is what Kubernetes is for; 61% disk is a ticket at most.
You are paged at 3 a.m. Errors began four minutes after a deploy. What is the correct first substantive action?
Mitigate before you diagnose. The recent change is the prime suspect, and rolling back is fast, reversible and stops user pain immediately — after which you can read the diff calmly. Scaling only multiplies a bug that is not caused by saturation, and waiting simply spends more error budget.