TOPIC 01What happens when you type a URL
Everything else in this chapter is a detail of this sequence. Learn the order and you can bisect any failure.
1. DNS shop.example.com → 203.0.113.42 "what is the address?" 2. route packets find their way to that address "can I reach it?" 3. TCP three-way handshake on port 443 "is anyone answering?" 4. TLS certificate check, keys agreed "is it really them?" 5. HTTP GET /checkout, headers, cookies "here's my request" 6. proxy nginx picks a healthy backend "who handles this?" 7. app your code runs, maybe queries a DB "do the work" 8. response status code + body travel back "here's the answer"
Posting a letter. DNS is looking the address up in a directory; routing is the postal network moving it; TCP is the recipient signing for delivery; TLS is a tamper-proof sealed envelope only they can open; HTTP is the language the letter is written in; the reverse proxy is the mailroom that decides which desk gets it. When a letter goes missing you do not check all steps at once — you check them in order.
TOPIC 02IP addresses and subnets
An IP address identifies a network interface. IPv4 looks like 10.0.3.14; IPv6 looks like 2001:db8::1. The part that actually matters day to day is the difference between public and private, and what the /24 means.
- Public addresses are reachable from the internet. Private ranges are not routable on the internet and are reused inside every network on earth:
10.0.0.0/8,172.16.0.0/12,192.168.0.0/16. - CIDR notation — the
/n— says how many leading bits are fixed as the network part. The rest is available for hosts. 127.0.0.1(localhost) is the machine talking to itself.0.0.0.0, when a server binds to it, means "listen on every interface" — which is why an app bound to127.0.0.1works locally but cannot be reached from outside. That single misunderstanding accounts for a huge share of "my container isn't responding" tickets.
| CIDR | Addresses | Typically used for |
|---|---|---|
/32 | 1 | One exact host, e.g. a firewall rule for your office IP |
/24 | 256 | One subnet — a tidy default for an app tier |
/20 | 4,096 | A large subnet in a cloud VPC |
/16 | 65,536 | A whole VPC, e.g. 10.0.0.0/16 |
A VPC is a gated township (10.0.0.0/16). Subnets are its lanes (10.0.1.0/24, 10.0.2.0/24). Public subnets have a gate to the main road; private subnets do not, which is exactly where you put your database. A NAT gateway is the township's shared courier: private houses can send parcels out, but no stranger can walk in.
TOPIC 03DNS: the phone book
DNS translates names to addresses. You will meet these record types constantly:
- A / AAAA
- Name → IPv4 / IPv6 address.
api.example.com → 203.0.113.42 - CNAME
- Name → another name. Points your domain at a load balancer's hostname that may change IPs.
- MX
- Where email for this domain goes.
- TXT
- Free text. Used for domain verification and email anti-spoofing (SPF/DKIM).
- NS
- Which nameservers are authoritative for this domain.
TTL (time to live) is how many seconds resolvers may cache an answer. It is the reason DNS changes seem not to work: your browser, your OS, your router and your ISP may all be holding the old answer. Before a planned migration, drop the TTL to 60 seconds a day in advance; afterwards, raise it again.
dig api.example.com +short # just the answer dig api.example.com # full answer, with TTL dig api.example.com @8.8.8.8 # ask a specific resolver — is it just MY cache? dig +trace api.example.com # follow the delegation from the root down dig -t MX example.com +short host api.example.com # shorter output getent hosts api.example.com # what the OS resolves, /etc/hosts included cat /etc/hosts # local overrides — check this before panicking
TOPIC 04Ports and sockets
An IP address gets you to the machine; a port gets you to the right program on it. A connection is identified by four things — source IP, source port, destination IP, destination port — and that tuple is a socket.
| Port | Service | Port | Service |
|---|---|---|---|
| 22 | SSH | 5432 | PostgreSQL |
| 80 | HTTP | 3306 | MySQL |
| 443 | HTTPS | 6379 | Redis |
| 53 | DNS | 9090 | Prometheus |
| 25 / 587 | SMTP | 3000 / 8080 | App defaults |
The IP is the building's street address; the port is the flat number. "Connection refused" means you reached the building and knocked on flat 8080, but nobody lives there. "Connection timed out" means your knock never even arrived — the gate guard (a firewall) silently dropped you. Those two errors point at completely different teams, which is why telling them apart is such a useful reflex.
ss -ltnp # listening TCP sockets + owning process
State Local Address:Port Process
LISTEN 0.0.0.0:80 users:(("nginx",pid=812))
LISTEN 127.0.0.1:3000 users:(("node",pid=1440)) ← local only!
ss -tn state established # current connections
nc -zv api.internal 5432 # can I even open that port? (great first test)
curl -v telnet://db-01:5432 # same idea without nc installed
TOPIC 05TCP vs UDP
| TCP | UDP | |
|---|---|---|
| Connection | Handshake first (SYN, SYN-ACK, ACK) | None — just send |
| Guarantees | Ordered, retransmitted, no duplicates | None; packets may vanish or arrive out of order |
| Cost | Slightly slower to start | Minimal overhead |
| Used by | HTTP, SSH, databases — anything that must be correct | DNS, video calls, game state, metrics push |
TCP is a registered parcel with signature on delivery: slower to set up, but you know it arrived. UDP is shouting across a noisy room: if a word is lost, you carry on. For a video call that is the right trade — a re-sent frame from two seconds ago is useless. For a bank transfer it very much is not.
TOPIC 06HTTP, properly
A request is a method, a path, headers and (sometimes) a body. A response is a status code, headers and a body. That is the whole protocol.
curl -i https://api.example.com/health # show response headers + body
curl -v https://api.example.com/health # show the whole conversation
curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' https://example.com
200 0.184s
# POST some JSON with a token
curl -X POST https://api.example.com/orders \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer $TOKEN" \
-d '{"item":"tea","qty":2}'
curl -I https://example.com # headers only (HEAD request)
curl --resolve api.example.com:443:10.0.3.14 https://api.example.com/
# ↑ test a specific server behind a DNS name — perfect during a migration
Status codes as a triage tool
- 2xx — worked.
200OK,201created,204no content. - 3xx — go elsewhere.
301permanent,302temporary,304not modified (your cache is fine). - 4xx — the client's fault.
400malformed,401not authenticated,403authenticated but not allowed,404no such thing,422validation failed,429rate limited. - 5xx — the server's fault, and yours to fix.
500unhandled error in your code,502bad gateway (the proxy got nothing usable from the app),503unavailable/overloaded,504gateway timeout (the app was too slow).
502 = your app is not answering at all: process dead, wrong port, crashed on boot. 504 = your app answered too slowly: a hung database query, an external API with no timeout, a thread pool exhausted. Same red page for the user, completely different investigation. Ninety seconds of reading the proxy's error log tells you which one you have.
TOPIC 07TLS and certificates
HTTPS is HTTP inside TLS. TLS gives you three things: encryption (nobody can read it), integrity (nobody can change it), and identity (you are talking to who you think). Identity is the part that breaks in production.
A certificate says "this public key belongs to shop.example.com", signed by a Certificate Authority your operating system already trusts. Browsers reject it when it has expired, when the name does not match, or when the signing chain is incomplete.
# when does it expire? (put this in a monitor — expiry outages are pure own-goals) echo | openssl s_client -connect example.com:443 2>/dev/null \ | openssl x509 -noout -dates -subject notBefore=Jun 2 00:00:00 2026 GMT notAfter=Aug 31 23:59:59 2026 GMT curl -vI https://example.com 2>&1 | grep -E 'subject|issuer|expire' openssl s_client -connect example.com:443 -servername example.com # -servername matters: one IP can serve many certs (SNI)
In practice you will use Let's Encrypt with certbot, or terminate TLS at a cloud load balancer / ingress controller and let it handle renewal. Either way: certificates are renewed automatically, and you alert on "expires in under 14 days", because a human calendar reminder is not a control.
TOPIC 08Reverse proxies and load balancers
A reverse proxy sits in front of your application and speaks to clients on its behalf. It is where you put the things that should not be your app's problem: TLS, compression, caching, rate limits, redirects, and routing multiple apps onto one domain. A load balancer is a reverse proxy that also spreads requests across several instances and stops sending traffic to unhealthy ones.
upstream notes_api {
server 127.0.0.1:3000; # add more lines to load balance
server 127.0.0.1:3001;
}
server {
listen 80;
server_name api.example.com;
location /health { access_log off; proxy_pass http://notes_api; }
location / {
proxy_pass http://notes_api;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_connect_timeout 3s;
proxy_read_timeout 30s; # the difference between 502 and 504
}
}
sudo nginx -t # syntax check FIRST nginx: configuration file /etc/nginx/nginx.conf test is successful sudo systemctl reload nginx # reload = no dropped connections sudo tail -f /var/log/nginx/error.log
Balancing strategies you should be able to name: round robin (in turn), least connections (to the least busy — better for uneven request times), IP hash (same client to the same server, for sticky sessions). And the feature that matters more than any of them: health checks, so a dead instance is removed from rotation automatically. That is what makes zero-downtime deploys possible in chapter 09.
TOPIC 09Firewalls and security groups
A firewall decides which traffic is allowed. On a single Linux host that is ufw or nftables; in a cloud it is a security group attached to the instance. The model to internalise: default deny, then allow the minimum.
sudo ufw default deny incoming sudo ufw default allow outgoing sudo ufw allow 22/tcp # do this BEFORE enabling, or lock yourself out sudo ufw allow 443/tcp sudo ufw allow from 10.0.1.0/24 to any port 5432 # db: app subnet only sudo ufw enable sudo ufw status numbered
Load balancer: 80/443 from anywhere. App servers: 3000 from the load balancer only. Database: 5432 from the app subnet only. SSH: from a bastion or your VPN, never from 0.0.0.0/0. If a database is reachable from the internet, that is not a hardening task for later — it is the incident that has not happened yet.
TOPIC 10The debugging ladder
Memorise this order. It turns "the site is down" from panic into six checks, and each rung eliminates a whole class of cause.
1. does the name resolve? dig api.example.com +short ✗ → DNS: wrong record, expired domain, /etc/hosts override, TTL cache 2. can I reach the host? ping 203.0.113.42 (ICMP may be blocked — not proof) ✗ → routing, VPN, wrong subnet, instance stopped 3. is the port open? nc -zv api.example.com 443 refused → nothing listening: app crashed, wrong port, bound to 127.0.0.1 timeout → something is dropping packets: firewall / security group 4. does TLS complete? curl -vI https://api.example.com ✗ → expired cert, name mismatch, missing chain, SNI 5. what does HTTP say? curl -i https://api.example.com/health 4xx → auth, path, or client problem 5xx → your app or its proxy 6. what do the logs say? proxy log first, then app log, then DB 502 → app not answering 504 → app too slow 500 → read the stack trace
A tap with no water. You do not start by dismantling the tap — you check the street supply, then the building tank, then the pipe to your flat, then the tap. Same ladder, same reason: each step is cheap and rules out everything below it. Engineers who look slow and methodical during an incident are usually the ones who finish first.
- Request order: DNS → route → TCP → TLS → HTTP → proxy → app. Debug in that order.
- Refused = nothing listening. Timeout = something is dropping the packets (firewall).
- Binding to
127.0.0.1makes a service unreachable from outside;0.0.0.0listens everywhere. - 502 = app not answering; 504 = app too slow; 500 = app threw an error.
- DNS changes are delayed by TTL caching — lower the TTL before a migration, not during one.
- Firewalls: default deny, then allow the minimum per tier. Databases never face the internet.
LABPut nginx in front of an app
Build the shape that runs most of the web, then break it in four ways
- Run something on port 3000. Anything:
python3 -m http.server 3000is enough. Confirm withcurl -i localhost:3000. - Install nginx and add a server block that proxies
/tohttp://127.0.0.1:3000, with the fourproxy_set_headerlines from above. sudo nginx -t, thensudo systemctl reload nginx. Nowcurl -i localhostshould return your app through port 80. You have just built a reverse proxy.- Break 1 — 502: stop the app on 3000 and curl again. Read the exact line nginx writes in
error.log. Learn to recognise "connect() failed (111: Connection refused)". - Break 2 — 504: set
proxy_read_timeout 2s, and point it at something slow (nc -l 3000and never reply). Watch a gateway timeout appear after exactly two seconds. - Break 3 — the bind trap: start your app bound to
127.0.0.1only, and try to reach it from another machine or from your host OS. Then rebind to0.0.0.0and try again. This is the single most common container networking bug. - Break 4 — the firewall:
sudo ufw allow 22/tcp, thensudo ufw enable. Try to reach port 80 from elsewhere and note that it now times out rather than being refused. Addsudo ufw allow 80/tcpand watch it recover. - Stretch: add a second app on 3001, list both in the
upstreamblock, reload, and runcurl -s localhost/whoamiin a loop to see requests alternate. You have now built a load balancer.
Why this lab matters: steps 4–6 are the three failures you will actually meet, and having produced each one on purpose is what lets you recognise them in seconds later.
CHECKCheck yourself
curl to your API hangs for 30 seconds and then reports a timeout. A colleague's curl to the same URL is refused instantly. What does that difference tell you?
A silent drop produces a timeout; an active rejection produces "connection refused" immediately. So the two of you are hitting different obstacles: you are blocked before the host (security group, ufw, VPN), and your colleague is getting through to a host where the process is not listening. Two symptoms, two different fixes — and if DNS were the problem you would see a resolution error instead.
Your app runs fine with curl localhost:3000 on the server, but nginx in front of it returns 502. Which cause fits best?
502 means nginx could not get a usable response from the upstream, so the fault is between proxy and app. Check the exact upstream address in the config against ss -ltnp output, then read error.log — it names the failing address and the OS error. Certificate and DNS problems fail earlier, before HTTP ever happens.
You updated an A record 10 minutes ago. Some users get the new server, some the old one. Why?
Every resolver in the chain — browser, OS, router, ISP — may hold the answer for its TTL, so a 3600-second TTL means up to an hour of split traffic. That is why you lower the TTL to 60 seconds a day before a migration, and why both old and new servers must serve correctly during the overlap.