devopsdiary
Next chapter

diary / chapters / 03

chapter 03 · ~40 min · terminal lab

Networking without the textbook

You do not need to memorise the seven-layer OSI model. You need to be able to answer one question quickly, under pressure: where exactly is this request dying? This chapter builds that skill from the ground up.

TOPIC 01What happens when you type a URL

Everything else in this chapter is a detail of this sequence. Learn the order and you can bisect any failure.

https://shop.example.com/checkout
1. DNS      shop.example.com → 203.0.113.42        "what is the address?"
2. route    packets find their way to that address  "can I reach it?"
3. TCP      three-way handshake on port 443         "is anyone answering?"
4. TLS      certificate check, keys agreed          "is it really them?"
5. HTTP     GET /checkout, headers, cookies         "here's my request"
6. proxy    nginx picks a healthy backend           "who handles this?"
7. app      your code runs, maybe queries a DB      "do the work"
8. response status code + body travel back          "here's the answer"
real life

Posting a letter. DNS is looking the address up in a directory; routing is the postal network moving it; TCP is the recipient signing for delivery; TLS is a tamper-proof sealed envelope only they can open; HTTP is the language the letter is written in; the reverse proxy is the mailroom that decides which desk gets it. When a letter goes missing you do not check all steps at once — you check them in order.

TOPIC 02IP addresses and subnets

An IP address identifies a network interface. IPv4 looks like 10.0.3.14; IPv6 looks like 2001:db8::1. The part that actually matters day to day is the difference between public and private, and what the /24 means.

  • Public addresses are reachable from the internet. Private ranges are not routable on the internet and are reused inside every network on earth: 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16.
  • CIDR notation — the /n — says how many leading bits are fixed as the network part. The rest is available for hosts.
  • 127.0.0.1 (localhost) is the machine talking to itself. 0.0.0.0, when a server binds to it, means "listen on every interface" — which is why an app bound to 127.0.0.1 works locally but cannot be reached from outside. That single misunderstanding accounts for a huge share of "my container isn't responding" tickets.
CIDRAddressesTypically used for
/321One exact host, e.g. a firewall rule for your office IP
/24256One subnet — a tidy default for an app tier
/204,096A large subnet in a cloud VPC
/1665,536A whole VPC, e.g. 10.0.0.0/16
real life

A VPC is a gated township (10.0.0.0/16). Subnets are its lanes (10.0.1.0/24, 10.0.2.0/24). Public subnets have a gate to the main road; private subnets do not, which is exactly where you put your database. A NAT gateway is the township's shared courier: private houses can send parcels out, but no stranger can walk in.

TOPIC 03DNS: the phone book

DNS translates names to addresses. You will meet these record types constantly:

A / AAAA
Name → IPv4 / IPv6 address. api.example.com → 203.0.113.42
CNAME
Name → another name. Points your domain at a load balancer's hostname that may change IPs.
MX
Where email for this domain goes.
TXT
Free text. Used for domain verification and email anti-spoofing (SPF/DKIM).
NS
Which nameservers are authoritative for this domain.

TTL (time to live) is how many seconds resolvers may cache an answer. It is the reason DNS changes seem not to work: your browser, your OS, your router and your ISP may all be holding the old answer. Before a planned migration, drop the TTL to 60 seconds a day in advance; afterwards, raise it again.

terminal · asking DNS directly
dig api.example.com +short          # just the answer
dig api.example.com                 # full answer, with TTL
dig api.example.com @8.8.8.8        # ask a specific resolver — is it just MY cache?
dig +trace api.example.com          # follow the delegation from the root down
dig -t MX example.com +short
host api.example.com                # shorter output
getent hosts api.example.com        # what the OS resolves, /etc/hosts included
cat /etc/hosts                      # local overrides — check this before panicking

TOPIC 04Ports and sockets

An IP address gets you to the machine; a port gets you to the right program on it. A connection is identified by four things — source IP, source port, destination IP, destination port — and that tuple is a socket.

PortServicePortService
22SSH5432PostgreSQL
80HTTP3306MySQL
443HTTPS6379Redis
53DNS9090Prometheus
25 / 587SMTP3000 / 8080App defaults
real life

The IP is the building's street address; the port is the flat number. "Connection refused" means you reached the building and knocked on flat 8080, but nobody lives there. "Connection timed out" means your knock never even arrived — the gate guard (a firewall) silently dropped you. Those two errors point at completely different teams, which is why telling them apart is such a useful reflex.

terminal · who is listening
ss -ltnp                    # listening TCP sockets + owning process
State  Local Address:Port   Process
LISTEN 0.0.0.0:80           users:(("nginx",pid=812))
LISTEN 127.0.0.1:3000       users:(("node",pid=1440))   ← local only!

ss -tn state established   # current connections
nc -zv api.internal 5432    # can I even open that port? (great first test)
curl -v telnet://db-01:5432 # same idea without nc installed

TOPIC 05TCP vs UDP

TCPUDP
ConnectionHandshake first (SYN, SYN-ACK, ACK)None — just send
GuaranteesOrdered, retransmitted, no duplicatesNone; packets may vanish or arrive out of order
CostSlightly slower to startMinimal overhead
Used byHTTP, SSH, databases — anything that must be correctDNS, video calls, game state, metrics push
real life

TCP is a registered parcel with signature on delivery: slower to set up, but you know it arrived. UDP is shouting across a noisy room: if a word is lost, you carry on. For a video call that is the right trade — a re-sent frame from two seconds ago is useless. For a bank transfer it very much is not.

TOPIC 06HTTP, properly

A request is a method, a path, headers and (sometimes) a body. A response is a status code, headers and a body. That is the whole protocol.

terminal · curl is your microscope
curl -i https://api.example.com/health      # show response headers + body
curl -v https://api.example.com/health      # show the whole conversation
curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' https://example.com
200 0.184s

# POST some JSON with a token
curl -X POST https://api.example.com/orders \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"item":"tea","qty":2}'

curl -I https://example.com                 # headers only (HEAD request)
curl --resolve api.example.com:443:10.0.3.14 https://api.example.com/
# ↑ test a specific server behind a DNS name — perfect during a migration

Status codes as a triage tool

  • 2xx — worked. 200 OK, 201 created, 204 no content.
  • 3xx — go elsewhere. 301 permanent, 302 temporary, 304 not modified (your cache is fine).
  • 4xx — the client's fault. 400 malformed, 401 not authenticated, 403 authenticated but not allowed, 404 no such thing, 422 validation failed, 429 rate limited.
  • 5xx — the server's fault, and yours to fix. 500 unhandled error in your code, 502 bad gateway (the proxy got nothing usable from the app), 503 unavailable/overloaded, 504 gateway timeout (the app was too slow).
the 502 / 504 distinction is worth money

502 = your app is not answering at all: process dead, wrong port, crashed on boot. 504 = your app answered too slowly: a hung database query, an external API with no timeout, a thread pool exhausted. Same red page for the user, completely different investigation. Ninety seconds of reading the proxy's error log tells you which one you have.

TOPIC 07TLS and certificates

HTTPS is HTTP inside TLS. TLS gives you three things: encryption (nobody can read it), integrity (nobody can change it), and identity (you are talking to who you think). Identity is the part that breaks in production.

A certificate says "this public key belongs to shop.example.com", signed by a Certificate Authority your operating system already trusts. Browsers reject it when it has expired, when the name does not match, or when the signing chain is incomplete.

terminal · certificate checks
# when does it expire? (put this in a monitor — expiry outages are pure own-goals)
echo | openssl s_client -connect example.com:443 2>/dev/null \
  | openssl x509 -noout -dates -subject
notBefore=Jun  2 00:00:00 2026 GMT
notAfter=Aug 31 23:59:59 2026 GMT

curl -vI https://example.com 2>&1 | grep -E 'subject|issuer|expire'
openssl s_client -connect example.com:443 -servername example.com
# -servername matters: one IP can serve many certs (SNI)

In practice you will use Let's Encrypt with certbot, or terminate TLS at a cloud load balancer / ingress controller and let it handle renewal. Either way: certificates are renewed automatically, and you alert on "expires in under 14 days", because a human calendar reminder is not a control.

TOPIC 08Reverse proxies and load balancers

A reverse proxy sits in front of your application and speaks to clients on its behalf. It is where you put the things that should not be your app's problem: TLS, compression, caching, rate limits, redirects, and routing multiple apps onto one domain. A load balancer is a reverse proxy that also spreads requests across several instances and stops sending traffic to unhealthy ones.

/etc/nginx/sites-available/notes-api
upstream notes_api {
    server 127.0.0.1:3000;      # add more lines to load balance
    server 127.0.0.1:3001;
}

server {
    listen 80;
    server_name api.example.com;

    location /health { access_log off; proxy_pass http://notes_api; }

    location / {
        proxy_pass http://notes_api;
        proxy_set_header Host              $host;
        proxy_set_header X-Real-IP         $remote_addr;
        proxy_set_header X-Forwarded-For   $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
        proxy_connect_timeout 3s;
        proxy_read_timeout   30s;   # the difference between 502 and 504
    }
}
terminal · never reload a config you haven't tested
sudo nginx -t                     # syntax check FIRST
nginx: configuration file /etc/nginx/nginx.conf test is successful
sudo systemctl reload nginx       # reload = no dropped connections
sudo tail -f /var/log/nginx/error.log

Balancing strategies you should be able to name: round robin (in turn), least connections (to the least busy — better for uneven request times), IP hash (same client to the same server, for sticky sessions). And the feature that matters more than any of them: health checks, so a dead instance is removed from rotation automatically. That is what makes zero-downtime deploys possible in chapter 09.

TOPIC 09Firewalls and security groups

A firewall decides which traffic is allowed. On a single Linux host that is ufw or nftables; in a cloud it is a security group attached to the instance. The model to internalise: default deny, then allow the minimum.

terminal · ufw basics
sudo ufw default deny incoming
sudo ufw default allow outgoing
sudo ufw allow 22/tcp              # do this BEFORE enabling, or lock yourself out
sudo ufw allow 443/tcp
sudo ufw allow from 10.0.1.0/24 to any port 5432  # db: app subnet only
sudo ufw enable
sudo ufw status numbered
the pattern for every tier

Load balancer: 80/443 from anywhere. App servers: 3000 from the load balancer only. Database: 5432 from the app subnet only. SSH: from a bastion or your VPN, never from 0.0.0.0/0. If a database is reachable from the internet, that is not a hardening task for later — it is the incident that has not happened yet.

TOPIC 10The debugging ladder

Memorise this order. It turns "the site is down" from panic into six checks, and each rung eliminates a whole class of cause.

climb from the bottom, stop where it breaks
1. does the name resolve?   dig api.example.com +short
   ✗ → DNS: wrong record, expired domain, /etc/hosts override, TTL cache

2. can I reach the host?    ping 203.0.113.42   (ICMP may be blocked — not proof)
   ✗ → routing, VPN, wrong subnet, instance stopped

3. is the port open?        nc -zv api.example.com 443
   refused → nothing listening: app crashed, wrong port, bound to 127.0.0.1
   timeout → something is dropping packets: firewall / security group

4. does TLS complete?       curl -vI https://api.example.com
   ✗ → expired cert, name mismatch, missing chain, SNI

5. what does HTTP say?      curl -i https://api.example.com/health
   4xx → auth, path, or client problem      5xx → your app or its proxy

6. what do the logs say?    proxy log first, then app log, then DB
   502 → app not answering    504 → app too slow    500 → read the stack trace
real life

A tap with no water. You do not start by dismantling the tap — you check the street supply, then the building tank, then the pipe to your flat, then the tap. Same ladder, same reason: each step is cheap and rules out everything below it. Engineers who look slow and methodical during an incident are usually the ones who finish first.

remember this much
  • Request order: DNS → route → TCP → TLS → HTTP → proxy → app. Debug in that order.
  • Refused = nothing listening. Timeout = something is dropping the packets (firewall).
  • Binding to 127.0.0.1 makes a service unreachable from outside; 0.0.0.0 listens everywhere.
  • 502 = app not answering; 504 = app too slow; 500 = app threw an error.
  • DNS changes are delayed by TTL caching — lower the TTL before a migration, not during one.
  • Firewalls: default deny, then allow the minimum per tier. Databases never face the internet.

LABPut nginx in front of an app

40 minutes · one Linux box (or WSL)

Build the shape that runs most of the web, then break it in four ways

  1. Run something on port 3000. Anything: python3 -m http.server 3000 is enough. Confirm with curl -i localhost:3000.
  2. Install nginx and add a server block that proxies / to http://127.0.0.1:3000, with the four proxy_set_header lines from above.
  3. sudo nginx -t, then sudo systemctl reload nginx. Now curl -i localhost should return your app through port 80. You have just built a reverse proxy.
  4. Break 1 — 502: stop the app on 3000 and curl again. Read the exact line nginx writes in error.log. Learn to recognise "connect() failed (111: Connection refused)".
  5. Break 2 — 504: set proxy_read_timeout 2s, and point it at something slow (nc -l 3000 and never reply). Watch a gateway timeout appear after exactly two seconds.
  6. Break 3 — the bind trap: start your app bound to 127.0.0.1 only, and try to reach it from another machine or from your host OS. Then rebind to 0.0.0.0 and try again. This is the single most common container networking bug.
  7. Break 4 — the firewall: sudo ufw allow 22/tcp, then sudo ufw enable. Try to reach port 80 from elsewhere and note that it now times out rather than being refused. Add sudo ufw allow 80/tcp and watch it recover.
  8. Stretch: add a second app on 3001, list both in the upstream block, reload, and run curl -s localhost/whoami in a loop to see requests alternate. You have now built a load balancer.

Why this lab matters: steps 4–6 are the three failures you will actually meet, and having produced each one on purpose is what lets you recognise them in seconds later.

CHECKCheck yourself

curl to your API hangs for 30 seconds and then reports a timeout. A colleague's curl to the same URL is refused instantly. What does that difference tell you?

A silent drop produces a timeout; an active rejection produces "connection refused" immediately. So the two of you are hitting different obstacles: you are blocked before the host (security group, ufw, VPN), and your colleague is getting through to a host where the process is not listening. Two symptoms, two different fixes — and if DNS were the problem you would see a resolution error instead.

Your app runs fine with curl localhost:3000 on the server, but nginx in front of it returns 502. Which cause fits best?

502 means nginx could not get a usable response from the upstream, so the fault is between proxy and app. Check the exact upstream address in the config against ss -ltnp output, then read error.log — it names the failing address and the OS error. Certificate and DNS problems fail earlier, before HTTP ever happens.

You updated an A record 10 minutes ago. Some users get the new server, some the old one. Why?

Every resolver in the chain — browser, OS, router, ISP — may hold the answer for its TTL, so a 3600-second TTL means up to an hour of split traffic. That is why you lower the TTL to 60 seconds a day before a migration, and why both old and new servers must serve correctly during the overlap.

saved in this browser only — no account needed