devopsdiary

diary / cheat sheet

The commands you'll forget

Everything from the twelve chapters, on one page, in the order you reach for it. Print it, bookmark it, or keep it open in a second tab during the labs.

The three ladders

If you memorise nothing else on this page, memorise these. Each one turns panic into an ordered set of cheap checks.

a request is failing (ch 03)
1 dig name +short        resolves?
2 ping / traceroute      reachable?
3 nc -zv host port       refused = nothing listening
                          timeout = firewall
4 curl -vI https://…     TLS ok?
5 curl -i  https://…     4xx client / 5xx server
6 proxy log → app log → db
   502 not answering
   504 too slow
   500 threw an error
a pod is unhappy (ch 07)
1 kubectl get pods       read STATUS
2 kubectl describe pod X read EVENTS
3 kubectl logs X --previous
4 kubectl get endpoints S
   empty = label/port mismatch
5 kubectl top pods       OOM? 137?

Pending      → not scheduled
ImagePull…   → tag / registry
CrashLoop…   → logs --previous
0/1 ready    → readiness probe
you are on call (ch 10)
1 ACKNOWLEDGE the page
2 ASSESS user impact (dashboard)
3 DECLARE + name a commander
4 MITIGATE  rollback / flag off
5 VERIFY    real request, all regions
6 DIAGNOSE  now, calmly
7 DOCUMENT  within 48h

mitigation BEFORE diagnosis,
every single time.

Linux & shell

Orientation

whoami · hostname · pwd
Who, where, which machine
ls -la
All files, long format
which cmd · echo $PATH
Which binary runs
history | tail -20
What did I just do?
cat /etc/os-release
Which distro
man cmd · cmd --help
The manual

Files & disk

tail -f -n 100 file
Follow a log
less file
Page through (q quits)
find /etc -name "*.conf" -mtime -1
Changed in 24h
du -sh /var/* | sort -h
What's eating the disk
df -h
How full each filesystem is
ls -l /proc/*/fd | grep deleted
Deleted-but-open file

Permissions

chmod +x deploy.sh
Make executable
chmod 750 file
owner rwx, group rx, other none
chmod 600 ~/.ssh/id_ed25519
Required for private keys
chown user:group file
Change ownership
sudo -u deploy cmd
Run as another user

Pipes that answer questions

awk '{print $1}' f | sort | uniq -c | sort -rn
Top values in a column
grep -rn "TODO" src/
Recursive, with line numbers
grep -c ' 500 ' access.log
Count matches
sed 's/old/new/g' f
Find and replace
cmd | tee out.txt
Print and save
cmd 2>/dev/null
Discard errors

Processes & services

ps aux | grep nginx
Find a process
kill PID · kill -9 PID
SIGTERM (polite) · SIGKILL (not)
top · htop · docker stats
Live resource use
systemctl status|restart|enable --now X
Manage a service
journalctl -u X -n 100 -f
That service's logs
free -h · nproc · uptime
Memory, CPUs, load

SSH & scripts

ssh-keygen -t ed25519
Make a key pair
ssh-copy-id user@host
Install the public key
scp file host:/tmp/
Copy one file
rsync -avz ./dir/ host:/dir/
Sync a folder
ssh -L 9090:localhost:9090 host
Tunnel a remote port
set -euo pipefail
First line of every script

Networking

DNS

dig name +short
Just the answer
dig name @8.8.8.8
Ask another resolver — is it my cache?
dig +trace name
Follow the delegation
getent hosts name
What the OS resolves
cat /etc/hosts
Local overrides

Ports & connections

ss -ltnp
Listening sockets + process
ss -tn state established
Current connections
nc -zv host 5432
Is that port open?
0.0.0.0 vs 127.0.0.1
All interfaces vs local only

HTTP & TLS

curl -i URL
Headers + body
curl -v URL
The whole conversation
curl -s -o /dev/null -w '%{http_code} %{time_total}\n' URL
Status + timing
curl --resolve host:443:IP URL
Test one server behind a name
openssl s_client -connect host:443
Certificate details

nginx & firewall

nginx -t
Test config before reloading
systemctl reload nginx
Reload with no dropped connections
ufw allow 22/tcp · ufw enable
Allow SSH before enabling
ufw status numbered
Current rules

Git

Daily

git status
Run it constantly
git add -p
Stage selected chunks
git commit -am "msg"
Stage tracked + commit
git log --oneline --graph -10
Shape of history
git diff · git diff --staged
Unstaged · staged changes

Branches

git switch -c feat/x
Create and switch
git fetch origin
Download, change nothing
git pull --rebase
Tidier than a merge pull
git rebase origin/main
Replay your commits on top
git log --oneline origin/main..HEAD
What's mine and unpushed

Undo

git restore file
Discard local edits
git restore --staged file
Unstage, keep edits
git commit --amend
Fix the last commit (unpushed)
git reset --soft HEAD~1
Undo commit, keep changes
git revert SHA
The only safe undo for pushed work
git reflog
Find "lost" commits
git stash push -m "wip" · git stash pop
Park work

Releases

git tag -a v2.3.0 -m "…" && git push origin v2.3.0
Mark a release
git describe --tags
Where am I relative to the last tag
git bisect start / bad / good
Binary-search for the bad commit

Docker & Compose

Run

docker run -d -p 8080:3000 --name api img
host:container ports
docker run -it --rm ubuntu bash
Throwaway shell
docker ps -a
Including stopped, with exit codes
docker logs -f api
Follow output
docker exec -it api sh
Shell inside a running container
docker run -it --rm --entrypoint sh img
When it exits instantly

Build

docker build -t api:v2 .
Don't forget the context dot
docker build --no-cache -t api:v2 .
Cold build, for comparison
docker history api:v2
Layer sizes — find the fat
docker images · docker system df
What's on disk
docker system prune -a
Reclaim space (careful)

Data & network

-v pgdata:/var/lib/postgresql/data
Named volume (databases)
-v "$PWD:/app"
Bind mount (local dev)
docker network create appnet
Then reach containers by name
docker inspect api
Env, mounts, network, command
docker diff api
What it wrote

Compose & registry

docker compose up -d --build
Build and start the stack
docker compose logs -f api
One service's logs
docker compose down
Stop; -v also deletes data
docker tag img ghcr.io/u/img:SHA
Immutable tag
docker push ghcr.io/u/img:SHA
Never deploy :latest

kubectl

Look

kubectl get pods -n ns -o wide
With node and IP
kubectl get pods -w
Watch state changes live
kubectl describe pod X
Events at the bottom = the answer
kubectl logs -f X · --previous
Live · the crashed instance
kubectl get endpoints svc
Empty = selector/port mismatch
kubectl get events --sort-by=.lastTimestamp
Recent cluster activity
kubectl top pods · nodes
Actual usage

Change

kubectl apply -f k8s/
Declarative, whole folder
kubectl set image deploy/api api=img:SHA
Deploy a version
kubectl rollout status deploy/api
Wait and watch
kubectl rollout undo deploy/api
The 3 a.m. command
kubectl rollout restart deploy/api
Pick up new config
kubectl scale deploy/api --replicas=5
More copies

Reach in

kubectl exec -it X -- sh
Shell in a pod
kubectl port-forward svc/api 8080:80
Test from your laptop
kubectl cp X:/path ./local
Copy a file out
kubectl config get-contexts
Check this before any change

Helm

helm upgrade --install name ./chart -f values.prod.yaml
Idempotent deploy
helm template ./chart -f values.yaml
See the YAML first
helm rollback name 3
Back to revision 3
helm list -A
Everything installed

Terraform & Ansible

Terraform loop

terraform init
Providers + backend
terraform fmt -recursive · validate
Format and check
terraform plan -out=tfplan
Read every line
terraform apply tfplan
Apply what you reviewed
terraform destroy
Your wallet's best friend
terraform plan -detailed-exitcode
Exit 2 = drift (nightly CI)

Plan symbols

+
Create
~
Update in place
-/+
Destroy and recreate — stop and read
-
Destroy

Terraform state

terraform state list · show ADDR
What's managed
terraform import ADDR ID
Adopt an existing resource
terraform state rm ADDR
Forget (does not delete)
terraform state mv A B
After renaming in code
Remote · locked · encrypted · versioned · per environment
The five rules

Ansible

ansible all -m ping
Can I reach every host?
ansible-playbook p.yml --check --diff
Dry run with diffs
ansible-playbook p.yml --limit web-01
One host first, always
changed=0 on a second run
Proof of idempotency

PromQL & alerting

the six queries that cover most days
# requests per second by route
sum(rate(http_requests_total[5m])) by (route)

# error percentage — the number that belongs in an alert
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
    / sum(rate(http_requests_total[5m]))

# p95 latency (never alert on averages)
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route))

# saturation: memory used vs limit
container_memory_working_set_bytes / container_spec_memory_limit_bytes

# is anything down?
min(up{job="notes-api"})

# restarts in the last hour — a leading indicator
increase(kube_pod_container_status_restarts_total[1h]) > 0

# rules of thumb
always wrap counters in rate()
range ≥ 4 × scrape interval (15s scrape → [5m])
every alert needs: a duration (for:), a severity, an impact line, a runbook link

Reference tables

HTTP status codes that matter

401Not authenticated
403Authenticated, not allowed
404No such thing
429Rate limited
500Your code threw
502App not answering the proxy
503Unavailable / overloaded
504App answered too slowly

Exit codes & signals

0Clean exit (in a container: the main process finished)
1Application error
126 / 127Not executable / not found
137SIGKILL — usually OOMKilled
143SIGTERM — asked to stop
SIGTERM"Finish up and exit" — handle this
SIGKILLCannot be handled; no cleanup

Availability budgets

SLOAllowed per 30 days
99%≈ 7h 12m
99.9%≈ 43m
99.95%≈ 21m 30s
99.99%≈ 4m 19s
99.999%≈ 26s

Common ports

22 / 53 / 80 / 443SSH / DNS / HTTP / HTTPS
3306 / 5432 / 6379MySQL / PostgreSQL / Redis
27017 / 9200MongoDB / Elasticsearch
9090 / 9093 / 3000Prometheus / Alertmanager / Grafana
6443 / 10250Kubernetes API / kubelet
the six rules behind all of it
  • Small batches. A change nobody can review is a change nobody can roll back.
  • Build once, promote the same artifact. Config changes; bytes do not.
  • Know the undo command before you run the do command.
  • Alert on what users feel, never on what is merely measurable.
  • Anything done by hand twice belongs in code — and in Git.
  • Mitigate first, diagnose second.

back to the chapters · go beat the arcade