devopsdiary
Next chapter

diary / chapters / 02

chapter 02 · ~45 min · terminal lab

Linux and the shell you live in

Almost every server you will ever be responsible for runs Linux, and almost every DevOps tool is a command-line program. This chapter is the one that turns the terminal from a scary black window into the fastest room in your house.

TOPIC 01Getting a Linux to play with

You need a Linux you are allowed to break. Pick whichever is easiest for you — none of them cost money:

  • Windows: install WSL2 — open PowerShell and run wsl --install -d Ubuntu. You get a real Ubuntu terminal inside Windows.
  • macOS: the built-in terminal is Unix and 90% of this chapter works as-is. For the rest, run Ubuntu in Docker: docker run -it --rm ubuntu bash.
  • Any machine: a virtual machine in VirtualBox, or a container as above. A cloud VM works too, but you do not need one yet.
habit to build now

Type the commands. Do not copy-paste them. Muscle memory for ls, cd, grep and tail is the difference between debugging an outage in three minutes and thirty. Every command below is safe on a throwaway machine.

TOPIC 02What a shell actually is

A shell is a program that reads what you type, finds a matching program, runs it, and shows you the output. bash and zsh are shells. The window they run inside is a terminal. Anything you type is either a program on disk or a builtin the shell handles itself.

real life

The shell is the counter clerk at a government office. You slide a written request across the counter; the clerk finds the right department, hands over your papers, and slides the reply back. PATH is the clerk's list of which departments exist and in which corridor. "Command not found" means the clerk searched every corridor on the list and found no such department.

terminal · orientation
whoami          # which user am I?
hostname        # which machine am I on? (ask before you break things)
pwd             # print working directory — where am I standing?
ls -la          # what is here, hidden files included
cd /var/log     # move; `cd -` jumps back to where you were
which python3   # which file on PATH actually runs?
echo $PATH      # the corridors the clerk searches, in order
man tail        # the manual. `q` to quit. `/word` to search inside it.
tail --help     # faster than man when you just want the flags
history | tail -20 # what did I just do? (invaluable during incidents)

Four keyboard habits that will save you hours:

  • Tab completes paths and commands. Press it constantly — it also proves a path exists before you commit to typing it.
  • Ctrl+R searches your command history. Type a fragment, keep pressing to cycle.
  • Ctrl+C stops the running program. Ctrl+D means "no more input" and exits the shell.
  • Ctrl+A / Ctrl+E jump to the start and end of the line.

TOPIC 03The filesystem tour

Linux has one tree starting at /. There are no drive letters; extra disks are mounted into the tree as folders. You only need to know a handful of destinations:

PathWhat lives thereWhy you'll go there
/etcConfiguration files, all plain textnginx configs, ssh config, systemd units
/var/logLogsFirst stop in almost every outage
/var/libData that programs ownDatabase files, docker's storage
/home/youYour files, your dotfiles~/.ssh, ~/.bashrc
/usr/bin, /usr/local/binInstalled programsWhere your PATH points
/tmpScratch space, wiped on rebootSafe place to make a mess
/proc, /sysLive kernel information as fake filescat /proc/meminfo — the machine describing itself
/mnt, /mediaMounted disksExtra volumes on cloud servers
real life

A hospital building. /etc is the policy binder at reception, /var/log is the patient notes trolley, /var/lib is the records room, /usr/bin is the equipment cupboard, and /tmp is the whiteboard that gets wiped every night. Knowing the building means you never search room by room.

terminal · files
cat /etc/os-release      # which distro and version is this?
less /etc/ssh/sshd_config # page through a long file (q quits)
head -20 access.log      # first 20 lines
tail -f access.log       # follow live — the single most-used flag in ops
cp app.conf app.conf.bak # ALWAYS back up before editing a live config
mv old.log archive/
rm -i junk.txt           # -i asks first. rm has no undo, ever.
mkdir -p deploy/scripts  # -p makes parents as needed
find /etc -name "*.conf" -mtime -1  # .conf files changed in last 24h
du -sh /var/log/*        # what is eating the disk?
df -h                    # how full is each filesystem?

TOPIC 04Permissions and sudo

Every file has an owner, a group, and three sets of permissions: read (r=4), write (w=2), execute (x=1). ls -l shows them as ten characters:

terminal · reading permissions
ls -l deploy.sh
-rwxr-x---  1 deploy  ops   412 Jul 30 09:14 deploy.sh
 │└┬┘└┬┘└┬┘     └──┬─┘ └┬┘
 │ │  │  │        │    └── group: ops
 │ │  │  │        └─────── owner: deploy
 │ │  │  └──────────────── others: no access at all
 │ │  └─────────────────── group:  read + execute (5)
 │ └────────────────────── owner:  read + write + execute (7)
 └──────────────────────── type:   - file, d directory, l symlink

# so this file is mode 750
chmod 750 deploy.sh       # numeric form
chmod +x deploy.sh        # or just "make it runnable"
chown deploy:ops deploy.sh # change owner and group
chmod 600 ~/.ssh/id_ed25519 # private keys: owner only, or ssh refuses
real life

Three keys to a shared flat: yours (owner), your flatmates' (group), and the building's (everyone else). 755 means you can rearrange the furniture, flatmates and visitors may walk in and look. 777 means the front door is propped open with a brick — which is why "just chmod 777 it" is a red flag in a code review, not a fix.

sudo runs one command as another user, normally root. Prefer sudo <command> over becoming root permanently: it keeps an audit trail in the logs and stops one careless rm from taking the machine with it. On a directory, x means "may enter", not "may execute" — a directory with r but no x lets you list names and open nothing.

TOPIC 05Pipes: the superpower

The Unix idea: many small programs that each do one thing, joined together with |. The output of the left becomes the input of the right. Once this clicks, you can answer questions about a server that no dashboard was built to answer.

terminal · pipes in anger
# the top 10 IP addresses hitting your site
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10
   4821 203.0.113.44
    912 198.51.100.7

# how many 500s today, and on which paths?
grep ' 500 ' access.log | awk '{print $7}' | sort | uniq -c | sort -rn

# the 5 biggest files under /var, human readable
du -ah /var 2>/dev/null | sort -h | tail -5

# is anything listening on 8080?
ss -ltnp | grep 8080

# count lines / find in many files
wc -l *.log
grep -rn "TODO" src/ | head       # -n prints line numbers

The pieces worth memorising: grep (filter lines), awk '{print $N}' (pull out column N), sort (order; -n numeric, -h human sizes, -r reverse), uniq -c (count duplicates — needs sorted input), wc -l (count lines), cut -d: -f1 (split on a delimiter), sed 's/old/new/g' (find and replace), xargs (turn a list into arguments), jq (the same idea for JSON).

redirection, in one line each

> overwrite a file · >> append · 2> redirect errors only · &> both · 2>/dev/null throw errors away · | pipe to another program · tee out.txt print and save.

TOPIC 06Processes and signals

A running program is a process with a numeric ID (PID). Ops work is often just: find the process, see what it is doing, and ask it to stop politely.

terminal · processes
ps aux | grep nginx     # find processes by name
top                     # live view; press M to sort by memory, q to quit
htop                    # nicer, if installed
pgrep -a node           # PIDs of node processes, with their command line

kill 4821               # SIGTERM: "please finish and exit" — the default
kill -9 4821            # SIGKILL: "die now" — no cleanup, last resort
pkill -f "python worker" # kill by matching the full command line

uptime                  # load average: 1, 5, 15 minute windows
free -h                 # memory. "available" is the number that matters
nproc                   # how many CPUs (load 4.0 on 4 CPUs = fully busy)

# run something that outlives your ssh session
nohup ./long-job.sh > job.log 2>&1 &
real life

SIGTERM is telling a shopkeeper "we're closing, please finish with this customer and lock up" — they save the till and switch off the lights. SIGKILL is cutting the power to the building. Both empty the shop; only one leaves the accounts in a usable state. This is exactly why applications need to handle SIGTERM properly — it is how Kubernetes and Docker ask them to stop during every single deploy.

TOPIC 07Packages and versions

Package managers install software and its dependencies. Debian/Ubuntu use apt; RHEL/Rocky/Amazon Linux use dnf/yum; Alpine (common in containers) uses apk.

terminal · packages
sudo apt update                 # refresh the catalogue (not an upgrade!)
sudo apt install -y nginx jq curl
apt list --installed | grep nginx
sudo apt remove nginx
nginx -v                        # always confirm the version you got

# RHEL family / Alpine equivalents
sudo dnf install -y nginx
apk add --no-cache curl         # --no-cache keeps images small (ch 06)
the version trap

"Install the latest" is how environments drift apart. Your laptop gets Node 22, the server has Node 18, and a subtle behaviour difference eats your afternoon. Pin versions — apt install nginx=1.24.*, node:20-alpine, terraform ~> 1.9. Chapter 06 solves this properly by shipping the whole environment together.

TOPIC 08systemd: services that stay up

nohup ./app & dies on reboot and never restarts after a crash. systemd is the supervisor that starts services at boot, restarts them when they fall over, and collects their logs. Writing a unit file is a genuinely useful skill — and it makes the Kubernetes chapter feel familiar, because Kubernetes is this same idea across many machines.

/etc/systemd/system/notes-api.service
[Unit]
Description=Notes API
After=network.target

[Service]
User=deploy
WorkingDirectory=/srv/notes-api
ExecStart=/usr/bin/node server.js
Restart=always              # bring it back if it dies
RestartSec=3
Environment=NODE_ENV=production
EnvironmentFile=/etc/notes-api.env   # secrets live here, mode 600

[Install]
WantedBy=multi-user.target  # start at boot
terminal · managing services
sudo systemctl daemon-reload          # after editing any unit file
sudo systemctl enable --now notes-api # start now + start at boot
systemctl status notes-api            # running? since when? last exit code?
sudo systemctl restart notes-api
journalctl -u notes-api -n 100 --no-pager  # last 100 log lines
journalctl -u notes-api -f            # follow live
journalctl -u notes-api --since "10 min ago" -p err  # errors only

TOPIC 09Logs and disk detective work

Two failures account for a large share of real-world "the site is down" pages, and both are found in this section.

The disk is full

terminal · the classic 3 a.m. investigation
df -h                       # /var at 100%? there's your outage
du -sh /var/* | sort -h     # which directory is the pig?
du -sh /var/log/* | sort -h | tail
journalctl --disk-usage
sudo journalctl --vacuum-size=200M   # trim systemd logs
ls -l /proc/*/fd 2>/dev/null | grep deleted | head
# ↑ a deleted-but-still-open file: space returns only when
#   the holding process is restarted. Catches everyone once.

Something in the log knows why

terminal · narrowing down
sudo tail -f /var/log/nginx/error.log
grep -c ' 502 ' /var/log/nginx/access.log   # how bad, in numbers
grep ' 502 ' access.log | tail -3            # what do the bad ones look like
journalctl --since "2025-07-30 03:00" --until "03:20"  # the incident window
dmesg -T | tail -20   # kernel messages: OOM kills, disk errors
real life

Logs are CCTV footage. Nobody watches it live; you scrub to the timestamp where the trouble started. That is why the two habits that matter are knowing when it started and being able to narrow by time — and why chapter 10 spends so long on making logs searchable instead of merely voluminous.

TOPIC 10SSH and keys

SSH gives you an encrypted shell on a remote machine. Use keys, never passwords: a key pair is a private key that never leaves your laptop and a public key you can hand out freely.

terminal · ssh
ssh-keygen -t ed25519 -C "you@laptop"   # creates ~/.ssh/id_ed25519{,.pub}
ssh-copy-id deploy@203.0.113.10        # put the PUBLIC key on the server
ssh deploy@203.0.113.10
ssh -i ~/.ssh/work.pem ubuntu@10.0.1.5 # specific key

scp report.tar.gz web-01:/tmp/         # copy a file up
rsync -avz --delete ./site/ web-01:/srv/site/  # sync a folder (resumable)

# tunnel a private port to your laptop — priceless for dashboards
ssh -L 9090:localhost:9090 web-01      # then open localhost:9090
~/.ssh/config — stop typing IP addresses
Host web-01
    HostName 203.0.113.10
    User deploy
    IdentityFile ~/.ssh/id_ed25519

Host *.internal
    ProxyJump bastion       # reach private hosts through a bastion
rules that are not optional

Private key permissions must be 600. Never commit a key or a .pem to Git (chapter 11 explains what to do if you already did). Disable password login on real servers (PasswordAuthentication no). And if SSH is your only way to change a server, you are one chapter away from a better answer — that is what chapter 08 is for.

TOPIC 11Bash scripting that survives contact

A script is a file of commands with a shebang line. The first three lines below are the difference between a script that fails loudly and one that quietly corrupts something.

backup.sh
#!/usr/bin/env bash
set -euo pipefail          # e: exit on error · u: error on unset var
IFS=$'\n\t'                # pipefail: a failing pipe stage fails the line

# --- config, overridable from the environment ---
SRC="${1:-/srv/app}"                    # $1 or a default
DEST="${BACKUP_DIR:-/var/backups}"
STAMP="$(date +%Y%m%d-%H%M%S)"
ARCHIVE="$DEST/app-$STAMP.tar.gz"

log() { printf '[%s] %s\n' "$(date -Is)" "$*" >&2; }

# --- checks before actions ---
[[ -d "$SRC" ]]  || { log "ERROR: $SRC does not exist"; exit 1; }
mkdir -p "$DEST"

log "archiving $SRC → $ARCHIVE"
tar -czf "$ARCHIVE" -C "$(dirname "$SRC")" "$(basename "$SRC")"

# keep the 7 most recent, delete the rest
ls -1t "$DEST"/app-*.tar.gz | tail -n +8 | xargs -r rm --

log "done: $(du -h "$ARCHIVE" | cut -f1)"

Points that matter more than the syntax:

  • Always quote your variables: "$SRC". Unquoted, a path with a space becomes two arguments and your rm hits the wrong thing.
  • set -euo pipefail makes failures stop the script. Without -e, a failed tar is followed cheerfully by the deletion step.
  • Log to stderr (>&2) so the script's real output can still be piped.
  • Check inputs first, act second. A five-line guard is cheaper than a restore.
  • Run shellcheck yourscript.sh. It catches the quoting bugs you cannot see yet, and it is the same class of tool as the linters you will wire into CI in chapter 05.
terminal · the control flow you need
# if / test
if [[ -f /etc/nginx/nginx.conf ]]; then echo present; else echo missing; fi
if curl -fsS localhost:3000/health > /dev/null; then echo healthy; fi

# loops
for host in web-01 web-02 web-03; do
  ssh "$host" 'systemctl is-active nginx' || echo "$host DOWN"
done

# retry with backoff — the pattern every deploy script needs
for i in 1 2 3 4 5; do
  curl -fsS localhost:3000/health && break
  echo "not ready, retry $i"; sleep $(( i * 2 ))
done

TOPIC 12Scheduling with cron and timers

crontab -e
# ┌─ minute (0-59)
# │ ┌─ hour (0-23)
# │ │ ┌─ day of month (1-31)
# │ │ │ ┌─ month (1-12)
# │ │ │ │ ┌─ day of week (0-6, Sunday = 0)
# * * * * *  command

0 2 * * *      /usr/local/bin/backup.sh >> /var/log/backup.log 2>&1
*/5 * * * *    /usr/local/bin/health-check.sh
0 9 * * 1      /usr/local/bin/weekly-report.sh   # Mondays 09:00

Three things cron beginners get wrong, every time: cron runs with a minimal environment (use absolute paths, and set variables inside the script); output that is not redirected disappears or turns into unread mail; and the timezone is the server's, not yours. On modern systems, systemd timers do the same job with better logging (journalctl -u backup.timer) and can catch up on missed runs after a reboot.

remember this much
  • One tree from /: config in /etc, logs in /var/log, data in /var/lib.
  • Permissions are owner/group/other × read/write/execute. 750 good, 777 a smell, 600 for private keys.
  • grep | awk | sort | uniq -c | sort -rn answers most "what is happening?" questions.
  • SIGTERM asks nicely and SIGKILL does not — which is why your app must handle SIGTERM to deploy cleanly.
  • Long-running things belong in a systemd unit with Restart=always, not in nohup.
  • Every script starts with set -euo pipefail and quotes its variables.

LABYour first real automation

45 minutes · on a throwaway VM, WSL or container

A health checker that runs itself, logs properly, and survives a reboot

  1. Install nginx: sudo apt update && sudo apt install -y nginx. Confirm with curl -I localhost — you want HTTP/1.1 200 OK.
  2. Write /usr/local/bin/health-check.sh: it should curl localhost, and append a timestamped OK or FAIL line to /var/log/health.log. Start with set -euo pipefail.
  3. chmod +x it and run it twice. Check the log with tail /var/log/health.log.
  4. Add a cron entry to run it every minute. Wait three minutes and confirm three lines appeared.
  5. Break it on purpose: sudo systemctl stop nginx. Your log should now record FAIL. This is the moment the script proves it works — a monitor that has never seen a failure is not a monitor.
  6. Now find the failure from the other side: systemctl status nginx, then journalctl -u nginx -n 20. Read what "inactive (dead)" looks like so you recognise it later.
  7. Start nginx again and make it survive reboots: sudo systemctl enable --now nginx.
  8. Stretch: make the script exit non-zero on failure, add a retry loop with backoff, and count the failures in the log with grep -c FAIL /var/log/health.log.

What you just built is the seed of everything in chapter 10: a probe, a record, and a way to tell healthy from broken without asking a human to look.

CHECKCheck yourself

df -h reports /var at 100%. You delete a 4 GB log file, but df still shows 100%. What is the most likely explanation?

On Linux, unlinking a file only removes the name. The blocks are freed when the last open file descriptor closes, so a process still writing to it keeps the space. Find it with ls -l /proc/*/fd | grep deleted or lsof | grep deleted, then restart that service. Rebooting works too — it is just the most expensive version of the same fix.

Why should a deployment stop a service with SIGTERM rather than SIGKILL?

SIGTERM is a request the program can catch: stop accepting new work, finish what's in flight, flush buffers, exit. SIGKILL is executed by the kernel with no notification — half-written files, dropped requests, locks left behind. Docker and Kubernetes both send SIGTERM first and SIGKILL after a grace period, so an app that ignores SIGTERM drops user requests on every single deploy.

Which line correctly counts how many times each HTTP status appears in access.log, where status is field 9?

uniq only collapses adjacent duplicate lines, so it must always be fed sorted input — that is the trap in the second option. The final sort -rn puts the most frequent status at the top, which is what you actually want at 3 a.m.

saved in this browser only — no account needed