TOPIC 01Getting a Linux to play with
You need a Linux you are allowed to break. Pick whichever is easiest for you — none of them cost money:
- Windows: install WSL2 — open PowerShell and run
wsl --install -d Ubuntu. You get a real Ubuntu terminal inside Windows. - macOS: the built-in terminal is Unix and 90% of this chapter works as-is. For the rest, run Ubuntu in Docker:
docker run -it --rm ubuntu bash. - Any machine: a virtual machine in VirtualBox, or a container as above. A cloud VM works too, but you do not need one yet.
Type the commands. Do not copy-paste them. Muscle memory for ls, cd, grep and tail is the difference between debugging an outage in three minutes and thirty. Every command below is safe on a throwaway machine.
TOPIC 02What a shell actually is
A shell is a program that reads what you type, finds a matching program, runs it, and shows you the output. bash and zsh are shells. The window they run inside is a terminal. Anything you type is either a program on disk or a builtin the shell handles itself.
The shell is the counter clerk at a government office. You slide a written request across the counter; the clerk finds the right department, hands over your papers, and slides the reply back. PATH is the clerk's list of which departments exist and in which corridor. "Command not found" means the clerk searched every corridor on the list and found no such department.
whoami # which user am I? hostname # which machine am I on? (ask before you break things) pwd # print working directory — where am I standing? ls -la # what is here, hidden files included cd /var/log # move; `cd -` jumps back to where you were which python3 # which file on PATH actually runs? echo $PATH # the corridors the clerk searches, in order man tail # the manual. `q` to quit. `/word` to search inside it. tail --help # faster than man when you just want the flags history | tail -20 # what did I just do? (invaluable during incidents)
Four keyboard habits that will save you hours:
- Tab completes paths and commands. Press it constantly — it also proves a path exists before you commit to typing it.
- Ctrl+R searches your command history. Type a fragment, keep pressing to cycle.
- Ctrl+C stops the running program. Ctrl+D means "no more input" and exits the shell.
- Ctrl+A / Ctrl+E jump to the start and end of the line.
TOPIC 03The filesystem tour
Linux has one tree starting at /. There are no drive letters; extra disks are mounted into the tree as folders. You only need to know a handful of destinations:
| Path | What lives there | Why you'll go there |
|---|---|---|
/etc | Configuration files, all plain text | nginx configs, ssh config, systemd units |
/var/log | Logs | First stop in almost every outage |
/var/lib | Data that programs own | Database files, docker's storage |
/home/you | Your files, your dotfiles | ~/.ssh, ~/.bashrc |
/usr/bin, /usr/local/bin | Installed programs | Where your PATH points |
/tmp | Scratch space, wiped on reboot | Safe place to make a mess |
/proc, /sys | Live kernel information as fake files | cat /proc/meminfo — the machine describing itself |
/mnt, /media | Mounted disks | Extra volumes on cloud servers |
A hospital building. /etc is the policy binder at reception, /var/log is the patient notes trolley, /var/lib is the records room, /usr/bin is the equipment cupboard, and /tmp is the whiteboard that gets wiped every night. Knowing the building means you never search room by room.
cat /etc/os-release # which distro and version is this? less /etc/ssh/sshd_config # page through a long file (q quits) head -20 access.log # first 20 lines tail -f access.log # follow live — the single most-used flag in ops cp app.conf app.conf.bak # ALWAYS back up before editing a live config mv old.log archive/ rm -i junk.txt # -i asks first. rm has no undo, ever. mkdir -p deploy/scripts # -p makes parents as needed find /etc -name "*.conf" -mtime -1 # .conf files changed in last 24h du -sh /var/log/* # what is eating the disk? df -h # how full is each filesystem?
TOPIC 04Permissions and sudo
Every file has an owner, a group, and three sets of permissions: read (r=4), write (w=2), execute (x=1). ls -l shows them as ten characters:
ls -l deploy.sh -rwxr-x--- 1 deploy ops 412 Jul 30 09:14 deploy.sh │└┬┘└┬┘└┬┘ └──┬─┘ └┬┘ │ │ │ │ │ └── group: ops │ │ │ │ └─────── owner: deploy │ │ │ └──────────────── others: no access at all │ │ └─────────────────── group: read + execute (5) │ └────────────────────── owner: read + write + execute (7) └──────────────────────── type: - file, d directory, l symlink # so this file is mode 750 chmod 750 deploy.sh # numeric form chmod +x deploy.sh # or just "make it runnable" chown deploy:ops deploy.sh # change owner and group chmod 600 ~/.ssh/id_ed25519 # private keys: owner only, or ssh refuses
Three keys to a shared flat: yours (owner), your flatmates' (group), and the building's (everyone else). 755 means you can rearrange the furniture, flatmates and visitors may walk in and look. 777 means the front door is propped open with a brick — which is why "just chmod 777 it" is a red flag in a code review, not a fix.
sudo runs one command as another user, normally root. Prefer sudo <command> over becoming root permanently: it keeps an audit trail in the logs and stops one careless rm from taking the machine with it. On a directory, x means "may enter", not "may execute" — a directory with r but no x lets you list names and open nothing.
TOPIC 05Pipes: the superpower
The Unix idea: many small programs that each do one thing, joined together with |. The output of the left becomes the input of the right. Once this clicks, you can answer questions about a server that no dashboard was built to answer.
# the top 10 IP addresses hitting your site
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10
4821 203.0.113.44
912 198.51.100.7
# how many 500s today, and on which paths?
grep ' 500 ' access.log | awk '{print $7}' | sort | uniq -c | sort -rn
# the 5 biggest files under /var, human readable
du -ah /var 2>/dev/null | sort -h | tail -5
# is anything listening on 8080?
ss -ltnp | grep 8080
# count lines / find in many files
wc -l *.log
grep -rn "TODO" src/ | head # -n prints line numbers
The pieces worth memorising: grep (filter lines), awk '{print $N}' (pull out column N), sort (order; -n numeric, -h human sizes, -r reverse), uniq -c (count duplicates — needs sorted input), wc -l (count lines), cut -d: -f1 (split on a delimiter), sed 's/old/new/g' (find and replace), xargs (turn a list into arguments), jq (the same idea for JSON).
> overwrite a file · >> append · 2> redirect errors only · &> both · 2>/dev/null throw errors away · | pipe to another program · tee out.txt print and save.
TOPIC 06Processes and signals
A running program is a process with a numeric ID (PID). Ops work is often just: find the process, see what it is doing, and ask it to stop politely.
ps aux | grep nginx # find processes by name top # live view; press M to sort by memory, q to quit htop # nicer, if installed pgrep -a node # PIDs of node processes, with their command line kill 4821 # SIGTERM: "please finish and exit" — the default kill -9 4821 # SIGKILL: "die now" — no cleanup, last resort pkill -f "python worker" # kill by matching the full command line uptime # load average: 1, 5, 15 minute windows free -h # memory. "available" is the number that matters nproc # how many CPUs (load 4.0 on 4 CPUs = fully busy) # run something that outlives your ssh session nohup ./long-job.sh > job.log 2>&1 &
SIGTERM is telling a shopkeeper "we're closing, please finish with this customer and lock up" — they save the till and switch off the lights. SIGKILL is cutting the power to the building. Both empty the shop; only one leaves the accounts in a usable state. This is exactly why applications need to handle SIGTERM properly — it is how Kubernetes and Docker ask them to stop during every single deploy.
TOPIC 07Packages and versions
Package managers install software and its dependencies. Debian/Ubuntu use apt; RHEL/Rocky/Amazon Linux use dnf/yum; Alpine (common in containers) uses apk.
sudo apt update # refresh the catalogue (not an upgrade!) sudo apt install -y nginx jq curl apt list --installed | grep nginx sudo apt remove nginx nginx -v # always confirm the version you got # RHEL family / Alpine equivalents sudo dnf install -y nginx apk add --no-cache curl # --no-cache keeps images small (ch 06)
"Install the latest" is how environments drift apart. Your laptop gets Node 22, the server has Node 18, and a subtle behaviour difference eats your afternoon. Pin versions — apt install nginx=1.24.*, node:20-alpine, terraform ~> 1.9. Chapter 06 solves this properly by shipping the whole environment together.
TOPIC 08systemd: services that stay up
nohup ./app & dies on reboot and never restarts after a crash. systemd is the supervisor that starts services at boot, restarts them when they fall over, and collects their logs. Writing a unit file is a genuinely useful skill — and it makes the Kubernetes chapter feel familiar, because Kubernetes is this same idea across many machines.
[Unit] Description=Notes API After=network.target [Service] User=deploy WorkingDirectory=/srv/notes-api ExecStart=/usr/bin/node server.js Restart=always # bring it back if it dies RestartSec=3 Environment=NODE_ENV=production EnvironmentFile=/etc/notes-api.env # secrets live here, mode 600 [Install] WantedBy=multi-user.target # start at boot
sudo systemctl daemon-reload # after editing any unit file sudo systemctl enable --now notes-api # start now + start at boot systemctl status notes-api # running? since when? last exit code? sudo systemctl restart notes-api journalctl -u notes-api -n 100 --no-pager # last 100 log lines journalctl -u notes-api -f # follow live journalctl -u notes-api --since "10 min ago" -p err # errors only
TOPIC 09Logs and disk detective work
Two failures account for a large share of real-world "the site is down" pages, and both are found in this section.
The disk is full
df -h # /var at 100%? there's your outage du -sh /var/* | sort -h # which directory is the pig? du -sh /var/log/* | sort -h | tail journalctl --disk-usage sudo journalctl --vacuum-size=200M # trim systemd logs ls -l /proc/*/fd 2>/dev/null | grep deleted | head # ↑ a deleted-but-still-open file: space returns only when # the holding process is restarted. Catches everyone once.
Something in the log knows why
sudo tail -f /var/log/nginx/error.log grep -c ' 502 ' /var/log/nginx/access.log # how bad, in numbers grep ' 502 ' access.log | tail -3 # what do the bad ones look like journalctl --since "2025-07-30 03:00" --until "03:20" # the incident window dmesg -T | tail -20 # kernel messages: OOM kills, disk errors
Logs are CCTV footage. Nobody watches it live; you scrub to the timestamp where the trouble started. That is why the two habits that matter are knowing when it started and being able to narrow by time — and why chapter 10 spends so long on making logs searchable instead of merely voluminous.
TOPIC 10SSH and keys
SSH gives you an encrypted shell on a remote machine. Use keys, never passwords: a key pair is a private key that never leaves your laptop and a public key you can hand out freely.
ssh-keygen -t ed25519 -C "you@laptop" # creates ~/.ssh/id_ed25519{,.pub}
ssh-copy-id deploy@203.0.113.10 # put the PUBLIC key on the server
ssh deploy@203.0.113.10
ssh -i ~/.ssh/work.pem ubuntu@10.0.1.5 # specific key
scp report.tar.gz web-01:/tmp/ # copy a file up
rsync -avz --delete ./site/ web-01:/srv/site/ # sync a folder (resumable)
# tunnel a private port to your laptop — priceless for dashboards
ssh -L 9090:localhost:9090 web-01 # then open localhost:9090
Host web-01
HostName 203.0.113.10
User deploy
IdentityFile ~/.ssh/id_ed25519
Host *.internal
ProxyJump bastion # reach private hosts through a bastion
Private key permissions must be 600. Never commit a key or a .pem to Git (chapter 11 explains what to do if you already did). Disable password login on real servers (PasswordAuthentication no). And if SSH is your only way to change a server, you are one chapter away from a better answer — that is what chapter 08 is for.
TOPIC 11Bash scripting that survives contact
A script is a file of commands with a shebang line. The first three lines below are the difference between a script that fails loudly and one that quietly corrupts something.
#!/usr/bin/env bash
set -euo pipefail # e: exit on error · u: error on unset var
IFS=$'\n\t' # pipefail: a failing pipe stage fails the line
# --- config, overridable from the environment ---
SRC="${1:-/srv/app}" # $1 or a default
DEST="${BACKUP_DIR:-/var/backups}"
STAMP="$(date +%Y%m%d-%H%M%S)"
ARCHIVE="$DEST/app-$STAMP.tar.gz"
log() { printf '[%s] %s\n' "$(date -Is)" "$*" >&2; }
# --- checks before actions ---
[[ -d "$SRC" ]] || { log "ERROR: $SRC does not exist"; exit 1; }
mkdir -p "$DEST"
log "archiving $SRC → $ARCHIVE"
tar -czf "$ARCHIVE" -C "$(dirname "$SRC")" "$(basename "$SRC")"
# keep the 7 most recent, delete the rest
ls -1t "$DEST"/app-*.tar.gz | tail -n +8 | xargs -r rm --
log "done: $(du -h "$ARCHIVE" | cut -f1)"
Points that matter more than the syntax:
- Always quote your variables:
"$SRC". Unquoted, a path with a space becomes two arguments and yourrmhits the wrong thing. set -euo pipefailmakes failures stop the script. Without-e, a failedtaris followed cheerfully by the deletion step.- Log to stderr (
>&2) so the script's real output can still be piped. - Check inputs first, act second. A five-line guard is cheaper than a restore.
- Run
shellcheck yourscript.sh. It catches the quoting bugs you cannot see yet, and it is the same class of tool as the linters you will wire into CI in chapter 05.
# if / test if [[ -f /etc/nginx/nginx.conf ]]; then echo present; else echo missing; fi if curl -fsS localhost:3000/health > /dev/null; then echo healthy; fi # loops for host in web-01 web-02 web-03; do ssh "$host" 'systemctl is-active nginx' || echo "$host DOWN" done # retry with backoff — the pattern every deploy script needs for i in 1 2 3 4 5; do curl -fsS localhost:3000/health && break echo "not ready, retry $i"; sleep $(( i * 2 )) done
TOPIC 12Scheduling with cron and timers
# ┌─ minute (0-59) # │ ┌─ hour (0-23) # │ │ ┌─ day of month (1-31) # │ │ │ ┌─ month (1-12) # │ │ │ │ ┌─ day of week (0-6, Sunday = 0) # * * * * * command 0 2 * * * /usr/local/bin/backup.sh >> /var/log/backup.log 2>&1 */5 * * * * /usr/local/bin/health-check.sh 0 9 * * 1 /usr/local/bin/weekly-report.sh # Mondays 09:00
Three things cron beginners get wrong, every time: cron runs with a minimal environment (use absolute paths, and set variables inside the script); output that is not redirected disappears or turns into unread mail; and the timezone is the server's, not yours. On modern systems, systemd timers do the same job with better logging (journalctl -u backup.timer) and can catch up on missed runs after a reboot.
- One tree from
/: config in/etc, logs in/var/log, data in/var/lib. - Permissions are owner/group/other × read/write/execute.
750good,777a smell,600for private keys. grep | awk | sort | uniq -c | sort -rnanswers most "what is happening?" questions.- SIGTERM asks nicely and SIGKILL does not — which is why your app must handle SIGTERM to deploy cleanly.
- Long-running things belong in a systemd unit with
Restart=always, not innohup. - Every script starts with
set -euo pipefailand quotes its variables.
LABYour first real automation
A health checker that runs itself, logs properly, and survives a reboot
- Install nginx:
sudo apt update && sudo apt install -y nginx. Confirm withcurl -I localhost— you wantHTTP/1.1 200 OK. - Write
/usr/local/bin/health-check.sh: it should curllocalhost, and append a timestamped OK or FAIL line to/var/log/health.log. Start withset -euo pipefail. chmod +xit and run it twice. Check the log withtail /var/log/health.log.- Add a cron entry to run it every minute. Wait three minutes and confirm three lines appeared.
- Break it on purpose:
sudo systemctl stop nginx. Your log should now record FAIL. This is the moment the script proves it works — a monitor that has never seen a failure is not a monitor. - Now find the failure from the other side:
systemctl status nginx, thenjournalctl -u nginx -n 20. Read what "inactive (dead)" looks like so you recognise it later. - Start nginx again and make it survive reboots:
sudo systemctl enable --now nginx. - Stretch: make the script exit non-zero on failure, add a retry loop with backoff, and count the failures in the log with
grep -c FAIL /var/log/health.log.
What you just built is the seed of everything in chapter 10: a probe, a record, and a way to tell healthy from broken without asking a human to look.
CHECKCheck yourself
df -h reports /var at 100%. You delete a 4 GB log file, but df still shows 100%. What is the most likely explanation?
On Linux, unlinking a file only removes the name. The blocks are freed when the last open file descriptor closes, so a process still writing to it keeps the space. Find it with ls -l /proc/*/fd | grep deleted or lsof | grep deleted, then restart that service. Rebooting works too — it is just the most expensive version of the same fix.
Why should a deployment stop a service with SIGTERM rather than SIGKILL?
SIGTERM is a request the program can catch: stop accepting new work, finish what's in flight, flush buffers, exit. SIGKILL is executed by the kernel with no notification — half-written files, dropped requests, locks left behind. Docker and Kubernetes both send SIGTERM first and SIGKILL after a grace period, so an app that ignores SIGTERM drops user requests on every single deploy.
Which line correctly counts how many times each HTTP status appears in access.log, where status is field 9?
uniq only collapses adjacent duplicate lines, so it must always be fed sorted input — that is the trap in the second option. The final sort -rn puts the most frequent status at the top, which is what you actually want at 3 a.m.