TOPIC 01The problem with clicking
You create a server in a web console: pick a size, a network, a disk, three security rules, tick two boxes. It works. Six months later:
- Nobody knows why one of those boxes is ticked, so nobody dares untick it.
- The staging server was built by someone else on a different afternoon and behaves differently, which is why the bug "only happens in staging".
- The person who built it has left, and the reasoning left with them.
- Recreating it in another region means repeating forty clicks from memory.
- There is no record of who changed the firewall rule last Tuesday, or why.
Infrastructure as code means the servers, networks, databases, DNS records and permissions are described in text files, kept in Git, reviewed in pull requests, and applied by a tool. Infrastructure gains everything code already has: history, review, diffs, rollback and repeatability.
Two ways to build the same kitchen. One: instruct the carpenter verbally, day by day, adjusting as you go — beautiful, and utterly unrepeatable. Two: a measured drawing. The drawing can be checked before anything is cut, handed to a different carpenter in another city and produce the same kitchen, and when you want the shelf moved you change the drawing rather than arguing about what was agreed in March.
TOPIC 02Declarative vs imperative
| Imperative (a script) | Declarative (Terraform) | |
|---|---|---|
| You write | The steps to take | The end state you want |
| Example | "create a server, then attach a disk" | "a server with this disk exists" |
| Run it twice | Two servers (unless you wrote guards) | No change — it already matches |
| Partial failure | Unclear state; rerunning may double up | Rerun; it completes what is missing |
Declarative tools compare desired state with actual state and work out the difference themselves. That property — running it again is always safe — is idempotency, and it is the same reconciliation idea as the Kubernetes thermostat in chapter 07. Once you spot the pattern you will see it everywhere in this field.
TOPIC 03Terraform in five files
Terraform (and its fork OpenTofu) creates and manages infrastructure across providers — AWS, Azure, GCP, Cloudflare, GitHub, even Kubernetes. The language is HCL. Here is a complete, realistic little stack:
terraform {
required_version = "~> 1.9"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.60" # allow patches, not surprises
}
}
backend "s3" { # shared state, see topic 05
bucket = "acme-tfstate"
key = "prod/web/terraform.tfstate"
region = "ap-south-1"
use_lockfile = true # prevents two people applying at once
}
}
provider "aws" {
region = var.region
default_tags {
tags = { # tag everything: this is how you
Environment = var.environment # answer "what is this and whose bill?"
ManagedBy = "terraform"
Repo = "acme/infra"
}
}
}
# look up the newest Ubuntu image instead of hard-coding an AMI id
data "aws_ami" "ubuntu" {
most_recent = true
owners = ["099720109477"]
filter {
name = "name"
values = ["ubuntu/images/hvm-ssd-gp3/ubuntu-noble-24.04-amd64-server-*"]
}
}
resource "aws_security_group" "web" {
name = "${var.environment}-web"
description = "HTTP/HTTPS in, everything out"
vpc_id = var.vpc_id
ingress {
description = "HTTPS from the internet"
from_port = 443
to_port = 443
protocol = "tcp"
cidr_blocks = ["0.0.0.0/0"]
}
ingress {
description = "SSH from the office VPN only"
from_port = 22
to_port = 22
protocol = "tcp"
cidr_blocks = [var.admin_cidr] # never 0.0.0.0/0 (chapter 03)
}
egress {
from_port = 0
to_port = 0
protocol = "-1"
cidr_blocks = ["0.0.0.0/0"]
}
}
resource "aws_instance" "web" {
count = var.instance_count
ami = data.aws_ami.ubuntu.id
instance_type = var.instance_type
subnet_id = var.subnet_ids[count.index % length(var.subnet_ids)]
vpc_security_group_ids = [aws_security_group.web.id] # ← the dependency
root_block_device {
volume_size = 20
encrypted = true
}
tags = { Name = "${var.environment}-web-${count.index + 1}" }
}
Notice aws_security_group.web.id inside the instance. That reference is how Terraform builds a dependency graph and works out the order by itself — you never write "create the security group first". Four block types cover almost everything: resource (something to create), data (something to look up), variable (input), output (result).
TOPIC 04plan, apply, destroy
terraform init # download providers, connect the backend terraform fmt -recursive # canonical formatting (put this in CI) terraform validate # syntax and type checking terraform plan -out=tfplan # the dry run. READ IT. every line. Plan: 3 to add, 1 to change, 0 to destroy. terraform apply tfplan # apply exactly the plan you reviewed terraform show # what exists now, per state terraform output alb_dns # a value another system needs terraform destroy # tear it all down (your wallet's best friend)
plan is the whole safety mechanism. Learn to read the symbols: + create, ~ update in place, -/+ destroy and recreate, - destroy. The dangerous one is -/+: changing an attribute that cannot be modified in place quietly means replacing a live resource. On a database, that is data loss written in a diff you skimmed. If a plan proposes destroying something you did not expect, stop and find out why before typing yes.
TOPIC 05State: the crown jewels
Terraform keeps a state file mapping your code to the real resource IDs it created — "the instance named web-1 is i-0abc123". Without it, Terraform has no idea what it owns, and a plan would propose creating everything again.
State is the property register. The code is your plan for the neighbourhood; the register records which house on the ground corresponds to which plot on the plan. Lose the register and the government cannot tell your house from anyone else's — it is not that your house vanished, it is that ownership can no longer be proven. Terraform behaves the same way: no state, no idea, so it builds a second house.
Rules that separate professionals from people about to have a bad week:
- Never keep state only on a laptop, and never commit it to Git. Use a remote backend (S3, Azure Blob, GCS, Terraform Cloud).
- Always enable locking. Two engineers applying at once with no lock is how state gets corrupted.
- Treat state as a secret. It contains resource attributes, and sometimes generated passwords, in plain text. Encrypt the bucket, restrict access.
- Version the bucket so you can recover a previous state file.
- Separate state per environment. One state file for prod and staging together means a staging mistake can plan a production change.
terraform state list # everything Terraform manages terraform state show aws_instance.web[0] terraform import aws_instance.web i-0abc123 # adopt an existing resource terraform state rm aws_instance.old # forget it (does NOT delete it) terraform state mv aws_instance.a aws_instance.b # after renaming in code terraform force-unlock LOCK_ID # only when a lock is genuinely stale
TOPIC 06Variables and outputs
variable "environment" {
description = "Deployment environment name"
type = string
validation {
condition = contains(["dev", "staging", "prod"], var.environment)
error_message = "environment must be dev, staging or prod."
}
}
variable "instance_count" {
description = "How many web servers"
type = number
default = 2
}
variable "admin_cidr" {
description = "CIDR allowed to SSH"
type = string
default = "10.8.0.0/24"
}
variable "db_password" {
description = "Database password — supplied from a secret store, never a file"
type = string
sensitive = true # keeps it out of plan/apply output
}
output "web_private_ips" {
value = aws_instance.web[*].private_ip
description = "Private IPs of the web servers"
}
# ---- supplying values, in order of precedence ----
# 1. -var on the command line
# 2. -var-file=prod.tfvars
# 3. terraform.tfvars / *.auto.tfvars
# 4. TF_VAR_environment=prod (how CI usually does it)
# 5. the default in the variable block
sensitive = true is worth a habit: it stops a password appearing in your terminal, in CI logs, and in the pull-request comment where a plan was posted. It does not encrypt it in state — that is what the backend encryption is for.
TOPIC 07Modules and environments
A module is a folder of Terraform you can call with different inputs — a function for infrastructure. It is how you stop copy-pasting a VPC definition into three environments and then patching two of them.
infra/
├── modules/
│ ├── network/ # vpc, subnets, routing
│ ├── web-service/ # asg or instances + sg + load balancer
│ └── database/
└── environments/
├── staging/
│ ├── main.tf # calls the modules with small numbers
│ └── terraform.tfvars
└── prod/
├── main.tf # same modules, bigger numbers, own state
└── terraform.tfvars
module "network" {
source = "../../modules/network"
environment = "prod"
cidr_block = "10.0.0.0/16"
}
module "web" {
source = "../../modules/web-service"
environment = "prod"
vpc_id = module.network.vpc_id
subnet_ids = module.network.private_subnet_ids
instance_count = 6 # staging passes 1
instance_type = "t3.medium" # staging passes t3.micro
}
The win: staging and production are the same code with different numbers, so a test in staging genuinely tells you something. The moment they are two separate copies, they start drifting and staging stops being a rehearsal.
TOPIC 08Drift and the manual fix
Drift is when reality no longer matches the code — usually because someone fixed something by hand during an incident. It is not a moral failing; at 3 a.m. you do what stops the bleeding. What matters is what happens next.
03:14 outage. someone widens a security group in the console. site recovers.
09:00 nobody writes it down.
now: the code says one thing, reality says another
14:00 a colleague applies an unrelated change…
14:01 …and terraform helpfully "corrects" the drift — the outage returns.
# detection
terraform plan -detailed-exitcode # exit 2 = drift found (run this nightly in CI)
# the fix, in order
1. reproduce the manual change IN CODE
2. plan → confirm it now proposes no changes
3. commit it, with the incident number in the message
Emergency console changes are allowed. Leaving them undocumented is not. Any hand-fix gets a follow-up commit the same day, or the next unrelated apply becomes an outage — and it will be one nobody can explain, because the code looked fine.
TOPIC 09Ansible: configuring what exists
Terraform creates infrastructure. Ansible configures it: install packages, write config files, manage services, deploy code. It works over plain SSH with no agent to install, and its tasks are written to be idempotent — "ensure this package is present", not "run apt install".
- name: Configure web servers
hosts: web
become: true # use sudo
vars:
app_port: 3000
tasks:
- name: Install packages
ansible.builtin.apt:
name: [nginx, curl, jq]
state: present # "present", not "install" — declarative
update_cache: true
- name: Deploy nginx site config
ansible.builtin.template:
src: templates/site.conf.j2
dest: /etc/nginx/sites-available/app.conf
mode: "0644"
notify: reload nginx # only reloads IF this task changed something
- name: Ensure nginx is running and enabled
ansible.builtin.service:
name: nginx
state: started
enabled: true
handlers:
- name: reload nginx
ansible.builtin.service:
name: nginx
state: reloaded
ansible-inventory --list ansible all -m ping # can I reach every host? ansible-playbook playbook.yml --check --diff # dry run, show the diffs ansible-playbook playbook.yml --limit web-01 # one host first, always ansible-playbook playbook.yml PLAY RECAP: web-01 : ok=6 changed=2 failed=0 # ← run again: changed=0
That last detail is the whole idea: a second run reports changed=0, because everything is already as declared. If a playbook reports changes on every run, some task is not idempotent — and that is a bug, because you can no longer tell "nothing to do" from "something moved".
TOPIC 10Immutable infrastructure
Two philosophies for updating a server:
| Mutable (configure in place) | Immutable (replace) | |
|---|---|---|
| Update process | SSH or Ansible changes the running server | Build a new image, boot new servers, remove old ones |
| Rollback | Try to undo the change | Boot the previous image |
| Drift risk | Accumulates over years | Near zero — nothing lives long enough to drift |
| Tooling | Ansible, Chef, Puppet | Packer + Terraform, or containers (chapters 06–07) |
Repairing a rented scooter versus swapping it for another from the rack. Repairs accumulate: this one has a soft brake, that one a rattling mirror, and eventually no two are alike — the "snowflake server" problem. Swapping means every rider gets a known machine, and a bad batch is withdrawn rather than nursed. Containers made this cheap enough to be the default.
You will meet both. Mutable is normal in long-lived enterprise estates; immutable is the direction of travel and what containers give you almost for free. The habit to carry either way: servers are cattle, not pets — anything you would be reluctant to delete and rebuild is a risk you have not measured.
TOPIC 11IaC in a pipeline
on:
pull_request:
paths: ['infra/**']
push:
branches: [main]
paths: ['infra/**']
jobs:
plan:
runs-on: ubuntu-latest
permissions:
id-token: write # OIDC → short-lived cloud creds, no static keys
pull-requests: write
steps:
- uses: actions/checkout@v4
- uses: hashicorp/setup-terraform@v3
- run: terraform fmt -check -recursive
- run: terraform init
- run: terraform validate
- run: terraform plan -no-color -out=tfplan
- run: tfsec . || true # security scan (chapter 11)
# post the plan as a PR comment so a human reviews the diff
apply:
needs: plan
if: github.ref == 'refs/heads/main'
environment: production # requires a manual approval
runs-on: ubuntu-latest
steps: [ ... terraform apply ... ]
The shape that works: plan on the pull request, apply only from main, with a human approving production. The plan output in the PR is the review artefact — it is far easier to spot "this will destroy the database" in a diff than in a terminal at the end of a long day.
- IaC gives infrastructure what code already has: review, history, diffs, repeatability.
- Declarative = describe the end state; re-running is safe. That property is idempotency.
- Terraform creates; Ansible configures. Both belong in Git.
- State maps code to real resources. Remote, locked, encrypted, versioned, one per environment.
- Read every plan.
-/+means destroy and recreate. - Modules keep staging and production the same code with different numbers.
- Manual console fixes are fine in an emergency and must be written back into code the same day.
LABInfrastructure you can destroy
Learn the loop with the Docker provider, then read a cloud plan
- No-cost start: install Terraform and use the
kreuzwerker/dockerprovider. Write a config that pullsnginx:alpineand runs a container on port 8080. Runinit,plan,apply, then curl it. You have just used Terraform for real. - Run
terraform applyagain. Note it reports no changes — this is idempotency, and it is the whole point. - Change the published port to 8081 and plan. Read the symbols carefully: does it propose
~or-/+? Explain to yourself why. - Create drift: stop the container by hand with
docker stop, then runterraform plan. Watch Terraform notice reality has diverged and offer to fix it. Apply. - Inspect state:
terraform state list, thenterraform state showon your container. Open the state file and read it — knowing what is in there is why you protect it. - Variables and modules: extract the image name and port into variables with defaults, add an output for the container name, then move the whole thing into
modules/web/and call it twice with different ports. Apply, and confirm two containers. terraform destroy. Confirm withdocker psthat nothing is left. The confidence to destroy is the real deliverable of this chapter.- If you have a cloud free tier: write the security group and one small instance from topic 03, and run
planonly. Read all of it. Then, if you apply, set a calendar reminder todestroy— a forgotten resource is the most common first cloud bill shock. - Stretch: add the plan-on-PR workflow from topic 11 to a repo, with
terraform fmt -checkas a gate.
CHECKCheck yourself
A teammate deletes their local terraform.tfstate and runs terraform apply against production. What happens?
Terraform does not discover existing resources; it trusts state. With an empty state, desired minus actual equals "create everything", so you get duplicates — and names that must be unique will fail halfway, leaving a mess. This is exactly why state belongs in a versioned, locked remote backend rather than on anyone's laptop. Recovery means restoring the state file or terraform import for each resource, one at a time.
Your plan output includes -/+ destroy and then create replacement for a database instance. What should you do?
Some attributes cannot be changed in place, so Terraform replaces the resource — and it does not move your data. The plan names the culprit as "forces replacement". Options include changing the approach, using lifecycle { prevent_destroy = true } on critical resources as a guardrail, or planning a proper migration with snapshots. This is the single most expensive mistake in the chapter, and reading the plan is the entire defence.
An Ansible playbook reports changed=4 every single time it runs, even with no edits. Why does that matter?
Idempotent tasks report changed only when they actually alter something, which makes the recap a useful signal — and makes notify handlers fire only when needed. Perpetual changes usually come from raw command/shell tasks without a creates: or when: guard. Fix them, or you lose the ability to spot drift in the noise.