devopsdiary
Next chapter

diary / chapters / 08

chapter 08 · ~55 min · terraform lab

Infrastructure as code

Clicking through a cloud console to create a server works exactly once. Writing that server down as code means you can review it, repeat it, destroy it, rebuild it identically, and know six months later why it exists.

TOPIC 01The problem with clicking

You create a server in a web console: pick a size, a network, a disk, three security rules, tick two boxes. It works. Six months later:

  • Nobody knows why one of those boxes is ticked, so nobody dares untick it.
  • The staging server was built by someone else on a different afternoon and behaves differently, which is why the bug "only happens in staging".
  • The person who built it has left, and the reasoning left with them.
  • Recreating it in another region means repeating forty clicks from memory.
  • There is no record of who changed the firewall rule last Tuesday, or why.

Infrastructure as code means the servers, networks, databases, DNS records and permissions are described in text files, kept in Git, reviewed in pull requests, and applied by a tool. Infrastructure gains everything code already has: history, review, diffs, rollback and repeatability.

real life

Two ways to build the same kitchen. One: instruct the carpenter verbally, day by day, adjusting as you go — beautiful, and utterly unrepeatable. Two: a measured drawing. The drawing can be checked before anything is cut, handed to a different carpenter in another city and produce the same kitchen, and when you want the shelf moved you change the drawing rather than arguing about what was agreed in March.

TOPIC 02Declarative vs imperative

Imperative (a script)Declarative (Terraform)
You writeThe steps to takeThe end state you want
Example"create a server, then attach a disk""a server with this disk exists"
Run it twiceTwo servers (unless you wrote guards)No change — it already matches
Partial failureUnclear state; rerunning may double upRerun; it completes what is missing

Declarative tools compare desired state with actual state and work out the difference themselves. That property — running it again is always safe — is idempotency, and it is the same reconciliation idea as the Kubernetes thermostat in chapter 07. Once you spot the pattern you will see it everywhere in this field.

TOPIC 03Terraform in five files

Terraform (and its fork OpenTofu) creates and manages infrastructure across providers — AWS, Azure, GCP, Cloudflare, GitHub, even Kubernetes. The language is HCL. Here is a complete, realistic little stack:

versions.tf — pin everything
terraform {
  required_version = "~> 1.9"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.60"          # allow patches, not surprises
    }
  }

  backend "s3" {                   # shared state, see topic 05
    bucket       = "acme-tfstate"
    key          = "prod/web/terraform.tfstate"
    region       = "ap-south-1"
    use_lockfile = true            # prevents two people applying at once
  }
}

provider "aws" {
  region = var.region
  default_tags {
    tags = {                       # tag everything: this is how you
      Environment = var.environment  # answer "what is this and whose bill?"
      ManagedBy   = "terraform"
      Repo        = "acme/infra"
    }
  }
}
main.tf — the resources
# look up the newest Ubuntu image instead of hard-coding an AMI id
data "aws_ami" "ubuntu" {
  most_recent = true
  owners      = ["099720109477"]
  filter {
    name   = "name"
    values = ["ubuntu/images/hvm-ssd-gp3/ubuntu-noble-24.04-amd64-server-*"]
  }
}

resource "aws_security_group" "web" {
  name        = "${var.environment}-web"
  description = "HTTP/HTTPS in, everything out"
  vpc_id      = var.vpc_id

  ingress {
    description = "HTTPS from the internet"
    from_port   = 443
    to_port     = 443
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }

  ingress {
    description = "SSH from the office VPN only"
    from_port   = 22
    to_port     = 22
    protocol    = "tcp"
    cidr_blocks = [var.admin_cidr]     # never 0.0.0.0/0 (chapter 03)
  }

  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }
}

resource "aws_instance" "web" {
  count                  = var.instance_count
  ami                    = data.aws_ami.ubuntu.id
  instance_type          = var.instance_type
  subnet_id              = var.subnet_ids[count.index % length(var.subnet_ids)]
  vpc_security_group_ids = [aws_security_group.web.id]   # ← the dependency

  root_block_device {
    volume_size = 20
    encrypted   = true
  }

  tags = { Name = "${var.environment}-web-${count.index + 1}" }
}

Notice aws_security_group.web.id inside the instance. That reference is how Terraform builds a dependency graph and works out the order by itself — you never write "create the security group first". Four block types cover almost everything: resource (something to create), data (something to look up), variable (input), output (result).

TOPIC 04plan, apply, destroy

terminal · the daily loop
terraform init              # download providers, connect the backend
terraform fmt -recursive    # canonical formatting (put this in CI)
terraform validate          # syntax and type checking

terraform plan -out=tfplan  # the dry run. READ IT. every line.
Plan: 3 to add, 1 to change, 0 to destroy.

terraform apply tfplan      # apply exactly the plan you reviewed

terraform show              # what exists now, per state
terraform output alb_dns    # a value another system needs
terraform destroy           # tear it all down (your wallet's best friend)
read the plan, especially the word "destroy"

plan is the whole safety mechanism. Learn to read the symbols: + create, ~ update in place, -/+ destroy and recreate, - destroy. The dangerous one is -/+: changing an attribute that cannot be modified in place quietly means replacing a live resource. On a database, that is data loss written in a diff you skimmed. If a plan proposes destroying something you did not expect, stop and find out why before typing yes.

TOPIC 05State: the crown jewels

Terraform keeps a state file mapping your code to the real resource IDs it created — "the instance named web-1 is i-0abc123". Without it, Terraform has no idea what it owns, and a plan would propose creating everything again.

real life

State is the property register. The code is your plan for the neighbourhood; the register records which house on the ground corresponds to which plot on the plan. Lose the register and the government cannot tell your house from anyone else's — it is not that your house vanished, it is that ownership can no longer be proven. Terraform behaves the same way: no state, no idea, so it builds a second house.

Rules that separate professionals from people about to have a bad week:

  • Never keep state only on a laptop, and never commit it to Git. Use a remote backend (S3, Azure Blob, GCS, Terraform Cloud).
  • Always enable locking. Two engineers applying at once with no lock is how state gets corrupted.
  • Treat state as a secret. It contains resource attributes, and sometimes generated passwords, in plain text. Encrypt the bucket, restrict access.
  • Version the bucket so you can recover a previous state file.
  • Separate state per environment. One state file for prod and staging together means a staging mistake can plan a production change.
terminal · state surgery (rarely, carefully)
terraform state list                          # everything Terraform manages
terraform state show aws_instance.web[0]
terraform import aws_instance.web i-0abc123   # adopt an existing resource
terraform state rm aws_instance.old           # forget it (does NOT delete it)
terraform state mv aws_instance.a aws_instance.b  # after renaming in code
terraform force-unlock LOCK_ID                # only when a lock is genuinely stale

TOPIC 06Variables and outputs

variables.tf
variable "environment" {
  description = "Deployment environment name"
  type        = string
  validation {
    condition     = contains(["dev", "staging", "prod"], var.environment)
    error_message = "environment must be dev, staging or prod."
  }
}

variable "instance_count" {
  description = "How many web servers"
  type        = number
  default     = 2
}

variable "admin_cidr" {
  description = "CIDR allowed to SSH"
  type        = string
  default     = "10.8.0.0/24"
}

variable "db_password" {
  description = "Database password — supplied from a secret store, never a file"
  type        = string
  sensitive   = true      # keeps it out of plan/apply output
}
outputs.tf + how values get in
output "web_private_ips" {
  value       = aws_instance.web[*].private_ip
  description = "Private IPs of the web servers"
}

# ---- supplying values, in order of precedence ----
# 1. -var on the command line
# 2. -var-file=prod.tfvars
# 3. terraform.tfvars / *.auto.tfvars
# 4. TF_VAR_environment=prod  (how CI usually does it)
# 5. the default in the variable block

sensitive = true is worth a habit: it stops a password appearing in your terminal, in CI logs, and in the pull-request comment where a plan was posted. It does not encrypt it in state — that is what the backend encryption is for.

TOPIC 07Modules and environments

A module is a folder of Terraform you can call with different inputs — a function for infrastructure. It is how you stop copy-pasting a VPC definition into three environments and then patching two of them.

layout that scales without pain
infra/
├── modules/
│   ├── network/          # vpc, subnets, routing
│   ├── web-service/      # asg or instances + sg + load balancer
│   └── database/
└── environments/
    ├── staging/
    │   ├── main.tf       # calls the modules with small numbers
    │   └── terraform.tfvars
    └── prod/
        ├── main.tf       # same modules, bigger numbers, own state
        └── terraform.tfvars
environments/prod/main.tf
module "network" {
  source      = "../../modules/network"
  environment = "prod"
  cidr_block  = "10.0.0.0/16"
}

module "web" {
  source         = "../../modules/web-service"
  environment    = "prod"
  vpc_id         = module.network.vpc_id
  subnet_ids     = module.network.private_subnet_ids
  instance_count = 6            # staging passes 1
  instance_type  = "t3.medium"  # staging passes t3.micro
}

The win: staging and production are the same code with different numbers, so a test in staging genuinely tells you something. The moment they are two separate copies, they start drifting and staging stops being a rehearsal.

TOPIC 08Drift and the manual fix

Drift is when reality no longer matches the code — usually because someone fixed something by hand during an incident. It is not a moral failing; at 3 a.m. you do what stops the bleeding. What matters is what happens next.

the drift cycle, and how to close it
03:14  outage. someone widens a security group in the console. site recovers.
09:00  nobody writes it down.
      now: the code says one thing, reality says another
14:00  a colleague applies an unrelated change…
14:01  …and terraform helpfully "corrects" the drift — the outage returns.

# detection
terraform plan -detailed-exitcode   # exit 2 = drift found (run this nightly in CI)

# the fix, in order
1. reproduce the manual change IN CODE
2. plan → confirm it now proposes no changes
3. commit it, with the incident number in the message
the rule to write on the wall

Emergency console changes are allowed. Leaving them undocumented is not. Any hand-fix gets a follow-up commit the same day, or the next unrelated apply becomes an outage — and it will be one nobody can explain, because the code looked fine.

TOPIC 09Ansible: configuring what exists

Terraform creates infrastructure. Ansible configures it: install packages, write config files, manage services, deploy code. It works over plain SSH with no agent to install, and its tasks are written to be idempotent — "ensure this package is present", not "run apt install".

playbook.yml
- name: Configure web servers
  hosts: web
  become: true                    # use sudo

  vars:
    app_port: 3000

  tasks:
    - name: Install packages
      ansible.builtin.apt:
        name: [nginx, curl, jq]
        state: present            # "present", not "install" — declarative
        update_cache: true

    - name: Deploy nginx site config
      ansible.builtin.template:
        src: templates/site.conf.j2
        dest: /etc/nginx/sites-available/app.conf
        mode: "0644"
      notify: reload nginx        # only reloads IF this task changed something

    - name: Ensure nginx is running and enabled
      ansible.builtin.service:
        name: nginx
        state: started
        enabled: true

  handlers:
    - name: reload nginx
      ansible.builtin.service:
        name: nginx
        state: reloaded
terminal · running it
ansible-inventory --list
ansible all -m ping                        # can I reach every host?
ansible-playbook playbook.yml --check --diff # dry run, show the diffs
ansible-playbook playbook.yml --limit web-01 # one host first, always
ansible-playbook playbook.yml
PLAY RECAP: web-01 : ok=6 changed=2 failed=0          # ← run again: changed=0

That last detail is the whole idea: a second run reports changed=0, because everything is already as declared. If a playbook reports changes on every run, some task is not idempotent — and that is a bug, because you can no longer tell "nothing to do" from "something moved".

TOPIC 10Immutable infrastructure

Two philosophies for updating a server:

Mutable (configure in place)Immutable (replace)
Update processSSH or Ansible changes the running serverBuild a new image, boot new servers, remove old ones
RollbackTry to undo the changeBoot the previous image
Drift riskAccumulates over yearsNear zero — nothing lives long enough to drift
ToolingAnsible, Chef, PuppetPacker + Terraform, or containers (chapters 06–07)
real life

Repairing a rented scooter versus swapping it for another from the rack. Repairs accumulate: this one has a soft brake, that one a rattling mirror, and eventually no two are alike — the "snowflake server" problem. Swapping means every rider gets a known machine, and a bad batch is withdrawn rather than nursed. Containers made this cheap enough to be the default.

You will meet both. Mutable is normal in long-lived enterprise estates; immutable is the direction of travel and what containers give you almost for free. The habit to carry either way: servers are cattle, not pets — anything you would be reluctant to delete and rebuild is a risk you have not measured.

TOPIC 11IaC in a pipeline

.github/workflows/terraform.yml (abridged)
on:
  pull_request:
    paths: ['infra/**']
  push:
    branches: [main]
    paths: ['infra/**']

jobs:
  plan:
    runs-on: ubuntu-latest
    permissions:
      id-token: write        # OIDC → short-lived cloud creds, no static keys
      pull-requests: write
    steps:
      - uses: actions/checkout@v4
      - uses: hashicorp/setup-terraform@v3
      - run: terraform fmt -check -recursive
      - run: terraform init
      - run: terraform validate
      - run: terraform plan -no-color -out=tfplan
      - run: tfsec . || true             # security scan (chapter 11)
      # post the plan as a PR comment so a human reviews the diff

  apply:
    needs: plan
    if: github.ref == 'refs/heads/main'
    environment: production              # requires a manual approval
    runs-on: ubuntu-latest
    steps: [ ... terraform apply ... ]

The shape that works: plan on the pull request, apply only from main, with a human approving production. The plan output in the PR is the review artefact — it is far easier to spot "this will destroy the database" in a diff than in a terminal at the end of a long day.

remember this much
  • IaC gives infrastructure what code already has: review, history, diffs, repeatability.
  • Declarative = describe the end state; re-running is safe. That property is idempotency.
  • Terraform creates; Ansible configures. Both belong in Git.
  • State maps code to real resources. Remote, locked, encrypted, versioned, one per environment.
  • Read every plan. -/+ means destroy and recreate.
  • Modules keep staging and production the same code with different numbers.
  • Manual console fixes are fine in an emergency and must be written back into code the same day.

LABInfrastructure you can destroy

55 minutes · no cloud account needed for steps 1–5

Learn the loop with the Docker provider, then read a cloud plan

  1. No-cost start: install Terraform and use the kreuzwerker/docker provider. Write a config that pulls nginx:alpine and runs a container on port 8080. Run init, plan, apply, then curl it. You have just used Terraform for real.
  2. Run terraform apply again. Note it reports no changes — this is idempotency, and it is the whole point.
  3. Change the published port to 8081 and plan. Read the symbols carefully: does it propose ~ or -/+? Explain to yourself why.
  4. Create drift: stop the container by hand with docker stop, then run terraform plan. Watch Terraform notice reality has diverged and offer to fix it. Apply.
  5. Inspect state: terraform state list, then terraform state show on your container. Open the state file and read it — knowing what is in there is why you protect it.
  6. Variables and modules: extract the image name and port into variables with defaults, add an output for the container name, then move the whole thing into modules/web/ and call it twice with different ports. Apply, and confirm two containers.
  7. terraform destroy. Confirm with docker ps that nothing is left. The confidence to destroy is the real deliverable of this chapter.
  8. If you have a cloud free tier: write the security group and one small instance from topic 03, and run plan only. Read all of it. Then, if you apply, set a calendar reminder to destroy — a forgotten resource is the most common first cloud bill shock.
  9. Stretch: add the plan-on-PR workflow from topic 11 to a repo, with terraform fmt -check as a gate.

CHECKCheck yourself

A teammate deletes their local terraform.tfstate and runs terraform apply against production. What happens?

Terraform does not discover existing resources; it trusts state. With an empty state, desired minus actual equals "create everything", so you get duplicates — and names that must be unique will fail halfway, leaving a mess. This is exactly why state belongs in a versioned, locked remote backend rather than on anyone's laptop. Recovery means restoring the state file or terraform import for each resource, one at a time.

Your plan output includes -/+ destroy and then create replacement for a database instance. What should you do?

Some attributes cannot be changed in place, so Terraform replaces the resource — and it does not move your data. The plan names the culprit as "forces replacement". Options include changing the approach, using lifecycle { prevent_destroy = true } on critical resources as a guardrail, or planning a proper migration with snapshots. This is the single most expensive mistake in the chapter, and reading the plan is the entire defence.

An Ansible playbook reports changed=4 every single time it runs, even with no edits. Why does that matter?

Idempotent tasks report changed only when they actually alter something, which makes the recap a useful signal — and makes notify handlers fire only when needed. Perpetual changes usually come from raw command/shell tasks without a creates: or when: guard. Fix them, or you lose the ability to spot drift in the noise.

saved in this browser only — no account needed