TOPIC 01Why shift left
The traditional model: build for six months, then a security team audits for two weeks and returns a spreadsheet of 200 findings, most of which are now expensive to fix because they are baked into the design.
DevSecOps moves those checks into the pipeline, where they run on every pull request and fail in seconds. The economics are stark:
in your editor seconds a linter squiggle on a pull request minutes a red check, fixed before review in staging hours a ticket, a re-test in production days an incident, a rollback, a rushed patch after a breach months disclosure, regulators, lost trust
Airport security placed at the gate rather than the entrance. If you screen everyone at the door, one person with a prohibited item is turned away in five seconds. Screen at the gate and you must clear a full aircraft, delay every passenger and re-run the whole process. Same check, wildly different cost — and the only variable is where you put it.
The DevSecOps mindset in one line: make the secure path the easy path. A template repository with scanning already wired in, a secret store that is simpler to use than an env file, a base image that is already hardened. People do not bypass security because they are careless; they bypass it because the alternative was faster.
TOPIC 02Secrets management
A secret is anything that grants access: passwords, API keys, tokens, private keys, connection strings. The rules, in order of importance:
- Never in Git. Not in code, not in config, not in a "temporary" commit. Git history is forever and repositories get cloned, forked and mirrored.
- Never in a container image.
ENVandARGvalues are visible indocker historyto anyone who can pull the image. - Inject at runtime as environment variables or mounted files, from a store the platform controls.
- Different secret per environment. A staging credential must not open production.
- Rotate on a schedule, and immediately on any suspicion. Rotation you have never practised is rotation that will fail when you need it.
- Prefer short-lived credentials. A token that expires in an hour is dramatically less valuable to an attacker than a key that never expires.
| Where | Use | Notes |
|---|---|---|
| CI/CD | The provider's secret store; OIDC for cloud access | OIDC removes static cloud keys entirely — worth the setup |
| Kubernetes | Secrets + RBAC, or External Secrets Operator | Base64 is not encryption (chapter 07) |
| Cloud | AWS Secrets Manager, GCP/Azure equivalents | Automatic rotation, audit logs, IAM-controlled |
| Self-hosted | HashiCorp Vault, Infisical, SOPS + age | SOPS lets you keep encrypted secrets in Git safely |
| Local dev | .env, gitignored, with a committed .env.example | Never real production values |
TOPIC 03You leaked a credential
This happens to good engineers. What separates a non-event from a breach is the order of the next ten minutes.
1. ROTATE FIRST. Create a new credential, deploy it, disable the old one.
Assume the secret is public the instant it was pushed —
bots scan new commits on public repos within seconds.
2. ASSESS ACCESS. What could that credential reach? Which accounts?
3. CHECK THE LOGS. Any use from an unexpected IP, region or time?
Cloud audit logs, database logs, provider dashboards.
4. THEN CLEAN HISTORY. git filter-repo / BFG, force-push, ask forks to re-clone.
This is hygiene, not containment.
5. PREVENT. Add a secret scanner to pre-commit AND to CI.
6. WRITE IT UP. Blameless postmortem (chapter 10). The action item is
"a secret could reach a commit", not "X was careless".
Deleting the commit first and rotating later — or not at all. Rewriting history does not un-publish anything: the value may already be in someone's clone, a CI log, a cache or a scraper's database. Only rotation actually revokes access, and it takes about the same effort.
TOPIC 04Least privilege and IAM
Least privilege: every human, service and pipeline gets exactly the permissions it needs, and nothing more. It is the control that limits how bad any single compromise can be.
# ✗ what people write when they are in a hurry
{ "Effect": "Allow", "Action": "*", "Resource": "*" }
# a leak here loses the entire cloud account
# ✓ what the pipeline actually needs
{
"Effect": "Allow",
"Action": ["s3:PutObject", "s3:GetObject"],
"Resource": "arn:aws:s3:::acme-web-assets/*"
}
# a leak here lets someone write to one bucket prefix
The practical habits: start from nothing and add permissions as things fail (rather than starting from admin and never trimming); use roles rather than long-lived users; scope by resource, not just by action; put MFA on every human account; and give production a separate account or project from staging so a mistake cannot cross the boundary. In Kubernetes the same idea is RBAC — a ServiceAccount per workload, with the narrowest role that works.
Hotel keycards. The cleaner's card opens rooms on one floor during working hours; it does not open the safe, the server room or the accounts office. Nobody argues this is distrust — it is simply the design that means a lost card is an inconvenience rather than a crisis.
TOPIC 05The scanner alphabet
| Type | What it looks at | Catches | Tools |
|---|---|---|---|
| SCA | Your dependencies | Known vulnerable library versions | Dependabot, npm audit, Trivy, Snyk |
| SAST | Your source code, statically | Injection, unsafe deserialisation, weak crypto | Semgrep, CodeQL, Bandit |
| Secret scanning | Commits and history | Keys and tokens about to be published | gitleaks, trufflehog, provider built-ins |
| Container scanning | Image layers | Vulnerable OS packages, bad configuration | Trivy, Grype, Docker Scout |
| IaC scanning | Terraform / K8s manifests | Public buckets, open security groups, no encryption | tfsec, Checkov, kube-score |
| DAST | The running application | Runtime issues: headers, auth flaws, XSS | OWASP ZAP, Nuclei |
You do not need all six on day one. The highest value for the least effort, in order: secret scanning (prevents the worst mistake), SCA (most real vulnerabilities arrive through dependencies), then IaC scanning (catches the open-to-the-world bucket before it exists).
A scanner that reports 400 findings gets muted, and a muted scanner is worse than none. Fail the build only on high and critical severity in code paths you actually ship; report the rest to a dashboard or a weekly ticket. Then tune. A gate people respect is worth more than a gate that is comprehensive.
TOPIC 06Container and IaC scanning
# what is vulnerable in my image? (and how much comes from the base) trivy image --severity HIGH,CRITICAL notes-api:v1 Total: 4 (HIGH: 3, CRITICAL: 1) # fail a build only on fixable criticals trivy image --exit-code 1 --severity CRITICAL --ignore-unfixed notes-api:v1 # misconfigurations in terraform tfsec ./infra Result #1 CRITICAL Security group rule allows ingress from 0.0.0.0/0 to port 22 # kubernetes manifests trivy config ./k8s # runAsNonRoot, no limits, privileged, etc. # secrets, in the whole history gitleaks detect --source . --redact
Most image findings come from the base image, which makes the fix delightfully simple: use a smaller base and rebuild regularly. Moving from a full distribution image to -slim, -alpine or distroless routinely removes dozens of vulnerabilities in packages your app never used — which is the same multi-stage build argument from chapter 06, now with a security number attached.
TOPIC 07Supply chain and SBOMs
Your application is mostly other people's code. A typical Node or Python service pulls in hundreds of transitive dependencies, any one of which could be compromised or abandoned. Supply chain security is about knowing and controlling what you ship.
- Pin and lock. Commit the lockfile; pin base images to a version or digest. Reproducibility is a security property, not just a convenience.
- Generate an SBOM — a Software Bill of Materials listing every component and version. When the next widely-exploited vulnerability lands, "are we affected?" becomes a one-minute query instead of a two-day audit.
- Verify provenance. Sign your images (
cosign) and, in a mature setup, let the cluster admit only signed images. That closes the "someone pushed a different image to that tag" gap. - Pin third-party CI actions to a commit SHA, not a moving tag. A pipeline action runs with your secrets — treat it as production code.
- Beware typosquats. Malicious packages with names one letter from a popular one are a standing attack. Copy names from the registry, do not retype them.
trivy image --format cyclonedx -o sbom.json notes-api:v1 syft notes-api:v1 -o spdx-json > sbom.spdx.json cosign sign ghcr.io/acme/notes-api@sha256:9f2a… cosign verify ghcr.io/acme/notes-api@sha256:9f2a… --certificate-identity-regexp '.*'
TOPIC 08Hardening checklists
Containers
- Run as a non-root user; set
runAsNonRoot: truein Kubernetes. - Read-only root filesystem where possible; drop all Linux capabilities you do not need.
- Never
privileged: true, and never mount the Docker socket into a container you did not write. - Set memory and CPU limits — resource exhaustion is a denial-of-service vector, not just an operations concern.
- Multi-stage build: no compilers, shells or package managers in the runtime image.
Network and hosts
- Default deny; open the minimum per tier. Databases are never internet-facing (chapter 03).
- SSH by key only, from a bastion or VPN. No password authentication, no root login.
- TLS everywhere, with automated renewal and an alert at 14 days to expiry.
- In Kubernetes, use NetworkPolicies so a compromised pod cannot talk to everything in the cluster.
Application
- Validate input at the boundary; use parameterised queries, never string-concatenated SQL.
- Security headers: HSTS, a content security policy,
X-Content-Type-Options. - Rate-limit authentication endpoints; log authentication failures (without logging the credentials).
- Patch on a schedule you can actually keep. An unpatched known vulnerability is the most commonly exploited thing there is.
TOPIC 09A secure pipeline
on: [pull_request]
permissions:
contents: read # start from the minimum, add what you need
jobs:
guardrails:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 } # full history for secret scanning
- name: Secret scan
uses: gitleaks/gitleaks-action@v2 # pin to a SHA in real life
- name: Dependency audit (SCA)
run: npm audit --audit-level=high
- name: Static analysis (SAST)
run: semgrep ci --config auto
- name: Build image
run: docker build -t app:${{ github.sha }} .
- name: Image scan
run: trivy image --exit-code 1 --severity CRITICAL
--ignore-unfixed app:${{ github.sha }}
- name: IaC scan
run: trivy config ./infra ./k8s
- name: SBOM
run: trivy image --format cyclonedx -o sbom.json app:${{ github.sha }}
- uses: actions/upload-artifact@v4
with: { name: sbom, path: sbom.json }
Note the placement: every one of these runs before the deploy stage, which is exactly the ordering the Pipeline Builder game tests. A scan after the deploy tells you what your users were already exposed to.
TOPIC 10Compliance, briefly
You will meet these acronyms, and DevOps work is usually what satisfies them:
- SOC 2 / ISO 27001
- Prove you have controls and that you follow them. Your pull-request reviews, access logs, change history and alerting are the evidence.
- GDPR / DPDP-style data laws
- Know what personal data you hold, where it lives, how long you keep it, and be able to delete it. Log hygiene matters here — see topic 03 of chapter 10.
- PCI DSS
- If you touch card data: network segmentation, encryption, strict access control, retained audit trails.
The reframe that makes this bearable: compliance is mostly automation plus evidence. A team already doing code review, IaC, immutable artifacts, least privilege and centralised logging is most of the way there, and the audit becomes an exercise in exporting what you already produce.
- Move checks left: the same finding costs seconds on a PR and weeks after a breach.
- Secrets: never in Git or images, injected at runtime, per environment, rotated, short-lived where possible.
- Leaked a credential? Rotate first. History cleanup is hygiene, not containment.
- Least privilege limits the damage of any single compromise. Roles over static keys; OIDC over stored cloud keys.
- Start with three scanners: secrets, dependencies (SCA), IaC.
- Fail builds only on high/critical, or people will mute the scanner.
- Smaller base images plus regular rebuilds remove most image findings for free.
- Keep an SBOM so "are we affected?" takes a minute.
LABFind your own holes
Scan what you built, then fix the worst of it
- Secret scan your history:
docker run --rm -v "$PWD:/repo" zricethezav/gitleaks:latest detect --source /repo --redact. Read every finding, including test fixtures. - Add a pre-commit hook that runs the same scan on staged changes. Then try to commit a fake AWS key (
AKIAIOSFODNN7EXAMPLE) and confirm you are blocked. - Scan your image:
trivy image --severity HIGH,CRITICAL yourapp:v1. Note the total. - Fix it the easy way: switch the base image from a full distribution to
-alpine,-slimor distroless, rebuild, and scan again. Write down both numbers — a 90% reduction from a one-line change is a genuinely good interview story. - Harden the container: add a non-root user, drop capabilities, and set a read-only root filesystem if your app allows it. Re-run
trivy configon your Kubernetes manifests and clear the warnings. - Scan your infrastructure: run
trivy config ./infra(ortfsec) against the chapter-08 Terraform. Deliberately add a security group open to0.0.0.0/0on port 22 first, and confirm the scanner catches it. - Least privilege drill: take any IAM policy or Kubernetes Role you have written and remove permissions one at a time until something breaks. Add back only what was needed. That is the whole discipline.
- Wire it into CI: add the security workflow from topic 09 and make one job fail on purpose so you can see the gate work.
- Generate an SBOM and store it as a build artifact. Then answer this out loud: if a critical vulnerability in your HTTP library were announced tonight, how would you know whether you ship it?
CHECKCheck yourself
A Kubernetes Secret holds your database password. A colleague says "it's fine, it's encrypted." What is the accurate correction?
Base64 is an encoding anyone can reverse in one command. What actually protects a Secret is who can read it (RBAC), whether etcd is encrypted at rest, and keeping the real value outside Git — for example with the External Secrets Operator pulling from a cloud secret manager.
A pull request from a fork triggers your pipeline. Why should that run not have production deploy credentials?
A pipeline executes code from the pull request, so an untrusted PR that gets secrets can simply print or upload them. That is why CI providers withhold secrets from fork runs by default, and why deploy jobs should be gated on a trusted branch or an environment with manual approval — never on an event a stranger can trigger.
Your image scanner reports 380 findings, so the team stops reading them. What is the best next step?
Signal quality decides whether a control survives contact with a deadline. Gate on the small set that matters and is actually fixable, report the rest without blocking, and attack the root cause — most findings live in base-image packages your application never calls, so a slimmer base deletes them wholesale.