Blast Radius — Policy-as-Code, Secrets & CI Applies
Terraform outages are usually applies that were allowed to do too much. This page covers policy on plan JSON, secret hygiene, least-privilege CI roles, and gates sized to blast radius.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What makes a plan's blast radius large?
Answer
Shared network scope, data-store replace, a role that can assume other accounts, and a delete you cannot restore.
L2
Why split the plan role from the apply role?
Answer
A pull request, especially from an untrusted branch, must not be able to mutate prod. Plan needs read APIs.
L3
Where do OPA and Checkov sit?
Answer
Checkov reads configuration. OPA or Conftest reads plan JSON. Both can fail the pull request.
L4
What is the rule for secret values?
Answer
Do not commit them. Terraform creates the container and the IAM binding. A pipeline or operator sets the value.
L5
Can OPA stop a laptop apply?
Answer
Not if the laptop holds raw cloud admin. Pair policy with permission boundaries and a rule against local prod applies.
L6
A sensitive output appeared in CI. Now what?
Answer
Treat the log as compromised and rotate. Then find the nested value or the unredacted show command.
L7
How do feature flags relate to Terraform?
Answer
App flags and progressive delivery live on the rollout pages. A Terraform enable flag is a coarse infra switch, not that system.
Failure modes
Admin on the plan role
A speculative pull request can create and delete because someone wanted CI to be simple.
One state file and a role that assumes every account
A single apply is the blast radius of the whole org.
Policy disabled until things calm down
The exception becomes the path, and destroys ship on Friday.
Secrets in tfvars shared in chat
The value is in git history, in the ticket, and in state.
A rubber-stamp apply approval
The gate exists and nobody reads the replace lines.
Misconceptions
Policy-as-code replaces IAM.
OPA fails a pipeline. The apply role and the org permission boundary are what stop a laptop.
sensitive = true keeps a value out of state.
It redacts some CLI output. Providers can still store the attribute. Restrict state pull anyway.
Terraform enable flags are feature flags.
They toggle infrastructure. Progressive delivery of an application lives on the Kubernetes rollout pages.
Interviewer traps
Giving the plan job AdministratorAccess so the graph can refresh.
Refresh needs read APIs. Creates and deletes belong to the apply role in that environment.
Designing a new feature-flag product in the Terraform answer.
One sentence to the existing rollout and rollback pages, then return to roles and plan policy.
Design scenario
Same prompt for every reader.
Requirements
OIDC plan and apply roles, policy on plan JSON and on HCL, no plaintext secrets in git, and a prod environment protection rule. Break-glass is time-boxed and ticketed.
Traffic / scale
Dozens of pull-request plans a day. A handful of prod applies inside a change window.
Latency
Policy fails the job in the pull request. The prod apply waits for reviewers, not for a laptop.
Consistency
A destroy of a typed data resource requires an explicit label and two reviewers. Tags such as owner are required.
Availability
Break-glass can still fix an outage in minutes, with an audit trail and an import ticket inside a day.
Failure assumptions
- A fork pull request runs the plan workflow.
- A provider writes a secret into a non-sensitive attribute.
- Someone applies prod from a laptop with personal admin keys.
Constraints
- Plan cannot create or delete.
- Prod apply uses a different OIDC subject than dev.
- Secret values are not in .tf or .tfvars.
Prompt
A GitHub Actions workflow uses a long-lived admin key to plan and apply every root, including prod, on merge. OPA was turned off during an incident and never restored. A database password was pasted into a tfvars file.
API
What identity does the plan job assume, and what does apply assume?
Data
Which plan actions are an automatic deny, and which require a label?
Architecture
Where do OPA, the apply environment protection, and the secret container sit?
A prod pull request can open a security group and replace a database
Prefer
Policy plus a narrow apply role
Static checks and plan JSON fail closed on known denies. The prod role cannot touch another account, and a human still reads the replace.
- The plan job uses OIDC and describe APIs only.
- A world-open SSH rule fails before merge.
- A database destroy needs a label and a second reviewer.
Alternative
One admin key for plan and apply
Every pull request workflow holds prod mutation rights, and policy is optional.
- A fork or a bad plan can delete data.
- Secrets in tfvars are one commit away from the log.
- Friday auto-apply has no change window.
Shrink what an apply is allowed to do
The saved plan is the previous page. Here you decide who may execute it and which diffs are illegal.
- 1
Name the radius
Which accounts, which data, whose on call, and whether the delete can be restored. - 2
Fail closed in CI
Checkov on the HCL. OPA or Sentinel on the plan JSON. Tune false positives or people skip the check. - 3
Separate identities
Plan reads. Dev apply is prefix-scoped. Prod apply is a different OIDC subject behind environment protection. - 4
Break glass on purpose
Time-boxed elevation, a ticket, and an import within a day. Personal admin keys are not the plan role.
Overview
Most Terraform outages are applies that were allowed to do too much, sometimes while leaking secrets into state or logs. Policy, roles, and human gates have to be sized to that radius.
Dimensions
| Dimension | Question |
|---|---|
| Scope | One application security group, or the shared network root? |
| Data | Does a replace touch a database, a bucket, or a volume? |
| Identity | Can this role assume other accounts? |
| Irreversibility | Deletes without a backup, or a DNS change with a long TTL? |
| Audience | Who is awake if this fails on Friday afternoon? |
Not every pull request needs an executive. Every database replace needs a grown-up.
Policy placement
Decisions
- 1
1. terraform plan JSON
- next3. OPA or Sentinel
- 2
3. OPA or Sentinel
- next4. Deny rule hit?
- 3
2. Checkov on HCL
- next3. OPA or Sentinel
- ?
4. Deny rule hit?
- yes5. Fail the pull request
- no6. Human apply gate
- 5
5. Fail the pull request
- 6
6. Human apply gate
Lesson map
Blast Radius — Policy-as-Code, Secrets & CI Applies
Terraform outages are usually applies that were allowed to do too much. This page covers policy on plan JSON, secret hygiene, least-privilege CI roles, and gates sized to blast radius.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB plan["1. terraform plan JSON"] hcl["2. Checkov on HCL"] opa["3. OPA or Sentinel"] deny["4. Deny rule hit?"] plan -->|1. terraform plan JSON| opa hcl -->|2. Checkov on HCL| opa opa -->|3. OPA or Sentinel| deny
High-value rules:
- Deny
0.0.0.0/0on sensitive ports. - Deny unencrypted buckets and volumes.
- Deny destroys of typed data resources unless an
allow-destroy-datalabel is present. - Require owner and cost-center tags.
- Cap huge
for_eachexpansions in prod.
Policies are code. Review them like code. A noisy pack trains people to skip it.
| Tool | Strength | Cost |
|---|---|---|
| Sentinel | Tight fit with Terraform Cloud or Enterprise | You are coupled to that control plane |
| OPA and Conftest | Open, and excellent on plan JSON | You own the policy UX and the library |
| Checkov, tfsec, or Trivy | Fast static checks and many built-in rules | Noise until you tune them |
A practical stack is a static scanner on the pull request, OPA on the plan JSON, and organization permission boundaries outside Terraform as the hard backstop. The plan file itself was built on Plan, apply, drift, and import.
Secrets
- Never commit plaintext secrets in
.tf,.tfvars, or the pull-request description. - Prefer a cloud secret store. Terraform creates the container and the IAM binding. A pipeline or an operator sets the value.
- Mark sensitive variables and outputs, and still assume state may store provider-returned secrets.
- Encrypt backends, restrict state pull, and scrub CI logs.
- Rotate through the secret-store process. Pasting a new string into tfvars is not rotation.
Envelope encryption and rotation procedure live on Secrets and KMS. Store selection lives on secret stores. If a provider forces a secret into state, document it, limit the ACL, and look for an alternative resource.
CI roles
| Role | Allowed |
|---|---|
| Plan on a pull request | Read state and describe APIs. No create or delete |
| Apply in dev | One account and a resource prefix |
| Apply in prod | Narrower still, a separate OIDC subject, a protected environment |
| Break-glass human | Time-boxed IAM, a ticket, and dual control where the org requires it |
Decisions
- 1
1. Plan role is read-only
- next2. Short-lived OIDC token
- 2
2. Short-lived OIDC token
- next4. Prod apply?
- 3
3. Apply role is env-scoped
- next2. Short-lived OIDC token
- ?
4. Prod apply?
- yes5. Required reviewers
- no6. Dev may auto-apply
- 5
5. Required reviewers
- 6
6. Dev may auto-apply
Short-lived OIDC from GitHub Actions or GitLab workload identity beats a static access key in repository secrets. The pipeline product around those jobs lives on CI/CD pipelines.
Gates that reviewers actually use
- Path filters. Changes under a prod directory require the stronger rule set.
- A plan comment. Post replace and destroy summaries from the JSON, not a wall of log text.
- Environment protection. A wait timer and required reviewers on the prod apply job.
- Freeze windows. Block applies during peak unless an override label is recorded.
A gate that nobody reads is a rubber stamp. Train reviewers to open replaces first.
Break-glass
- Declare the incident and open a ticket.
- Elevate with a time box, or use a pre-baked break-glass role with MFA.
- Fix from the CLI or the console if minutes matter.
- Within a day, import the result or record an intentional exception.
- In the retro, name which policy would have caught the non-emergency path.
Git history for the follow-up revert or forward fix lives on Git. Network objects the module emits live on private networking. Application rollout after the infrastructure exists lives on Kubernetes workloads and rolling, blue-green, and canary.
Progressive delivery and feature-flag cutover already have a home on rollback and feature flags. A Terraform enable_x variable is a coarse infrastructure switch.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Anti-patterns
- Admin on the plan role so CI is simple.
- One state file for every account, with a role that can assume all of them.
- Auto-apply of prod on Friday with no destroy policy.
- A tfvars secret passed around in chat.
- Disabling OPA or Sentinel "temporarily" for months.
Interview Q&A
Why bother splitting the plan role and the apply role?
Answer
Untrusted pull requests must not mutate prod. Even a trusted plan job should not hold AdminAccess when it only needs read APIs to refresh the graph.
Can OPA stop a bad apply from a laptop?
Answer
Not if that laptop has raw cloud credentials. Pair pipeline policy with service-control policies or permission boundaries, and with a rule that prod is not applied locally.
A sensitive output showed up in CI. What happened?
Answer
A nested value, terraform show without redaction, or a provider writing into an attribute that was not marked sensitive. Treat the log as compromised and rotate. Then fix the redaction.
How do feature flags relate to this cluster?
Answer
Application flags and progressive cutover live on the Kubernetes rollout page and the rollback page. Terraform enable flags are coarse. They are not a second flag system.
Sentinel, OPA, or Checkov?
Answer
Sentinel when you already run Terraform Cloud and want policy next to remote runs. OPA when you want an open policy on plan JSON. Checkov when you want fast feedback on the HCL itself. Most teams run a static scanner and OPA, and keep IAM as the backstop.
What does a useful prod gate look like?
Answer
Environment protection, a required reviewer, a comment that lists replaces and destroys, and a freeze window. The reviewer opens the replace lines first.
Walk break-glass in five beats.
Answer
Incident and ticket. Time-boxed role. Fix. Import or document within a day. Retro on which policy should have caught a non-emergency.
What should you say in the first minute?
Answer
Radius is scope, data, identity, and irreversibility. Plan reads, apply writes, OIDC over static keys. Policy on the plan JSON. Secrets are references. Feature flags stay on the rollout pages.
Production checklist
A plan replaces one database, opens SSH to the world on a bastion, and tags a bucket. Say which line policy denies with no human, which line needs two reviewers, which role is allowed to run the apply, and where an application feature flag would live instead.