Infrastructure as Code — Terraform State, Modules & Safe Change
Infrastructure as code turns console clicks into reviewable change with a known blast radius. This hub is the Terraform decision map for remote state, modules, saved plans, drift, and policy, with a light contrast to Pulumi, CloudFormation, and Crossplane.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What three things does a Terraform plan compare?
Answer
Configuration in git, remote state, and refreshed reality from the provider APIs.
L2
When is local state acceptable?
Answer
Solo prototypes and throwaway sandboxes. Any shared environment needs a remote backend and a lock.
L3
Workspaces or separate roots for prod versus staging?
Answer
Separate roots, often separate accounts, when blast radius and approval differ. Workspaces fit many identical short-lived previews.
L4
Why apply a saved plan file instead of running apply with no file?
Answer
A bare apply can re-plan and pick up a merge or a drift that landed after review.
L5
Where do secret values belong?
Answer
In a secret store. Terraform should create the container and an IAM binding. The value is not a committed tfvars string.
L6
Terraform or Crossplane for the same cluster add-on?
Answer
Terraform bootstraps the cluster, IAM, and network. Crossplane reconciles in-cluster if the platform team owns Kubernetes as the control plane. One owner per cloud object.
L7
What is the production baseline before the first shared apply?
Answer
Remote state with locking, plan and apply roles split, a saved plan artifact, module pins, encrypted state, and a drift plan on a read-only role.
Failure modes
Local state for a shared environment
Two laptops write two truths. The next apply creates duplicates or deletes the wrong object.
Apply that re-plans
Reviewers approved one diff. The apply job computed another after a merge or a console change.
Secrets in tfvars and in state
Anyone who can read the backend or the CI log can read the credential.
Three tools owning one VPC
Terraform, CloudFormation, and Crossplane fight over the same IDs and the plan never stays empty.
Floating module main
A provider or module commit on main replaces prod resources the pull request never showed.
Misconceptions
Infrastructure as code means the YAML is the system.
The system is config plus state plus the cloud. A plan that ignores state will lie.
ClickOps is always forbidden.
Break-glass with a ticket, time-boxed credentials, and a same-day import is a valid incident path. Steady-state ClickOps is the smell.
State is a backup of the .tf files.
Git is the config source of truth. The backend is the binding from addresses to real IDs.
Crossplane replaces Terraform for the whole company on day one.
Crossplane wants a solid Kubernetes control plane and a single owner per cloud object.
Interviewer traps
Recommending workspaces as the isolation boundary between prod and staging.
Say separate roots and accounts when IAM and approvals differ. Workspaces are a soft name in one backend.
Redesigning the CI pipeline, the VPC, or the KMS envelope when the question was the plan.
Name the neighboring cluster in one sentence. Stay on state, the graph, modules, and the saved plan.
Design scenario
Same prompt for every reader.
Requirements
One owner per cloud object. Remote state with locking. Prod applies use a reviewed plan file. Secrets are references. Module versions are pinned.
Traffic / scale
A handful of pull requests a day across network, data, and app roots. Prod applies are gated. Dev roots may auto-apply.
Latency
A plan comment lands on the pull request before review. Prod apply waits for the saved artifact and an approval, not for a laptop.
Consistency
The next plan after a successful apply is empty, except for fields an owner marked ignore_changes.
Availability
A stuck lock is cleared only after proving no apply is running. Break-glass changes are imported within a day.
Failure assumptions
- Two applies can read the same state serial if locking is off.
- A console change lands between plan and apply.
- A module ref on main moves under a prod root.
Constraints
- Do not store prod state on a laptop.
- Do not let Terraform and Crossplane both own the VPC.
- Do not commit secret values.
Prompt
A payments team provisions AWS and one SaaS vendor. Staging drifted from prod after a Friday console change. The platform group also runs Kubernetes and is considering Crossplane for the same network.
API
What does the pull request check publish, and what does the apply job consume?
Data
What lives in state, and who is allowed to read it?
Architecture
Where do the remote backend, the module pins, and the prod apply role sit?
A multi-cloud team wants reviewable infrastructure, and a platform group already lives in Kubernetes
Prefer
Terraform for provisioning, one owner per object
Pull requests carry a plan. Remote state is locked. Crossplane is a later control plane only for objects the platform team agrees to own.
- The plan JSON is the review artifact across AWS, GCP, and SaaS providers.
- Module and provider versions are pinned, including the lockfile.
- Prod apply uses the saved plan and a narrower role than plan.
- Break-glass is ticketed and imported the same day.
Alternative
ClickOps plus three tools on the same VPC
The console, a CloudFormation stack, and a Crossplane claim all try to be the source of truth.
- Staging and prod diverge with no diff to read.
- The next plan wants to delete what another tool just created.
- Secrets land in shell history and in a laptop state file.
From a desired change to a gated apply
The diagrams below are the same loop. Pick the tool first, then refuse a path that cannot be reviewed.
- 1
Pick one owner
Terraform for portable provisioning. Pulumi when the team already ships a language runtime in CI. CloudFormation when the org is AWS-only and already invested. Crossplane when Kubernetes is the control plane. - 2
Bind config to remote state
Git holds the .tf files. A locked backend holds IDs, dependencies, and attribute cache. - 3
Publish the plan artifact
fmt, validate, and a saved plan. Reviewers read replaces and destroys before merge. - 4
Apply those bytes
The prod job applies the reviewed file. A later drift plan catches console edits.
Overview
Infrastructure as code is how a senior team turns "someone clicked in the console" into change you can review, repeat, and bound. This hub centers on Terraform, the usual multi-cloud default, and says when Pulumi, CloudFormation, or Crossplane is the better fit.
Interviews probe whether you treat infrastructure as software with state:
- Can you contrast the four tools without reciting marketing?
- Can you explain remote state, locks, and workspaces?
- Can you read a dependency graph and a replace?
- Can you version a module and apply a saved plan?
- Can you keep secrets and blast radius inside a policy?
Ask this out loud: two engineers apply the same root at once, and a third person edited a security group in the console an hour ago. What does each of them think is true?
Why teams adopt it
Without a declared source of truth you inherit:
- Snowflake environments. Staging is not prod, so the incident only exists in one account.
- Unreviewable change. A console click has no pull request and no plan artifact.
- Secret sprawl. Access keys land in local tfvars and then in git.
- Friday edits. Someone "just updated the security group" and locked the fleet out.
- Audit gaps. Compliance asks who changed the subnet group, and the answer is a Slack shrug.
Tool choice
| Dimension | Terraform (HCL) | Pulumi | CloudFormation or CDK | Crossplane |
|---|---|---|---|---|
| Language | HCL and JSON | TypeScript, Python, Go, and others | YAML, JSON, or CDK | Kubernetes CRDs |
| Multi-cloud | Provider ecosystem | Provider ecosystem | AWS-first | Providers on a Kubernetes control plane |
| State | Explicit file plus backend | Service or a DIY backend | Stacks managed by AWS | Desired state in etcd |
| Drift | Plan and refresh | Preview and refresh | Stack drift detection | Reconcile loop |
| Review surface | Plan JSON and policy | Preview and policy packs | Change sets and stack policies | Compositions, RBAC, sync waves |
| Best fit | Portable, PR-driven provisioning | App engineers who want real libraries | Deep AWS shops already on CFN | Platform teams whose product is Kubernetes |
| Step away when | You need a continuous reconciler as the control plane | CI cannot standardize one language runtime | You are multi-cloud and refuse AWS lock-in | You do not already run a solid Kubernetes control plane |
Pick Terraform for portable pull-request provisioning. Pick Pulumi when abstractions should be ordinary libraries. Pick CloudFormation or CDK for an AWS-only org whose guardrails already assume stacks. Pick Crossplane when the platform product is Kubernetes and you want continuous reconciliation. One cloud object has one owner.
Decisions
- ?
1. K8s control plane?
- yes2. Crossplane
- no3. AWS-only CFN shop?
- 2
2. Crossplane
- ?
3. AWS-only CFN shop?
- yes4. CloudFormation or CDK
- no5. Real language in CI?
- 4
4. CloudFormation or CDK
- ?
5. Real language in CI?
- yes6. Pulumi
- no7. Terraform
- 6
6. Pulumi
- 7
7. Terraform
Lesson map
Infrastructure as Code — Terraform State, Modules & Safe Change
Infrastructure as code turns console clicks into reviewable change with a known blast radius. This hub is the Terraform decision map for remote state, modules, saved plans, drift, and policy, with a light contrast to Pulumi, CloudFormation, and Crossplane.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB cp["1. K8s control plane?"] xp["2. Crossplane"] aws["3. AWS-only CFN shop?"] cfn["4. CloudFormation or CDK"] cp -->|yes| xp cp -->|no| aws aws -->|yes| cfn
When a console click is still honest
Decisions
- 1
1. Change infrastructure
- next2. Will it repeat?
- ?
2. Will it repeat?
- yes3. Model it in IaC
- no4. Outage right now?
- 3
3. Model it in IaC
- next7. Secrets in play?
- ?
4. Outage right now?
- yes5. Break-glass with a ticket
- no3. Model it in IaC
- 5
5. Break-glass with a ticket
- next6. Import within a day
- 6
6. Import within a day
- ?
7. Secrets in play?
- yes8. Store a reference
- no9. Plan then gated apply
- 8
8. Store a reference
- next9. Plan then gated apply
- 9
9. Plan then gated apply
- next10. Alert on real drift
- 10
10. Alert on real drift
Break-glass is acceptable when the blast radius is an active outage, the credentials are time-boxed, and someone owns the backfill. It is a smell when it becomes the delivery path, when prod and staging diverge in silence, or when the import ticket never exists.
Cluster map
| Page | You should be able to | Ring |
|---|---|---|
| Hub (this) | Choose a tool and describe safe change | Blast radius to State |
| State | Lock a remote backend and place prod in its own root | Hub to Resources |
| Resources | Read the graph, lifecycle, and for_each | State to Modules |
| Modules | Pin a small interface and ship a version | Resources to Plan |
| Plan and import | Apply the reviewed file and adopt drift | Modules to Blast radius |
| Blast radius | Gate destroys, roles, and secret values | Plan to Hub |
Config, state, reality
Flow
- 1
1. Config in git
- next4. Plan the diff
- 2
4. Plan the diff
- next5. Apply the saved file
- 3
2. Remote state
- next4. Plan the diff
- 4
3. Cloud reality
- next4. Plan the diff
- 5
5. Apply the saved file
terraform plan refreshes reality, diffs config against state, and proposes create, update, or delete. terraform apply should execute that plan file. If state lies, the plan lies. If the lock is missing, two applies race. If secrets sit in state, every backend reader is a secret reader.
Production baseline
- Remote state and a lock. S3 with a lock table or native lock, GCS, Azure Blob leases, or Terraform Cloud. Shared environments do not use a laptop file.
- Split roles. The plan role can read. The apply role can change one environment. Prefer OIDC to long-lived keys.
- Plan artifacts. Upload the plan and a JSON summary on the pull request. Apply uses that file.
- Apply gates. Required reviewers for destroy and replace. Environment protection on prod.
- Pins. Provider constraints, a committed lockfile, and a module version or git SHA. Prod does not track
main. - Secret references. Terraform creates the container. Values arrive through the secret store. Encrypt the backend anyway, because providers sometimes write sensitive attributes back.
- Drift alerts. A scheduled read-only plan pages a human when a critical root is dirty.
Dependency order, in miniature
Terraform topologically sorts a DAG. Implicit edges come from attribute references such as subnet_id = aws_subnet.a.id. Explicit depends_on is for couplings the references cannot see, such as IAM's eventual consistency. The full graph lesson is next after state.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
What stays on neighboring pages
A pipeline builds and deploys what this cluster provisions. Stage graphs, digests, and supply-chain checks live on CI/CD pipelines. Workload rollouts live on Kubernetes workloads and rolling, blue-green, and canary. History hygiene for a .tf pull request lives on Git. Envelope encryption and secret stores live on Secrets and KMS and secret stores. The VPC, NAT, and ingress those modules emit live on private networking.
Interview Q&A
Local state versus remote state. When is local acceptable?
Answer
Solo prototypes and throwaway sandboxes. Any shared environment, including staging, needs a remote backend and a lock. Local state invites an overwrite race and makes a laptop the source of truth.
Workspaces versus separate directories?
Answer
Workspaces are a soft namespace in one backend. Prefer separate directories, state files, and often accounts when prod and non-prod differ in blast radius, IAM, or approval. Workspaces fit many similar ephemeral environments with the same blast radius.
Terraform versus Crossplane for the same cluster add-ons?
Answer
Terraform for bootstrap: cluster, IAM, and network. Crossplane or an operator for in-cluster reconcile when the platform team owns Kubernetes as the control plane. Two controllers for one AWS object is a footgun.
Where do secrets go?
Answer
Not in git, and not in plaintext tfvars that get committed. Create the secret container and the IAM binding with Terraform. Put values in the secret store. Encrypt the backend and limit who can pull state. The Secrets and KMS pages own envelope encryption.
Why can a plan be green locally and noisy in CI?
Answer
Different provider builds, different variable files, a different workspace, or a missing lockfile. Align the backend, the var files, and .terraform.lock.hcl before you debug the diff.
What does apply without a plan file risk?
Answer
It may re-plan. A merge or a cloud change that landed after review becomes part of the apply. Prod should apply the exact artifact reviewers saw.
Is ClickOps ever the right call?
Answer
During a prod outage, with logged time-boxed credentials and a ticket. The follow-up is an import or a deliberate recreate inside a day. ClickOps as the weekly delivery path is the failure mode.
What should you say in the first minute?
Answer
Config, state, and reality. One owner per object. Remote state with a lock. Pinned modules. A saved plan. Split plan and apply roles. Secret references. Drift on a schedule. Neighboring pages own pipelines, rollouts, Git, KMS, and the VPC.
Pitfalls
- Treating HCL as the whole system and forgetting the state file.
- Running Terraform, CloudFormation, and Crossplane against one VPC.
- Tracking a module
ref=mainfrom a prod root. - Using one admin cloud role for both the pull-request plan and the prod apply.
- Explaining pipeline stages or KMS envelope details when the question was the plan artifact.
A staff engineer asks how you would stop a Friday console edit from becoming the architecture. Name the three pictures of the world, the artifact reviewers must see, the role split, and which neighboring page owns the pipeline versus the VPC.