State, Backends, Locking & Workspaces
Terraform state maps configuration addresses to real resource IDs. This page covers remote backends, locking, encryption, and why prod usually gets its own root instead of a workspace.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does state store that git does not?
Answer
Addresses mapped to provider IDs, dependency order, cached attributes, outputs, serial, and lineage.
L2
Why is the bucket encrypted if the account is already private?
Answer
State often contains sensitive attributes. Encryption, tight IAM, and access logs are defense in depth.
L3
What breaks if two applies skip the lock?
Answer
Both read serial N and write divergent N+1. The backend has split-brain.
L4
When is force-unlock allowed?
Answer
After you prove no apply is running, you read the lock metadata, and you pull a backup. Then you plan before any new apply.
L5
Why not a workspace named prod in the same backend as dev?
Answer
A mis-selected workspace applies with the same IAM surface. Prod wants its own root, role, and usually its own account.
L6
What is a safe remote-state output?
Answer
A narrow value such as a subnet ID. Whole objects and circular reads deadlock later applies.
L7
What is lineage?
Answer
A UUID for this state stream. Copying a file across backends without care trips the lineage check, which is the signal that the streams diverged.
Failure modes
State committed to git
The file races, leaks attributes, and teaches the team that a laptop commit is the backend.
Force-unlock during a live apply
A second writer lands while the first writer still believes it holds the lease.
Workspace typo on prod
The config is shared, the blast radius is not, and apply hits the wrong serial.
Circular remote state
Root A reads root B and root B reads root A. Neither apply can finish.
Stale lease after a dead runner
The lock looks held. An old runner whose token expired must not still be able to write.
Misconceptions
State is a backup of the Terraform files.
Git is the config source of truth. The backend is the binding source of truth.
A private account means a plaintext state bucket is fine.
Sensitive attributes, a stolen laptop role, and an audit all want encryption plus access logs.
Workspaces are environments.
Workspaces are names in one state stream. Isolation of IAM and approvals is a different root.
Interviewer traps
Approving one workspace for prod and staging because the config is identical.
Identical config can still have different blast radius. Separate the prod backend and role.
Explaining KMS envelope rotation when the question was the state bucket ACL.
Say encrypt the backend and restrict state pull. Point envelope design at the Secrets and KMS page.
Design scenario
Same prompt for every reader.
Requirements
Versioned encrypted backend, a lock, separate prod root and role, and a written force-unlock check. No state files in git.
Traffic / scale
Several plans an hour on pull requests. A few prod applies a day, each holding the lock for minutes.
Latency
A contended lock fails the job immediately. Humans do not wait on a silent overwrite.
Consistency
A successful write advances serial N to N+1 only while the fencing token matches.
Availability
A dead runner leaves a lock that a human can clear after proving the apply is gone. Backend versioning can restore the previous object.
Failure assumptions
- A runner dies after acquiring the lock and before release.
- Someone selects the prod workspace by habit.
- An expired lease holder retries the state write.
Constraints
- Prod does not share an apply role with staging.
- force-unlock requires proof that no apply is running.
- Remote-state outputs stay narrow.
Prompt
A platform team stores staging and prod in one S3 backend using workspaces. Two CI jobs overlapped and the state serial went backwards. A developer wants to force-unlock and keep going.
API
Which commands read lock metadata, pull a backup, and clear a stale lock?
Data
What do serial and lineage protect?
Architecture
Where do the bucket, the lock table, and the prod apply role sit?
Prod and staging share one configuration shape and very different blast radius
Prefer
Separate roots, backends, and apply roles
Each environment has its own serial, its own lock, and its own IAM boundary.
- A staging apply cannot select prod by changing a workspace name.
- The prod bucket policy can exclude the pull-request plan role from writes.
- State files stay small enough to review a destroy.
Alternative
One backend and a workspace switch
The same credentials and the same bucket hold every environment.
- A wrong workspace is a legendary outage class.
- Lock contention couples unrelated applies.
- A state pull in CI prints prod attributes into a shared log.
One writer advances the serial
The diagram is the same lease. An expired holder must not still be able to write.
- 1
Read serial N
The job loads the remote object and remembers the serial and lineage. - 2
Acquire the lease
The lock records who, when, and a fencing token. If it is held, fail fast. - 3
Refresh and write N+1
The backend accepts the write only when the token and the serial still match. - 4
Release, or prove the holder is dead
A stuck lock is cleared only after CI, the audit log, and a teammate agree the apply is gone.
Overview
Lose state integrity and you get duplicate resources, phantom deletes, or an apply that cannot proceed. This page is the operational contract for the backend: what the file holds, how the lock works, and how environments are separated.
The hub chose the tool. Here the binding has to stay boring.
What the file actually holds
- Resource addresses mapped to provider IDs, such as
aws_instance.apptoi-0abc. - Dependency edges used to order destroy.
- An attribute cache. Some providers write sensitive values back into that cache.
- A serial and a lineage so two writers cannot both claim the next version.
- Outputs that other roots may read. That read is another coupling surface.
Git remains the config source of truth. The backend is the binding source of truth.
Backends teams actually run
| Backend | Lock | What you still configure |
|---|---|---|
| S3, with DynamoDB or a native S3 lock | Yes | Versioning, encryption, block public access, a narrow IAM policy, access logs |
| GCS | Object generation or a lock | CMEK when the org requires a customer-managed key |
| Azure Blob | Blob leases | Azure AD auth instead of a static key |
| Terraform Cloud or Enterprise | Built-in | Remote runs so state need not sit on a CI disk |
| HTTP, Consul, and similar | Varies | An owner. An unowned backend is a future incident |
The AWS-shaped baseline is a versioned encrypted bucket, public access blocked, a lock, IAM limited to CI roles plus break-glass admins, and access logging.
The lock
Decisions
- 1
1. Read serial N
- next2. Lock free?
- ?
2. Lock free?
- no3. Fail the job
- yes4. Refresh and apply
- 3
3. Fail the job
- 4
4. Refresh and apply
- next5. Write serial N+1
- 5
5. Write serial N+1
- next6. Release the lock
- 6
6. Release the lock
Lesson map
State, Backends, Locking & Workspaces
Terraform state maps configuration addresses to real resource IDs. This page covers remote backends, locking, encryption, and why prod usually gets its own root instead of a workspace.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB read["1. Read serial N"] held["2. Lock free?"] fail["3. Fail the job"] run["4. Refresh and apply"] read -->|1. Read serial N to 2. Lock free?| held held -->|no| fail held -->|yes| run
Without the lock, two applies read serial N and write two different N+1 objects. That is split-brain state.
Stale locks
Locks stick when a runner dies mid-apply. The recovery order is:
- Confirm no apply is actually running. Check the CI UI, the cloud audit log, and a teammate.
- Read the lock metadata: who, which operation, and when.
- Pull a state backup, or rely on backend versioning you have already tested.
- Only then run
terraform force-unlockwith the lock ID, or the Terraform Cloud equivalent. - Plan before any new apply.
Force-unlock during a live apply is how you corrupt state. Treat it like killing a database migration.
Workspaces versus roots
| Approach | Gain | Cost |
|---|---|---|
| Workspaces such as dev, stage, and prod | One config, fast switching | Easy to point at the wrong name. Shared IAM. Weak isolation |
| A directory or root per environment | Clear blast radius and different roles | More boilerplate, so modules have to carry the repeated shape |
| A root per slice: network, data, app | Smaller state and parallel applies | Cross-root contracts via remote state or a published output |
The senior default is a separate root, and often a separate account, for prod. Workspaces remain reasonable for many short-lived previews that share one blast radius.
Decisions
- ?
1. Prod isolation required?
- no2. Workspace is enough
- yes3. Separate state root
- 2
2. Workspace is enough
- 3
3. Separate state root
- next4. Separate account and role
- 4
4. Separate account and role
Readers of remote state
terraform_remote_state, or a published output in a registry or parameter store, couples roots. Publish narrow outputs such as subnet IDs, not entire objects. A cycle, where A reads B and B reads A, deadlocks both applies.
Encryption of the values themselves, rotation, and envelope layout live on Secrets and KMS. This page only requires that the backend is encrypted and that state pull is rare and audited.
A lease with a fencing token
Backends differ. The sketch is the idea: after expiry, the old holder cannot write.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Fat state versus many small states
| One fat state | Many small states |
|---|---|
| A single apply sees the whole graph | Plans are faster and destroys are smaller |
| Simple while the org is tiny | Roots need an output contract |
| Slow plans and scary destroys | A rotten contract leaves orphaned dependencies |
Interview Q&A
Why encrypt the state bucket if the account is already private?
Answer
Defense in depth and compliance. State often contains sensitive attributes. Pair bucket encryption with tight IAM, access logs, and a customer-managed key when duties must be separated.
Can I commit terraform.tfstate for staging only?
Answer
No. It races, it leaks, and it trains the next environment to copy the habit. Staging gets a remote backend too.
Would you approve one workspace for prod and staging?
Answer
Only when isolation requirements are genuinely weak. The interview-safe design is a separate account and backend for prod. A mis-selected workspace plus apply is a known outage class.
What is state lineage?
Answer
A UUID that identifies this state stream. Copying a file across backends incorrectly trips lineage checks. That failure is a feature: the streams are not the same history.
Two applies both read serial N. What did you forget?
Answer
The lock. Both writers believe they may publish N+1. The second upload wins or the object is corrupt, and Terraform's view no longer matches the cloud.
A runner died and the lock is stuck. What are the steps?
Answer
Prove the apply is not running. Read who held the lock. Pull a backup or confirm versioning. Force-unlock. Plan before the next apply.
How should root A pass a subnet ID to root B?
Answer
A narrow output, read via remote state or a published parameter. Do not export the whole VPC object. Do not let B's state be an input to A.
Who may pull state?
Answer
The plan role can read. The apply role can write. Humans are break-glass. CI logs should not print sensitive values. Scrubbing lives with the blast-radius page. This page decides the ACL.
Pitfalls
.gitignorethat forgets*.tfstate*and a backup file that still gets committed.- A plan role that can also
PutObjecton the state bucket. - Remote-state outputs that include entire resource objects, so every consumer breaks when an attribute is renamed.
- Treating backend versioning as optional. It is the restore path for a bad serial.
- Explaining envelope encryption in depth when the question was the lock.
Operational checklist
Two jobs read serial 14. One holds a lock token that expires in 60 seconds and then crashes. The other starts at second 70. Say which write succeeds, which token must be rejected, and what you check before force-unlock if the first job is actually still refreshing.