Distributed systems
Part 3 of 6 · Two-Phase Commit — Protocol, Coordinator & ParticipantsFailures & Recovery - Blocking, Coordinator Crash & Participant Crash
Coordinator and participant crashes, blocking, termination rules, and heuristic decisions.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
A participant crashes before PREPARED. What does recovery do?
Answer
Ordinary local abort. If asked, it votes NO. The coordinator will not commit.
L2
A participant crashes after PREPARED and before a decision. What is forbidden?
Answer
The usual uncommitted-transactions-abort rule. PREPARED outranks that default. Locks come back. The participant asks the coordinator.
L3
The coordinator crashes with a start record and no decision. What is the outcome?
Answer
Abort. Votes that lived only in RAM do not count. Send ABORT to anyone who might have prepared.
L4
COMMIT is durable and one participant never saw it. What changes on restart?
Answer
Nothing about the outcome. Resend COMMIT. A participant that already committed acks again and must not apply the business effects twice.
L5
All prepared participants can see each other but not the coordinator. May they commit?
Answer
No. We all voted YES is not a commit record. The coordinator log might say COMMIT or ABORT.
L6
What do cooperative termination rules actually remove?
Answer
Blocking when some reachable member already committed, aborted, or can still abort. They do not remove the all-prepared, coordinator-dead case.
L7
How does a replicated coordinator log change the blocker?
Answer
COMMIT in a quorum lets a new leader finish phase 2. You still block if you cannot get a quorum. The blocker moves from one machine to no quorum.
Failure modes
PREPARED rolled back on restart
Treating a prepared transaction like an in-progress local transaction breaks anyone who already committed.
New coordinator, old disk gone
A replacement process without the log does not create a decision. Prepared participants still block.
Partition read as a crash
Aborting because the coordinator seems down diverges from participants that already received COMMIT.
Heuristic unlock
ROLLBACK PREPARED on one side while the other side already committed is an atomicity break that needs a human and an incident id.
Misconceptions
A timeout is a decision record.
A timeout tells you a message did not arrive. It does not tell you what the other side's log contains.
If every participant voted YES, they may commit without the coordinator.
They knew that when they entered prepared. The missing piece is the decision, which may already be abort on the far side of a partition.
Heuristic commit means recovery succeeded.
It means an operator forced an outcome the log did not prove. The resource manager must remember it until the transaction manager acknowledges the hazard.
Interviewer traps
Healthy participants imply a healthy decision.
Participants can be up and still blocked because the only decision copy is on a dead disk.
Re-teaching Raft elections as the recovery protocol.
Say that a Raft group can be the quorum log behind a coordinator, then stay on the crash matrix.
Design scenario
Same prompt for every reader.
Requirements
Do not commit one side to finish the user request. Do not abort one side because a health check failed.
Traffic / scale
One prepared transaction, locks held on both resource managers.
Latency
User latency is now restart time or operator time, not a round trip.
Consistency
One outcome, matching the log if the log still exists.
Availability
Conflicting work on those keys waits.
Failure assumptions
- The coordinator disk may be intact, or it may be gone.
- One participant may already have applied COMMIT while its ack was lost.
Constraints
- No pure function from prepared to commit or abort.
- A heuristic choice needs an incident id and a reconciliation plan.
Prompt
Both participants voted YES. The coordinator process is gone. You do not know whether a COMMIT record was forced.
API
Who is allowed to answer an inquiry, and what do they answer from?
Data
Which log records make abort safe, and which make commit mandatory?
Architecture
What changes if the decision log is a quorum instead of one disk?
What recovery is allowed to do
Prefer
Replay the log
No record, or a start without a decision, aborts. A decision record is resent. PREPARED with no path to that record waits.
- Memory that was not forced is neither a vote nor a decision.
- Duplicate COMMIT is an ack, not a second business effect.
- Termination uses a peer only when that peer already knows the outcome.
Alternative
Guess, because the coordinator looks down
Each unilateral choice is safe for only one hidden decision. The disk that knew is gone, or it is on the other side of a partition.
- Health checks are not a commit record.
- A new process without the log does not create an outcome.
- Heuristic unlock is an incident, not success.
Look up the log before you touch a lock
The same timeout can be a safe abort, a resend, or a block. The record decides.
- 1
No PREPARED on the participant
Abort local work. Vote NO if anyone still asks. The coordinator is not allowed to commit this cohort. - 2
PREPARED, decision unknown
Re-acquire locks, stay prepared, and ask. Do not run the ordinary uncommitted-abort rule on this transaction. - 3
Decision record exists
Resend COMMIT or ABORT. Participants that already applied it ack again. Do not flip the outcome because an ack was lost. - 4
Everyone prepared, coordinator silent
Block. A peer protocol cannot invent the contents of a log it cannot read. A human guess is a heuristic and needs reconciliation.
Overview
2PC survives crashes only by replaying forced log records. It does not survive a missing decision. A prepared participant with no path to the decision record blocks, holding locks, until the coordinator, or a termination protocol with enough information, speaks. Logging of the happy path is on prepare and commit.
Crash matrix
| Who died | What is on their stable log | Recovery action | Others |
|---|---|---|---|
| Participant before PREPARED | No prepared record | Abort local work, vote NO if asked | Coordinator will not commit |
| Participant after PREPARED, before decision | PREPARED | Re-acquire locks, stay prepared, ask the coordinator | The cohort waits on this resource manager for acks only if the decision already exists |
| Participant after the decision is applied | Commit or abort record | Answer retries with the same ack; forget when told | Idempotent |
| Coordinator before any decision record | Start, maybe partial votes in memory only | Abort. Votes in RAM do not count. Send ABORT to anyone who might have prepared | Prepared resource managers unblock by aborting |
| Coordinator after COMMIT is forced, before all acks | COMMIT | Resend COMMIT. Do not abort | Participants that already committed ack again |
| Coordinator after ABORT is forced | ABORT | Resend ABORT | Same rule |
| Coordinator after END | Forgotten | On inquiry, presumed abort answers ABORT. Presumed commit can answer COMMIT. Know which protocol you run | See the optimizations page |
| Coordinator disk and a prepared resource manager both gone | Decision gone, PREPARED remains | Block. A heuristic decision is a human breaking atomicity on purpose | The nightmare case |
Memory that was not forced is not a vote and not a decision. Recovery code that "remembers" a YES from a heap dump is wrong.
Coordinator crash
Case A. Crash while sending PREPARE, COMMIT record absent. Some participants may be prepared. Some may not have seen PREPARE. Restart logs ABORT, which is what basic 2PC and presumed abort both do here, and broadcasts it. Participants that never prepared abort locally if they have anything open. Participants that prepared apply ABORT. Atomicity holds because nobody was allowed to commit.
Case B. Crash after COMMIT is forced, before participant B receives it. A may have committed. Restart reads COMMIT and sends COMMIT to A and B. B commits. A treats the duplicate as a success ack. The user-visible pause is the restart time, not a split outcome.
Case C. The log disk is gone and there is no replica. Prepared participants have a PREPARED record and nobody can answer. They block. Buying a new coordinator process without the log does not create a decision. That is the operational argument for putting the transaction-manager log in a quorum, which is the Paxos Commit and Raft transaction-manager discussion on tradeoffs. Raft is the quorum log behind a coordinator, not a second copy of Raft consensus.
Participant crash
Before the force of PREPARED, undo on restart is ordinary single-node recovery. The coordinator, if it is still waiting, times out and aborts. That is safe.
After PREPARED, restart must not run the usual "uncommitted transactions abort" rule on this transaction. The PREPARED record is a promise that outranks the default abort-on-recovery rule. Engines surface these as prepared transactions that survive restart: Postgres pg_prepared_xacts, MySQL XA RECOVER. Locks come back with them. If recovery treats PREPARED like an in-progress local transaction and rolls it back, you have broken 2PC.
Cooperative termination
If the coordinator is down, prepared participants can ask each other. Bernstein's rules, in interview form:
- If any cohort member has a commit record, the decision is commit. Tell the others.
- If any cohort member has an abort record, or never prepared and aborted, the decision is abort.
- If any cohort member has not voted yet, it can still abort, and then the decision is abort. It must actually abort, not merely be undecided.
- If every reachable member is only PREPARED, and the coordinator is unreachable, block. Nobody has evidence of the decision, and the dead coordinator's log might say COMMIT.
Decisions
- 1
1. Prepared, coordinator unreachable
- next2. A peer committed?
- ?
2. A peer committed?
- yes3. Commit and inform peers
- no4. A peer aborted or can abort?
- 3
3. Commit and inform peers
- ?
4. A peer aborted or can abort?
- yes5. Abort and make it durable
- no6. All prepared: block
- 5
5. Abort and make it durable
- 6
6. All prepared: block
Lesson map
Recovery asks, then blocks
The coordinator is crashed with no decision record. Both stores are locked on PREPARED.
Architecture. Coordinator Crashed. Store A Locked · YES. Store B Locked · YES
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB coordinator["Coordinator Crashed"] store_a["Store A Locked YES"] store_b["Store B Locked YES"] store_a -->|Inquiry| store_b store_b -->|PREPARED only| store_a store_a -->|Lock held| coordinator store_b -->|Lock held| coordinator coordinator -->|No decision| store_a store_a -->|Heuristic| store_b
Termination protocols reduce blocking when someone already knows the outcome. They do not remove the all-prepared, coordinator-dead case. That case is why papers say atomic commit cannot be non-blocking in an asynchronous network that can partition. 3PC tries to change the state machine so a new coordinator can decide. Under partition it can decide twice. That comparison stays on the tradeoffs page. It is not a fix you should ship from this page.
Partitions
A partition is not a crash. The coordinator may be alive on the other side and may already have committed. A prepared participant that aborts because the coordinator seems down will diverge from participants that received COMMIT. Health checks are not a decision record.
The safe partition behavior of classic 2PC is to stall the prepared side. Availability of those locked keys is what you lose. You do not get to claim the locked keys stay available and the commit stays atomic at the same time.
Heuristic decisions
When operators force a prepared branch to commit or abort without the transaction manager, XA calls that a heuristic decision.
| Name | Meaning |
|---|---|
| Heuristic commit | This resource manager committed while the transaction manager might have aborted |
| Heuristic rollback | This resource manager aborted while the transaction manager might have committed |
| Heuristic mixed | Inside one resource manager, some work was forced one way and some the other |
| Heuristic hazard | The resource manager does not know whether its heuristic was mixed |
These are not recovery success. They are a recorded atomicity break so a human can reconcile. The participant must remember the heuristic until the transaction manager acknowledges it. Silently unlocking a prepared Postgres transaction with ROLLBACK PREPARED while another database already ran COMMIT PREPARED is heuristic rollback. Do it only with a reconciliation plan and an incident id. Compensation, if you need it after the fact, is the existing saga contract on Two-Phase Commit vs Sagas, not a silent unlock.
Sandbox
After both vote YES, the coordinator dies before a decision record. Each unilateral choice is safe only for one hidden decision.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
If you can write a pure function from "I am prepared" to commit or abort, the function is wrong for one of the two coordinator logs.
What to monitor
- Age of the oldest prepared transaction. Page before lock queues and vacuum or purge stall.
- Count of prepared transactions after a coordinator deploy. A leak means the transaction manager crashed after prepare and is not replaying, or the app started prepares and never finished.
- Log-disk latency. Happy-path latency is cross-node round trip plus at least two forces, participant prepare and coordinator decision, often more.
- Heuristic counters. Any non-zero value is an incident, not a graph decoration.
Pitfalls
- Heap memory of a YES is not a vote after a crash.
- A replacement coordinator without the disk is not recovery.
- Duplicate COMMIT must be idempotent at the participant.
Interview Q&A
The coordinator crashes before writing COMMIT. One participant voted YES. The other had not answered. On restart, what is the decision?
Answer
Abort. The YES is real, so that participant must be told ABORT. The silent participant, if it later wakes prepared, must also abort. If it never prepared, it aborts locally anyway.
Why can't a prepared participant use a timeout to abort?
Answer
Because the coordinator may have force-logged COMMIT and delivered it to someone else during the partition you are interpreting as a timeout. Timeout-abort of a prepared resource manager is a heuristic rollback.
A participant committed and then the ack was lost. The coordinator retries COMMIT. What does the participant do?
Answer
Ack success again. It must not apply the business effects twice. The commit record makes the retry idempotent.
All participants are prepared and can all see each other, but not the coordinator. Can they commit?
Answer
No. The coordinator may have aborted and its ABORT record is on the other side of the partition, or it may have committed and not told them yet. "We all voted YES" is not a commit record. They already knew that when they entered prepared.
How does replicating the coordinator log change this page?
Answer
If COMMIT is committed to a quorum, a new leader reads it and finishes phase 2. You still block if you cannot get a quorum, and a prepared shard still waits until it learns the decision. You have moved the blocking condition from one machine to no quorum. Details of Raft itself stay on Raft consensus. Use it only as the quorum log behind a coordinator.
What is wrong with rolling back every uncommitted transaction on restart?
Answer
It is right for local transactions that never prepared. It is wrong for a transaction whose log contains PREPARED. That record is a promise. Engines keep those rows visible in pg_prepared_xacts or XA RECOVER and put the locks back.
A peer has an abort record and another peer is only prepared. What does termination do?
Answer
Abort. Someone already knows the outcome. Tell the prepared peer and make the abort durable there. Termination is stuck only when every reachable peer is merely prepared.
When is a heuristic decision the least-bad operational move?
Answer
When the decision disk is gone, locks are taking down the system, and a human accepts a recorded atomicity break. Compare both ledgers, compensate or repair with an incident id, and do not pretend the commit is still atomic. The saga pages, starting at saga failure modes, own reconciliation after the fact.
Draw three coordinator disks: empty except a start record, COMMIT forced, and disk missing. For each disk, write what a prepared participant must do. Only one of the three rows is "wait with locks held."
Go deeper
- Bernstein, Hadzilacos, and Goodman on termination protocols.
- Skeen, Nonblocking Commit Protocols, including why partitions still hurt.
- Gray and Lamport on why a failed transaction manager blocks 2PC, and Paxos Commit as the fix.
- Kleppmann on the coordinator crash that leaves everyone stuck.
- Postgres pg_prepared_xacts: prepared transactions survive crash and hold locks.