Distributed systems
Part 2 of 6 · Two-Phase Commit — Protocol, Coordinator & ParticipantsPrepare & Commit - Votes, Logging & the Decision
Prepare and commit votes, force-before-send logging, and the decision record.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why is prepare a separate phase?
Answer
So every updater has promised that both commit and abort are still possible before anyone is told the outcome.
L2
What does a READ-ONLY vote mean?
Answer
The participant has no redo. It is not a soft YES, and it is left out of phase 2.
L3
Which record must name the participants?
Answer
The coordinator start record, or the collecting record under presumed commit. Recovery has to know who may be prepared.
L4
Is waiting for acks a third phase?
Answer
No. Acks do not change the decision. They let the coordinator forget log records. 3PC adds a pre-commit round that changes what participants may assume.
L5
A YES packet is lost and the coordinator aborts. Is that already inconsistency?
Answer
Not if the prepared participant later applies ABORT. Inconsistency starts if that participant heuristically commits.
L6
What counts as a force?
Answer
A write that survives process crash and power loss, such as a WAL fsync or a quorum log commit. A write into the page cache is not a force.
L7
When is one-phase commit legal?
Answer
When there is a single participant. That resource manager's local commit is the global commit. It is not 2PC with a step skipped across two resource managers.
Failure modes
COMMIT sent before the force
A crash recovers as no decision and aborts, while a fast participant already committed.
Participant names only in RAM
Restart forgets a participant that is holding locks, or fails to send COMMIT to one of them.
Prepared work aborted by a local timer
That timer is a heuristic decision. It can break atomicity if the coordinator already committed.
READ-ONLY treated as a veto
A read-only vote does not stop a commit of the updaters. It only drops that participant from phase 2.
Misconceptions
NO and YES are symmetric, so both can be forgotten immediately.
NO already applied the only legal local outcome. YES promised to accept either outcome, so PREPARED stays until the global decision arrives.
Group commit lets you send YES before that transaction's record is flushed.
Many PREPARED records may share one fsync. That transaction's record must be in the flushed batch before its YES is sent.
A lost ack means the decision is unknown.
Acks can be lost. The coordinator retries the same decision until END is logged. Participants answer from the log.
Interviewer traps
Calling the ack wait three-phase commit.
3PC changes the state machine with pre-commit. Acks only garbage-collect the log.
Using a cloud disk ack as if it were a quorum.
A volume that acknowledges before replication can lose the record with the machine. A serious transaction manager puts the decision in a quorum.
Design scenario
Same prompt for every reader.
Requirements
Stock is already aborted. Orders must not stay committed. A retransmitted ABORT is a no-op on stock.
Traffic / scale
One distributed transaction, two votes.
Latency
One force on the YES path, one force for the ABORT decision, then a message to the prepared participant.
Consistency
The cohort aborts. No updater commits.
Availability
Orders holds locks only until ABORT arrives.
Failure assumptions
- The NO packet and the YES packet can arrive in either order.
- The coordinator can crash after the ABORT force and before orders applies it.
Constraints
- Do not commit orders because its vote was YES.
- Do not ask stock to prepare again after it has forgotten the transaction.
Prompt
Two participants. Orders can commit. Stock hits a constraint and must vote NO. The coordinator has not written a decision yet.
API
Which participants receive ABORT, and which may ignore a duplicate?
Data
What is on the coordinator log before ABORT is sent?
Architecture
Where do participant names live so restart can still find orders?
One phase or two
Prefer
Prepare, then one decision
Participants promise both outcomes are still possible. Only then does the coordinator pick one outcome for everyone.
- YES is a durable promise, not a guess.
- NO has already aborted and can be forgotten.
- The decision record outranks a late vote.
Alternative
Please commit, and hope the other side can
If A is told to commit before B votes, there is no distributed undo when B cannot.
- Legal only for a single participant.
- The X/Open transaction manager uses that one-phase shortcut.
- Skipping prepare across two resource managers is a bug.
Log order on the YES path and the NO path
Basic 2PC, before presumed-abort shortcuts. Each force completes before the message that depends on it.
- 1
Coordinator records the names
Force a start record that names everyone you will ask. If you crash and the names are not durable, you cannot finish. - 2
YES voters force PREPARED
Finish the local work. Force transaction id, locks or write-set, undo, and redo. Only then send YES. - 3
NO voters abort first
Write abort, undo, release locks, send NO, and forget. A later ABORT is a no-op. - 4
Seal one decision
All required votes YES: force COMMIT, then send it. Any NO or timeout while no COMMIT record exists: force ABORT and tell anyone who might have prepared.
Overview
Phase 1 collects a vote that is safe to rely on later. Phase 2 applies the single decision the coordinator has force-logged. This page is the message sequence, the vote vocabulary, and the exact log records. Crash cases live on failures and recovery. Presumed-abort shortcuts live on optimizations.
Vote vocabulary
| Vote | Local state before the reply | Log | Locks | In phase 2? |
|---|---|---|---|---|
| YES | Can commit and can abort | Force PREPARED, with undo and redo | Held until the decision | Yes |
| NO | Already aborted | Abort record, then forget | Released | No. The coordinator must abort the global transaction if it has not committed |
| READ-ONLY | No updates to redo | Often nothing under presumed abort | Released at the end of prepare, with the isolation caveat below | No |
READ-ONLY is not a soft YES. The participant has no redo to apply and must not be in the commit set. If it still holds read locks that serializability depends on, dropping it from phase 2 early can admit anomalies. Engines that use snapshot isolation often let read-only participants exit because their snapshot is already fixed. Know which isolation story you are in before you copy the optimization.
Logging rule
Participant, YES path
- Finish the local work in the transaction.
- Force a PREPARED record: transaction id, locks or write-set, undo, redo.
- Only then send YES.
- On COMMIT: write commit (often forced), release locks, ack.
- On ABORT: apply undo, write abort, release locks, ack.
Participant, NO path
- Force or write abort, undo, and release.
- Send NO.
- Forget. A retransmission of ABORT is a no-op.
Coordinator
- Force a start record that names the participants you will ask.
- Send PREPARE.
- Collect votes until the set is complete or a timeout fires.
- All YES: force COMMIT listing who must be told, then send COMMIT.
- Any NO or timeout while no COMMIT record exists: force ABORT, and send ABORT to anyone who might have prepared. Skip participants who answered NO and already forgot, or send it anyway. It is a no-op for them.
- Wait for acks from participants that were prepared. Then write END and forget.
If you send COMMIT before the COMMIT record hits stable storage, a crash lets you recover as "no decision" and abort, while a fast participant already committed. That is the bug the force-before-send rule exists to forbid.
Flow
- 1
1. Force start, names included
- next2. PREPARE both participants
- 2
2. PREPARE both participants
- next3. A force-logs PREPARED, votes YES
- 3
3. A force-logs PREPARED, votes YES
- next4. B aborts locally and votes NO
- 4
4. B aborts locally and votes NO
- next5. Coordinator force-logs ABORT
- 5
5. Coordinator force-logs ABORT
- next6. A undoes and releases locks
- 6
6. A undoes and releases locks
Lesson map
Votes and the decision record
Store A is locked on YES. Store B has already aborted and voted NO. The coordinator still has only a start record.
Architecture. Coordinator Collecting. Store A Locked · YES. Store B Aborted · NO
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB coordinator["Coordinator Collecting"] store_a["Store A Locked YES"] store_b["Store B Aborted NO"] coordinator -->|PREPARE| store_a coordinator -->|PREPARE| store_b store_a -->|YES| coordinator store_b -->|NO| coordinator coordinator -->|ABORT| store_a store_a -->|ACK| coordinator store_a -->|Lock held| coordinator coordinator -->|Early COMMIT| store_a
The decision record is the truth
After the coordinator force-logs COMMIT, the transaction is committed even if every participant is currently unreachable. Recovery's job is to drive them to that record, not to reopen the vote. After ABORT is force-logged, the transaction is aborted even if a late YES arrives. Votes do not outrank a durable decision.
Before either record exists:
- The coordinator may still abort, on a timeout or a NO.
- It must not commit unless every required vote is YES.
- Participants that have not voted may still abort locally: crash before PREPARED, deadlock victim, or constraint failure.
Idempotence matters. COMMIT delivered twice is still commit. ABORT delivered to a NO voter is a no-op. Acks can be lost. The coordinator retries until END is logged. Participants should answer from their log, not from memory of the socket.
What stable has to mean
A force is a write that survives process crash and power loss: a WAL fsync, or a replicated log commit in a quorum. write() into the kernel page cache is not a force. A cloud disk that acknowledges before replication might still lose the record together with the machine. That is why a serious transaction manager puts the decision in a quorum. The tradeoff is on 2PC versus consensus commit. Raft, when you use it here, is only the quorum log behind a coordinator. The election lesson stays on Raft consensus.
Group commit is allowed. Many PREPARED records can share one fsync. The order constraint is per transaction: that transaction's PREPARED must be in a flushed batch before its YES is sent. Batches do not let you reorder a vote ahead of its record.
Timeout placement
| Where the timer fires | Legal outcome |
|---|---|
| Coordinator waiting for votes, no decision record yet | Abort. Tell anyone who already said YES |
| Coordinator waiting for acks, decision record exists | Resend the same decision. Do not flip it |
| Participant waiting for PREPARE | Abort locally. You never promised |
| Participant prepared, waiting for the decision | Block. Ask again. Do not flip |
The last row is why phase 2 is a wait, not a local timer. A local timer that aborts prepared work is a heuristic decision. It is an operational escape hatch on the XA page, and it can break atomicity.
Sandbox
Phase-1 votes become a phase-2 decision. An existing decision record is sticky. None means no record yet.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
A read-only cohort can commit with nothing to redo. Presumed abort often writes no commit record for that case. The optimization page owns the presumption. This function only shows that read-only is not a veto, and that a durable decision ignores a later timeout.
Pitfalls
- Logging the participant list only in memory leaks a prepared resource manager after restart.
- Releasing read locks at prepare time can break serializability even when atomicity holds. Say which isolation level you mean.
- One-phase commit is a single-participant optimization. It is not optimism across two resource managers.
Interview Q&A
Can the coordinator commit if one participant is read-only and the others vote YES?
Answer
Yes. Read-only is not a veto. That participant is omitted from phase 2. The COMMIT record lists only participants that must apply redo.
A YES packet is lost and the coordinator times out and aborts. The participant is prepared. Is the system inconsistent?
Answer
Not if the coordinator reaches that participant with ABORT and the participant obeys. The participant is blocked only while ABORT is unknown. It is not free to commit. Inconsistency starts if the participant heuristically commits.
Why log participant names in the coordinator start record?
Answer
Recovery after a crash needs to know who may be prepared. If the name list was only in RAM, a restarted coordinator can forget a participant that is holding locks forever, or fail to send COMMIT to one of them.
Is the ack phase a third phase?
Answer
No. Acks do not change the decision. They let the coordinator garbage-collect log records. 3PC adds a pre-commit round that changes what participants are allowed to assume. Do not call "wait for acks" three-phase commit. The contrast is on tradeoffs.
Why can NO be forgotten immediately but YES cannot?
Answer
NO already applied the only legal local outcome, abort. YES promised to accept either outcome, so the participant must retain PREPARED until the global decision arrives.
Does a page-cache write count as a force?
Answer
No. A force survives process crash and power loss. Group commit may batch several records into one fsync. It may not send a vote ahead of that transaction's record.
Where does a participant timer become illegal?
Answer
Before PREPARE, aborting locally is safe because you never promised. After PREPARED, a timer that aborts is a heuristic rollback. The coordinator may already have force-logged COMMIT and delivered it to someone else.
A late YES arrives after ABORT is durable. What changes?
Answer
Nothing. The decision record wins. The late vote does not reopen the ballot. Send or resend ABORT if that participant might be prepared.
Write the six coordinator steps on one card. Cross out any step that sends a message before the matching force returns. Then add a timeout to each of the four rows in the timeout table and say the only legal outcome.
Go deeper
- Gray, Notes on Data Base Operating Systems on logging and commit.
- Bernstein, Hadzilacos, and Goodman on atomic commitment.
- Fowler, Two-Phase Commit.
- Kleppmann, two-phase commit and CMU 15-445.