Distributed systems
Part 6 of 6 · Two-Phase Commit — Protocol, Coordinator & Participants2PC vs 3PC vs Consensus Commit vs Sagas
Why 3PC splits under partition, Paxos Commit as F>0 2PC, and when to use the existing saga series.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does classic 2PC do under a partition?
Answer
Prepared participants stall. They do not commit on one side and abort on the other.
L2
Why can 3PC lose atomicity?
Answer
One side can see pre-commit and commit, while the other side never sees pre-commit, times out, and aborts.
L3
What is Paxos Commit with F=0?
Answer
Two-phase commit. One acceptor, the coordinator. If that acceptor is down, prepared participants wait.
L4
What does just use Raft fail to do?
Answer
It does not atomically commit two independent resource managers. Either both writes share one consensus group, or you run 2PC across groups and store the decision in one of them.
L5
What does consensus commit still block on?
Answer
Loss of the quorum that knows the decision. A prepared shard must not guess. You moved the blocker from one host to a quorum loss.
L6
When is a saga the right replacement?
Answer
A participant cannot prepare, lock time across a round trip is unacceptable, or the product accepts pending and undo as visible states.
L7
A design uses 2PC between orders and a card API. What do you change?
Answer
The card API will not sit in prepare holding your locks. Make the charge a local idempotent step and the rest a saga. Point at the existing comparison page.
Failure modes
3PC under partition
Pre-commit reaches one side only. That side commits. The other side aborts. Atomicity is gone.
Electing a new transaction manager by timeout
Two processes can both decide they are the new coordinator. Electing one is itself consensus, done badly.
One Raft group described as a substitute for cross-shard commit
A single group orders one log. It does not prepare two independent resource managers.
Saga used where partial commit is an integrity incident
Compensations are visible and fallible. Finance that needs all-or-nothing still wants atomic commit inside one system.
Misconceptions
3PC is strictly safer than 2PC.
It aims to be non-blocking for fail-stop crashes. Under partition, 2PC preserves atomicity by stalling. 3PC can commit and abort.
Paxos Commit is a different business protocol.
It is fault-tolerant commit. Classic 2PC is the special case with one acceptor.
Multi-shard databases that use Raft never block.
A prepared shard that cannot learn the decision from a quorum must not guess. Consensus removed the single-process coordinator, not the uncertainty window.
Interviewer traps
Proposing 3PC as the production fix for 2PC blocks.
Know the third phase, then refuse it. The production fix for a single coordinator disk is a replicated decision, or a saga if atomic commit was the wrong contract.
Restating orchestration, outbox, and idempotency keys on this page.
Name the boundary and link the existing saga series. Those pages already own the long form.
Design scenario
Same prompt for every reader.
Requirements
The ledger pair is atomic or it does not happen. The card charge must not hold database locks. A single coordinator machine must not be the only copy of a shard decision.
Traffic / scale
Shard transactions are synchronous. The card step can be asynchronous.
Latency
Cross-shard commit may pay a quorum round trip. The card call must not sit inside prepare.
Consistency
Shard pair is atomic commit. Card plus shipment is eventual, with compensations.
Availability
A shard quorum loss blocks that prepared transaction. The card path stays up.
Failure assumptions
- One coordinator process can die after prepare.
- The network can partition the cohort.
- The card API cannot prepare.
Constraints
- Do not propose 3PC.
- Do not put the card API inside XA.
Prompt
Orders and inventory are two shards you operate. A card charge is an HTTP API. Partial inventory decrement is acceptable only with an explicit undo. A partial ledger pair is not.
API
Which boundary is 2PC, and which boundary is a saga step?
Data
Where is the cross-shard decision record stored?
Architecture
What still blocks if the decision log is a quorum?
The repair has to match the failure
Prefer
Replicate the decision, or stop promising atomic commit
If you need all-or-nothing and the coordinator must survive one machine, put the decision in a quorum. If a participant cannot prepare, write a saga.
- 2PC inside one data system is a yes when locks fit the budget.
- Paxos Commit, or 2PC on a Raft transaction-manager log, keeps one outcome while a quorum is up.
- Sagas make intermediate states visible on purpose.
Alternative
Add a phase and hope partitions behave
3PC's proof sits in a fail-stop model with reliable failure detectors. Real networks delay and split.
- Pre-commit on one side and timeout on the other is two outcomes.
- Electing a new transaction manager is a consensus problem.
- The extra phase is on the success path, and almost nobody shipped it.
Choose the contract before the protocol
Atomic commit, fault-tolerant atomic commit, or compensations. 3PC is not the middle option.
- 1
Can every participant prepare and recover?
If not, do not fake XA around an HTTP API. The existing saga series is the design. - 2
Is lock time across the round trip acceptable?
If the product cannot hold locks that long, you already chose compensations, even if both sides speak prepare. - 3
Must the coordinator survive one machine?
No: classic 2PC or presumed abort. Yes: Paxos Commit, or 2PC whose decision is the quorum log behind a coordinator. - 4
Refuse a split decision
A partition is not allowed to commit and abort. That rules out 3PC. Stall, or replicate the decision. Do not decide twice.
Overview
Classic 2PC blocks if the coordinator's decision is unreachable. 3PC tries to let survivors finish after a crash and can commit on one side of a partition while aborting on the other. Paxos Commit (Gray and Lamport) makes the decision a consensus problem so any majority of acceptors can finish, with more messages and the same disk delay in the failure-free case. Sagas do not solve atomic commit at all. They replace it with compensations.
The saga series already exists. Read Two-Phase Commit vs Sagas and Sagas and Distributed Transactions. This page only draws the boundary so you do not "fix" 2PC by rewriting that material.
What you are actually choosing
| Approach | Atomic commit | Survives coordinator process crash | Survives partition | Extra cost | Ship it? |
|---|---|---|---|---|---|
| 2PC, single TM log | Yes | No. Prepared participants block | Stalls, does not split | 2 RTTs, log forces | Yes, inside one data system |
| 3PC | Intended yes | Claims non-blocking for fail-stop | No. Can decide both ways | Another phase | Essentially no |
| Paxos Commit | Yes | Yes, if F+1 of 2F+1 acceptors are up | Yes, while a quorum remains | More messages. F=0 is exactly 2PC | Yes, when the TM must be fault tolerant |
| 2PC whose TM log is a Raft group | Yes | Yes, while the quorum is up | Same as that quorum | Shard RTT plus cross-shard 2PC | Yes, multi-shard databases |
| Saga or outbox | No | Not applicable. Each step already committed | Business-level anomalies | Compensations, idempotency | Yes, across services that cannot prepare |
3PC in one screen
Skeen's non-blocking commit adds a round so participants do not enter a state where only the old coordinator knows the outcome.
Sketch of centralized 3PC:
- Prepare, as in 2PC. Participants that vote YES are willing, but not yet pre-committed.
- If all vote YES, the coordinator sends PRE-COMMIT. A participant that records pre-commit knows every participant voted YES. It still has not committed.
- The coordinator sends COMMIT. A participant that times out and has seen pre-commit may commit, because the protocol promised no abort after pre-commit. A participant that has not seen pre-commit may abort if the coordinator dies, because the protocol promised no commit before pre-commit was acknowledged.
That state split is the point. It is also the bug under partition.
Flow
- 1
1. Both sides vote YES
- next2. PRE-COMMIT reaches A only
- 2
2. PRE-COMMIT reaches A only
- next3. Partition hides side B
- 3
3. Partition hides side B
- next4. Side A commits
- 4
4. Side A commits
- next5. Side B aborts on timeout
- 5
5. Side B aborts on timeout
- next6. Two outcomes, atomicity lost
- 6
6. Two outcomes, atomicity lost
Lesson map
3PC splits under partition
Both sides are willing and locked. Neither has a pre-commit record yet.
Architecture. Coordinator All YES. Side A Willing. Side B Willing
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB coordinator["Coordinator All YES"] side_a["Side A Willing"] side_b["Side B Willing"] side_a -->|YES| coordinator side_b -->|YES| coordinator coordinator -->|PRE-COMMIT| side_a coordinator -->|PRE-COMMIT| side_b side_a -->|Still locked| coordinator side_b -->|Still locked| coordinator coordinator -->|Partition| side_b side_a -->|Saga undo| side_b
Fail-stop crashes with reliable failure detectors are the model where the 3PC proof sits. Real networks partition and delay. A detector that marks the coordinator dead on one side and alive on the other is a split brain with extra steps. Gray and Lamport's critique is sharper: published 3PC variants often do not even specify what to do when two processes both claim to be the new transaction manager. Electing one is itself a consensus problem. You did not avoid consensus. You did it badly.
Latency is a third phase on the success path. Interview answer: know it, do not propose it as the production fix for "2PC blocks."
Paxos Commit
Gray and Lamport, Consensus on Transaction Commit: run a consensus instance for each resource manager's vote, with that resource manager as the distinguished proposer for its own yes or no. The decision is commit only if every resource manager's instance chose yes. With 2F+1 acceptors, any F+1 working acceptors can make progress. Classic 2PC is the special case F=0: one acceptor, the single coordinator, and losing it blocks.
Failure-free cost can match 2PC's stable-storage delay and message delay. The message count is higher because each vote is a consensus instance, not one RPC to one transaction manager. You pay messages to avoid a blocked prepared state when one coordinator machine dies.
Do not confuse this with "we run Paxos, so we do not need 2PC." Inside one shard, Raft or Paxos already orders a local commit. Across shards, you still need atomic commit unless you put everything in one consensus group, which couples failure domains and does not scale the write path. The usual industrial shape is Raft or Multi-Paxos per shard, 2PC across shards, and the 2PC decision record in a consensus group so the coordinator is not one box. Spanner-class and Cockroach-class systems are in this family, with extra machinery for isolation that 2PC does not provide.
Raft election and log matching stay on Raft consensus. A Raft group, in this cluster, is only the quorum log behind a coordinator.
What consensus commit still will not do
- If a shard quorum is down, its prepared transaction still cannot learn or apply a decision. You moved the blocker from one host to a quorum loss.
- It is still lock-holding atomic commit, not a saga.
- It is still the wrong tool for a participant that is a REST call to a vendor.
Sagas, only the boundary
A saga is a sequence of local transactions with a compensation for each, coordinated by orchestration or choreography. Each local transaction commits for real. There is no prepare and no uncertainty window. There is also no atomicity. Other clients can see the intermediate state, and a failure runs compensations that are themselves fallible. They must be idempotent, and some actions cannot be compensated.
Use a saga when:
- Some participant cannot prepare.
- You will not hold database locks across a network round trip and a human retry.
- The product owner accepts pending and undo as user-visible states.
Use 2PC when:
- Both sides are resource managers you control, usually shards of one system.
- Partial commit is a financial or integrity incident, not a state on a screen.
- You can operate prepared-transaction age as a page.
The catalog already has the long form: Two-Phase Commit vs Sagas, inside Sagas and Distributed Transactions. Follow those pages for orchestration versus choreography, timeouts and retries, and poison steps and reconciliation. This cluster does not restate them.
Decisions
- 1
1. Cross-system write
- next2. Can every side prepare?
- ?
2. Can every side prepare?
- no3. Saga or outbox
- yes4. Is cross-RTT lock time ok?
- 3
3. Saga or outbox
- ?
4. Is cross-RTT lock time ok?
- no3. Saga or outbox
- yes5. Must the TM survive one loss?
- ?
5. Must the TM survive one loss?
- no6. Classic 2PC or presumed abort
- yes7. Paxos Commit or a Raft TM log
- 6
6. Classic 2PC or presumed abort
- 7
7. Paxos Commit or a Raft TM log
The Trade-offs view on the map above is this choice. A side that cannot prepare, or that cannot hold locks across the round trip, is a saga. Classic 2PC is enough when both sides can prepare and one log is acceptable. If that log must survive one machine, the decision belongs in a quorum, not in a third phase.
Sandbox
The function never returns 3PC. A participant that cannot prepare, or that cannot hold locks, is a saga. Fault-tolerant atomic commit is a replicated decision, not a third phase.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Is 3PC strictly safer than 2PC?
Answer
No. It is an attempt to be non-blocking when sites crash in a fail-stop model. Under partition it can commit and abort. 2PC under the same partition stalls, which preserves atomicity. Safer for atomicity is 2PC. More available for the locked keys is not automatically 3PC.
Someone says just use Raft instead of 2PC. What do you answer?
Answer
Raft, or Paxos, gives you a total order and a fault-tolerant log for one group. It does not by itself atomically commit two independent resource managers. You either put both writes in one Raft group, or you run 2PC across groups whose decision is stored in Raft. Raft consensus is the first half, and only as the quorum log behind a coordinator. This series is the second half.
Paxos Commit with F=0 is what?
Answer
Two-phase commit. One acceptor, the coordinator. If that acceptor is down, prepared participants wait. The paper's point is that fault-tolerant commit is 2PC with the decision replicated, not a different business protocol.
A design doc proposes 2PC between the orders service and a card API. What do you change?
Answer
The card API will not sit in prepare holding your database locks until the transaction manager recovers. The charge is a local step with idempotency, and the rest is a saga or outbox. Point the author at Two-Phase Commit vs Sagas. Do not invent a prepare on an HTTP 200.
Why do multi-shard databases still block some transactions if they use Raft?
Answer
Because a prepared shard that cannot talk to a quorum that knows the decision must not guess. Consensus removed the single-process coordinator. It did not make unilateral commit of a prepared shard safe.
What did 3PC hope the pre-commit round would prove?
Answer
That every participant had voted YES, and that no abort would follow once pre-commit was in place. The hope depends on failure detectors that agree. A partition makes one side see pre-commit and the other side see a dead coordinator. The protocol then allows both outcomes.
When do you still ship classic 2PC?
Answer
Inside one data system, few resource managers, locks inside the latency budget, and an operator who will page on prepared-transaction age. Presumed abort is the usual engine shape of that choice. The log rules are on optimizations. If the coordinator disk is the outage you cannot accept, replicate that log instead of adding a third phase.
What does this page refuse to teach in full?
Answer
Orchestration versus choreography, idempotent undo, saga deadlines, and poison-step reconciliation. Those lessons already exist. Link them and stop. Rewriting them here would be a second saga cluster.
Take one cross-shard transfer and one checkout that calls a card API. Write the protocol name on each arrow. The transfer may say 2PC on a quorum log. The card arrow may not say prepare. If you wrote 3PC anywhere, replace it and say what split you just avoided.
Go deeper
- Dale Skeen, Nonblocking Commit Protocols.
- Gray and Lamport, Consensus on Transaction Commit and the TODS 2006 PDF.
- Bernstein, Hadzilacos, and Goodman on atomic commit versus non-blocking commit.
- Kleppmann on 2PC, then a total order of votes.
- CMU 15-445 on why 3PC is ignored in practice.
- Fowler, Two-Phase Commit, and the saga patterns on the same site for the boundary only.