Distributed systems
Part 2 of 6 · Sagas & Distributed TransactionsTwo-Phase Commit vs Sagas — Why 2PC Breaks at Scale
Why 2PC blocks at scale, how that compares with sagas, and when a shared-database ACID transaction still wins.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
All-or-nothing across two stores
Prefer
Saga: short local commits, explicit undo
Each resource manager commits quickly. Intermediate states are visible. A later failure runs a compensation, not a distributed rollback.
- Autonomy stays with the service that owns the schema.
- Lock time is the local transaction, not the cross-region round trip.
- You still owe observability and a reconciler.
Alternative
2PC: prepare, then commit
The coordinator asks every participant to vote. A yes vote persists an intent and holds resources until commit or abort.
- Familiar atomic story when every participant really speaks the protocol.
- A crashed coordinator stalls everyone who voted yes.
- Operators sometimes force a decision the logs cannot prove.
How 2PC moves
Phase 1 is the vote. The coordinator sends PREPARE. Each resource manager writes a durable vote and, on yes, keeps the locks it will need to commit. Phase 2 is the decision. If every vote is yes, the coordinator sends COMMIT. If any vote is no, it sends ABORT and participants release locks.
Sequence
- 1
Coordinator → Store A
PREPARE
- 2
Coordinator → Store B
PREPARE
- 3
Store A → Coordinator
YES locked
- 4
Store B → Coordinator
YES locked
- 5
Coordinator → Store A
COMMIT
- 6
Coordinator → Store B
COMMIT
- 7
Store A → Coordinator
ACK
- 8
Store B → Coordinator
ACK
Lesson map
Why the prepare lock breaks checkout
Both stores are locked on YES. The coordinator has votes and no decision. This is the window checkout cannot afford.
Architecture. Coordinator No decision. Store A Locked · YES. Store B Locked · YES
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB coordinator["Coordinator No decision"] store_a["Store A Locked YES"] store_b["Store B Locked YES"] coordinator -->|PREPARE| store_a coordinator -->|PREPARE| store_b store_a -->|YES| coordinator store_b -->|YES| coordinator store_a -->|Lock held| coordinator store_b -->|Lock held| coordinator coordinator -->|No decision| store_a store_a -->|Heuristic| store_b
The abort path is the same shape with a different decision. Nobody stays prepared.
Flow
- 1
1. Coordinator sends PREPARE
- next2. A participant votes NO
- 2
2. A participant votes NO
- next3. Coordinator sends ABORT
- 3
3. Coordinator sends ABORT
- next4. Participants release locks
- 4
4. Participants release locks
The dangerous window is a crash after a yes vote and before a durable decision everyone has applied. Prepared participants sit on locks until recovery finishes. That is the availability hit. Participants need a durable vote log so they can answer “what did I promise?” after they restart. The coordinator needs its own log so it can finish commit or abort. Dual failures in that window are why this protocol has an operations chapter, not just a diagram.
Where the two designs differ
| Dimension | 2PC / XA | Saga |
|---|---|---|
| Atomicity | Attempted across resource managers | Business-level eventual outcome |
| Isolation during the protocol | Locks held from prepare through decision | Each local transaction is short |
| Failure mode | Block, or a heuristic decision | Compensate, then reconcile |
| Autonomy | Low; everyone speaks one protocol | High; each service commits locally |
| Latency | Extra round trips plus lock hold | A pipeline of local commits |
| Operations | Coordinator HA and XA drivers | Saga state store and a dead-letter path |
| Polyglot microservices | Poor fit | The usual fit |
Why 2PC breaks at scale
Flow
- 1
1. Participants vote YES and lock
- next2. Coordinator crash stalls those locks
- 2
2. Coordinator crash stalls those locks
- next3. Many stores do not speak real XA
- 3
3. Many stores do not speak real XA
- next4. Cross-region round trips hit p99
- 4
4. Cross-region round trips hit p99
- next5. Heuristic outcomes still need reconcile
- 5
5. Heuristic outcomes still need reconcile
- next6. Checkout wants short local commits
- 6
6. Checkout wants short local commits
- Blocking. Prepared participants cannot release locks until commit or abort. A network blip becomes an availability incident for every other transaction that needs those rows.
- Coordinator failure. High availability means a recovery log, fencing, and a story for “both sides thought they were coordinator.” That is real work, not a checkbox.
- Heterogeneous stores. Not every database, cache, or third-party API implements XA. Many “we use XA” stories are a driver flag with a gap behind it.
- Latency. Prepare plus commit across regions adds round trips on the user path and holds locks for that whole interval.
- Heuristic outcomes. When logs disagree, an operator forces commit or abort. Correctness then depends on offline repair. You paid for 2PC and still built a saga-like cleanup.
Raft Consensus — Leader Election, Log Replication & Safety is not a substitute. Raft agrees on one replicated log inside a cluster. It does not atomically commit an inventory row and a payment row in two independent databases. Say that when the interview offers consensus as a distributed transaction.
When 2PC or one database still wins
| Situation | Prefer |
|---|---|
| One Postgres or MySQL database owns every row in the action | Local ACID |
| Two co-located stores, a short transaction, one team, real XA | Maybe 2PC or one shared transaction manager |
| Checkout across inventory, payments, and shipping | Saga |
| Independent deploy and schema ownership | Saga |
| A rule of “never double-charge” | Saga, an idempotent ledger, and reconcile — not a naive XA wrapper |
Try-Confirm-Cancel (TCC) is a cousin, not classic XA. Try places a hold, confirm finalizes it, cancel releases it. That is a semantic reservation, closer to a saga step than to a prepare vote that freezes arbitrary SQL locks. Compensations in the later lesson cover void-versus-refund in that family.
Regulatory language sometimes sounds like it demands XA. The invariant is usually “money moves once, and we can prove it.” An idempotency key at the payment provider plus a ledger reconcile meets that bar more often than a coordinator that blocks checkout.
Lock-hold sketch (run this)
The model is deliberately small. 2PC holds locks from the prepare round trip through the commit round trip. A saga holds locks only for the local transaction. Same-AZ looks cheap. Cross-region does not.
Press Run. Snippets must be self-contained — no network, files, or native modules.
165 ms of lock hold for 5 ms of real work is the interview picture. It ignores coordinator failover, which only makes the 2PC number worse.
Vote aggregation (run this)
A single timeout is not an abort. The coordinator does not know whether that participant prepared. Treating timeout as no releases locks on one side and can commit on the other if you guess. The safe toy result is BLOCKED until recovery.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
What are the two phases?
Answer
Prepare: each participant durably votes and, on yes, holds the resources required to finish. Commit or abort: the coordinator’s decision, which participants apply and then acknowledge. Crash recovery reads those durable votes. There is no legal “forget the vote and continue” step.
Why is blocking dangerous?
Answer
Other transactions cannot touch the locked rows. If the coordinator is dead, the prepared fleet stalls even though the business work itself already finished locally. User requests queue behind locks that nobody can release until recovery.
What is a heuristic decision?
Answer
An operator or resource manager forces commit or abort when the protocol cannot finish safely. The forced side may disagree with another participant. You repair that with reconciliation, which is the cost you hoped 2PC would avoid.
Does 2PC give you serializability across microservices?
Answer
Only if every participant implements the protocol and its isolation rules correctly. Polyglot stores rarely hand you that for free. Assume you do not have cross-service serializability until you can name the resource managers and the isolation each one actually provides.
Why do sagas win for checkout?
Answer
Inventory, payments, and shipping are different teams and failure domains. A refund or a released hold is a better availability trade than holding XA locks across a payment RPC. Customers can see “payment failed, stock released” if you design that state on purpose.
Is Try-Confirm-Cancel the same as 2PC?
Answer
No. TCC reserves with a try, then confirms or cancels that reservation. It is a semantic hold, in the saga family. Classic XA prepare is a database protocol that pins locks until the global decision.
Can Raft replace 2PC?
Answer
No. Raft replicates a log so one cluster agrees on order. It does not commit two independent databases together. Link the consensus lesson for elections and quorums; keep business atomicity on this page.
What is the trap answer?
Answer
“We will wrap the whole flow in a distributed transaction.” The follow-up is blocking, who runs the coordinator, what a heuristic outcome means, and who owns compensation if you leave XA behind.
Pitfalls
Take a checkout you know. Write prepare RTT, local work, and commit RTT for same-AZ and for cross-region. Compare that hold time with the local transaction alone. Then say which participant would block the catalog if the coordinator died, and what compensation you would owe if you refused 2PC.