Distributed systems
Part 5 of 5 · Raft consensusRaft vs Multi-Paxos vs Zab — When to Choose What
Same CFT RSM goal, different models/ops/ecosystems; prefer Raft/etcd/Consul greenfield; keep ZK when watches own estate; custom Multi-Paxos only with deep expertise.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Greenfield metadata: what you actually ship
Prefer
Managed Raft (etcd / Consul) unless ZK already owns the app
Understandable CFT log, boring ops, Kubernetes gravity. Consul if you already live in the Hashi stack.
- Interview default for new strongly consistent metadata.
- Leases, watches, small KV — not a general database.
- KRaft shows the industry moving off extra ZK ensembles.
- You can still say Paxos in the same breath: same safety class.
Alternative
Custom Multi-Paxos or a second ZooKeeper
Right for a storage engine with experts, or an existing ZK recipe surface. Wrong as a fashion statement.
- Paxos is more academic ≠ stronger.
- Two ensembles for the same metadata = two failure domains.
- fsync is the latency floor on all three.
- AP multi-region active-active needs a higher-level design.
Overview
Raft, Multi-Paxos, and Zab all implement replicated state machine consensus under crash faults, but they optimize for different mental models, ecosystems, and operational defaults.
- Raft wins on understandability and etcd/Consul/TiKV gravity.
- Multi-Paxos remains the flexible ancestor for custom low-latency stores.
- Zab powers ZooKeeper's primary-backup broadcast with zxid ordering and watch-centric APIs.
Choose based on tooling, latency path, ops skill, and API semantics — not brand loyalty.
Decision table (interview gold)
| Criterion | Raft | Multi-Paxos | Zab / ZooKeeper |
|---|---|---|---|
| Mental model | Strong leader, terms, explicit log | Ballots, promises, per-slot choose | Primary + epoch/zxid broadcast |
| Ops complexity | Medium (well-trodden etcd) | High if self-built | Medium (ZK ensemble lore) |
| Typical latency path | Leader RTT + majority disk | Can pipeline / batch aggressively | Primary + majority ack |
| Tooling gravity | etcd, Consul, TiKV, Cockroach bits | Spanner-like / proprietary | ZooKeeper, old Hadoop/Kafka (legacy) |
| Membership changes | Joint consensus / single-add | Implementation-defined | Reconfig via ZK protocols |
| Best fit | K8s control plane, service discovery, metadata KV | Custom DB consensus with experts | Existing ZK watches / recipes |
| Weak fit | WAN multi-master writes without hierarchy | Teams without Paxos veterans | Greenfield if etcd already standard |
What fails if you choose wrong
- Build a custom Multi-Paxos for a tiny lease service when etcd would do → months of safety bugs.
- Standardize on ZooKeeper in 2026 for new K8s-adjacent metadata while the platform already runs etcd → dual ensembles, dual failure domains.
- Expect Raft/ZK to give AP multi-region active-active without higher-level design → you get CP majority behavior.
Comparative teaching: leadership & logs
Architecture
Step 1 — Client write
- 1
Client
- nextPropose command
- 2
Propose command
- nextRaft leader append plus majority
- nextStable proposer Phase2
- nextPrimary assigns zxid
Step 2 — Protocol path
- 3
Raft leader append plus majority
- nextRaft followers ack
- 4
Stable proposer Phase2
- nextPaxos acceptors
- 5
Primary assigns zxid
- nextZab secondaries ack
Step 3 — Learners and followers
- 6
Raft followers ack
- nextApply in log order
- 7
Paxos acceptors
- nextChosen values in slot order
- 8
Zab secondaries ack
- nextDeliver by zxid
Step 4 — Deliver
- 9
Apply in log order
- 10
Chosen values in slot order
- 11
Deliver by zxid
Lesson map
Raft vs Multi-Paxos vs Zab — When to Choose What
Same CFT RSM goal, different models/ops/ecosystems; prefer Raft/etcd/Consul greenfield; keep ZK when watches own estate; custom Multi-Paxos only with deep expertise.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["Client"] cmd["Propose command"] rl["Raft leader append plus majority"] pp["Stable proposer Phase2"] c -->|Client to Propose command| cmd cmd -->|Propose command to Raft leader append plus majority| rl cmd -->|Propose command to Stable proposer Phase2| pp
Ecosystem context
- etcd (Raft): Kubernetes heart; lease + watch + MVCC keyspace. Prefer when you need small strongly consistent metadata and K8s alignment.
- Consul (Raft): Service discovery + KV; similar ops to etcd with the HashiCorp ecosystem.
- ZooKeeper (Zab): Mature recipes (locks, barriers) and watches; still common in older big-data stacks. Kafka left ZK for KRaft (Raft-based metadata quorum).
- Custom Multi-Paxos: Justified inside databases where you co-design batching, parallel commits, and storage engines (and staff the expertise).
Latency & throughput levers
| Lever | Raft practice | Multi-Paxos practice | Zab practice |
|---|---|---|---|
| Batching | Leader batches AppendEntries | Phase2 pipelining | Primary batches proposals |
| Read path | ReadIndex / lease | Lease / quorum read | Sync read vs local |
| Slow followers | Snapshot install | Same class of catch-up | SNAP sync |
| Disk | fsync policy dominates | Same | Same |
Disk is the honest floor. Protocol brand will not beat a slow fsync.
Sandbox: decision helper (Python)
Encodes the table for study — not production routing.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same advisor (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Is Raft better than Paxos?
Answer
Not asymptotically. Raft prioritizes understandability and a strong-leader decomposition; Paxos is more abstractly flexible. Same CFT RSM safety class when implemented correctly.
Why did Kafka move from ZK to KRaft?
Answer
Remove an external ensemble; simplify ops; use a Raft-based metadata quorum inside Kafka. The data log is still Kafka's partitioned log — KRaft is the controller quorum.
When is ZooKeeper still justified?
Answer
A large existing ZK estate, or apps deeply tied to ZK recipes/watches — not greenfield by default in 2026 if etcd is already on the platform.
Does Multi-Paxos require a leader?
Answer
A stable leader/proposer is common for performance; classical Paxos allows different proposers with higher ballots. Production Multi-Paxos almost always sticks a leader.
Shared safety goal?
Answer
Agreement on a total order of commands for a deterministic state machine under majority crash faults.
Ops complexity winner?
Answer
Usually managed etcd/Consul over self-built Paxos; ZK if you already mastered it. Complexity is people and pages, not big-O.
Cross-region active-active?
Answer
None of these give magic AP multi-master. You need higher-level partitioning or you accept CP majority across regions.
etcd vs Consul?
Answer
Both Raft; choose by ecosystem (K8s vs Hashi service mesh/discovery) and operational familiarity.
Pitfalls
(1) New Kubernetes-adjacent lease store. (2) Hadoop-era app with ZK locks in production. (3) New OLTP engine with a consensus-specialist staff. (4) "We need multi-master writes in three regions." Write the protocol, the product, and the trap.