Database Replication — Leader-Follower, Multi-Leader & Leaderless Quorums
Replication keeps the same data on several machines for durability, read scale, or multi-region latency. The three shapes are leader-follower, multi-leader, and leaderless. Who may accept a write and when the client hears committed set lag, RPO, failover risk, and whether you must resolve conflicts.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Replication means keeping a copy of the same data on several machines. Teams do it for three different reasons: durability (survive a disk or node loss), read scale (serve reads from replicas), and latency/availability across regions. There are only three basic shapes. In leader-follower (single leader), one node accepts writes and ships its log to the others. In multi-leader, several nodes accept writes and replicate to each other. In leaderless (Dynamo style), the client writes to many replicas and counts acks against a quorum. Every interview question about replication comes down to two choices: who is allowed to accept a write, and when the client gets told "done". Those two choices decide your lag, your data-loss window (RPO), your failover story and whether you have to resolve conflicts.
The three models compared
| Model | Who accepts writes | Ordering | Conflicts | Failover | Typical systems | Best for |
|---|---|---|---|---|---|---|
| Leader-follower (async) | 1 leader | Total order (leader's log) | None | Promote a follower; may lose the tail of acked writes | Postgres streaming replicas, MySQL replicas, MongoDB replica sets, Redis replicas | OLTP with read replicas, most apps |
| Leader-follower (sync / semi-sync) | 1 leader | Total order | None | Promote the sync follower; zero acked-write loss | Postgres synchronous_standby_names, MySQL semi-sync, Aurora (quorum storage) | Money, inventory, anything with RPO = 0 |
| Consensus-replicated leader | 1 leader elected by majority | Total order, majority-committed | None | Automatic and safe (Raft/Paxos term) | etcd, CockroachDB, Spanner, TiDB, YugabyteDB | Strong consistency plus automatic failover |
| Multi-leader / active-active | Several leaders (often one per region) | Per-leader only | Yes: must detect and merge | Each region keeps writing; no global failover | Postgres BDR/pglogical, MySQL Group Replication multi-primary, CouchDB, DynamoDB Global Tables | Multi-region local writes, offline clients |
| Leaderless quorum | Any replica (client or coordinator fans out) | None globally; per-key versions | Yes: siblings, LWW or CRDT | No failover step; quorums tolerate node loss | Cassandra, ScyllaDB, Riak, DynamoDB (internally) | High write availability, large key-value data |
Pick a replication model from the write path
Who accepts writes and when the client hears committed decide lag, RPO, failover, and conflicts.
- 1
Who accepts writes?
One leader, several regional leaders, or any replica in a quorum. - 2
Leader-follower
Total order on the leader log. Add sync or semi-sync when RPO must be zero. - 3
Multi-leader
Local writes in every region. You must detect and merge concurrent updates. - 4
Leaderless quorum
Client or coordinator fans out; success is W of N acks. - 5
Measure lag and RPO
Session guarantees, failover fencing, and repair close the loop.
Flow
- 1
1. Who accepts writes?
- One node2. Leader-follower
- Several regions3. Multi-leader
- Any replica4. Leaderless quorum
- 2
2. Leader-follower
- next5. Pick ack mode: async / semi-sync / sync
- 3
3. Multi-leader
- next5. Plan conflict merge: LWW / CRDT / home region
- 4
4. Leaderless quorum
- next5. Tune N / R / W and repair
- 5
5. Pick ack mode: async / semi-sync / sync
- next6. Measure lag and failover RPO
- 6
5. Plan conflict merge: LWW / CRDT / home region
- next6. Measure lag and failover RPO
- 7
5. Tune N / R / W and repair
- next6. Measure lag and failover RPO
- 8
6. Measure lag and failover RPO
Lesson map
Database Replication — Leader-Follower, Multi-Leader & Leaderless Quorums
Replication keeps the same data on several machines for durability, read scale, or multi-region latency. The three shapes are leader-follower, multi-leader, and leaderless. Who may accept a write and when the client hears committed set lag, RPO, failover risk, and whether you must resolve conflicts.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB q["1. Who accepts writes?"] lf["2. Leader-follower"] ml["3. Multi-leader"] ll["4. Leaderless quorum"] q -->|One node| lf q -->|Several regions| ml q -->|Any replica| ll
Press Run. Snippets must be self-contained — no network, files, or native modules.
Output:
leader-follower (async) ack= 0ms visible: A@0ms, B@5ms, C@40ms
-> fast ack, followers lag; leader loss can drop acked writes
leader-follower (semi-sync) ack= 10ms visible: A@0ms, B@5ms, C@40ms
-> one extra RTT buys 'acked write survives one node loss'
multi-leader ack= 0ms visible: A@0ms, B@5ms, C@40ms
-> local writes everywhere, but concurrent writes CONFLICT
leaderless (N=3, W=2) ack= 10ms visible: A@0ms, B@5ms, C@40ms
-> no failover step; reads must also hit R replicasWhy single leader is the default, and what happens if you choose something else
Single leader wins by default because a single total order of writes makes everything downstream simple: there are no conflicts, unique constraints work, transactions work, CDC streams are ordered, and "the latest value" has one meaning. You pay for it in three places. All writes go through one node, so write throughput is capped. Failover is a dangerous moment (see the failover page). And if the leader is in one region, users far away get cross-region write latency.
- If you choose multi-leader to "get writes in every region": you now own conflict resolution forever. Two regions can both accept
UPDATE users SET email=...for the same user, and something has to decide. Last-writer-wins drops data silently. Unique constraints such as "usernames must be unique" can't be enforced across leaders without coordination. Most teams that try generic active-active SQL end up partitioning instead: each user gets a home region that owns their writes. That is still single leader, just per key. - If you choose leaderless for "no failover": you get great write availability, but reads have to do more work (R replicas, read repair), there is no transaction across keys, and "R + W > N" still isn't linearizable under sloppy quorums, concurrent writes or partial failures. It fits key-value and time-series workloads, not a ledger.
- If you choose consensus (Raft/Paxos) everywhere: you get correctness and automatic failover, but every write pays a majority round trip. Across regions that is tens of milliseconds per commit, and throughput is bounded by the leader per range. That's why systems like CockroachDB or Spanner split data into many ranges, each with its own leader.
// Pick a replication model from requirements (a tiny decision helper).
// Sandbox-runnable: tsc --strict then node. No dependencies.
type Needs = {
writeRegions: 1 | "many"; // must users in several regions write locally?
tolerateConflicts: boolean; // can the data model merge concurrent writes?
maxRpoSeconds: number; // how much acked data may we lose on failover?
wantNoFailoverStep: boolean; // availability over strict ordering?
};
function choose(n: Needs): string {
// 1. Multiple write regions only make sense if conflicts are mergeable.
if (n.writeRegions === "many") {
return n.tolerateConflicts
? "multi-leader / active-active with CRDT or app-level merge"
: "single leader per KEY (partitioned home region) - route writes to owner";
}
// 2. Leaderless trades ordering for no-failover availability (Dynamo style).
if (n.wantNoFailoverStep && n.tolerateConflicts) {
return "leaderless quorum (Cassandra/Dynamo style, R+W>N)";
}
// 3. Default: single leader; RPO decides sync vs async.
return n.maxRpoSeconds === 0
? "single leader + synchronous/semi-sync (or consensus: Raft-based SQL)"
: "single leader + async read replicas";
}
const cases: Array<[string, Needs]> = [
["payments ledger", { writeRegions: 1, tolerateConflicts: false, maxRpoSeconds: 0, wantNoFailoverStep: false }],
["product catalog", { writeRegions: 1, tolerateConflicts: false, maxRpoSeconds: 5, wantNoFailoverStep: false }],
["shopping cart, 3 regions", { writeRegions: "many", tolerateConflicts: true, maxRpoSeconds: 5, wantNoFailoverStep: true }],
["user profiles, 3 regions", { writeRegions: "many", tolerateConflicts: false, maxRpoSeconds: 1, wantNoFailoverStep: false }],
["IoT time series", { writeRegions: 1, tolerateConflicts: true, maxRpoSeconds: 60, wantNoFailoverStep: true }],
];
for (const [name, needs] of cases) console.log(`${name.padEnd(26)} -> ${choose(needs)}`);Output:
payments ledger -> single leader + synchronous/semi-sync (or consensus: Raft-based SQL)
product catalog -> single leader + async read replicas
shopping cart, 3 regions -> multi-leader / active-active with CRDT or app-level merge
user profiles, 3 regions -> single leader per KEY (partitioned home region) - route writes to owner
IoT time series -> leaderless quorum (Cassandra/Dynamo style, R+W>N)The knobs that matter in every model
| Knob | Question it answers | Leader-follower | Multi-leader | Leaderless |
|---|---|---|---|---|
| Ack policy | When is the client told "committed"? | async / semi-sync / sync | local commit, then async | W of N acks |
| Read policy | Which copy can a read use? | leader, any replica, or a caught-up replica (LSN token) | local leader | R of N replicas, newest version wins |
| Lag visibility | How stale can a read be? | replay_lag, seconds-behind-source | replication queue depth | hinted handoff backlog, repair age |
| Conflict policy | What if two writes race? | n/a (one order) | LWW, version vectors, CRDTs, app merge | LWW timestamps, siblings, CRDTs |
| Failure handling | What if a node dies? | detect, promote, fence, repoint | region keeps writing | sloppy quorum, hinted handoff |
Pros and cons at a glance
Leader-follower. Pros: simple mental model, one order, transactions and constraints work, mature tooling (Patroni, Orchestrator, RDS Multi-AZ). Cons: write bottleneck, failover risk, async lag leads to stale reads, and async loses acked writes on failover.
Multi-leader. Pros: local-latency writes in every region, survives a region partition, works for offline-first apps. Cons: conflicts, no global constraints, hard to reason about, and the topology (ring, star, all-to-all) has its own ordering bugs.
Leaderless. Pros: no single point of failure, tunable per request (QUORUM, ONE, LOCAL_QUORUM), linear scale with consistent hashing. Cons: weaker semantics than they look, repair work (read repair, anti-entropy), tombstones and clock-based LWW.
Interview Q&A
Q1. You're designing a global app. Users in the US, EU and APAC all write. Which replication model? Start by asking what the data is. For user-owned data (profile, settings, orders), give each user a home region with a single leader per partition, route writes there, and serve local follower reads. You get no conflicts and only the owner's writes cross regions. For data that is truly shared and mergeable (counters, carts, presence), use multi-leader with CRDTs. For strongly consistent global state (inventory, money), use a consensus database with leaders placed near the main writers and accept the cross-region commit latency.
Q2. What's the difference between replication and sharding? Replication puts the same data on several nodes, for durability and read scale. Sharding puts different data on different nodes, for write scale and capacity. Real systems do both: each shard is a replica set (for example a Cassandra token range has RF=3, and a Vitess shard has a primary plus replicas).
Q3. Why don't we just make all replication synchronous? Synchronous replication to all followers means the slowest or dead follower blocks every commit, so availability gets worse as you add replicas. Real systems use semi-sync: wait for one or a quorum, which gives RPO = 0 for a single failure without waiting on stragglers. Consensus does the same thing with a majority.
Q4. Is Raft "leader-follower replication"? Yes. Raft is leader-follower with two additions: a safe election (terms and log-up-to-date checks), and majority commit, so failover never loses committed entries. Plain Postgres streaming replication has neither built in. An external tool such as Patroni with etcd supplies the election, and you choose sync settings for durability.
Q5. A product manager says reads from replicas show old data. What are your options? Read your own writes from the leader for a short window, or route with an LSN/GTID token. Pin users to one replica for monotonic reads. Expose lag and stop routing to replicas past a lag threshold. Or accept the staleness explicitly for that screen. The session-guarantees page covers each of these.
Q6. Which model does DynamoDB use? Inside a region, each partition is a replication group with an elected leader (Multi-Paxos), and a write is acknowledged once a quorum of replicas has persisted it. Global Tables add multi-region, multi-active replication, which by default resolves conflicts with last-writer-wins (AWS has since added a multi-region strong consistency option). So "Dynamo-style leaderless" describes the 2007 Dynamo paper and Cassandra/Riak better than today's DynamoDB.
Go Deeper
- PostgreSQL docs: High Availability, Load Balancing, and Replication
- MySQL 8.4: Semisynchronous Replication
- Amazon Dynamo paper (SOSP 2007)
- DynamoDB paper (USENIX ATC 2022)
- Designing Data-Intensive Applications (Kleppmann), chapter 5: Replication
- Jepsen analyses of real database replication bugs
Related
- Prev: Leaderless Replication - N/R/W Quorums, Hinted Handoff, Read Repair & Anti-Entropy (
leaderless-replication-quorums-read-repair-anti-entropy) - Next: Sync vs Async vs Semi-Sync Replication - Durability, RPO & Physical vs Logical Logs (
replication-sync-async-semi-sync-physical-logical)
This series:
- Database Replication for Engineers - Leader-Follower, Multi-Leader & Leaderless Quorums (
database-replication-leader-follower-multi-leader-leaderless) - this page - Sync vs Async vs Semi-Sync Replication - Durability, RPO & Physical vs Logical Logs (
replication-sync-async-semi-sync-physical-logical) - Replication Lag & Session Guarantees - Read-Your-Writes, Monotonic Reads & Consistent Prefix (
replication-lag-read-your-writes-session-guarantees) - Failover & Split Brain - Detection, Promotion, Fencing & Lost Writes (
replication-failover-split-brain-fencing) - Multi-Leader & Active-Active - Conflict Detection, LWW, CRDTs & Home Regions (
multi-leader-replication-conflicts-active-active) - Leaderless Replication - N/R/W Quorums, Hinted Handoff, Read Repair & Anti-Entropy (
leaderless-replication-quorums-read-repair-anti-entropy)
Existing Study pages (cross-link only, not rewritten here):
- Raft consensus - leader election & log replication (
raft-consensus-leader-election-log-replication) - Database sharding - partition keys & rebalancing (
db-sharding-partitioning-keys-rebalancing) - Consistent hashing - rings, vnodes & replicas (
consistent-hashing-rings-vnodes-replicas) - CRDTs - conflict-free replicated data types (
crdts-conflict-free-replicated-data) - Storage engines - WAL, B-trees & LSM-trees (
database-storage-engines-wal-b-trees-lsm-trees)