Multi-Leader & Active-Active — Conflict Detection, LWW, CRDTs & Home Regions
Multi-leader and active-active give local writes in every region and create conflicts. Last-writer-wins can silently drop data; version vectors, CRDTs, and home-region routing are the real tools.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Multi-leader (active-active) replication lets several nodes accept writes, usually one per region or per offline device, and replicates changes between them asynchronously. Users get local write latency everywhere, and each region keeps working through an inter-region partition. The price is write conflicts: two leaders can change the same row before hearing about each other. You must detect concurrency (version vectors, or a conflict on the same key within the replication window) and resolve it with last-writer-wins, merge rules, CRDTs or a human. Constraints that need a global view, such as unique usernames or "balance must stay at or above 0", can't be enforced locally. The senior answer is often to avoid conflicts by giving each record a home region, and to use true multi-leader only for data you can merge.
When multi-leader makes sense
| Use case | Why multi-leader | Conflict story |
|---|---|---|
| Multi-region apps with local writes | Avoid ~70-150 ms cross-region write RTT | Route each user to a home region; true conflicts are rare |
| Offline-first / mobile / collaborative editing | Every device is a leader while offline | CRDTs or OT (Figma, Notion, CouchDB/PouchDB, Automerge) |
| Region-partition tolerance | Keep accepting writes when the inter-region link is down | Must merge on heal |
| Shared counters, carts, likes, presence | Natural merge semantics | CRDT counters and sets |
| Ledgers, inventory, unique constraints | Bad fit | Use a single leader per key or consensus |
Accept multi-leader only with a merge story
Local writes are easy. Concurrent updates are the product.
- 1
Local write in a region
Ack without waiting on the far region. - 2
Replicate asynchronously
Ship changes to peer leaders. - 3
Detect conflicts
Version vectors or causal metadata beat wall-clock LWW alone. - 4
Merge or home-region
CRDTs / app merge, or route each key to one home leader.
Sequence
- 1
US client → US leader
1. add lamp to cart 42
- 2
EU client → EU leader
1. add mug to cart 42
- 3
US leader → US client
2. OK locally
- 4
EU leader → EU client
2. OK locally
- 5
US leader → EU leader
3. replicate version us2 eu0
- 6
EU leader → US leader
3. replicate version us1 eu1
- 7
US leader
4. neither version dominates, so writes are concurrent
- 8
US leader → US leader
5. resolve by LWW, union merge or CRDT
- 9
EU leader → EU leader
5. same deterministic rule so both regions converge
Lesson map
Multi-Leader & Active-Active — Conflict Detection, LWW, CRDTs & Home Regions
Multi-leader and active-active give local writes in every region and create conflicts. Last-writer-wins can silently drop data; version vectors, CRDTs, and home-region routing are the real tools.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["US client"] lu["US leader"] le["EU leader"] e["EU client"] u -->|1. add lamp to cart 42| lu e -->|1. add mug to cart 42| le lu -->|2. OK locally| u le -->|2. OK locally| e lu -->|3. replicate version us2 eu0| le le -->|3. replicate version us1 eu1| lu
Conflict resolution strategies compared
| Strategy | How | Pros | Cons | Seen in |
|---|---|---|---|---|
| Avoidance (home region / per-key leader) | All writes for a key go to its owner region | No conflicts; constraints work per key | Cross-region writes for travelling users; failing over a home region is a migration | Most large SaaS, CockroachDB REGIONAL BY ROW, Spanner leader placement |
| Last-writer-wins (LWW) | Highest timestamp (plus tie-break node id) wins | Simple, always converges | Silently loses concurrent writes; clock skew picks the "wrong" winner | Cassandra, DynamoDB Global Tables, many BDR configs |
| Version vectors plus siblings | Detect concurrency and keep both versions; app merges on read | No silent loss | App must write merge logic; siblings can pile up | Riak, CouchDB conflicts, Dynamo paper |
| Custom merge handlers | DB calls your function on conflict | Domain-correct | Code runs inside replication; hard to test | Postgres BDR/pgactive conflict handlers, Oracle GoldenGate |
| CRDTs | Data types whose merge is commutative, associative, idempotent | Automatic convergence | Limited types; invariants (stock at or above 0) not enforced; metadata growth | Redis Enterprise Active-Active, Riak data types, Automerge/Yjs |
| Operational transform | Transform concurrent ops against each other via a server | Great for text | Needs central server ordering; complex | Google Docs |
Press Run. Snippets must be self-contained — no network, files, or native modules.
Output:
LWW result : ['book', 'lamp'] <- the mug is silently lost
version vectors : CONCURRENT -> conflict
union merge : ['book', 'lamp', 'mug'] vv = {'us': 2, 'eu': 1}// A PN-counter CRDT: each region increments its own slot; merge = per-slot max.
// Concurrent updates in different regions converge without coordination.
// Sandbox-runnable: tsc --strict then node.
type Slots = Record<string, number>;
class PNCounter {
p: Slots = {};
n: Slots = {};
constructor(readonly region: string) {}
inc(by = 1) { this.p[this.region] = (this.p[this.region] ?? 0) + by; }
dec(by = 1) { this.n[this.region] = (this.n[this.region] ?? 0) + by; }
value() {
const sum = (s: Slots) => Object.values(s).reduce((a, b) => a + b, 0);
return sum(this.p) - sum(this.n);
}
merge(o: PNCounter) {
// Commutative, associative, idempotent: replay or reorder merges freely.
for (const [k, v] of Object.entries(o.p)) this.p[k] = Math.max(this.p[k] ?? 0, v);
for (const [k, v] of Object.entries(o.n)) this.n[k] = Math.max(this.n[k] ?? 0, v);
}
}
// Seats remaining for an event, edited in two regions during a partition.
const us = new PNCounter("us"), eu = new PNCounter("eu");
us.inc(100); // initial capacity loaded in us
eu.merge(us);
us.dec(3); eu.dec(2); // concurrent bookings
console.log("before sync us =", us.value(), " eu =", eu.value());
us.merge(eu); eu.merge(us); eu.merge(us); // duplicate merge is harmless
console.log("after sync us =", us.value(), " eu =", eu.value(), "(both converge to 95)");
// Caveat: a counter can go below zero - CRDTs converge, they do NOT enforce invariants.Output:
before sync us = 97 eu = 98
after sync us = 95 eu = 95 (both converge to 95)What happens if you choose differently
- Generic multi-master SQL for "write anywhere": autoincrement IDs collide (use UUIDs or offset sequences), unique constraints pass in both regions, foreign keys reference rows the other region deleted, and LWW drops updates without warning. Teams usually find out from customer support.
- LWW with wall clocks: a region with a clock 2 seconds fast wins every race for 2 seconds. Hybrid logical clocks (HLC) shrink the problem but don't remove the loss: concurrent writes still collapse to one.
- Single leader instead: writes from far regions pay cross-region latency, and a leader-region outage stops writes until failover. In return you get no conflicts, constraints, and simpler code. For many products, 100 ms extra on writes only is perfectly acceptable.
- Consensus across regions instead (Spanner/CockroachDB global tables): strongly consistent writes, but every commit pays a cross-region quorum RTT. It's good for low-write, high-value data.
Topologies
| Topology | Shape | Risk |
|---|---|---|
| All-to-all | Every leader sends to every other | Causality: an update can arrive before the insert it depends on, through different paths. Needs version vectors |
| Star / hub | Leaves replicate through a hub | Hub is a single point of failure and a bottleneck |
| Ring | Each forwards to the next | One node failure breaks the ring; loop prevention by node-id tagging |
Pros and cons
Pros: local write latency in every region, survives a region partition, enables offline clients, no global failover event. Cons: conflicts and their silent-loss modes, no global constraints or transactions, harder testing (you have to simulate concurrent cross-region writes), schema changes must be compatible across all leaders, and operational complexity.
Interview Q&A
Q1. Two regions update the same user's email concurrently. What happens under LWW, and how would you do better? Under LWW the write with the higher timestamp wins and the other disappears, even if it was "later" in real time on a skewed clock. Better options: avoid the conflict with a home region per user, or detect it with version vectors and resolve by a rule (for example, keep the change that went through verification), or surface it to the user.
Q2. How do you enforce unique usernames with multi-leader?
You can't safely do it locally. Options: route username claims to a single authority (a leader-per-key or consensus store), reserve namespaces per region (ajay@eu), or claim optimistically and resolve conflicts afterwards by renaming one claimant, which is a bad user experience. Most teams choose the single authority.
Q3. What makes a CRDT merge safe? The merge must be commutative, associative and idempotent (a join-semilattice). Replicas can then apply merges in any order, any number of times, and still converge. Examples include G-counters, PN-counters, OR-sets and LWW-registers. Convergence is not the same as business correctness: a PN-counter can still go negative.
Q4. How does CockroachDB offer "multi-region writes" without multi-leader conflicts?
Each range has a single Raft leaseholder. REGIONAL BY ROW tables place each row's leaseholder in that row's home region, so local users get local writes. GLOBAL tables replicate for fast reads everywhere with slower writes. Consensus keeps a single order per range, so there are no conflicts to resolve.
Q5. Why is multi-leader a natural fit for offline mobile apps? Each device is effectively a leader while offline: it accepts writes and syncs later. Conflicts are expected, so you design data as CRDTs or mergeable operations from day one (Automerge, Yjs, CouchDB revision trees).
Go Deeper
- crdt.tech: CRDT resources and papers
- CockroachDB docs: Multi-region capabilities overview
- AWS docs: DynamoDB Global Tables
- Redis docs: Active-Active geo-distribution (CRDT-based)
- Martin Kleppmann: CRDTs, the hard parts (video)
Related
- Prev: Failover & Split Brain - Detection, Promotion, Fencing & Lost Writes (
replication-failover-split-brain-fencing) - Next: Leaderless Replication - N/R/W Quorums, Hinted Handoff, Read Repair & Anti-Entropy (
leaderless-replication-quorums-read-repair-anti-entropy)
This series:
- Database Replication for Engineers - Leader-Follower, Multi-Leader & Leaderless Quorums (
database-replication-leader-follower-multi-leader-leaderless) - Sync vs Async vs Semi-Sync Replication - Durability, RPO & Physical vs Logical Logs (
replication-sync-async-semi-sync-physical-logical) - Replication Lag & Session Guarantees - Read-Your-Writes, Monotonic Reads & Consistent Prefix (
replication-lag-read-your-writes-session-guarantees) - Failover & Split Brain - Detection, Promotion, Fencing & Lost Writes (
replication-failover-split-brain-fencing) - Multi-Leader & Active-Active - Conflict Detection, LWW, CRDTs & Home Regions (
multi-leader-replication-conflicts-active-active) - this page - Leaderless Replication - N/R/W Quorums, Hinted Handoff, Read Repair & Anti-Entropy (
leaderless-replication-quorums-read-repair-anti-entropy)
Existing Study pages (cross-link only, not rewritten here):
- CRDTs - conflict-free replicated data types (
crdts-conflict-free-replicated-data) - Database sharding - partition keys & rebalancing (
db-sharding-partitioning-keys-rebalancing)