Reliability & Disaster Recovery
Part 5 of 6 · Disaster Recovery & Multi-RegionActive-Active vs Active-Passive Multi-Region - Data, Write Routing & Home Regions
Active-passive vs home-region vs multi-writer merge vs consensus topologies; runnable home-region routing with read-your-writes tokens; runnable LWW lost-update on balances; per-data-type choices; failover behavior and fencing in each topology.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Who writes in active-passive?
Answer
The primary region only. The standby may serve stale reads. RPO is the replication lag, and failover is a promotion.
L2
Design a global profile service.
Answer
Give each user a home region near them. Writes go home and replicate asynchronously. Reads are local with a read-your-writes token. On region loss, re-home and accept the lag as RPO.
L3
Why is last-writer-wins dangerous for a balance?
Answer
Two concurrent debits both read 100. LWW keeps one result. The snippet keeps 70 from us-east and drops the eu-west debit of 50. The expected balance is 20.
L4
What does fencing mean?
Answer
The old primary must be unable to accept writes before you promote: revoke credentials, cut the network, or require a fencing epoch that storage rejects.
L5
When do you pay for Spanner-style consensus?
Answer
When cross-region correctness matters more than tens of milliseconds, such as a ledger or a global uniqueness constraint, and write volume is moderate.
L6
Which data can merge?
Answer
A shopping cart can union adds. A display name can take the newest value. A balance and an oversell limit cannot.
L7
What is split brain in active-passive?
Answer
The old primary keeps accepting writes after the standby is promoted because nothing fenced it. Those writes conflict or are lost.
Failure modes
LWW on a money path
Lost updates show up later as reconciliation tickets. The snippet's missing debit is the shape.
Consensus for every write
Each write pays a cross-region round trip, often 60 to 150 ms.
Promotion without fencing
The old primary still accepts writes. That is split brain.
No plan for travelers
A user far from home gets slow writes unless you have a re-homing rule.
Misconceptions
Active-active means every region can write every key.
That is one topology. Home region and consensus are the other workable answers.
The later timestamp is a merge.
On a balance it is a lost update. The snippet result is 70, not 20.
Zero RPO from consensus means the service always takes writes.
If the failure removes the majority, writes stop until quorum returns.
Interviewer traps
Putting a ledger on DynamoDB-style LWW because the product is multi-region.
Say the newest value is not the sum of the debits.
Forgetting to fence before promote.
Name the epoch or the credential revoke before traffic moves.
Design scenario
Same prompt for every reader.
Requirements
Show the read-your-writes behavior after a profile write, and the lost update if the balance uses LWW.
Traffic / scale
The home-region snippet is one profile write. The LWW snippet is two debits on a starting balance of 100.
Latency
Home-region writes are local for users in that region and cross-region for travelers. Consensus adds a cross-region round trip per write.
Consistency
The balance must not go negative, and both debits must apply.
Availability
Home-region failover accepts lag for the failed home. Consensus stops writes if the majority is lost.
Failure assumptions
- Replicas lag.
- Clocks differ by a millisecond across regions.
- A partition can leave the old primary reachable.
Constraints
- Do not use LWW for the balance.
- Do not promote a standby that can still accept writes.
Prompt
A profile, a shopping cart, and an account balance must run in us-east and eu-west. Pick a write rule for each.
API
What does the client carry after a write, and when does eu-west forward the read?
Data
What balance does LWW keep, and what does the delta merge compute?
Architecture
Where is the single writer for the ledger, and what fences the old primary?
Who is allowed to write the record
Prefer
One home region for anything with an invariant
Balances, ledgers, and inventory with an oversell limit have a single writer. Other regions serve reads and forward when the token is ahead of the replica.
- The example token v1 falls back to us-east until eu-west has applied it.
- After replication, the same read is served in eu-west.
- Re-homing on region loss accepts that home's lag as RPO, or marks those users read-only.
Alternative
Last-writer-wins on a balance
Two debits, minus 30 and minus 50, both start from 100. The later timestamp keeps 70 and the other debit vanishes.
- Fine for a display name or a last-seen time.
- Wrong for counters, inventory, and non-negative balances.
- Commutative deltas reach 20 and still cannot enforce the invariant across regions.
Route the write, then the read that must see it
The sequence is a user in Europe whose home is us-east.
- 1
Forward the write home
The nearest edge accepts the call and sends the write to the home region. - 2
Return a version token
The commit version goes back to the client with the OK. - 3
Read locally only if caught up
If eu-west has applied the version, it serves. If not, it forwards home. - 4
Do not LWW a debit
A second region writing the same balance will drop one update when timestamps reconcile.
Overview
The compute tier is easy to run in two regions; the data is what makes multi-region hard. In active-passive, one region takes all writes and replicates to a standby that is promoted on failure: simple, but RPO equals the replication lag and failover needs a promotion step. In active-active, more than one region serves traffic. You then have to decide who may write which data. The three workable answers are: partition writes by home region (each user or tenant is owned by one region), accept concurrent writes and merge them (last-writer-wins or CRDTs, only for data where that is safe), or pay for consensus across regions (Spanner-style synchronous replication with higher write latency). Most real systems mix all three per data type.
Topologies compared
| Topology | Who writes | Read path | RPO on region loss | Write latency | Typical systems |
|---|---|---|---|---|---|
| Active-passive (async) | Primary region only | Primary, or stale reads from standby | Replication lag | Local | Aurora Global Database, cross-region read replicas |
| Active-active, home region | Each record's home region | Local, with stale reads elsewhere | Lag for that home's data | Local for local users, cross-region for travelers | Per-tenant routing, CockroachDB regional-by-row tables |
| Active-active, multi-writer merge | Any region | Local | Near zero, but conflicts | Local | DynamoDB global tables (last-writer-wins), Cassandra multi-DC |
| Active-active, consensus | Any region via quorum | Local or leader | Zero | Cross-region round trip per write | Google Spanner multi-region, CockroachDB global |
Home-region routing with read-your-writes (runnable)
# Home-region writes: each user's data has ONE writable region; other regions serve
# (possibly stale) reads. A version token gives read-your-writes after a write.
class Region:
def __init__(self, name):
self.name, self.data, self.applied_version = name, {}, 0
class Cluster:
def __init__(self):
self.regions = {"us-east": Region("us-east"), "eu-west": Region("eu-west")}
self.log = [] # replication log owned by the home region
def write(self, home, key, value):
r = self.regions[home]
self.log.append((key, value))
r.data[key] = value
r.applied_version = len(self.log)
return r.applied_version # token returned to the client
def replicate(self, region, upto):
r = self.regions[region]
for key, value in self.log[r.applied_version:upto]:
r.data[key] = value
r.applied_version = max(r.applied_version, upto)
def read(self, region, key, min_version=0, home="us-east"):
r = self.regions[region]
if r.applied_version < min_version: # replica too stale for this client: go to home region
return self.regions[home].data.get(key), f"{home} (fallback, replica at v{r.applied_version})"
return r.data.get(key), region
c = Cluster()
token = c.write("us-east", "profile:7", "new avatar")
print("eu read, no token :", c.read("eu-west", "profile:7"))
print("eu read, token v%d:" % token, c.read("eu-west", "profile:7", min_version=token))
c.replicate("eu-west", upto=token)
print("after replication :", c.read("eu-west", "profile:7", min_version=token))Output:
eu read, no token : (None, 'eu-west')
eu read, token v1: ('new avatar', 'us-east (fallback, replica at v0)')
after replication : ('new avatar', 'eu-west')The token is the important trick: after a write, the client carries the version it wrote, and any region that has not caught up to that version forwards the read to the home region instead of returning stale data.
Sequence
- 1
User in Europe (home: us-east) → eu-west region
1. Write profile (routed by nearest edge)
- 2
eu-west region → us-east region (home)
2. Forward write to home region
- 3
us-east region (home) → eu-west region
3. Commit at version v41
- 4
eu-west region → User in Europe (home: us-east)
4. OK + token v41
- 5
User in Europe (home: us-east) → eu-west region
5. Read profile with token v41
- 6
eu-west region → User in Europe (home: us-east)
6a. Serve locally
- 7
eu-west region → us-east region (home)
6b. Forward read to home
- 8
us-east region (home) → User in Europe (home: us-east)
7. Fresh value
Lesson map
Active-Active vs Active-Passive Multi-Region - Data, Write Routing & Home Regions
Active-passive vs home-region vs multi-writer merge vs consensus topologies; runnable home-region routing with read-your-writes tokens; runnable LWW lost-update on balances; per-data-type choices; failover behavior and fencing in each topology.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["User in Europe (home: us-east)"] eu["eu-west region"] us["us-east region (home)"] u -->|1. Write profile (routed by nearest edge)| eu eu -->|2. Forward write to home region| us us -->|3. Commit at version v41| eu eu -->|4. OK + token v41| u u -->|5. Read profile with token v41| eu eu -->|6a. Serve locally| u eu -->|6b. Forward read to home| us us -->|7. Fresh value| u
Why last-writer-wins can silently lose money (runnable)
// Active-active with last-writer-wins (LWW): two regions accept a write to the same key
// during a partition. When they reconcile, the "later" timestamp silently wins.
type Write = { region: string; value: number; ts: number };
// Balance starts at 100. Two concurrent debits happen in different regions.
const usEast: Write = { region: "us-east", value: 100 - 30, ts: 1_000_002 }; // -30
const euWest: Write = { region: "eu-west", value: 100 - 50, ts: 1_000_001 }; // -50 (clock 1 ms behind)
const lww = usEast.ts >= euWest.ts ? usEast : euWest;
console.log(`LWW result: ${lww.value} from ${lww.region} (expected 20 after both debits)`);
console.log(`lost update: the ${lww === usEast ? "eu-west" : "us-east"} debit vanished`);
// Fix options: route the account to a home region (no concurrent writers),
// or model the balance as commutative operations (a CRDT-style counter of deltas).
const deltas = [-30, -50];
const merged = 100 + deltas.reduce((a, b) => a + b, 0);
console.log(`delta merge (commutative): ${merged}`);
console.log(`but deltas alone cannot enforce "balance >= 0" across regions -> needs a single writer or consensus`);Output:
LWW result: 70 from us-east (expected 20 after both debits)
lost update: the eu-west debit vanished
delta merge (commutative): 20
but deltas alone cannot enforce "balance >= 0" across regions -> needs a single writer or consensusExpectedLWW result: 70 from us-east (expected 20 after both debits) lost update: the eu-west debit vanished delta merge (commutative): 20 but deltas alone cannot enforce "balance >= 0" across regions -> needs a single writer or consensus
Press Run. Snippets must be self-contained — no network, files, or native modules.
Last-writer-wins is fine for data where the newest value really is the right one (a user's display name, a device's last-seen time). It is wrong for anything that accumulates (balances, inventory, counters) or has invariants (no overselling, no negative balance). For those, either give the record a single writer (home region) or use consensus.
Choosing per data type
| Data | Recommended approach | Why |
|---|---|---|
| User profile, preferences | Multi-writer with LWW, or home region | Newest value wins is acceptable |
| Shopping cart | CRDT-style merge (union of adds) | Merging is better than losing items |
| Account balance, ledger | Home region or consensus | Invariants need one writer |
| Inventory with oversell limits | Home region per SKU or a reservation service | Concurrent decrements break the limit |
| Session tokens | Replicate everywhere, short TTL | Reads dominate, staleness is bounded |
| Analytics events | Write locally, ship asynchronously | Order rarely matters, volume is huge |
Failover in each topology
- Active-passive: detect, fence the old primary so it cannot accept writes, promote the standby, repoint applications, shift traffic. Writes in flight during the lag window are lost or must be recovered later from the old primary.
- Home region: re-home the failed region's users to a surviving region. Their data there is only as fresh as replication allowed, so you accept that RPO for those users or mark them read-only until the original region returns.
- Multi-writer merge: route everyone to surviving regions. When the failed region returns, its unreplicated writes merge in, and conflicts are resolved by the merge rule.
- Consensus: if a majority of replicas survives, nothing to do. If the failure removes the majority, writes stop until quorum returns, which is the price of zero RPO.
What happens if you choose otherwise
- Active-active everywhere with LWW: lost updates in money paths that surface weeks later as reconciliation tickets.
- Consensus for everything: every write pays a cross-region round trip (often 60 to 150 ms), which hurts latency-sensitive paths.
- Active-passive with no fencing: during a network partition, the old primary keeps accepting writes after the standby is promoted. That is split brain, and those writes conflict or are lost.
- No plan for travelers: a user whose home region is far away gets slow writes, so you need re-homing rules.
Pros and cons
| Approach | Pros | Cons |
|---|---|---|
| Active-passive | Simple model, one writer, cheap standby | RPO equals lag, promotion step, idle capacity |
| Home region | No write conflicts, local latency for most users | Re-homing logic, cross-region writes for travelers |
| Multi-writer merge | Local writes everywhere, high availability | Conflicts, only safe for mergeable data |
| Consensus | Zero RPO, strong consistency | Write latency, cost, quorum loss blocks writes |
Interview Q&A
Design a global user profile service with users on three continents.
Answer
Assign each user a home region near them. Writes go to the home region and replicate asynchronously elsewhere. Reads are served locally with a read-your-writes token so a user always sees their own latest change. On region loss, re-home affected users to the nearest survivor and accept the lag window as RPO.
Why is last-writer-wins dangerous for a bank balance?
Answer
Two concurrent debits in different regions both read the same starting balance; LWW keeps one result and silently drops the other debit. Balances need a single writer or consensus.
What does fencing mean in an active-passive failover?
Answer
Making sure the old primary can no longer accept writes, for example by revoking its credentials, isolating it from the network, or using a fencing token or epoch that storage rejects, before promoting the standby.
When would you accept the write latency of Spanner-style consensus?
Answer
For data where correctness across regions matters more than tens of milliseconds, such as financial ledgers or global uniqueness constraints, and when write volume is moderate.
What does the read-your-writes token do?
Answer
After a write, the client carries the version. A replica that has not applied that version forwards the read to the home region instead of returning stale or missing data.
Why can deltas add to 20 and still be the wrong design?
Answer
Commutative deltas cannot enforce balance at least zero across regions. That still needs a single writer or consensus.
What is the RPO when a home region fails?
Answer
The replication lag for the users whose home failed, unless you mark them read-only until that region returns.
What happens to consensus writes if the failure removes the majority?
Answer
Writes stop until quorum returns. That is the price of zero RPO.
Check yourself
Add a third debit in the LWW snippet with an even later timestamp. State the LWW result and the delta-merge result, and say which data types on the page would still allow LWW.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Multi-leader conflicts, CRDTs, Raft consensus, Failover and fencing.