Reliability & Disaster Recovery
Part 4 of 6 · Disaster Recovery & Multi-RegionMulti-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static Stability
Host/AZ/region/cell failure domains; availability math and why correlated regional dependencies cap it; static stability and data plane vs control plane; cell-based architecture and shuffle sharding (runnable); when multi-region is worth it.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does an AZ protect that a host does not?
Answer
A data center power, cooling, or network event. Latency to the other zones is about 1 to 2 ms, so synchronous replication is practical.
L2
Why are three 99.9 percent AZs not nine nines?
Answer
The formula assumes independent failures. A shared regional control plane, DNS, or a bad deploy is correlated and sits in series.
L3
What is static stability?
Answer
The system keeps working during a failure without having to change anything. The classic shape is 9 instances when you need 6, so one lost zone still leaves full capacity.
L4
How do cells differ from shards?
Answer
Sharding usually splits the data layer. A cell is a full vertical slice, including compute, queues, and data, so a bad deploy cannot spread.
L5
What does the cell snippet report?
Answer
With 8 cells, the bad cell hits 124 of 1000 customers, 12.4 percent. Shuffle sharding fully hits 34 who share both nodes, and 432 share one node and can retry.
L6
When do you add a region?
Answer
When a regional outage's cost justifies it, when users are global and latency matters, or when data must reside in a specific region. The regions must not share the request path.
L7
What did December 2021 show about reacting?
Answer
Control plane APIs and monitoring in the region were impaired. Plans that needed those APIs to replace capacity struggled.
Failure modes
Single AZ to save cross-AZ transfer
One data center event takes the tier fully down.
Multi-region with a shared dependency
Both regions call the same global auth or config service and fail together.
No cells
A bad config push or poison message reaches 100 percent of customers before the alarm fires.
Too many tiny cells
The router, migrations, and cross-cell search become their own burden.
Misconceptions
Availability multiplies when you add zones.
Only if the failures are independent. The example regional dependency at 0.9995 leaves about 262.80 minutes down per year.
Auto-scaling is a disaster plan.
It needs the control plane, which may be the impaired component.
A cell is just a shard.
A cell includes the stack above the data, so the blast radius is the customer slice, not only a partition.
Interviewer traps
Quoting nine nines from three zones.
Put the shared dependency in series and recompute.
Calling multi-region done while both regions use one identity service.
The second region does not help that request path.
Design scenario
Same prompt for every reader.
Requirements
Use the example availability numbers. Say what a bad deploy does with and without cells.
Traffic / scale
Example zone availability 0.999. Example regional dependency 0.9995. 1000 customers, 8 cells.
Latency
Cross-AZ is about 1 to 2 ms. Cross-region is 10 to 200 ms or more, so replication is usually asynchronous.
Consistency
Each cell has its own database. Cross-cell queries are a product decision, not a free join.
Availability
Losing one zone must leave full capacity without a launch API call.
Failure assumptions
- The control plane can be down with the zone event.
- A deploy can be wrong.
- A noisy tenant can overload its nodes.
Constraints
- Do not rely on auto-scaling during the regional event.
- Do not share one global auth path across the regions you claim are independent.
Prompt
A stateless tier needs 6 instances of capacity. It runs in three AZs and also calls a regional IAM dependency. Decide the instance count and whether a second region helps.
API
How many instances do you run, and why not 6?
Data
What is the downtime math once the regional dependency is in series?
Architecture
What does a cell contain that a shard does not?
Survive the zone without asking the control plane
Prefer
Static stability at 150 percent
Need 6 instances, run 3 in each of 3 AZs. Losing one zone leaves 6 already running. No launch API call.
- Users stay up while the control plane is impaired.
- The data plane keeps the last known good configuration.
- The example 3-AZ math only holds if failures are independent.
Alternative
Reactive scale-out from 100 percent
Run 6 and ask auto-scaling to replace the lost zone. If the regional control plane is down, you stay at about 66 percent.
- The APIs you need are often impaired by the same event.
- Three independent 99.9 percent zones are not nine nines once a regional dependency is in series.
- A bad deploy with no cells reaches every customer.
Shrink the blast radius on purpose
The first diagram is static stability. The second is the cell router and a one-cell deploy wave.
- 1
Put the tier in three zones
Synchronous replication is practical at about 1 to 2 ms. A zone failure should not be a region failover. - 2
Pre-provision the spare
If one zone dies, the other two already hold full capacity. - 3
Isolate a cell
A cell is compute, queues, and its own data for a fixed set of customers. A bad deploy bakes in one cell first. - 4
Watch the shared dependency
IAM, DNS, or a global config service in series caps the AZ math and can couple two regions.
Overview
Cloud providers give you two nested failure domains. An availability zone (AZ) is one or more data centers with independent power, cooling and networking, a few milliseconds from the other zones in the same region. Spreading across AZs is cheap (synchronous replication is practical at that latency) and protects against the most common large failures: a data center power or network event. Spreading across regions protects against region-wide failures, such as a regional control plane or network problem, at the cost of tens to hundreds of milliseconds of latency, asynchronous data, and much more complexity. Inside either, cells and shuffle sharding shrink the blast radius of the failure you cause yourself most often: a bad deploy or a poison request.
Failure domains compared
| Failure domain | Typical cause | Latency between copies | Data replication | What it protects |
|---|---|---|---|---|
| Host or rack | Disk, NIC, power supply, top-of-rack switch | Sub-millisecond | Synchronous | One machine |
| Availability zone | Data center power, cooling, network, fire | About 1 to 2 ms | Synchronous is practical | A building or campus |
| Region | Regional network, control plane, widespread software bug | 10 to 200+ ms | Usually asynchronous | A metro area or provider region |
| Cell | Your own bad deploy, poison request, noisy tenant | Same as host or AZ | Independent per cell | A slice of customers |
| Provider or global service | Global DNS, BGP, identity, a provider-wide bug | n/a | n/a | Needs multi-provider or out-of-band paths |
Availability math, and where it lies (runnable)
# Availability math for zones and regions, and why correlated failures break the "multiply" rule.
def parallel(*avail: float) -> float:
# Service is up if ANY copy is up (assumes independent failures).
down = 1.0
for a in avail:
down *= (1 - a)
return 1 - down
def serial(*avail: float) -> float:
# Request needs EVERY dependency to be up.
up = 1.0
for a in avail:
up *= a
return up
az = 0.999 # one availability zone (example)
minutes_per_year = 525_600
for label, a in [("1 AZ", az), ("3 AZs (independent)", parallel(az, az, az))]:
print(f"{label:24} {a:.9f} -> {(1 - a) * minutes_per_year:8.2f} min down/yr")
# Reality check: a shared regional dependency (control plane, IAM, DNS) is in SERIES with the AZs.
regional_dep = 0.9995
multi_az = serial(parallel(az, az, az), regional_dep)
print(f"{'3 AZs + regional dep':24} {multi_az:.9f} -> {(1 - multi_az) * minutes_per_year:8.2f} min down/yr")
two_regions = parallel(multi_az, multi_az)
print(f"{'2 regions of that':24} {two_regions:.9f} -> {(1 - two_regions) * minutes_per_year:8.2f} min down/yr")Output:
1 AZ 0.999000000 -> 525.60 min down/yr
3 AZs (independent) 0.999999999 -> 0.00 min down/yr
3 AZs + regional dep 0.999499999 -> 262.80 min down/yr
2 regions of that 0.999999750 -> 0.13 min down/yrMultiplying independent AZ availabilities gives absurd numbers like nine nines. Real failures are correlated: a shared regional dependency is in series with every zone, so it caps you. That is why multi-region exists, and also why multi-region only helps if the regions truly share nothing in the request path.
Static stability
Static stability (an AWS Builders' Library term) means the system keeps working during a failure without needing to make any changes. The classic example: if you run in three AZs and need 6 instances of capacity, a statically stable design runs 9 (3 per zone) so that losing a zone leaves 6 still running. A design that runs 6 and relies on auto-scaling to replace the lost zone's capacity depends on the control plane, which may be the thing that is impaired.
Architecture
Statically stable: 150% pre-provisioned
- 1
1. AZ-a fails
- next2. AZ-b and AZ-c already hold full capacity
- 2
2. AZ-b and AZ-c already hold full capacity
- next3. No API calls needed, users unaffected
- 3
3. No API calls needed, users unaffected
Reactive: 100% provisioned
- 4
1. AZ-a fails
- next2. Auto-scaling must launch replacements
- 5
2. Auto-scaling must launch replacements
- next3. Control plane healthy?
- 6
3. Control plane healthy?
- yes4a. Recovers in minutes
- no, regional event4b. Stuck at 66% capacity
- 7
4a. Recovers in minutes
- 8
4b. Stuck at 66% capacity
Lesson map
Multi-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static Stability
Host/AZ/region/cell failure domains; availability math and why correlated regional dependencies cap it; static stability and data plane vs control plane; cell-based architecture and shuffle sharding (runnable); when multi-region is worth it.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["1. AZ-a fails"] r1["1. AZ-a fails"] s2["2. AZ-b and AZ-c already hold full capacity"] r2["2. Auto-scaling must launch replacements"] s3["3. No API calls needed, users unaffected"] r3["3. Control plane healthy?"] r4["4a. Recovers in minutes"] r5["4b. Stuck at 66% capacity"] s1 -->|continues| s2 s2 -->|continues| s3 r1 -->|continues| r2 r2 -->|continues| r3 r3 -->|yes| r4 r3 -->|no, regional event| r5
The same idea applies to data planes versus control planes: the data plane (serving requests, routing packets) should keep running on its last known good configuration even when the control plane (APIs that change configuration) is down.
Cell-based architecture and shuffle sharding (runnable)
A cell is a complete, independent copy of the service stack serving a fixed subset of customers. A thin routing layer maps each customer to a cell. Deploys roll out one cell at a time, so a bad change hurts one cell's customers before it is caught.
// Cell-based architecture: customers are pinned to one of N independent cells.
// A bad deploy or poison request takes out one cell, not the whole service.
// Shuffle sharding goes further: each customer gets a random PAIR of nodes, so two customers
// rarely share their whole failure domain.
function hash(s: string): number {
let h = 2166136261; // FNV-1a, small and deterministic
for (const c of s) { h ^= c.charCodeAt(0); h = Math.imul(h, 16777619) >>> 0; }
return h;
}
const CELLS = 8;
const customers = Array.from({ length: 1000 }, (_, i) => `cust-${i}`);
const cellOf = (c: string) => hash(c) % CELLS;
const badCell = 3; // a bad deploy lands on cell 3 first
const hit = customers.filter((c) => cellOf(c) === badCell).length;
console.log(`cells=${CELLS}: bad cell hits ${hit}/${customers.length} customers (${(100 * hit / customers.length).toFixed(1)}%)`);
// Shuffle sharding: 8 nodes, each customer assigned 2 distinct nodes -> C(8,2)=28 possible shards.
const shardOf = (c: string): [number, number] => {
const h = hash(c);
const a = h % 8;
const b = (a + 1 + ((h >>> 8) % 7)) % 8; // second node: different from a, picked from other hash bits
return [Math.min(a, b), Math.max(a, b)];
};
const noisy = shardOf("cust-42"); // a noisy customer overloads both of its nodes
const fullyHit = customers.filter((c) => {
const [x, y] = shardOf(c);
return x === noisy[0] && y === noisy[1];
}).length;
const partlyHit = customers.filter((c) => {
const [x, y] = shardOf(c);
return (x === noisy[0] || x === noisy[1] || y === noisy[0] || y === noisy[1]) && !(x === noisy[0] && y === noisy[1]);
}).length;
console.log(`shuffle sharding: ${fullyHit}/${customers.length} share BOTH nodes with the noisy customer (fully impacted)`);
console.log(`${partlyHit} share one node and still have a healthy second node to retry on`);Output:
cells=8: bad cell hits 124/1000 customers (12.4%)
shuffle sharding: 34/1000 share BOTH nodes with the noisy customer (fully impacted)
432 share one node and still have a healthy second node to retry onExpectedcells=8: bad cell hits 124/1000 customers (12.4%) shuffle sharding: 34/1000 share BOTH nodes with the noisy customer (fully impacted) 432 share one node and still have a healthy second node to retry on
Press Run. Snippets must be self-contained — no network, files, or native modules.
With 8 cells, a poisoned cell hits about 1/8 of customers. With shuffle sharding, each customer gets a random combination of nodes, so a noisy or poisonous customer fully impacts only those few who happen to share its exact combination; everyone else who overlaps on one node can retry on their other node.
Flow
- 1
1. Request with customer id
- next2. Cell router: lookup table or hash
- 2
2. Cell router: lookup table or hash
- next3a. Cell 1: full stack + own DB
- next3b. Cell 2: full stack + own DB
- next3c. Cell N: full stack + own DB
- 3
3a. Cell 1: full stack + own DB
- 5. bake time + health check passes3b. Cell 2: full stack + own DB
- 4
3b. Cell 2: full stack + own DB
- 5
3c. Cell N: full stack + own DB
- 6
4. Deploy wave 1 targets one cell only
- related3a. Cell 1: full stack + own DB
Where each pattern is used
- Multi-AZ is the default for every production database and stateless tier. Managed databases offer it as a checkbox (synchronous standby or quorum storage across zones).
- Multi-region is used for tier 0 journeys, latency to global users, and regulatory data residency.
- Cells are used by large SaaS and cloud providers themselves to bound the impact of deploys and noisy tenants.
- Shuffle sharding is used for multi-tenant front ends such as DNS and API gateways.
What happens if you choose otherwise
- Single AZ to save cross-AZ data transfer cost: a single data center event takes you fully down.
- Multi-region without isolating dependencies: both regions call the same global auth or config service, so you pay double and still fail together.
- No cells: a bad config push or poison message reaches 100% of customers before the alarm fires.
- Too many tiny cells: the router, migrations and capacity planning become their own operational burden, and cross-cell features (search across all customers) get hard.
Pros and cons
| Pattern | Pros | Cons |
|---|---|---|
| Multi-AZ | Cheap, synchronous data, transparent to the app | Does not survive regional or logical failures |
| Multi-region | Survives regional events, lower latency for global users | Async data, cost, complex failover |
| Static stability | Survives control plane failures | Pays for spare capacity all the time |
| Cells | Bounded blast radius for deploys and tenants | Routing layer, cross-cell queries, more units to operate |
| Shuffle sharding | Isolates noisy tenants cheaply | Needs client retries across nodes, harder to reason about |
Interview Q&A
Three AZs at 99.9% each gives nine nines. Why do real systems not get that?
Answer
The math assumes independent failures. Shared dependencies (regional control planes, DNS, a bad deploy pushed everywhere) are correlated and sit in series, so they dominate.
What is static stability and why does it matter during a regional event?
Answer
The system survives a failure without having to change anything, for example by pre-provisioning enough capacity in the remaining zones. It matters because the APIs you would use to react are often impaired by the same event.
How do cells differ from shards?
Answer
Sharding usually splits only the data layer. A cell is a full vertical slice of the stack, including compute, queues and data, so a failure or bad deploy in one cell cannot spread to others.
When would you choose multi-region over just multi-AZ?
Answer
When the cost of a regional outage justifies the extra spend and complexity, when users are global and latency matters, or when regulations require data to live in specific regions.
Why run 9 instances when you need 6?
Answer
Three AZs with 3 each still have 6 after one zone is gone, and you do not call auto-scaling to get there.
What is the data plane versus the control plane here?
Answer
The data plane serves requests on the last known good configuration. The control plane is the API that changes that configuration, and it may be down during the event.
What blast radius does shuffle sharding change?
Answer
A noisy customer fully hits only the others who share its exact node pair. Customers who overlap on one node can retry the other.
What fails if both regions call the same global auth service?
Answer
You pay for two regions and still fail together.
Check yourself
Re-run the availability snippet with the regional dependency removed, then with it set to 0.99. Write the minutes per year and say whether a second region still helps if both regions call that same dependency.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Resilience patterns, Consistent hashing, Private networking, Feature flags.