Reliability & Disaster Recovery
Part 1 of 6 · Disaster Recovery & Multi-RegionDisaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-Active
Interview hub: HA vs DR vs backup, RTO vs RPO, the four DR strategies (backup/restore, pilot light, warm standby, active-active) with cost tiers, RTO as a phase budget, hidden single-region dependencies; replication is not backup.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does RTO measure?
Answer
How long the business can tolerate the service being down.
L2
What does RPO measure?
Answer
How much recent data can be lost, as a time window, such as the last 5 minutes of writes.
L3
Why is replication not a backup?
Answer
A replica copies every change, including a DELETE without a WHERE. Only a backup or a PITR log stopped before that statement saves you.
L4
Is multi-AZ a disaster recovery plan?
Answer
No. It is high availability. It survives a zone, not a region-wide control plane event or logical corruption that replicates to every zone.
L5
How do you choose a strategy for a new service?
Answer
Price downtime and lost data, turn that into RTO and RPO, map dependencies, pick the cheapest strategy that meets both, and drill it.
L6
What is usually the biggest RTO win?
Answer
Cutting decision time with a pre-agreed rule, not shaving the promote script.
L7
What hidden dependency sets the real RTO?
Answer
Whatever the request path still needs in one region: identity, KMS, the control plane, the image registry, observability, or a single-region third party.
Failure modes
Replication copies the bad DELETE
Synchronous replicas apply the statement within milliseconds. There is no earlier point to stop.
Decision time blows the 30 minute target
The example manual plan totals 60 minutes. Warm standby removes scale-up and a pre-approved decision brings the sum to 30.
The recovery region cannot see its dependencies
Secrets, images, dashboards, or DNS APIs that live only in the failed region are unavailable during the event.
Misconceptions
Multi-AZ deployment is a disaster recovery plan.
It is high availability. Region loss and logical corruption are still open.
RTO is how fast the failover script runs.
The clock starts when users are hurt. Detect and decide are part of the budget.
A replica is a backup.
Replication copies the damage. Backups and PITR logs are how you return to a time before it.
Interviewer traps
Promising an RPO from the replication brochure and ignoring logical corruption.
State the lag window and the last restorable point separately.
Putting every service on active-active because it sounds safer.
Show the example cost tiers and say most services should not pay for a second full region.
Design scenario
Same prompt for every reader.
Requirements
Meet each service's RTO and RPO with the cheapest standing cost. Name one dependency that would invalidate the tier.
Traffic / scale
Example targets from the picker: wiki 1440/1440 minutes, order history 240/15, checkout 15/1, ledger 2/0.5.
Latency
Recovery time includes detect and decide, not only the promote command.
Consistency
Replication must not be treated as the backup for a bad delete.
Availability
A regional event has to land on a strategy that was already standing.
Failure assumptions
- One region can become unusable.
- A bad delete can replicate before anyone notices.
- Control-plane APIs in the failed region may be degraded.
Constraints
- Do not put the wiki on active-active.
- Do not leave the ledger on backup-and-restore.
Prompt
Four services share a company: an internal wiki, order history, checkout, and a payments ledger. Assign each a DR strategy.
API
Which phases are in the RTO budget before users are served again?
Data
What still protects you after a DELETE without a WHERE replicates?
Architecture
Where do identity, images, and dashboards have to exist before the drill?
What you are actually protecting
Prefer
A tier whose strategy meets both numbers
Price downtime and lost writes per journey, then take the cheapest strategy that meets that RTO and that RPO. Prove it with a drill.
- Most services do not need active-active.
- The payments ledger and the internal wiki should not share a strategy.
- A tier 0 path that calls a tier 3 dependency inherits the worse recovery.
Alternative
One strategy for the whole company
Active-active everywhere doubles spend and forces every team to solve conflicts. Backup-and-restore everywhere turns a region event into a day-long outage for checkout.
- Complexity from unused active-active paths causes its own outages.
- Restoring every service at once overwhelms the recovery region.
- The standing cost does not follow business impact.
From the failure to the number that matters
The hub map. Sibling pages hold the strategy math, backups, cells, data topology, and the failover runbook.
- 1
Name what broke
A node or AZ is high availability. A region is disaster recovery. Wrong data is a backup problem. - 2
Set RTO and RPO from the business
RTO is downtime. RPO is lost writes, as a time window. The restore script's runtime is only one phase. - 3
Pick the cheapest strategy that meets both
Backup/restore, pilot light, warm standby, then active-active. The example picker sends the wiki to backups and the ledger to active-active. - 4
Hunt the single-region dependency
Identity, the control plane, the registry, and dashboards in the failed region set the real RTO.
Overview
Disaster recovery (DR) is the plan for getting a service back when something bigger than a single server fails: a whole zone, a whole region, a corrupted database, a deleted bucket, ransomware, or a bad deploy that poisons data everywhere at once. Two numbers drive every DR decision. RTO (Recovery Time Objective) is how long the business can tolerate the service being down. RPO (Recovery Point Objective) is how much recent data the business can tolerate losing, measured in time ("we can lose at most the last 5 minutes of writes"). Lower RTO and RPO cost more, because they need more infrastructure running before the disaster happens.
The industry vocabulary comes in four standard strategies, ordered from cheapest and slowest to most expensive and fastest: backup and restore, pilot light, warm standby, and multi-site active-active. A senior engineer's job is not to pick the fanciest one. It is to tier services by business impact, pick the cheapest strategy that meets each tier's RTO and RPO, and then prove with drills that the plan actually works. A DR plan that has never been exercised is a hope, not a plan.
This cluster covers the whole picture:
- RTO, RPO and the four DR strategies, with the cost math behind them.
- Backups that actually restore: snapshots, point-in-time recovery, immutable copies and restore drills.
- Multi-AZ vs multi-region: blast radius, cell-based architecture and static stability.
- Active-active vs active-passive data: home regions, write routing, conflicts.
- Region failover in practice: runbooks, DNS and global load balancer cutover, failback and game days.
High availability vs disaster recovery vs backup
These three get mixed up constantly in interviews. They protect against different failures.
| Concern | Protects against | Typical mechanism | Does NOT protect against |
|---|---|---|---|
| High availability (HA) | One machine, disk or zone failing | Multi-AZ replicas, load balancers, auto-healing | Region loss, logical corruption, deletes |
| Disaster recovery (DR) | A whole region or site becoming unusable | Standby or active region, cross-region replication | Bad data that replicates everywhere |
| Backup | Data being destroyed or corrupted (bugs, humans, ransomware) | Snapshots, PITR logs, immutable offsite copies | Fast recovery on its own (restores are slow) |
The key insight: replication is not backup. If someone runs DELETE FROM orders without a WHERE, a synchronous replica faithfully deletes everything within milliseconds. Only a backup taken before the delete, or a point-in-time log you can stop replaying before the bad statement, saves you.
Decisions
- 1
1. Failure happens
- next2. What broke?
- ?
2. What broke?
- one node or AZ3a. HA: replicas in other AZs take over in seconds
- whole region3b. DR: shift traffic to another region
- data itself is wrong3c. Backup: restore to a point before the damage
- 3
3a. HA: replicas in other AZs take over in seconds
- next4. Users served again
- 4
3b. DR: shift traffic to another region
- next4. Users served again
- 5
3c. Backup: restore to a point before the damage
- next4b. Reconcile writes made after the damage
- 6
4. Users served again
- 7
4b. Reconcile writes made after the damage
- next4. Users served again
Lesson map
Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-Active
Interview hub: HA vs DR vs backup, RTO vs RPO, the four DR strategies (backup/restore, pilot light, warm standby, active-active) with cost tiers, RTO as a phase budget, hidden single-region dependencies; replication is not backup.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Failure happens"] b["2. What broke?"] c["3a. HA: replicas in other AZs take over in seconds"] d["3b. DR: shift traffic to another region"] e["3c. Backup: restore to a point before the damage"] f["4. Users served again"] g["4b. Reconcile writes made after the damage"] a -->|continues| b b -->|one node or AZ| c b -->|whole region| d b -->|data itself is wrong| e c -->|continues| f d -->|continues| f e -->|continues| g g -->|continues| f
The four DR strategies compared
| Strategy | What runs in the recovery region before a disaster | Typical RTO | Typical RPO | Standing cost |
|---|---|---|---|---|
| Backup and restore | Nothing but backups copied cross-region | Hours to a day | Hours (last backup) or minutes with PITR logs | Lowest |
| Pilot light | Data replicated continuously; core infra defined but scaled to zero or minimal | Tens of minutes to hours | Seconds to minutes | Low |
| Warm standby | A smaller but fully working copy of the stack, already serving health checks | Minutes | Seconds | Medium |
| Multi-site active-active | Full capacity in two or more regions, all serving live traffic | Near zero (seconds) | Near zero, depending on data model | Highest (often 2x or more) |
The ranges are planning guides, not guarantees. Your real RTO depends on detection time, decision time and how much of the runbook is automated.
Choosing the cheapest strategy that meets the targets (runnable)
# Pick the cheapest DR strategy that still meets a service's RTO and RPO targets.
# Numbers are illustrative planning ranges, not vendor quotes.
from dataclasses import dataclass
@dataclass
class Strategy:
name: str
rto_min: float # typical worst-case recovery time, minutes
rpo_min: float # typical worst-case data loss window, minutes
standing_cost: float # monthly cost as a multiple of one production region (1.0 = a full copy)
STRATEGIES = [
Strategy("backup-and-restore", rto_min=12 * 60, rpo_min=24 * 60, standing_cost=0.05),
Strategy("pilot-light", rto_min=60, rpo_min=5, standing_cost=0.15),
Strategy("warm-standby", rto_min=15, rpo_min=1, standing_cost=0.40),
Strategy("active-active", rto_min=1, rpo_min=0.1, standing_cost=1.00),
]
def pick(rto_target_min: float, rpo_target_min: float) -> Strategy:
# Keep only strategies that meet BOTH targets, then take the cheapest one.
ok = [s for s in STRATEGIES if s.rto_min <= rto_target_min and s.rpo_min <= rpo_target_min]
if not ok:
raise ValueError("no strategy meets these targets; revisit the requirement or the design")
return min(ok, key=lambda s: s.standing_cost)
services = {
"internal-wiki": (24 * 60, 24 * 60), # a day of downtime and a day of lost edits is tolerable
"order-history": (4 * 60, 15), # hours of downtime OK, but only minutes of data loss
"checkout": (15, 1), # revenue path: minutes, near-zero loss
"payments-ledger": (2, 0.5), # must survive a region loss almost transparently
}
for svc, (rto, rpo) in services.items():
s = pick(rto, rpo)
print(f"{svc:16} RTO<={rto:>6.0f}m RPO<={rpo:>6.1f}m -> {s.name:18} (~{s.standing_cost:.2f}x region cost)")Output:
internal-wiki RTO<= 1440m RPO<=1440.0m -> backup-and-restore (~0.05x region cost)
order-history RTO<= 240m RPO<= 15.0m -> pilot-light (~0.15x region cost)
checkout RTO<= 15m RPO<= 1.0m -> warm-standby (~0.40x region cost)
payments-ledger RTO<= 2m RPO<= 0.5m -> active-active (~1.00x region cost)Notice that most services in a typical company do not need active-active. Paying a full extra region for an internal wiki is waste; under-protecting the payments ledger is a business risk.
RTO is a budget, not a script runtime (runnable)
// RTO is not "how fast the failover script runs". It is the sum of every phase
// from the moment users are hurt until they are served again.
type Phase = { name: string; minutes: number; automatable: boolean };
const manualPlan: Phase[] = [
{ name: "detect (alerts fire, on-call paged)", minutes: 5, automatable: true },
{ name: "triage + decide to fail over", minutes: 20, automatable: false },
{ name: "promote DB replica in region B", minutes: 5, automatable: true },
{ name: "scale app tier in region B", minutes: 15, automatable: true },
{ name: "shift traffic (DNS TTL drain)", minutes: 5, automatable: true },
{ name: "validate (smoke tests, dashboards)", minutes: 10, automatable: false },
];
function total(plan: Phase[]): number {
return plan.reduce((sum, p) => sum + p.minutes, 0);
}
const rtoTarget = 30;
console.log(`manual plan total: ${total(manualPlan)} min (target ${rtoTarget})`);
// Warm standby removes the scale-up phase; a pre-approved runbook shrinks the decision.
const improved = manualPlan.map((p) =>
p.name.startsWith("scale") ? { ...p, minutes: 0 } :
p.name.startsWith("triage") ? { ...p, minutes: 5 } : p,
);
console.log(`warm standby + pre-approved decision: ${total(improved)} min`);
const biggest = [...improved].sort((a, b) => b.minutes - a.minutes)[0];
console.log(`largest remaining phase: ${biggest.name} (${biggest.minutes} min)`);Output:
manual plan total: 60 min (target 30)
warm standby + pre-approved decision: 30 min
largest remaining phase: validate (smoke tests, dashboards) (10 min)Expectedmanual plan total: 60 min (target 30) warm standby + pre-approved decision: 30 min largest remaining phase: validate (smoke tests, dashboards) (10 min)
Press Run. Snippets must be self-contained — no network, files, or native modules.
The biggest win usually comes from cutting decision time: a pre-agreed rule like "if region health checks fail for 10 minutes and the provider confirms a regional event, the incident commander fails over without further approval".
Why tiering beats one strategy for everything, and what happens if you choose otherwise
- Everything active-active: you pay for double capacity everywhere, every team must solve multi-region data conflicts, and the complexity itself causes outages. Most companies that try this end up with a few truly active-active paths and a lot of accidental single-region dependencies.
- Everything backup and restore: cheap, but a region outage becomes a day-long outage for revenue-critical paths, and restoring dozens of services at once overwhelms the recovery region and the on-call team.
- Tiered (for example tier 0 active-active, tier 1 warm standby, tier 2 pilot light, tier 3 backups only): spend follows business impact. The cost is discipline: every service needs an owner who knows its tier and its dependencies, because a tier 0 service that synchronously calls a tier 3 service inherits tier 3 recovery.
The dependency trap
A service's real RTO is the worst RTO of anything it needs in the request path. Common hidden single-region dependencies:
- Identity and secrets: the login service, IAM, KMS keys or a secrets manager that only exists in the primary region.
- Control planes: the DR plan needs to launch instances or change DNS, but those APIs are degraded during the very regional event you are escaping (AWS us-east-1, December 2021, is the classic example).
- CI/CD and container registries: you cannot deploy into the recovery region if images only live in the primary.
- Observability: dashboards and alerting hosted in the failed region go dark exactly when you need them.
- Third parties: a payment provider or SaaS API that is itself single-region.
Pros and cons at a glance
| Approach | Pros | Cons |
|---|---|---|
| Backup and restore | Cheap, simple, protects against logical corruption | Slow RTO, restores rarely tested, large RPO without PITR |
| Pilot light | Data is current, infra is code-defined, low cost | Scale-up time, capacity may be unavailable during a regional rush |
| Warm standby | Minutes of RTO, standby is continuously exercised by health checks | Paying for idle capacity, standby can drift from primary |
| Active-active | Near-zero RTO, capacity is proven daily | Highest cost, hard data consistency, cross-region latency |
Interview Q&A
What is the difference between RTO and RPO?
Answer
RTO is how long you can be down; RPO is how much data you can lose, measured as a time window. A nightly backup gives an RPO up to 24 hours regardless of how fast you restore. Synchronous replication can give RPO near zero but says nothing about how fast you can recover.
Is multi-AZ deployment a disaster recovery plan?
Answer
It is high availability, not full DR. It survives a data center or zone failure, but not a region-wide control plane event, a regional network failure, or logical corruption that replicates to every zone.
How would you choose a DR strategy for a new service?
Answer
Start with business impact: cost per minute of downtime and per minute of lost data, plus regulatory requirements. Turn that into RTO and RPO targets, map dependencies, then pick the cheapest strategy that meets the targets and schedule a drill to prove it.
Why is replication not a backup?
Answer
Replication copies every change, including the bad ones. Backups and point-in-time logs let you go back to before the damage.
What is the hardest part of DR in practice?
Answer
Not the technology. It is keeping the plan current as the system changes, finding hidden single-region dependencies, and actually rehearsing the failover so people trust it during a real incident.
When does the RTO clock start?
Answer
When users are hurt, not when the failover command runs. Detection and the decision to fail over are phases in the budget. In the example manual plan those two phases are 5 and 20 minutes before any promote step.
Why does a tier 0 service that calls a tier 3 dependency inherit the worse recovery?
Answer
The real RTO is the worst RTO of anything on the request path. A checkout path that synchronously needs a tier 3 service cannot come back faster than that dependency.
What did the December 2021 us-east-1 event show about control planes?
Answer
A plan that must launch instances or change DNS can stall when those APIs are degraded by the same regional event you are leaving.
Check yourself
Take one service you run. Write its tolerable downtime and its tolerable lost-write window. Run it through the four strategies in the picker and say which dependency would still pin it to the primary region.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Database replication, Global load balancing, Resilience patterns, SLOs and error budgets, Load, chaos, and production validation.