Reliability & Disaster Recovery
Part 2 of 6 · Disaster Recovery & Multi-RegionRTO, RPO & DR Strategies - Backup/Restore vs Pilot Light vs Warm Standby vs Active-Active
Business impact analysis to tiers with RTO/RPO targets; each DR strategy in depth; expected-yearly-cost math that picks the strategy; measuring real RPO from p99 replication lag instead of averages; capacity and decision-time pitfalls.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Where do RTO and RPO come from?
Answer
From the cost of a minute of downtime and a minute of lost writes, plus regulatory requirements.
L2
What runs in pilot light before the disaster?
Answer
The data replica is live. Compute is defined and scaled to zero or near zero.
L3
Why keep the data layer on in the cheap strategy?
Answer
Moving data is the slow part. A 10 TB restore at 500 MB/s is already more than 5 hours.
L4
What RPO do you promise when average lag is 300 ms?
Answer
Not 300 ms. Use p99 and the maximum over a long window, including batch jobs, then add the other async paths.
L5
When is active-active not worth it?
Answer
When the outage cost you avoid is smaller than the standing cost and the conflict complexity. Warm standby often captures most of the benefit.
L6
Why is warm standby more reliable in practice than pilot light?
Answer
The standby is already running and checked, so expired certificates and missing secrets show up before the disaster.
L7
How do you cut RTO without more infrastructure?
Answer
Faster detection, a pre-approved decision, an automated runbook, and a global load balancer or a short DNS TTL.
Failure modes
Targets set by engineers alone
Low-value services get gold-plated, or checkout is assumed to have an RPO of zero only after the outage.
RTO measured from the failover command
Detection and decision often exceed the time the script takes.
Pilot light with no capacity reservation
During a regional event, many customers scale up in the same neighboring region at once.
Misconceptions
Average replication lag is the RPO.
The example average is 4.06 seconds while p99 is 45 seconds.
Active-active is always the responsible choice.
At the example rates it is the most expensive yearly total and buys little beyond warm standby.
Managed cross-region lag is the whole RPO.
Queues, object replication, and third parties add their own windows.
Interviewer traps
Quoting the mean lag from the dashboard.
Show p99, the spike during vacuum or backfill, and the alert tied to the RPO.
Ignoring restore throughput for backup-and-restore.
A 10 TB database at 500 MB/s is hours before validation starts.
Design scenario
Same prompt for every reader.
Requirements
Use standing cost plus expected outage cost. State the RPO you would actually promise from lag samples, not from the mean.
Traffic / scale
Example inputs in the cost snippet. They are planning numbers, not a vendor quote.
Latency
Pilot light RTO is about 60 minutes in the example table. Warm standby is about 15.
Consistency
Async replication loses whatever had not shipped when the region died.
Availability
Tier 0 still needs a capacity story in the neighboring region.
Failure assumptions
- Lag spikes during backfill and vacuum.
- The neighboring region may be short on capacity during the same event.
- The business, not the brochure, owns the targets.
Constraints
- Do not promise the average lag.
- Do not skip the standing-cost term.
Prompt
A checkout journey loses about $2,000 per minute down and $5,000 per minute of lost writes. A region-level event is modeled at 5 percent per year. One region costs $40,000 a month. Pick a strategy.
API
Which tier table row fits checkout, and which strategy does that row name?
Data
What is the p99 of the sample lag, and how many samples breach 5 seconds?
Architecture
What is already running in the recovery region for pilot light?
Which number picks the strategy
Prefer
Expected yearly cost, then a drill
Standing cost plus the chance of a region event times the outage cost. The example inputs make pilot light the minimum. Change the cost per minute and the winner changes.
- Backup/restore is cheap to run and expensive in the year the disaster happens.
- Active-active buys little extra over warm standby at these example rates.
- p99 lag, not the average, is the RPO you can say out loud.
Alternative
The database brochure lag
A sub-second cross-region replica number ignores queues, write-behind caches, large object replication, and third parties. Average lag hides the spikes.
- The example samples average 4.06 seconds and p99 at 45 seconds.
- Two of fifteen samples breach a 5 second RPO.
- Spikes line up with backfills, migrations, vacuum, and congestion.
From the journey to the strategy
Targets come from the business. The strategy is the cheapest point that still meets both.
- 1
Price the journey
List checkout or login, not the service name. Price a minute of downtime and a minute of lost writes. - 2
Assign the tier
Tier 0 is under 5 minutes and near-zero RPO. Tier 3 can wait a day. - 3
Keep data moving even in pilot light
Restoring terabytes takes hours. Launching stateless compute from images takes minutes. - 4
Read the tail of the lag
Alert when lag exceeds the RPO. The average will still look fine.
Overview
RTO and RPO are business numbers first and engineering numbers second. You get them by asking what one minute of downtime costs (lost revenue, SLA credits, support load, reputation) and what one minute of lost writes costs (refunds, manual reconciliation, legal exposure). Once you have targets, the four standard strategies (backup and restore, pilot light, warm standby, active-active) are just points on a cost curve. This page goes deep on each strategy, how to measure the RPO you actually have rather than the one on the slide, and the expected-cost math that tells you when spending more on DR is worth it.
Getting the targets: a business impact analysis in four steps
- List the user journeys (checkout, login, search, reporting), not the services. Users experience journeys.
- Price the outage: revenue per minute, contractual SLA penalties, and the less visible costs like support tickets and churn.
- Price the data loss: can the lost writes be reconstructed (from payment provider records, logs, event streams) or are they gone forever?
- Assign a tier with explicit RTO and RPO, and record which services and data stores the journey depends on.
| Tier | Example journeys | RTO target | RPO target | Usual strategy |
|---|---|---|---|---|
| 0 | Payments, auth, checkout | Under 5 minutes | Near zero | Active-active or hot standby |
| 1 | Order history, notifications | Under 1 hour | Under 5 minutes | Warm standby |
| 2 | Search, recommendations | Under 4 hours | Under 1 hour | Pilot light |
| 3 | Internal tools, analytics | Under 24 hours | Under 24 hours | Backup and restore |
The four strategies in detail
Architecture
1. Backup and restore
- 1
Primary region serves all traffic
- nextBackups copied to region B
- 2
Backups copied to region B
- nextOn disaster: provision infra from code, restore data, then shift traffic
- 3
On disaster: provision infra from code, restore data, then shift traffic
2. Pilot light
- 4
Primary serves traffic
- nextDB replica live in region B, app tier at zero
- 5
DB replica live in region B, app tier at zero
- nextOn disaster: promote replica, scale app tier up, shift traffic
- 6
On disaster: promote replica, scale app tier up, shift traffic
3. Warm standby
- 7
Primary serves traffic
- nextSmall but complete stack live in region B
- 8
Small but complete stack live in region B
- nextOn disaster: promote DB, scale out, shift traffic
- 9
On disaster: promote DB, scale out, shift traffic
4. Active-active
- 10
Both regions serve live traffic
- nextData replicated or partitioned by home region
- 11
Data replicated or partitioned by home region
- nextOn disaster: stop routing to the failed region
- 12
On disaster: stop routing to the failed region
Lesson map
RTO, RPO & DR Strategies - Backup/Restore vs Pilot Light vs Warm Standby vs Active-Active
Business impact analysis to tiers with RTO/RPO targets; each DR strategy in depth; expected-yearly-cost math that picks the strategy; measuring real RPO from p99 replication lag instead of averages; capacity and decision-time pitfalls.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB br1["Primary region serves all traffic"] br2["Backups copied to region B"] br3["On disaster: provision infra from code, restore data, then shift traffic"] pl1["Primary serves traffic"] pl2["DB replica live in region B, app tier at zero"] pl3["On disaster: promote replica, scale app tier up, shift traffic"] ws1["Primary serves traffic"] ws2["Small but complete stack live in region B"] ws3["On disaster: promote DB, scale out, shift traffic"] aa1["Both regions serve live traffic"] aa2["Data replicated or partitioned by home region"] aa3["On disaster: stop routing to the failed region"] br1 -->|continues| br2 br2 -->|continues| br3 pl1 -->|continues| pl2 pl2 -->|continues| pl3 ws1 -->|continues| ws2 ws2 -->|continues| ws3 aa1 -->|continues| aa2 aa2 -->|continues| aa3
Backup and restore. Infrastructure is defined as code (Terraform, CloudFormation) and backups are copied to another region. Recovery means building the environment, restoring the data, and warming caches. RTO is dominated by restore throughput: restoring a 10 TB database at 500 MB/s takes over 5 hours before you even start validating.
Pilot light. The data layer is always on and replicating in the recovery region, because data is the slowest thing to move. Compute is defined but scaled to zero or near zero. The risk is capacity: during a large regional event, thousands of customers try to scale up in the same neighboring region at once.
Warm standby. A scaled-down but fully functional copy of the whole stack is running and receiving synthetic traffic or health checks. Because it is always running, you discover configuration drift before the disaster rather than during it.
Active-active. Every region serves real users. Failover is just "stop sending traffic to the broken region", which is why RTO approaches zero. The hard part is the data model, covered later in this cluster.
The expected-cost math (runnable)
# Compare DR strategies by expected yearly cost: standing cost + (chance of disaster x outage cost).
# This is the business-impact math behind "why not just go active-active everywhere?".
REGION_MONTHLY = 40_000 # cost of one production region per month, USD (example)
P_REGION_EVENT = 0.05 # chance per year of a region-level outage that needs DR (example)
COST_PER_DOWN_MIN = 2_000 # revenue + reputation cost per minute down (example)
COST_PER_LOST_MIN = 5_000 # cost per minute of lost writes (refunds, reconciliation)
strategies = {
# name: (standing cost multiple, RTO minutes, RPO minutes)
"backup-and-restore": (0.05, 720, 1440),
"pilot-light": (0.15, 60, 5),
"warm-standby": (0.40, 15, 1),
"active-active": (1.00, 1, 0.1),
}
rows = []
for name, (mult, rto, rpo) in strategies.items():
standing = mult * REGION_MONTHLY * 12
expected_outage = P_REGION_EVENT * (rto * COST_PER_DOWN_MIN + rpo * COST_PER_LOST_MIN)
rows.append((standing + expected_outage, name, standing, expected_outage))
for total, name, standing, outage in sorted(rows):
print(f"{name:18} standing ${standing:>9,.0f} expected outage ${outage:>9,.0f} total ${total:>9,.0f}")
print("cheapest overall:", min(rows)[1])Output:
pilot-light standing $ 72,000 expected outage $ 7,250 total $ 79,250
warm-standby standing $ 192,000 expected outage $ 1,750 total $ 193,750
backup-and-restore standing $ 24,000 expected outage $ 432,000 total $ 456,000
active-active standing $ 480,000 expected outage $ 125 total $ 480,125
cheapest overall: pilot-lightWith these example numbers pilot light wins: backup and restore is cheap to run but catastrophically expensive in the year a disaster happens, and active-active buys very little extra over warm standby. Change the cost per minute (a high-frequency trading desk, a hospital system) and the answer flips toward active-active. That is the point: the strategy falls out of the numbers.
Measuring your real RPO (runnable)
With asynchronous replication, the data you lose in a sudden region failure is whatever had not replicated yet, so the replication lag at the moment of failure is your RPO.
// Your real RPO under async replication is roughly the replication lag at the moment of failure.
// Estimate it from lag samples: the p99/max matters, not the average.
const lagSamplesSec = [0.2, 0.3, 0.25, 0.4, 0.3, 0.35, 12.0, 0.3, 0.28, 45.0, 0.31, 0.29, 0.33, 0.3, 0.27];
function percentile(xs: number[], p: number): number {
const s = [...xs].sort((a, b) => a - b);
const idx = Math.min(s.length - 1, Math.ceil((p / 100) * s.length) - 1);
return s[idx];
}
const avg = lagSamplesSec.reduce((a, b) => a + b, 0) / lagSamplesSec.length;
console.log(`avg lag ${avg.toFixed(2)} s <- looks great on a dashboard`);
console.log(`p50 lag ${percentile(lagSamplesSec, 50).toFixed(2)} s`);
console.log(`p99 lag ${percentile(lagSamplesSec, 99).toFixed(2)} s <- closer to the RPO you can promise`);
// If a region dies during a lag spike (bulk backfill, vacuum, network blip), you lose that window.
const rpoTargetSec = 5;
const breaches = lagSamplesSec.filter((x) => x > rpoTargetSec).length;
console.log(`samples breaching RPO ${rpoTargetSec}s: ${breaches}/${lagSamplesSec.length}`);Output:
avg lag 4.06 s <- looks great on a dashboard
p50 lag 0.30 s
p99 lag 45.00 s <- closer to the RPO you can promise
samples breaching RPO 5s: 2/15Expectedavg lag 4.06 s <- looks great on a dashboard p50 lag 0.30 s p99 lag 45.00 s <- closer to the RPO you can promise samples breaching RPO 5s: 2/15
Press Run. Snippets must be self-contained — no network, files, or native modules.
Lag spikes happen during bulk backfills, schema migrations, large transactions, vacuum and network congestion, which are exactly the moments failures like to coincide with. Alert on replication lag against your RPO, not just on replica health.
What happens if you get the targets wrong
- Targets set by engineers alone: you either gold-plate low-value services or discover during an outage that the business assumed checkout had an RPO of zero.
- RPO assumed from the database brochure: a managed database may advertise sub-second cross-region lag, but your RPO also includes queues, caches with write-behind, object storage replication (which can take minutes for large objects) and third-party systems.
- RTO measured from when the failover command runs: the clock starts when users are hurt. Detection and decision time often exceed execution time.
- No capacity reservation: pilot light assumes you can get instances in a neighboring region during a regional event. Consider reserved capacity or capacity reservations for tier 0 and tier 1.
Pros and cons
| Strategy | Best when | Watch out for |
|---|---|---|
| Backup and restore | Low business impact, large data that changes slowly | Untested restores, long restore times, IaC drift |
| Pilot light | Medium tier, data must be current but compute can wait | Scale-up time, capacity crunch, cold caches |
| Warm standby | Minutes of RTO with moderate spend | Paying for idle capacity, keeping versions in sync |
| Active-active | Very high cost per minute down | Data conflicts, cross-region latency, 2x cost |
Interview Q&A
Your database replicates asynchronously with an average lag of 300 ms. What RPO do you promise?
Answer
Not 300 ms. Look at the p99 and maximum lag over a long window, including during batch jobs and migrations, add every other async path (queues, object replication), and promise something with margin. Alert when lag exceeds it.
Why keep the data layer always on even in the cheap pilot light strategy?
Answer
Moving data is the slowest part of recovery. Restoring terabytes takes hours, while launching stateless compute from images takes minutes.
When is active-active not worth it?
Answer
When the expected outage cost avoided is smaller than the standing cost and complexity. Many services get most of the benefit from warm standby at a fraction of the price.
What makes warm standby more reliable than pilot light in practice?
Answer
The standby is continuously running and checked, so broken configuration, expired certificates or missing secrets show up before the disaster, not during it.
How do you reduce RTO without spending on more infrastructure?
Answer
Shorten detection with better alerts, pre-approve the failover decision, automate the runbook steps, and lower DNS TTLs or use a global load balancer so traffic shifts quickly.
What four steps turn a journey into a tier?
Answer
List the user journeys, price the outage, price the data loss, then write an explicit RTO and RPO plus the stores that journey depends on.
Why can a brochure lag number understate RPO?
Answer
Queues, write-behind caches, object replication, and third-party systems sit outside the database lag sample.
What does the example cost model say about backup-and-restore?
Answer
Standing cost is about $24,000 a year, but the expected outage cost is $432,000, so the yearly total is worse than pilot light for those example inputs.
Check yourself
Take the lag samples in the snippet and change the RPO target from 5 seconds to 60 seconds. Say whether you would still page, and which other async path you would add before you promise the number.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Sync, async, and semi-sync replication, SLOs and error budgets, Terraform state and safe change.