Reliability & Disaster Recovery
Part 6 of 6 · Disaster Recovery & Multi-RegionRegion Failover in Practice - Runbooks, DNS/GLB Cutover, Failback & Game Days
Automatic vs human-approved failover; runnable failover state machine with hysteresis; 9-step runbook; DNS TTL drain math vs global load balancer/anycast/client routing; failback risks; tabletop to region-evacuation drills; Facebook 2021, AWS 2021, GitLab 2017 lessons.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why not make region failover fully automatic?
Answer
Health checks fail for reasons a failover will not fix, and an automatic promotion during a partition can split-brain. Automate stateless traffic shifts. Keep a fast human approval for data promotion.
L2
Your DNS TTL is 60 seconds and traffic took 20 minutes. Why?
Answer
Some resolvers cache longer than the TTL, long-lived HTTP/2, gRPC, and WebSocket connections never re-resolve until they break, and clients hold endpoints in memory. A global load balancer does not wait on that cache.
L3
What does the drain snippet show at a 60 minute TTL after 15 minutes?
Answer
About 76 percent of traffic is still on the old answer, with a 3 percent sticky floor. At a 1 minute TTL the non-sticky cache is gone by 1 minute and the floor remains.
L4
How do you run a drill safely?
Answer
Write a hypothesis, start in staging, set abort criteria and a rollback, tell on-call, limit the blast radius to one service or one cell, time every phase against the RTO, and fix the runbook.
L5
What most often makes a failover fail?
Answer
Hidden dependencies on the failed region, such as auth, secrets, config, images, and DNS control planes, plus a runbook nobody has executed.
L6
What is different about failback?
Answer
You are returning to a region that just failed, and data written in the recovery region has to flow back before traffic does. Many teams leave the recovery region as primary and fail back later as a planned change.
L7
What should the drill's success signal be?
Answer
A business metric such as orders per minute, plus synthetic transactions. Green health checks are not enough.
Failure modes
First attempt during the incident
The runbook names scripts, dashboards, and people that no longer exist.
DNS drain assumed equal to the TTL
Sticky clients and long-lived connections remain. The snippet keeps a 3 percent floor.
Failback without a cutover
Writes from both sides interleave, and stranded primary writes are forgotten.
Tooling that lives in the failed network
Facebook in October 2021 and the us-east-1 control-plane impairment in December 2021 both punished teams whose recovery path was inside the outage.
Misconceptions
A health check is a sufficient failover trigger for the database.
Use it for stateless shift. Data promotion needs a person because a false positive can split-brain.
Failback is failover in reverse and therefore safer.
It is usually riskier. The recovery region has accepted writes that must move back first.
A tabletop proves the region can evacuate.
It proves people know the roles. Capacity, alerts, and hidden dependencies show up in a production game day.
Interviewer traps
Setting TTL to 60 seconds and declaring the RTO met.
Account for resolvers that ignore TTL and connections that do not re-resolve.
Scheduling failback at peak because the team wants the old primary back.
Wait for a quiet window, or make the recovery region the new primary and plan the return.
Design scenario
Same prompt for every reader.
Requirements
Use the hysteresis rules in the snippet. Name the runbook steps that must finish before the DNS or load balancer change.
Traffic / scale
The sample check series fails over at t=6 and failbacks at t=21 after approval. Sticky fraction in the DNS snippet is 0.03.
Latency
Anycast or a global load balancer moves traffic in seconds. DNS is bounded by TTL plus clients that ignore it.
Consistency
Failback waits until recovery-region writes have resynced. Stranded writes on the old primary are reviewed.
Availability
The traffic shift should not depend on the control plane of the region you are leaving.
Failure assumptions
- A single bad health check is a blip.
- The primary can look healthy and then flap.
- Some clients cache DNS past the TTL.
Constraints
- Do not promote the database on one failed check.
- Do not fail back before approval and a resync.
Prompt
Primary health blips, then fails, then flaps, then stays healthy. DNS TTL is 60 minutes. Decide when to fail over, how traffic moves, and when failback is allowed.
API
Which steps are automatic, and which need the incident commander?
Data
What has to be true of the replication stream before failback?
Architecture
Why can a global load balancer drain faster than the 60 minute TTL row?
Who is allowed to move the region
Prefer
Automatic traffic shift, approved data promotion
Stateless tiers move on health checks. Database promotion waits for one pre-authorized incident commander. Failback is manual and only after resync.
- Three consecutive bad checks avoid a single blip.
- Five healthy checks are not enough: the snippet still waits for approval.
- A flap sends the controller back to the secondary.
Alternative
A fully automatic promote, or a fully manual runbook
Automatic data promotion during a partition can split-brain. A manual runbook for every step adds the decision minutes you did not budget, and it rots if nobody drills it.
- Health checks fail for reasons a failover will not fix.
- DNS TTL is not the drain time when clients cache longer or hold connections.
- Failback during peak, or without a cutover point, causes the second outage.
Leave, prove, then come back on purpose
The state diagram is the controller. The runbook is what the humans do inside those states.
- 1
Declare and fence
Name an incident commander, confirm scope from outside the region, freeze deploys, and stop the old primary from accepting writes. - 2
Promote, scale, shift, validate
Promote data, scale if you were not already static, shift traffic, and watch orders per minute rather than only green health checks. - 3
Wait out the flap
A short healthy streak is not failback. The controller returns to the secondary when the primary blips again. - 4
Resync, then fail back
Writes from the recovery region flow back first. Stranded writes in the old primary are replayed or discarded on purpose.
Overview
A region failover is an incident procedure, not a button. It has a trigger (who decides, based on what signals), a sequence (fence, promote data, scale, shift traffic, validate), a traffic mechanism (DNS, anycast, a global load balancer, or client-side routing), and a return path (failback, which is usually riskier than failover because data has to be resynchronized in the opposite direction). The only way to know it works is to rehearse it: scheduled DR drills, game days and, at the extreme, deliberately evacuating a region with real traffic the way Netflix's Chaos Kong exercises did.
Automatic vs manual failover
| Mode | When it fits | Risk |
|---|---|---|
| Fully automatic | Stateless tiers, active-active traffic shifting, health checks that are clearly reliable | False positives move traffic during a blip; flapping |
| Automatic with human approval (one click) | Data promotion in active-passive, tier 0 journeys | Approval delay adds to RTO |
| Fully manual runbook | Rare, complex failovers with many dependencies | Slow, error-prone under stress |
A common pattern: traffic shifting for stateless tiers is automatic, database promotion needs one human approval from a pre-authorized incident commander, and failback is always manual.
A failover controller with hysteresis (runnable)
# Automated region failover with hysteresis: fail over only after N consecutive bad checks,
# fail back only after M consecutive good checks AND a human approval. Prevents flapping.
FAIL_AFTER, RECOVER_AFTER = 3, 5
def run(checks, approve_failback_at=None):
state, bad, good, log = "PRIMARY_ACTIVE", 0, 0, []
for t, healthy in enumerate(checks):
if state == "PRIMARY_ACTIVE":
bad = 0 if healthy else bad + 1
if bad >= FAIL_AFTER:
state, good = "FAILED_OVER", 0
log.append(f"t={t}: {bad} bad checks -> FAIL OVER to secondary")
elif state == "FAILED_OVER":
good = good + 1 if healthy else 0
if good >= RECOVER_AFTER:
state = "READY_TO_FAILBACK"
log.append(f"t={t}: primary healthy {good}x -> ready, waiting for approval")
elif state == "READY_TO_FAILBACK":
if not healthy:
state, good = "FAILED_OVER", 0
log.append(f"t={t}: primary flapped -> stay on secondary")
elif approve_failback_at is not None and t >= approve_failback_at:
state, bad = "PRIMARY_ACTIVE", 0
log.append(f"t={t}: approved -> FAIL BACK (after data resync)")
return state, log
# primary health over time: blip, real outage, flapping recovery, stable recovery
checks = [1, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1]
final, log = run([bool(c) for c in checks], approve_failback_at=21)
print("\n".join(log))
print("final state:", final)Output:
t=6: 3 bad checks -> FAIL OVER to secondary
t=12: primary healthy 5x -> ready, waiting for approval
t=13: primary flapped -> stay on secondary
t=18: primary healthy 5x -> ready, waiting for approval
t=21: approved -> FAIL BACK (after data resync)
final state: PRIMARY_ACTIVERequiring several consecutive bad checks avoids failing over on a single blip, and requiring a longer run of good checks plus approval before failback avoids bouncing back into a region that is still unstable.
States
- 1
Start → PrimaryActive
Start → PrimaryActive
- 2
PrimaryActive → FailedOver
PrimaryActive → FailedOver
1. N consecutive failed health checks
- 3
FailedOver → ReadyToFailback
FailedOver → ReadyToFailback
2. M consecutive healthy checks
- 4
ReadyToFailback → FailedOver
ReadyToFailback → FailedOver
3. Primary flaps again
- 5
ReadyToFailback → Resyncing
ReadyToFailback → Resyncing
4. Human approves failback
- 6
Resyncing → PrimaryActive
Resyncing → PrimaryActive
5. Data caught up, traffic shifted back
Lesson map
Region Failover in Practice - Runbooks, DNS/GLB Cutover, Failback & Game Days
Automatic vs human-approved failover; runnable failover state machine with hysteresis; 9-step runbook; DNS TTL drain math vs global load balancer/anycast/client routing; failback risks; tabletop to region-evacuation drills; Facebook 2021, AWS 2021, GitLab 2017 lessons.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB primaryactive["PrimaryActive"] failedover["FailedOver"] readytofailback["ReadyToFailback"] resyncing["Resyncing"] primaryactive -->|1. N consecutive failed health checks| failedover failedover -->|2. M consecutive healthy checks| readytofailback readytofailback -->|3. Primary flaps again| failedover readytofailback -->|4. Human approves failback| resyncing resyncing -->|5. Data caught up, traffic shifted back| primaryactive
The failover runbook, step by step
- Declare the incident and name an incident commander. Confirm the scope (provider status page, your own synthetic checks from outside the region).
- Freeze deploys and risky batch jobs everywhere.
- Fence the failing primary so it cannot keep accepting writes: revoke credentials, block its network path, or bump the fencing epoch.
- Promote data stores in the recovery region (database replicas, queue consumers, caches warmed from the database).
- Scale the recovery region to full capacity if it is not already statically provisioned.
- Shift traffic with the global load balancer, DNS or anycast change.
- Validate with synthetic transactions and business metrics (orders per minute), not just green health checks.
- Communicate status to customers and internal teams.
- Plan failback only after the primary is stable, data is resynchronized and there is a quiet window.
How traffic actually moves (runnable)
// DNS-based failover is not instant: resolvers and clients cache the old answer for up to the TTL,
// and some clients ignore TTLs entirely. Estimate how much traffic still hits the dead region.
function stillOnOld(minutesAfterChange: number, ttlMin: number, stickyFraction: number): number {
// Caches expire roughly uniformly over the TTL window; sticky clients (JVMs with infinite
// DNS cache, long-lived connections) stay until restarted or the connection breaks.
const caching = Math.max(0, 1 - minutesAfterChange / ttlMin) * (1 - stickyFraction);
return caching + stickyFraction;
}
for (const ttl of [1, 5, 60]) {
const row = [0, 1, 5, 15, 60].map((m) => `${m}m:${(100 * stillOnOld(m, ttl, 0.03)).toFixed(0)}%`);
console.log(`TTL ${String(ttl).padStart(2)} min -> ${row.join(" ")}`);
}
console.log("anycast / global load balancer health checks move traffic without waiting on client DNS caches");Output:
TTL 1 min -> 0m:100% 1m:3% 5m:3% 15m:3% 60m:3%
TTL 5 min -> 0m:100% 1m:81% 5m:3% 15m:3% 60m:3%
TTL 60 min -> 0m:100% 1m:98% 5m:92% 15m:76% 60m:3%
anycast / global load balancer health checks move traffic without waiting on client DNS cachesExpectedTTL 1 min -> 0m:100% 1m:3% 5m:3% 15m:3% 60m:3% TTL 5 min -> 0m:100% 1m:81% 5m:3% 15m:3% 60m:3% TTL 60 min -> 0m:100% 1m:98% 5m:92% 15m:76% 60m:3% anycast / global load balancer health checks move traffic without waiting on client DNS caches
Press Run. Snippets must be self-contained — no network, files, or native modules.
| Mechanism | Speed | How it works | Caveats |
|---|---|---|---|
| DNS failover | Bounded by TTL plus clients that ignore TTL | Health-checked DNS returns the healthy region | Long-lived connections and sticky resolvers lag behind |
| Global load balancer (anycast front end) | Seconds | One global IP, edge proxies route to healthy backends | Ties you to the provider's global control plane |
| Anycast (BGP) | Seconds to a minute | Same IP announced from many sites; withdraw from the failed one | Needs your own network expertise or a provider |
| Client-side routing | As fast as the client notices | Apps hold a list of regional endpoints and retry elsewhere | Must ship and update the logic in every client |
Prefer mechanisms whose data plane keeps working when a provider's control plane is impaired. Some providers offer dedicated recovery controls for this (for example AWS Route 53 Application Recovery Controller), designed so that flipping traffic does not depend on the region you are leaving.
Failback is the dangerous half
During failover you knew exactly what you wanted: get out. During failback you are moving traffic back to a region that just failed, and data written in the recovery region must flow back first. Common mistakes:
- Reversing replication without a clean cutover point, so writes from both sides interleave.
- Failing back during peak traffic because the team is eager to return to normal.
- Forgetting the writes stranded in the old primary during the lag window; they must be reviewed and replayed or discarded deliberately.
Many teams simply make the recovery region the new primary and fail back days later as a planned change.
Drills and game days
| Exercise | What it proves | Frequency |
|---|---|---|
| Tabletop | People know the runbook and roles | Quarterly |
| Restore drill | Backups restore within RTO | Monthly or quarterly per tier |
| Component failover in staging | Automation works end to end | Every release of the DR tooling |
| Production game day | Real dependencies, real capacity, real alerts | Quarterly or twice a year |
| Region evacuation with live traffic | The whole organization can do it | Only for mature active-active systems |
Google's DiRT program and Netflix's region evacuations are the well-known examples. Start small: one service, one dependency, a staging environment, a written hypothesis ("checkout keeps working when the primary DB is promoted"), and a scheduled time with everyone watching.
Lessons from real incidents
- Facebook, October 2021: a maintenance command accidentally disconnected their backbone network, their DNS servers then withdrew their BGP routes, and all services went offline for hours. Internal tools and even physical access depended on the same network, which slowed recovery. Lesson: keep out-of-band access and tooling that does not depend on production.
- AWS us-east-1, December 2021: an internal network event impaired control plane APIs and monitoring in the region for hours. Teams whose failover required calling those APIs struggled. Lesson: static stability and recovery paths outside the failing region.
- GitLab, January 2017: during an incident, an engineer ran a delete on the wrong database host, and several backup paths were broken. Lesson: guardrails on destructive commands, and tested restores.
What happens if you skip this
- The first real failover is also the first attempt, under stress, at 3 a.m.
- Runbooks reference dashboards, scripts and people that no longer exist.
- Capacity in the recovery region is not there when needed.
- Failback causes a second outage.
Pros and cons
| Practice | Pros | Cons |
|---|---|---|
| Automatic failover | Lowest RTO, no human delay | False positives, flapping, harder to reason about |
| Human-approved failover | Prevents unnecessary failovers | Adds decision minutes to RTO |
| Production game days | Finds real hidden dependencies | Risk of causing a real outage; needs maturity |
| Make recovery region the new primary | Avoids rushed failback | Region roles become dynamic, tooling must handle both |
Interview Q&A
Why not make region failover fully automatic?
Answer
Health checks can fail for reasons that a failover would not fix, and an automatic data promotion during a partition can cause split brain. Automate traffic shifting for stateless tiers; keep a fast human approval for data promotion.
Your DNS TTL is 60 seconds but traffic took 20 minutes to drain. Why?
Answer
Some resolvers and runtimes cache longer than the TTL, long-lived connections (HTTP/2, gRPC, WebSockets) never re-resolve until they break, and mobile apps may hold endpoints in memory. Close connections from the server side and prefer a global load balancer.
How do you run a DR drill safely?
Answer
Write a hypothesis, start in staging, define abort criteria and a rollback, schedule it with on-call aware, limit blast radius (one service, one cell), measure every phase against RTO, and turn findings into runbook and automation fixes.
What is the most common cause of failed failovers?
Answer
Hidden dependencies on the failed region (auth, secrets, config, container images, DNS control planes) and runbooks that drifted from reality because nobody exercised them.
Why does the controller wait for 3 bad checks and 5 good checks?
Answer
Three bad checks skip a single blip. Five good checks plus a human approval avoid failing back into a region that is still flapping. The sample fails over at t=6, rejects a flap at t=13, and returns at t=21.
What are the nine runbook steps?
Answer
Declare, freeze deploys, fence the primary, promote data, scale, shift traffic, validate with business metrics, communicate, and plan failback only after the primary is stable, data has resynced, and there is a quiet window.
Why is failback riskier than failover?
Answer
Writes accepted in the recovery region must flow back first. Reversing replication without a cutover point interleaves both sides, and writes stranded on the old primary need a deliberate replay or discard.
What did the October 2021 Facebook outage show about tooling?
Answer
A backbone change withdrew DNS routes, and internal tools plus physical access depended on that same network. Keep an out-of-band path that does not ride production.
Check yourself
Change the DNS snippet's sticky fraction from 0.03 to 0.20 and read the 60 minute column again. Say which mechanism on the page would move those clients without waiting for their resolver cache.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Global load balancing, L4 and L7 load balancing, Failover and fencing, Load, chaos, and production validation.