TopicsReliability & Disaster Recovery
Reliability & Disaster Recovery
RTO and RPO, backups and PITR, multi-region failover, and cells with static stability you can defend in interviews.
Common tags: disaster-recovery, rto, rpo, multi-region, failover
- Reliability & Disaster Recovery
RTO, RPO & DR Strategies - Backup/Restore vs Pilot Light vs Warm Standby vs Active-Active
Cluster · Disaster Recovery & Multi-Region
Business impact analysis to tiers with RTO/RPO targets; each DR strategy in depth; expected-yearly-cost math that picks the strategy; measuring real RPO from p99 replication lag instead of averages; capacity and decision-time pitfalls.
Open study →- reliability
- disaster-recovery
- rto
- rpo
- backups
- pitr
- multi-region
- failover
- availability-zones
- cells
- interview
- Reliability & Disaster Recovery
Region Failover in Practice - Runbooks, DNS/GLB Cutover, Failback & Game Days
Cluster · Disaster Recovery & Multi-Region
Automatic vs human-approved failover; runnable failover state machine with hysteresis; 9-step runbook; DNS TTL drain math vs global load balancer/anycast/client routing; failback risks; tabletop to region-evacuation drills; Facebook 2021, AWS 2021, GitLab 2017 lessons.
Open study →- reliability
- disaster-recovery
- rto
- rpo
- backups
- pitr
- multi-region
- failover
- availability-zones
- cells
- interview
- Reliability & Disaster Recovery
Multi-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static Stability
Cluster · Disaster Recovery & Multi-Region
Host/AZ/region/cell failure domains; availability math and why correlated regional dependencies cap it; static stability and data plane vs control plane; cell-based architecture and shuffle sharding (runnable); when multi-region is worth it.
Open study →- reliability
- disaster-recovery
- rto
- rpo
- backups
- pitr
- multi-region
- failover
- availability-zones
- cells
- interview
- Reliability & Disaster Recovery
Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-Active
Cluster · Disaster Recovery & Multi-Region
Interview hub: HA vs DR vs backup, RTO vs RPO, the four DR strategies (backup/restore, pilot light, warm standby, active-active) with cost tiers, RTO as a phase budget, hidden single-region dependencies; replication is not backup.
Open study →- reliability
- disaster-recovery
- rto
- rpo
- backups
- pitr
- multi-region
- failover
- availability-zones
- cells
- interview
- Reliability & Disaster Recovery
Backups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore Drills
Cluster · Disaster Recovery & Multi-Region
Snapshot vs incremental vs logical dump vs PITR vs object versioning; crash- vs application-consistent; runnable PITR replay stopping before a bad DELETE; 3-2-1-1-0 and immutable copies (Object Lock); GFS retention; GitLab 2017 and OVHcloud 2021; restore drills.
Open study →- reliability
- disaster-recovery
- rto
- rpo
- backups
- pitr
- multi-region
- failover
- availability-zones
- cells
- interview
- Reliability & Disaster Recovery
Active-Active vs Active-Passive Multi-Region - Data, Write Routing & Home Regions
Cluster · Disaster Recovery & Multi-Region
Active-passive vs home-region vs multi-writer merge vs consensus topologies; runnable home-region routing with read-your-writes tokens; runnable LWW lost-update on balances; per-data-type choices; failover behavior and fencing in each topology.
Open study →- reliability
- disaster-recovery
- rto
- rpo
- backups
- pitr
- multi-region
- failover
- availability-zones
- cells
- interview