Reliability & Disaster Recovery
Topic, then cluster, then study. Recently added is the short list at the top.
Recently added
- 5.Active-Active vs Active-Passive Multi-Region - Data, Write Routing & Home RegionsActive-passive vs home-region vs multi-writer merge vs consensus topologies; runnable home-region routing with read-your-writes tokens; runnable LWW lost-update on balances; per-data-type choices; failover behavior and fencing in each topology.
- 3.Backups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore DrillsSnapshot vs incremental vs logical dump vs PITR vs object versioning; crash- vs application-consistent; runnable PITR replay stopping before a bad DELETE; 3-2-1-1-0 and immutable copies (Object Lock); GFS retention; GitLab 2017 and OVHcloud 2021; restore drills.
- 1.Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-ActiveInterview hub: HA vs DR vs backup, RTO vs RPO, the four DR strategies (backup/restore, pilot light, warm standby, active-active) with cost tiers, RTO as a phase budget, hidden single-region dependencies; replication is not backup.
- 4.Multi-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static StabilityHost/AZ/region/cell failure domains; availability math and why correlated regional dependencies cap it; static stability and data plane vs control plane; cell-based architecture and shuffle sharding (runnable); when multi-region is worth it.
- 6.Region Failover in Practice - Runbooks, DNS/GLB Cutover, Failback & Game DaysAutomatic vs human-approved failover; runnable failover state machine with hysteresis; 9-step runbook; DNS TTL drain math vs global load balancer/anycast/client routing; failback risks; tabletop to region-evacuation drills; Facebook 2021, AWS 2021, GitLab 2017 lessons.
- 2.RTO, RPO & DR Strategies - Backup/Restore vs Pilot Light vs Warm Standby vs Active-ActiveBusiness impact analysis to tiers with RTO/RPO targets; each DR strategy in depth; expected-yearly-cost math that picks the strategy; measuring real RPO from p99 replication lag instead of averages; capacity and decision-time pitfalls.
Reliability & Disaster Recovery
RTO and RPO, backups and PITR, multi-region failover, and cells with static stability you can defend in interviews.
Disaster Recovery & Multi-Region
6 studies- 1.Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-ActiveInterview hub: HA vs DR vs backup, RTO vs RPO, the four DR strategies (backup/restore, pilot light, warm standby, active-active) with cost tiers, RTO as a phase budget, hidden single-region dependencies; replication is not backup.
- 2.RTO, RPO & DR Strategies - Backup/Restore vs Pilot Light vs Warm Standby vs Active-ActiveBusiness impact analysis to tiers with RTO/RPO targets; each DR strategy in depth; expected-yearly-cost math that picks the strategy; measuring real RPO from p99 replication lag instead of averages; capacity and decision-time pitfalls.
- 3.Backups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore DrillsSnapshot vs incremental vs logical dump vs PITR vs object versioning; crash- vs application-consistent; runnable PITR replay stopping before a bad DELETE; 3-2-1-1-0 and immutable copies (Object Lock); GFS retention; GitLab 2017 and OVHcloud 2021; restore drills.
- 4.Multi-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static StabilityHost/AZ/region/cell failure domains; availability math and why correlated regional dependencies cap it; static stability and data plane vs control plane; cell-based architecture and shuffle sharding (runnable); when multi-region is worth it.
- 5.Active-Active vs Active-Passive Multi-Region - Data, Write Routing & Home RegionsActive-passive vs home-region vs multi-writer merge vs consensus topologies; runnable home-region routing with read-your-writes tokens; runnable LWW lost-update on balances; per-data-type choices; failover behavior and fencing in each topology.
- 6.Region Failover in Practice - Runbooks, DNS/GLB Cutover, Failback & Game DaysAutomatic vs human-approved failover; runnable failover state machine with hysteresis; 9-step runbook; DNS TTL drain math vs global load balancer/anycast/client routing; failback risks; tabletop to region-evacuation drills; Facebook 2021, AWS 2021, GitLab 2017 lessons.