Reliability & Disaster Recovery
Part 3 of 6 · Disaster Recovery & Multi-RegionBackups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore Drills
Snapshot vs incremental vs logical dump vs PITR vs object versioning; crash- vs application-consistent; runnable PITR replay stopping before a bad DELETE; 3-2-1-1-0 and immutable copies (Object Lock); GFS retention; GitLab 2017 and OVHcloud 2021; restore drills.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is the difference between a snapshot and PITR?
Answer
A snapshot restores one instant. PITR restores a base, then replays logs to any second in the retention window.
L2
What is crash-consistent versus application-consistent?
Answer
Crash-consistent looks like power was pulled. The database recovers from its log, but multi-volume state may disagree. Application-consistent flushes or freezes so the captured state lines up.
L3
Walk a column corrupted 40 minutes ago.
Answer
Stop the job, restore the base into a new instance, replay WAL to just before the migration, diff the column, copy the good values back, and reconcile the last 40 minutes.
L4
How do you protect backups from ransomware?
Answer
A separate account, separate credentials, write-once object lock or a vault lock, and an alert on deletion. Then restore from that isolated copy in a drill.
L5
How do you know the backups work?
Answer
Scheduled restores of a random backup to a random time, integrity checks, alerts on a missing log shipment, and a recorded duration compared with the RTO.
L6
What did GitLab lose in January 2017?
Answer
Several backup paths were not working. They restored from a staging copy about six hours old.
L7
What does the retention snippet keep?
Answer
400 daily backups prune to 21 copies, oldest 2025-11-01 and newest 2026-10-06, under a 7 daily, 4 weekly, 12 monthly rule.
Failure modes
Backups in the same account
One compromised admin credential deletes production and the backups together.
No PITR
Logical corruption falls back to the last snapshot, which can be a full day.
No timed restore
You learn during the outage that the restore takes 14 hours or fails on a version mismatch.
The encryption key died with the account
An encrypted backup is useless if the key was deleted with the region. Drills have to include key access.
Misconceptions
Taking backups means you can restore.
The job can be silently broken for months. Only a measured restore proves it.
Replication covers a bad delete.
The delete reaches every replica in seconds. You need a log you can stop.
The newest backup is the one to drill.
Drill a random backup and a random target time, or you never see an old chain break.
Interviewer traps
Restoring over the damaged primary.
Use a new instance so you can diff and keep evidence.
Forgetting writes after the recovery target.
The snippet's u9 is a legitimate write that the t=129 copy does not contain.
Design scenario
Same prompt for every reader.
Requirements
Name the restore target, the instance you restore into, and what you do with writes after 14:31.
Traffic / scale
The runnable log is five statements from t=110 through t=140. Treat that as the shape, not the production volume.
Latency
Replay time grows with the log volume after the base. The drill records each step against the RTO.
Consistency
Application-consistent backups matter when more than one volume has to agree.
Availability
The immutable copy must remain if the production account is compromised.
Failure assumptions
- The log archive can have a gap.
- The encryption key might live only in the failed account.
- Legitimate writes continue after the bad statement.
Constraints
- Do not restore over the damaged instance.
- Do not replay past the bad DELETE if the goal is the prior rows.
Prompt
A bad DELETE without a WHERE hits at 14:31. Nightly base backups and a cross-region WAL archive exist. Recover the table.
API
What is the recovery target time relative to 14:31?
Data
Which rows exist at t=129, and which key needs a manual reconcile?
Architecture
Where does the immutable copy live, and who holds its credentials?
Where the restore point comes from
Prefer
Base backup plus a log you can stop
The nightly snapshot is the base. WAL replay stops at the target time, before the bad statement. Restore lands in a new instance.
- t=129 keeps alice, bobby, and carol.
- Replaying through t=140 re-applies the delete and leaves only zed.
- u9 was written after the incident and has to be reconciled by hand.
Alternative
The newest snapshot, restored over production
A snapshot-only copy jumps to the last backup, which may be a day of RPO. Restoring over the damaged instance destroys the evidence and the chance to diff.
- Crash-consistent copies can disagree across volumes.
- A gap in the log archive breaks the PITR chain.
- A copy in the same account dies with the admin credential.
Restore to a chosen second
The sequence in the diagram is the same path as the runnable replay.
- 1
Take a base and ship the log
The base is the nightly backup. WAL segments ship continuously to a cross-region immutable archive. - 2
Stop before the bad statement
Replay until 14:30:59, not through the 14:31 DELETE. - 3
Validate, then reconcile
Check row counts and checksums. Writes after the target are not in the copy. - 4
Keep the damaged instance
Swap or copy rows back. Do not restore on top of the only evidence.
Overview
Nobody cares whether you take backups; they care whether you can restore. Most backup disasters are restore disasters: the job silently failed for months, the backup was in the same region or account that got destroyed, the encryption key was gone, or nobody knew how long a restore takes. A solid backup design combines full snapshots with continuous logs for point-in-time recovery (PITR), keeps at least one copy immutable and isolated from production credentials, follows a retention scheme that balances cost against how far back you might need to go, and is proven by scheduled restore drills with measured timings.
Kinds of backups
| Type | What it captures | Restore granularity | Speed | Gotchas |
|---|---|---|---|---|
| Full physical / block snapshot | Disk or volume at an instant | Snapshot time only | Fast to take, restore depends on lazy loading | Crash-consistent unless you quiesce the app |
| Incremental snapshot | Only changed blocks since last snapshot | Snapshot time only | Cheap and fast | Restore depends on the chain being intact |
Logical dump (pg_dump, mysqldump) | Rows and schema as SQL or archive | Per table or object | Slow for large databases | Portable across versions; long dumps can stress the DB |
| PITR (base backup + WAL or binlog archive) | Every committed change since the base | Any second in the retention window | Restore base, then replay logs | Log archive gaps break the chain |
| Object versioning | Every version of every object | Per object version | Fast per object | Deletes become delete markers; costs grow with churn |
Crash-consistent vs application-consistent. A crash-consistent snapshot looks like the power was pulled: databases recover using their write-ahead log, but multi-volume or multi-service state may disagree. Application-consistent backups coordinate with the app (flush, freeze, or use the database's own backup API) so everything lines up.
Point-in-time recovery in miniature (runnable)
# Point-in-time recovery (PITR) in miniature: restore the last base snapshot,
# then replay the write-ahead log (WAL) up to just BEFORE the bad statement.
import copy
snapshot = {"taken_at": 100, "rows": {"u1": "alice", "u2": "bob"}} # nightly base backup
wal = [ # (timestamp, operation, key, value) archived continuously after the snapshot
(110, "put", "u3", "carol"),
(120, "put", "u2", "bobby"),
(130, "delete_all", None, None), # the incident: someone ran DELETE without WHERE
(140, "put", "u9", "zed"), # writes that landed after the incident
]
def restore(target_ts: int) -> dict:
db = copy.deepcopy(snapshot["rows"]) # 1. restore base backup into a NEW instance
for ts, op, key, val in wal: # 2. replay WAL in order
if ts > target_ts: # 3. stop at the recovery target time
break
if op == "put":
db[key] = val
elif op == "delete_all":
db.clear()
return db
print("restore to t=129 ->", restore(129)) # just before the bad delete
print("restore to t=140 ->", restore(140)) # replaying past it re-applies the damage
# Writes after the incident (u9) are NOT in the t=129 copy: you must reconcile them separately.
lost = {k: v for k, v in restore(140).items() if k not in restore(129)}
print("needs manual reconcile:", lost)Output:
restore to t=129 -> {'u1': 'alice', 'u2': 'bobby', 'u3': 'carol'}
restore to t=140 -> {'u9': 'zed'}
needs manual reconcile: {'u9': 'zed'}The last line is the part people forget: restoring to before the incident also throws away legitimate writes made after it. You need a plan to reconcile them, for example by replaying them from an event log or the original API requests.
Sequence
- 1
Production DB → Nightly base backup
1. Base backup at 02:00
- 2
Production DB → WAL archive (cross-region, immutable)
2. Ship WAL segments continuously
- 3
Production DB
3. 14:31 bad DELETE runs
- 4
New restored instance → Nightly base backup
4. Restore base backup into a NEW instance
- 5
New restored instance → WAL archive (cross-region, immutable)
5. Replay WAL up to 14:30:59
- 6
New restored instance → New restored instance
6. Validate row counts and checksums
- 7
New restored instance
7. Swap traffic or copy back the missing rows, then reconcile post-incident writes
Lesson map
Backups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore Drills
Snapshot vs incremental vs logical dump vs PITR vs object versioning; crash- vs application-consistent; runnable PITR replay stopping before a bad DELETE; 3-2-1-1-0 and immutable copies (Object Lock); GFS retention; GitLab 2017 and OVHcloud 2021; restore drills.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB db["Production DB"] arch["WAL archive (cross-region, immutable)"] snap["Nightly base backup"] new["New restored instance"] db -->|1. Base backup at 02:00| snap db -->|2. Ship WAL segments continuously| arch new -->|4. Restore base backup into a NEW instance| snap new -->|5. Replay WAL up to 14:30:59| arch
Restore into a new instance, never over the damaged one, so you can compare and you keep evidence.
The 3-2-1 rule and its modern extension
- 3 copies of the data (production plus two backups).
- 2 different storage types or services.
- 1 copy offsite (another region, and ideally another account).
- The modern 3-2-1-1-0 variant adds 1 immutable or offline copy and 0 errors on verified restore tests.
Immutability matters because ransomware and compromised admin credentials go after backups first. Options include object lock in write-once mode (S3 Object Lock in compliance mode, which even the root account cannot shorten), backup vault locks, and a separate backup account whose credentials production systems never hold.
Retention: how far back can you go? (runnable)
// Grandfather-father-son retention: keep 7 dailies, 4 weeklies (Sundays), 12 monthlies (1st of month).
// Shows which backups survive pruning on a given day, and how many copies you pay for.
function daysAgo(today: Date, n: number): Date {
const d = new Date(today);
d.setUTCDate(d.getUTCDate() - n);
return d;
}
function keep(today: Date, backups: Date[]): Set<string> {
const iso = (d: Date) => d.toISOString().slice(0, 10);
const sorted = [...backups].sort((a, b) => b.getTime() - a.getTime()); // newest first
const kept = new Set<string>();
sorted.slice(0, 7).forEach((d) => kept.add(iso(d))); // sons: last 7 days
sorted.filter((d) => d.getUTCDay() === 0).slice(0, 4).forEach((d) => kept.add(iso(d))); // fathers
sorted.filter((d) => d.getUTCDate() === 1).slice(0, 12).forEach((d) => kept.add(iso(d))); // grandfathers
return kept;
}
const today = new Date(Date.UTC(2026, 9, 6));
const all = Array.from({ length: 400 }, (_, i) => daysAgo(today, i)); // one backup per day for 400 days
const kept = [...keep(today, all)].sort();
console.log(`backups taken: ${all.length}, kept after pruning: ${kept.length}`);
console.log(`oldest kept: ${kept[0]} newest kept: ${kept[kept.length - 1]}`);
console.log(`a 2-day-old ransomware infection? the 7 dailies still hold clean copies before it`);Output:
backups taken: 400, kept after pruning: 21
oldest kept: 2025-11-01 newest kept: 2026-10-06
a 2-day-old ransomware infection? the 7 dailies still hold clean copies before itExpectedbackups taken: 400, kept after pruning: 21 oldest kept: 2025-11-01 newest kept: 2026-10-06 a 2-day-old ransomware infection? the 7 dailies still hold clean copies before it
Press Run. Snippets must be self-contained — no network, files, or native modules.
Grandfather-father-son retention keeps many recent copies and a few old ones. Pick the depth based on how long corruption can go unnoticed: a bug that silently writes bad prices for three weeks needs backups older than three weeks.
Real failures worth knowing
- GitLab, January 2017: an engineer deleted the primary database directory during an incident. Several of their backup mechanisms turned out not to be working, and they restored from a staging copy about six hours old, losing roughly six hours of data. Their public postmortem is required reading.
- OVHcloud Strasbourg fire, March 2021: a data center building burned. Customers who stored backups in the same site lost both production and backups.
- Lost encryption keys: encrypted backups are useless if the key was deleted with the account or region. Back up key material or use multi-region keys, and include key access in drills.
Restore drills
A restore drill is a scheduled, measured exercise:
- Pick a random backup (not the newest one) and a random target time.
- Restore into an isolated environment using only the documented runbook.
- Run integrity checks: row counts, checksums, application smoke tests.
- Record the elapsed time per step and compare to the RTO.
- Fix whatever was missing in the runbook, then repeat next quarter.
What happens if you skip these practices
- Backups in the same account: one compromised admin credential deletes production and backups together.
- No PITR: your RPO for logical corruption is the time since the last snapshot, which can be a full day.
- No restore tests: you learn the restore takes 14 hours, or fails on a version mismatch, during the outage.
- Replication treated as backup: the bad delete propagates to every replica within seconds.
Pros and cons
| Approach | Pros | Cons |
|---|---|---|
| Snapshots only | Simple, fast, cheap with incrementals | Coarse RPO, crash-consistent only |
| Logical dumps | Portable, table-level restore | Slow for big data, load on the database |
| PITR logs | Second-level recovery point | Chain fragility, replay time grows with log volume |
| Immutable offsite copy | Survives ransomware and account compromise | Storage cost, retention cannot be shortened by mistake or on purpose |
Interview Q&A
Someone ran a bad migration that corrupted a column 40 minutes ago. Walk through recovery.
Answer
Stop the bleeding first (pause the job or feature flag it off). Restore the latest base backup into a new instance and replay WAL to just before the migration. Diff the affected column against production, then copy the correct values back or swap traffic, and reconcile writes made in the last 40 minutes.
What is the difference between crash-consistent and application-consistent backups?
Answer
Crash-consistent captures disk state as if power failed; the database can recover but cross-volume or cross-service state may be inconsistent. Application-consistent coordinates with the application so the captured state is clean.
How do you protect backups from ransomware?
Answer
Keep a copy in a separate account with separate credentials, use write-once object lock or vault lock, and alert on deletion attempts. Test that you can restore from that isolated copy.
How do you know your backups work?
Answer
Scheduled automated restores with integrity checks, alerts when a backup or log shipment is missing, and timing every restore against the RTO.
Why restore into a new instance?
Answer
You can compare it with the damaged primary, and you keep the damaged copy as evidence.
What does 3-2-1-1-0 add?
Answer
One immutable or offline copy, and zero errors on verified restore tests, on top of three copies, two storage types, and one offsite copy.
Why does a three-week silent price bug need more than last night's snapshot?
Answer
Grandfather-father-son keeps a few old copies so you can go back past the unnoticed corruption, not only the last seven dailies.
What broke for customers in the OVHcloud Strasbourg fire?
Answer
Backups stored in the same site burned with production.
Check yourself
Change the recovery target in the PITR snippet from 129 to 120 and to 140. Write down which rows survive, and which post-incident write you would have to rebuild from the API log.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Object storage, Secrets and KMS, Sync, async, and semi-sync replication.