Write-Ahead Log — Durability, Checkpoints & Crash Recovery
Force log before data (ARIES-style); LSN, checkpoints, REDO/UNDO; sync vs group commit vs async; torn pages / doublewrite.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How hard you fsync the log
Prefer
Group commit (default mental model)
Many ready commits share one fsync. Acknowledged commits are as durable as sync-per-commit; throughput jumps; latency is batched by a few milliseconds.
- Ack only after the group fsync succeeds.
- Postgres and InnoDB both group in practice under load.
- Tune wait vs batch size — too little grouping is an fsync storm.
Alternative
Async WAL as the house default, or fsync every commit with no grouping
Async: client saw success, crash deletes the commit. Sync-every-commit with no grouping: durable TPS collapses to 1/fsync_ms.
- synchronous_commit=off / innodb_flush_log_at_trx_commit=2 are RPO knobs, not free lunch.
- Replication does not replace local WAL discipline; ack policy still matters.
- Checkpoints that are too rare make REDO the outage.
Commit, then crash
Force the log. Flush data pages later. Recovery starts at the last checkpoint.
- 1
Log the change + COMMIT
LSNs increase. Pages remember the LSN that dirtied them. - 2
Force WAL per sync policy
Sync, group commit, or async. Only durable LSN survives power loss. - 3
Ack the client
Never ack a durable commit before the policy's LSN is on stable media. - 4
Async page flush
Force-log: a dirty page may hit its final location only after its redo is durable. - 5
Recover: checkpoint → REDO → UNDO
Replay from last checkpoint; roll back losers if the engine is ARIES-style.
Overview
The write-ahead log (WAL / redo log) records changes before data pages are durable. Force the log first (ARIES-style), assign LSNs, checkpoint periodically, and on crash REDO from the last checkpoint. Sync policy (every commit vs group commit vs async) decides which recent commits survive power loss.
"Did we lose data?" is almost always a WAL/fsync question. Mis-tuned synchronous_commit, InnoDB innodb_flush_log_at_trx_commit, or RocksDB WAL sync turns into silent commit loss or multi-second durable latency. Checkpoints that are too rare make recovery slow; too frequent amplify writes.
Core mechanism
- Begin txn / mutate → generate log records describing the change.
- Assign LSN (log sequence number) — monotonically increases; pages remember the latest LSN that dirtied them.
- Force-log rule: before a dirty page may be written to its final location, its redo must be on durable WAL (STEAL/NO-STEAL variants exist; classic ARIES allows STEAL with UNDO).
- Commit: append a COMMIT record; durability depends on sync policy.
- Checkpoint: record a position plus dirty-page / active-txn metadata so recovery need not scan from genesis.
- Crash recovery: find last checkpoint → REDO forward to bring pages to latest committed state → UNDO loser txns (if needed).
Torn pages and doublewrite
A 16KB page write can partial-fail mid-sector. InnoDB's doublewrite buffer writes a coherent copy first so recovery can repair torn pages. Postgres logs full-page images (FPI) in WAL after checkpoint for similar reasons.
WAL protects logical durability and ordering. Doublewrite / FPI protect torn physical page writes. They are cousins, not substitutes.
Sync policy
| Policy | Pros | Cons |
|---|---|---|
| Sync per commit | No acknowledged-commit loss on crash | fsync dominates latency |
| Group commit | High durable TPS | Small wait to fill the batch |
| Async WAL | Lowest latency | Lose recent commits |
| Replication as durability | Geo survival | Still need local WAL discipline; replica ack policy matters |
Checkpoint frequency:
- Too rare → long REDO, long RTO.
- Too frequent → extra flushes / write amplification.
Group-commit mechanics (OS page cache, fdatasync, disks that lie): fsync.
Crash recovery flow
Sequence
- 1
Engine
1. Append redo + COMMIT
- 2
Client → Engine
COMMIT
- 3
Engine → WAL
2. Force WAL (sync policy)
- 4
WAL → Engine
3. Durable LSN
- 5
Engine → Client
4. Ack commit
- 6
Engine
5. Async flush dirty pages
- 7
Engine → Data pages
Write pages (LSN stamped)
- 8
Engine
CRASH
- 9
Engine → WAL
6. Read last checkpoint
- 10
Engine → WAL
7. REDO forward to end of WAL
- 11
Engine → Data pages
8. Apply missing page fixes / UNDO losers
- 12
Engine
9. Accept new traffic
Lesson map
Write-Ahead Log — Durability, Checkpoints & Crash Recovery
Force log before data (ARIES-style); LSN, checkpoints, REDO/UNDO; sync vs group commit vs async; torn pages / doublewrite.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["Client"] e["Engine"] w["WAL"] d["Data pages"] c -->|COMMIT| e e -->|2. Force WAL| w w -->|3. Durable LSN| e e -->|4. Ack commit| c e -->|Write pages (LSN| d e -->|6. Read last| w
Sandbox: WAL append + checkpoint (Python)
Educational mini-WAL. Unsynced commits vanish on crash; a checkpoint force-logs through its LSN.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Group-commit flavored WAL (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Force-log rule in one sentence?
Answer
Never write a dirty data page to its final disk location until the redo describing that change is durable in the WAL.
REDO vs UNDO roles?
Answer
REDO brings pages forward to reflect logged changes after crash. UNDO rolls back loser (incomplete) transactions. The analysis phase finds winners/losers in full ARIES.
Why full-page images in Postgres WAL?
Answer
After a checkpoint, the first modification to a page may log an FPI so REDO can restore a coherent page if a tear occurred.
What does a checkpoint actually buy you?
Answer
Bounds recovery: you start REDO near the checkpoint instead of the beginning of time. It may also flush dirty pages. Too aggressive → extra write amp and latency spikes.
Sync every commit vs group commit?
Answer
Same durability for acknowledged commits if the group fsync completes before ack. Group commit amortizes fsync across many txns (throughput up, latency slightly batched). Depth: fsync.
Async commit failure mode?
Answer
Client saw success but crash before WAL sync → that commit disappears. Fine for some caches/metrics; bad for payments.
Checkpoint too aggressive — what hurts?
Answer
Extra write amplification and latency spikes from flushing; WAL recycling pressure. RTO is not the only SLO.
How does InnoDB doublewrite relate to WAL?
Answer
WAL protects logical durability/ordering. Doublewrite protects against torn physical page writes during data-file updates. Buffer-pool dirty pages still need both: B-Tree internals.
Pitfalls
Draw COMMIT → force WAL → ack → async page flush. Mark a crash (a) before fsync, (b) after fsync but before page flush, (c) after checkpoint. For each, say which commits survive and whether REDO or UNDO runs. Then name the Postgres/InnoDB knob that picks (a) vs (b).