fsync, Group Commit & Durable Latency
OS page cache lie; fsync/fdatasync; group commit coalesces commits; latency vs throughput; disk-lies / BBWC.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
One fsync for the group
Prefer
Group commit with a bounded wait
Gather txns that became ready during the fsync wait. One fdatasync makes the whole group durable. Ack together. Throughput climbs; extra latency is the batch wait.
- No txn in the group is acked until the shared sync succeeds.
- Postgres and InnoDB already do this under load.
- Cap wait and batch size so p99 stays inside the SLO.
Alternative
fsync every commit with no grouping, or async as the default
Solo fsync: durable TPS ≈ 1000/fsync_ms. Async: the page-cache lie becomes lost acknowledged commits. Both are valid knobs — not silent defaults on payments.
- synchronous_commit and innodb_flush_log_at_trx_commit are RPO vs latency.
- Controllers that lie make fsync return early.
- Sync replica ack is extra RTT, not a replacement for local WAL discipline.
Two commits, one disk barrier
The OS cache is not durability. The group is not durable until the shared sync returns.
- 1
Append WAL bytes with write
May sit in the page cache. A following read sees them. Power loss does not. - 2
Gather ready commits
Group commit waits a bounded time or batch size. - 3
One fdatasync on the WAL fd
Data vs data+metadata. WAL appends often only need data. - 4
Ack the batch
All members become durable together — or none do if power dies first.
Overview
write() only hits the OS page cache — a lie until fsync / fdatasync (or equivalent) forces durable media. Databases batch many transaction commits into one group commit fsync to raise throughput while bounding durable latency. Replication flush/ack policies and hardware (battery-backed cache, disable-barriers myths) change what "durable" really means.
p99 commit latency is often fsync-bound, not SQL-bound. Misunderstanding synchronous_commit=off, InnoDB flush settings, or "disk controllers that lie" causes data-loss incidents. Group commit is the classic systems interview topic connecting OS I/O to database throughput.
Core mechanism
- Engine appends WAL bytes via
write→ may sit in page cache. - To acknowledge a durable commit, engine calls
fdatasync/fsyncon the WAL fd (semantics: data vs data+metadata). - Group commit: gather all txns that became ready during the wait; one sync makes all of them durable; then ack the batch.
- Background writers flush data pages separately; the force-log rule still applies.
- Replicas: sync/async replication and
synchronous_commitremote variants trade RPO vs latency.
OS page cache lie
After write, read sees data, but power loss loses it. Only sync / FUA / barriers push to stable storage — unless the device silently caches without power-fail protection.
Failure modes
- Disk lies / disabled barriers: fsync returns early → acknowledged commits vanish on power loss.
- Battery-backed write cache (BBWC): safe when the battery is healthy; treated as durable early.
- Too little grouping: fsync storms, low TPS.
- Too much wait: high commit latency to fill the batch.
Durability knobs
| Knob | Pros | Cons |
|---|---|---|
| fsync every commit | Minimal RPO | Worst durable latency |
| Group commit | High durable TPS | Small batching delay |
| Async commit | Lowest latency | Lose recent commits |
| Sync replication | Survive node loss | Network RTT in the commit path |
| BBWC / NVMe with capacitors | Fast durable writes | Hardware dependency / failure modes |
Group commit sequence
Sequence
- 1
Txn1 → Engine
1. COMMIT
- 2
Txn2 → Engine
2. COMMIT
- 3
Engine
3. Append both WAL records
- 4
Engine → OS Disk
4. One fdatasync for group
- 5
OS Disk
5. Power loss before sync: neither durable
- 6
OS Disk → Engine
6. Sync OK
- 7
Engine → Txn1
7. Ack
- 8
Engine → Txn2
8. Ack
Lesson map
fsync, Group Commit & Durable Latency
OS page cache lie; fsync/fdatasync; group commit coalesces commits; latency vs throughput; disk-lies / BBWC.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB t1["Txn1"] t2["Txn2"] e["Engine"] os["OS Disk"] t1 -->|1. COMMIT| e t2 -->|2. COMMIT| e e -->|4. One fdatasync| os os -->|6. Sync OK| e e -->|7. Ack| t1 e -->|8. Ack| t2
Sandbox: group-commit batcher (Python)
Coalesce commits into one sync. Count fsyncs vs txns.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Latency vs throughput (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
write vs fsync?
Answer
write buffers in the OS (and maybe device) cache. fsync waits until data is on stable storage (as the stack claims).
fsync vs fdatasync?
Answer
fdatasync flushes file data; it may skip non-essential metadata — often enough for WAL append workloads. fsync also pushes metadata (size, mtime).
How does group commit preserve durability?
Answer
No txn in the group is acked until the shared sync succeeds; all become durable together. Power loss before that sync loses the whole unacked group — which is correct.
Latency vs throughput tradeoff?
Answer
Waiting to build a group adds latency but amortizes fsync cost → higher TPS. Too little grouping: fsync storm. Too much wait: SLO miss.
Postgres synchronous_commit?
Answer
Can wait for local WAL flush, or remote apply/flush depending on setting — tunes RPO vs latency. off is the async-commit failure mode from the WAL lesson.
What if the disk lies?
Answer
The controller acknowledges before non-volatile media; power loss loses "fsynced" data. Disable write-back without battery; verify with power-pull tests. Community lore: fsync error handling (fsyncgate).
Does SSD change the story?
Answer
Faster syncs, still not free. Parallelism / queue depth and capacitors matter. Write amp still burns endurance.
Relation to replication?
Answer
Local fsync ≠ multi-AZ durability. Sync replica ack policies are the distributed cousin of group commit. You can still lose the node that acked locally if the replica was async.
Pitfalls
Assume 5ms fsync and 2 commits/ms. Compute durable TPS and added latency for groupSize=1 vs 8 vs 64. Then say which Postgres/InnoDB knob moves you toward async, and what a crash deletes in each case.