Sync vs Async vs Semi-Sync Replication — Durability, RPO & Physical vs Logical Logs
Async, semi-sync, quorum, and fully sync commit modes trade latency for RPO. Postgres synchronous_commit levels and physical versus logical logs (plus replication slots) decide what a failover can lose and what CDC can see.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
The commit mode decides what "committed" means. In async mode the leader acks after its own WAL fsync and ships the log afterwards: commits are fastest, but if the leader dies, acked writes that never left it are gone. In semi-sync mode the leader waits until at least one follower (or a quorum) has received or flushed the record: you pay one extra round trip and an acked write survives the loss of one node. In fully synchronous mode the leader waits for every follower: maximum durability, but the slowest follower sets your latency and a dead follower stops writes. Separately, the log you ship is either physical (WAL bytes and pages: exact copy, same major version) or logical (row-level changes: cross-version, selective, the basis of CDC).
Commit modes compared
| Mode | Client ack after | Extra latency | RPO on leader loss | Availability risk | Real knobs |
|---|---|---|---|---|---|
| Async | Leader local fsync | ~0 | Seconds of acked writes (whatever was in flight) | None from followers | Postgres default, MySQL default, Redis replicas |
| Semi-sync (receive) | 1 follower received in memory | 1 RTT to nearest follower | ~0 for single failure (not if both lose power) | Falls back to async on timeout (MySQL) | rpl_semi_sync_source_wait_point=AFTER_SYNC, Postgres synchronous_commit=remote_write |
| Semi-sync (flush) | 1 follower fsynced | 1 RTT + follower fsync | 0 for single failure | Blocks if no sync standby is available | Postgres synchronous_commit=on + synchronous_standby_names='ANY 1 (...)' |
| Quorum | Majority / k-of-n flushed | RTT to k-th fastest | 0 while a majority survives | Needs majority up | Raft/Paxos DBs, Aurora 4-of-6 storage writes, ANY 2 (a,b,c) |
| Fully sync | All followers | RTT to slowest | 0 | Any follower down = writes stop | Rarely used beyond 2 nodes |
| Remote apply | Follower has applied (visible to reads) | Plus replay time | 0 and read-your-writes on that follower | Highest latency | Postgres synchronous_commit=remote_apply |
Choose commit durability
Ack policy is the RPO knob. Physical vs logical is the log shape.
- 1
Write on the leader
Append to the WAL / binlog before replicas see it. - 2
Ship the log
Async ships and returns; semi-sync waits for one (or a quorum); sync waits for all named standbys. - 3
Replica applies
Physical replays bytes; logical replays row changes and enables CDC/slots. - 4
Client hears committed
That moment is your RPO promise on failover.
Sequence
- 1
Client → Leader
1. COMMIT
- 2
Leader → Leader
2. WAL fsync locally
- 3
Leader
async mode acks here and risks losing the tail
- 4
Leader → Sync follower
3. stream WAL record
- 5
Leader → Async follower
3. stream WAL record without waiting
- 6
Sync follower → Sync follower
4. write or flush record
- 7
Sync follower → Leader
5. ack at received or flushed LSN
- 8
Leader → Client
6. COMMIT OK in semi-sync mode
- 9
Client
acked write now lives on 2 machines
Lesson map
Sync vs Async vs Semi-Sync Replication — Durability, RPO & Physical vs Logical Logs
Async, semi-sync, quorum, and fully sync commit modes trade latency for RPO. Postgres synchronous_commit levels and physical versus logical logs (plus replication slots) decide what a failover can lose and what CDC can see.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["Client"] l["Leader"] s["Sync follower"] a["Async follower"] c -->|1. COMMIT| l l -->|3. stream WAL record| s l -->|3. stream WAL record without waiting| a s -->|5. ack at received or flushed LSN| l l -->|6. COMMIT OK in semi-sync mode| c
Press Run. Snippets must be self-contained — no network, files, or native modules.
Output:
async acked-then-lost= 2 p50 commit latency= 0.0ms
semi-sync acked-then-lost= 0 p50 commit latency= 4.5ms
sync-all acked-then-lost= 0 p50 commit latency= 60.5msThe numbers above come from a toy model, but the shape is real. Async loses only the writes that were in flight at crash time, and that number grows with throughput and lag. Semi-sync removes that loss for the price of one round trip to the nearest follower. Waiting for all followers ties every commit to the farthest one.
What happens if you choose differently
- Async for a payments table: a leader crash followed by promotion can roll back acknowledged payments. The client was told "paid", and the new leader never saw it. You'd find out during reconciliation. GitHub's 2018 incident is the classic example: after a cross-region failover, writes that existed only on the old primary had to be reconciled by hand.
- Fully sync with 3 followers: one slow EBS volume or GC pause on any follower shows up as p99 commit latency for everyone, and a follower crash halts writes until someone removes it from the sync list.
- Semi-sync that silently degrades: MySQL semi-sync falls back to async after
rpl_semi_sync_source_timeout. Without an alert onRpl_semi_sync_source_status=OFF, you think RPO is 0 when it isn't. Postgres takes the other side: it blocks commits instead of degrading, which is safer for data but bad for availability. - Sync replica in the same AZ only: RPO is 0 for node loss but not for an AZ loss. Put the sync standby in another AZ and keep cross-region replicas async.
Physical vs logical replication
| Physical (WAL / binary / page) | Logical (row changes / binlog ROW / decoding) | |
|---|---|---|
| What's shipped | Byte-level page changes | INSERT/UPDATE/DELETE with keys and values |
| Replica | Exact clone, read-only hot standby | Independent DB that can have extra indexes and tables |
| Version and schema | Same major version and architecture | Can cross major versions and schemas (upgrades) |
| Selectivity | Whole cluster | Per table or publication, row filters |
| DDL | Replicated automatically | Usually not replicated (Postgres); handle it separately |
| Uses | HA standby, read replicas, PITR | Zero-downtime upgrades, CDC to Kafka (Debezium), consolidation, partial copies |
| Gotchas | Replica conflicts with long queries (hot_standby_feedback) | Needs primary keys or replica identity; replication slots can fill the disk |
// Physical (WAL/page) vs logical (row-change) replication records.
// Shows why logical replication can cross versions/schemas and physical can't.
// Sandbox-runnable: tsc --strict then node.
type PhysicalRecord = { lsn: number; relfilenode: number; page: number; offset: number; bytes: string };
type LogicalRecord = { lsn: number; table: string; op: "INSERT" | "UPDATE" | "DELETE"; key: number; row?: Record<string, unknown> };
// The same UPDATE expressed two ways.
const physical: PhysicalRecord = { lsn: 0x16b3740, relfilenode: 16384, page: 42, offset: 128, bytes: "0a1f..ff" };
const logical: LogicalRecord = { lsn: 0x16b3740, table: "users", op: "UPDATE", key: 7, row: { id: 7, email: "a@x.io", plan: "pro" } };
// A subscriber on a NEWER schema: column "plan" was renamed to "tier".
function applyLogical(rec: LogicalRecord, rename: Record<string, string>) {
const row: Record<string, unknown> = {};
for (const [k, v] of Object.entries(rec.row ?? {})) row[rename[k] ?? k] = v;
return `${rec.op} ${rec.table} id=${rec.key} -> ${JSON.stringify(row)}`;
}
function applyPhysical(rec: PhysicalRecord, replicaMajorVersion: number, primaryMajorVersion: number) {
// Physical bytes only make sense on an identical on-disk layout.
if (replicaMajorVersion !== primaryMajorVersion) return "REJECT: on-disk format differs (needs same major version)";
return `write ${rec.bytes} to file ${rec.relfilenode} page ${rec.page}+${rec.offset}`;
}
console.log("physical, same version :", applyPhysical(physical, 16, 16));
console.log("physical, v16 -> v17 :", applyPhysical(physical, 17, 16));
console.log("logical, renamed col :", applyLogical(logical, { plan: "tier" }));Output:
physical, same version : write 0a1f..ff to file 16384 page 42+128
physical, v16 -> v17 : REJECT: on-disk format differs (needs same major version)
logical, renamed col : UPDATE users id=7 -> {"id":7,"email":"a@x.io","tier":"pro"}Statement-based vs row-based (MySQL): shipping SQL statements is compact, but non-deterministic statements (NOW(), UUID(), LIMIT without ORDER BY, triggers) diverge on replicas. Row-based (binlog_format=ROW) is the safe default today. Replication slots (Postgres): a slot makes the primary keep WAL until the consumer confirms it. A dead CDC consumer therefore fills the primary's disk. Alert on retained WAL size, and set max_slot_wal_keep_size.
Pros and cons
Async: pros: fastest, simplest, followers can be anywhere. cons: RPO greater than 0, and failover can lose acked writes. Semi-sync / quorum: pros: RPO = 0 for one failure at about one RTT. cons: latency depends on placement of the sync replica, plus fallback or blocking behavior you must monitor. Physical: pros: exact, low overhead, includes DDL. cons: same version only, all or nothing. Logical: pros: flexible, enables upgrades and CDC. cons: no DDL, more CPU, slot management, needs replica identity.
Interview Q&A
Q1. What does RPO = 0 actually require? Every acknowledged write must be durable on at least two failure domains before the ack: a semi-sync or quorum commit with the sync copy in a different AZ. You also need failover to promote a replica that has that write (the most caught-up sync standby), not just any replica.
Q2. Postgres synchronous_commit levels, from weakest to strongest?
off (no local fsync wait, may lose recent commits even on one node) → local → remote_write (standby received it and handed it to the OS) → on (standby flushed) → remote_apply (standby replayed, so reads there see it). You can set it per transaction, for example SET LOCAL synchronous_commit = off for an analytics insert and on for a payment.
Q3. Why might you use logical replication for a major-version upgrade? Physical replicas need the same major version. With logical replication you build a v17 subscriber from the v16 publisher, let it catch up, then switch application traffic in seconds. That gives you a near-zero-downtime upgrade with rollback possible, provided DDL and sequences are handled separately.
Q4. A CDC consumer has been down for two days and the primary's disk is at 95%. Why?
Its logical replication slot pins WAL. The primary can't recycle segments the slot hasn't confirmed. Fix it now by advancing or dropping the slot (accept a resnapshot), then set max_slot_wal_keep_size and alert on slot lag.
Q5. Semi-sync acked, then both leader and follower lost power. Is the write safe?
Only if the follower flushed the record, not just received it. AFTER_SYNC / remote_write mean it may be only in memory or the OS cache. Durable means fsync'd on a second machine (synchronous_commit=on), or in a separate failure domain such as another AZ.
Go Deeper
- PostgreSQL docs: synchronous_commit and synchronous replication settings
- PostgreSQL docs: Log-Shipping Standby Servers
- PostgreSQL docs: Logical Replication
- MySQL 8.4: Semisynchronous Replication
- GitHub: October 21 post-incident analysis
Related
- Prev: Database Replication for Engineers - Leader-Follower, Multi-Leader & Leaderless Quorums (
database-replication-leader-follower-multi-leader-leaderless) - Next: Replication Lag & Session Guarantees - Read-Your-Writes, Monotonic Reads & Consistent Prefix (
replication-lag-read-your-writes-session-guarantees)
This series:
- Database Replication for Engineers - Leader-Follower, Multi-Leader & Leaderless Quorums (
database-replication-leader-follower-multi-leader-leaderless) - Sync vs Async vs Semi-Sync Replication - Durability, RPO & Physical vs Logical Logs (
replication-sync-async-semi-sync-physical-logical) - this page - Replication Lag & Session Guarantees - Read-Your-Writes, Monotonic Reads & Consistent Prefix (
replication-lag-read-your-writes-session-guarantees) - Failover & Split Brain - Detection, Promotion, Fencing & Lost Writes (
replication-failover-split-brain-fencing) - Multi-Leader & Active-Active - Conflict Detection, LWW, CRDTs & Home Regions (
multi-leader-replication-conflicts-active-active) - Leaderless Replication - N/R/W Quorums, Hinted Handoff, Read Repair & Anti-Entropy (
leaderless-replication-quorums-read-repair-anti-entropy)
Existing Study pages (cross-link only, not rewritten here):
- Storage engines - WAL, B-trees & LSM-trees (
database-storage-engines-wal-b-trees-lsm-trees) - Change Data Capture hub (
cdc-for-engineers-hub) - Zero-downtime DB migrations - expand/contract (
zero-downtime-database-migrations-expand-contract)