Replication Lag & Session Guarantees — Read-Your-Writes, Monotonic Reads & Consistent Prefix
Async replicas introduce read-your-writes, monotonic-read, and consistent-prefix anomalies. Fix them with LSN/GTID tokens, sticky replicas, and bounded-staleness routing — measure lag correctly first.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Async followers are always a little behind. Usually it's milliseconds, but during a bulk load, a long query on the replica or a network hiccup it can be minutes. That's replication lag. Lag on its own is fine. The trouble is anomalies the user can see. The three that interviewers ask about come from the 1994 "session guarantees" work. Read-your-writes: after I save, I see my save. Monotonic reads: I never see time go backwards between two reads. Consistent prefix: I never see an answer before its question. Each has a cheap, targeted fix, so you rarely need to send every read to the leader.
Anomalies and fixes
| Anomaly | What the user sees | Cause | Targeted fix | Cost |
|---|---|---|---|---|
| Read-your-writes violation | "I changed my avatar and it reverted" | Write went to the leader, read hit a lagging replica | Read from the leader for N seconds after a write, or pass an LSN/GTID token and only use replicas past it | Some leader load, or a little extra router logic |
| Monotonic reads violation | Comment count 12, refresh, now 9 | Successive reads hit replicas with different lag | Sticky replica per user/session (hash user to replica), or "min LSN seen" token | Uneven load and rebalancing when a replica dies |
| Consistent prefix violation | Reply appears before the message it answers | Partitions/shards replicate independently | Keep causally related writes in one partition, or carry causal metadata | Constrains partition keys |
| Cross-device RYW | Saved on phone, laptop shows old | Token lives on one device | Store last-write LSN server-side per user (session store) | Extra lookup per request |
| Stale reads after failover | Data "disappears" | New leader was behind (async) | Semi-sync plus promote the most caught-up replica | See failover page |
Keep session reads sane under lag
Targeted routing beats "always read the primary" for every query.
- 1
Write on the primary
Capture an LSN / GTID token with the ack. - 2
Route the next read
Sticky replica, token-aware router, or primary for read-your-writes. - 3
Enforce monotonicity
Do not bounce a session to an older replica. - 4
Bound staleness
Stop sending traffic to replicas past a lag threshold.
Decisions
- 1
1. Client writes avatar
- next2. Leader commits at LSN 1007
- next4. Next read carries token
- 2
2. Leader commits at LSN 1007
- 3. returns token lsn 10071. Client writes avatar
- async WAL6a. Serve from replica A
- async WAL, lagging at 1002Replica B skipped
- 3
4. Next read carries token
- next5. Router: which replica replayed at least 1007?
- ?
5. Router: which replica replayed at least 1007?
- yes: replica A at 10096a. Serve from replica A
- none caught up6b. Wait up to 50ms or fall back to leader
- 5
6a. Serve from replica A
- 6
6b. Wait up to 50ms or fall back to leader
- 7
Replica B skipped
Lesson map
Replication Lag & Session Guarantees — Read-Your-Writes, Monotonic Reads & Consistent Prefix
Async replicas introduce read-your-writes, monotonic-read, and consistent-prefix anomalies. Fix them with LSN/GTID tokens, sticky replicas, and bounded-staleness routing — measure lag correctly first.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB w["1. Client writes avatar"] l["2. Leader commits at LSN 1007"] rq["4. Next read carries token"] rt["5. Router: which replica replayed at least 1007?"] w -->|1. Client writes avatar to 2. Leader commits at LSN 1007| l l -->|3. returns token lsn 1007| w w -->|1. Client writes avatar to 4. Next read carries token| rq rq -->|4. Next read carries token to 5. Router: which replica replayed at least 1007?| rt
Press Run. Snippets must be self-contained — no network, files, or native modules.
Output:
leader lsn=7, r1 replay=7, r2 replay=4
naive read -> ('old-avatar.png', 'r2') (stale: user sees old avatar)
token read -> ('new-avatar.png', 'r1') (read-your-writes holds)Strategies compared, and what happens if you pick another
| Strategy | How | Pros | Cons | Where it's used |
|---|---|---|---|---|
| All reads to leader | No replicas for app reads | Perfectly fresh | Leader becomes the read bottleneck; replicas are only for HA | Small apps, money paths |
| Time-window pinning | After a write, read from the leader for max_lag seconds (cookie) | Trivial to build | A guess: breaks when lag exceeds the window | Many Rails/Django setups |
| LSN / GTID token routing | Client or session stores the commit position; router checks replica replay position | Precise; most reads stay on replicas | Needs DB support and router integration | MySQL WAIT_FOR_EXECUTED_GTID_SET, Postgres pg_last_wal_replay_lsn(), MongoDB causal sessions, Aurora/ProxySQL |
| Sticky replica | Hash user to one replica | Monotonic reads for free | Hot replicas; when that replica dies, the user may jump backwards | Read-heavy feeds |
| Bounded staleness | DB guarantees reads are at most T seconds old | Clean contract | Reads can block or be rejected when lag is greater than T | Spanner bounded reads, Cosmos DB, CockroachDB follower reads |
| Synchronous apply | remote_apply to the replica you read from | RYW without routing | Write latency includes replay | Low-volume critical flows |
- If you only pin by time window: it works until the day lag spikes during a backfill. Then support tickets say "my settings don't save". Pair it with a lag circuit breaker: if a replica's lag exceeds the threshold, take it out of the read pool.
- If you send everything to the leader "to be safe": it's correct, but you've given up read scaling, and the leader fails under read load before write load.
- If you rely on sticky replicas only: monotonic reads hold, but your own writes still aren't visible until that replica catches up, so you still need RYW handling.
// Monotonic reads: random replica per request vs sticky replica per user.
// Random routing lets a user see a newer value, then an OLDER one ("time travel").
// Sandbox-runnable: tsc --strict then node. Deterministic PRNG.
let seed = 42;
const rand = () => ((seed = (seed * 1103515245 + 12345) % 2 ** 31) / 2 ** 31);
const replicaLag = [0, 2, 6]; // records behind leader for replicas 0..2
const REQUESTS = 2000;
function simulate(pick: (user: number) => number): number {
let timeTravel = 0;
const lastSeen = new Map<number, number>(); // user -> highest version seen
for (let t = 10; t < REQUESTS + 10; t++) {
const user = Math.floor(rand() * 50);
const r = pick(user);
const versionSeen = t - replicaLag[r]; // leader is at version t
const prev = lastSeen.get(user) ?? -1;
if (versionSeen < prev) timeTravel++; // went backwards!
lastSeen.set(user, Math.max(prev, versionSeen));
}
return timeTravel;
}
const randomPick = () => Math.floor(rand() * replicaLag.length);
const stickyPick = (user: number) => user % replicaLag.length; // hash(user) -> replica
console.log(`random replica per request : ${simulate(randomPick)} backwards reads`);
console.log(`sticky replica per user : ${simulate(stickyPick)} backwards reads`);Output:
random replica per request : 32 backwards reads
sticky replica per user : 0 backwards readsMeasuring lag correctly
- Postgres:
pg_stat_replicationon the primary showswrite_lag,flush_lag,replay_lagand the LSNs. On a replica,now() - pg_last_xact_replay_timestamp()is misleading when the primary is idle, so compare LSNs instead. - MySQL:
Seconds_Behind_Sourcemeasures the relay log, not the source, and reads 0 when the IO thread is stalled. A heartbeat table (pt-heartbeat) is more reliable: the primary writesnow()every second and you compare on the replica. - Causes of lag: single-threaded apply (use parallel apply such as
replica_parallel_workers), long replica queries conflicting with replay (Postgresmax_standby_streaming_delay), huge transactions, under-provisioned replica IO, and cross-region bandwidth. - Alert on lag in seconds and in bytes, and route around replicas that exceed the threshold.
Interview Q&A
Q1. How would you implement read-your-writes for a web app with read replicas?
On write, capture the commit position (Postgres pg_current_wal_lsn(), MySQL GTID) and store it in the user's session. On read, the router chooses a replica whose replay position is at least that token. If none qualifies, it waits briefly (MySQL WAIT_FOR_EXECUTED_GTID_SET with a timeout), then falls back to the leader. Keep the token server-side so it works across devices.
Q2. Why is "read from leader for 5 seconds after a write" fragile? It assumes lag stays under 5 seconds. Lag is unbounded in exactly the situations where the guarantee matters: incidents, bulk jobs, failovers. It also ignores other users' writes (a comment you were notified about) and cross-device reads.
Q3. What is consistent prefix and when does it bite?
Observers should see writes in an order consistent with causality. It bites in sharded or partitioned systems where a question and its answer live on different partitions that replicate at different speeds. Fix it by co-locating causally related data (same partition key, for example conversation_id) or by tracking causal dependencies.
Q4. Monotonic reads vs read-your-writes? RYW concerns your own writes being visible to you. Monotonic reads concerns any data never going backwards between your successive reads. Sticky routing gives monotonic reads. A token with your last write's LSN gives RYW. A token with the max LSN you've seen gives both.
Q5. Replica lag is 30 minutes and climbing. What do you check? Is the replica's apply single-threaded and CPU-bound? Is there a long-running transaction or huge batch on the primary? Is replay blocked by a long query on the replica (conflict wait)? Look at IO saturation and network throughput, and check whether a schema change or index build is replaying. In the meantime, pull it from the read pool.
Go Deeper
- PostgreSQL docs: Hot Standby (query conflicts, replay)
- PostgreSQL docs: pg_stat_replication view
- MySQL 8.4: WAIT_FOR_EXECUTED_GTID_SET and GTID functions
- MongoDB: Causal consistency and read/write concerns
- Werner Vogels: Eventually Consistent, Revisited
Related
- Prev: Sync vs Async vs Semi-Sync Replication - Durability, RPO & Physical vs Logical Logs (
replication-sync-async-semi-sync-physical-logical) - Next: Failover & Split Brain - Detection, Promotion, Fencing & Lost Writes (
replication-failover-split-brain-fencing)
This series:
- Database Replication for Engineers - Leader-Follower, Multi-Leader & Leaderless Quorums (
database-replication-leader-follower-multi-leader-leaderless) - Sync vs Async vs Semi-Sync Replication - Durability, RPO & Physical vs Logical Logs (
replication-sync-async-semi-sync-physical-logical) - Replication Lag & Session Guarantees - Read-Your-Writes, Monotonic Reads & Consistent Prefix (
replication-lag-read-your-writes-session-guarantees) - this page - Failover & Split Brain - Detection, Promotion, Fencing & Lost Writes (
replication-failover-split-brain-fencing) - Multi-Leader & Active-Active - Conflict Detection, LWW, CRDTs & Home Regions (
multi-leader-replication-conflicts-active-active) - Leaderless Replication - N/R/W Quorums, Hinted Handoff, Read Repair & Anti-Entropy (
leaderless-replication-quorums-read-repair-anti-entropy)
Existing Study pages (cross-link only, not rewritten here):
- MVCC & snapshot isolation (
mvcc-snapshot-isolation) - Database sharding - partition keys & rebalancing (
db-sharding-partitioning-keys-rebalancing)