Caching
Part 2 of 8 · Redis cacheWrite-Through vs Write-Behind Caching
Cache-aside invalidates after DB commit. Write-through warms cache on the write path. Write-behind acks from cache/queue and flushes later — faster writes, durability risk.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
The cache-aside hub covers the read path: miss, fill, stampede. This lesson is the write path. The question is not "which pattern is best." It is "what is true the instant the client gets 200, and what is true if the process dies one millisecond later."
By the end of this lesson you should be able to:
- Draw cache-aside, read-through, write-through, and write-behind paths, and say who fills the cache
- Say who still has the value after ack, after crash, and after flush
- Pick a pattern from read/write ratio, freshness, and failure domain — not from habit
- Explain why write-through still needs a TTL, and why write-behind still needs stampede control on miss
- Refuse write-behind for money unless the queue is durable and replayable
Payments: cache-aside + DEL vs write-behind
Prefer
Cache-aside, then DEL after DB commit
The ledger commits first. Redis is a hint. A crash after 200 cannot un-capture a payment, and a crash before DEL only leaves a bounded stale window.
- Durability lives in the primary (or a real WAL). Redis can vanish.
- Invalidation is simpler than keeping two writes in lockstep.
- TTL is the backstop if DEL is lost — not the happy path.
Alternative
Write-behind for the capture
Ack after cache or an in-process queue accepts. Flush to the ledger later. Faster p99 writes — until the process dies with the queue.
- A crash before flush loses the capture unless the queue is itself durable.
- Ops will not notice until flush-lag alerts fire. Silent is worse than slow.
- Read-your-writes looks correct in the happy path and lies after failover.
One client write, three possible truths
Vertical path. After step 2 the patterns diverge — that is the interview.
- 1
Client sends a mutation
Update a profile, mint a session, or capture a payment. The API must pick a write path before it acks. - 2
Cache-aside: commit DB, then DEL
Source of truth first. Cache is empty until the next read rebuilds. - 3
Write-through: commit DB, then SET
Same durability as cache-aside, but the cache is already warm. Writes pay extra latency and memory. - 4
Write-behind: SET + enqueue, then ack
The client is told success while the DB may still be empty. Flush is someone else's job. - 5
Crash before flush
In-memory queue is gone. If Redis died too, nobody has the value. That is a lost write — not a stale read.
The three write paths
Interview default for user-facing reads: cache-aside plus invalidate-on-write. The other two patterns exist because some writes have a freshness or latency constraint that invalidation does not meet.
| Pattern | Write path | What is true at ack | Strength | Cost |
|---|---|---|---|---|
| Cache-aside | App writes DB, then DEL | DB has it; cache may miss | Simple; memory only for hot keys | Miss latency + stampede risk |
| Write-through | App writes DB, then SET | DB and cache agree | Always warm after write | Writes slower; unused keys waste memory |
| Write-behind | App writes cache + queue; worker flushes | Cache has it; DB maybe not | Absorbs write bursts | Durability risk; harder correctness |
Decisions
- 1
Step 1 Client write
- nextStep 2 Pattern
- ?
Step 2 Pattern
- cache-asideStep 3a Commit DB, then DEL the key
- write-throughStep 3b Commit DB, then SET cache in the same request
- write-behindStep 3c SET cache, enqueue durably, ack the client
- 3
Step 3a Commit DB, then DEL the key
- DEL fails after commitFailure path - stale until TTL; retry or outbox
- 4
Step 3b Commit DB, then SET cache in the same request
- two writers interleaveFailure path - cache keeps the older value; version the SET
- 5
Step 3c SET cache, enqueue durably, ack the client
- nextStep 4 Worker flushes to DB, coalescing per key
- crash before flushFailure path - acked write lost unless the queue is durable
- 6
Step 4 Worker flushes to DB, coalescing per key
- 7
Failure path - stale until TTL; retry or outbox
- 8
Failure path - cache keeps the older value; version the SET
- 9
Failure path - acked write lost unless the queue is durable
Lesson map
Write-Through vs Write-Behind Caching
Cache-aside invalidates after DB commit. Write-through warms cache on the write path. Write-behind acks from cache/queue and flushes later — faster writes, durability risk.
Architecture. Step 1 Client write Ready. Step 2 Pattern Ready. Step 3a Commit DB, then DEL the key Ready. Step 3b Commit DB, then SET cache in the same request Ready. Step 3c SET cache, enqueue durably, ack the client Ready. Step 4 Worker flushes to DB, coalescing per key Ready. Failure path - stale until TTL; retry or outbox Ready. Failure path - cache keeps the older value; version the SET Ready. Failure path - acked write lost unless the queue is durable Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB W["Step 1 Client write Ready"] P["Step 2 Pattern Ready"] CA["Step 3a Commit DB, then DEL the key Ready"] WT["Step 3b Commit DB, then SET cache in the same request Ready"] WB["Step 3c SET cache, enqueue durably, ack the client Ready"] FL["Step 4 Worker flushes to DB, coalescing per key Ready"] F1["Failure path - stale until TTL retry or outbox Ready"] F2["Failure path - cache keeps the older value version the SET Ready"] F3["Failure path - acked write lost unless the queue is durable Ready"] W -->|continues| P P -->|cache-aside| CA P -->|write-through| WT P -->|write-behind| WB WB -->|continues| FL CA -->|DEL fails after commit| F1 WT -->|two writers interleave| F2 WB -->|crash before flush| F3
The mermaid is the whole lesson: ack moves. In cache-aside and write-through, ack is after the primary. In write-behind, ack is before the primary.
Who fills the cache
Write path is only half the interview. The other half is who loads on miss.
| Strategy | Who fills cache | Write path | Best when |
|---|---|---|---|
| Cache-aside (lazy) | App on miss | App writes DB, then DEL (or SET) | Most web/API reads; policy stays in one service |
| Read-through | Cache library / proxy loads DB | Same as aside, or paired write-through | Shared loading logic across services |
| Write-through | Cache on write | App → cache → DB (sync) | Must be warm immediately after write |
| Write-behind (write-back) | Cache on write | App → cache; DB async | High write QPS and a durable queue |
Read-through does not delete stampede races. It only centralizes TTL, serialization, and metrics. Pair it with single-flight. It also does not keep multi-region caches coherent — that is near-cache consistency.
Cache-aside recap (the default)
The app owns policy. Redis is a shared L2, not the source of truth.
- Write the primary and wait for
COMMIT DELthe cache key- The next read misses and rebuilds
Do not SET the cache before commit. Readers must never see a row that rolls back. If DEL is lost, retry it (outbox / queue); TTL bounds how long the stale copy can live. That backstop is why this cluster still has a TTL lesson.
Forget invalidate on all write paths — admin tools, batch jobs, and migrations silently leave stale data. Mixing "this service write-throughs, that job writes SQL directly" is how you get a week of ghost profile bios.
Sequence
- 1
Client
Step1 read
- 2
Client → App
GET /item/42
- 3
App → Cache
GET item:42
- 4
Cache → App
value
- 5
App → Client
200
- 6
App
Step2 load DB
- 7
Cache → App
nil
- 8
App → DB
SELECT
- 9
DB → App
row
- 10
App
Step3 populate
- 11
App → Cache
SET item:42 TTL
- 12
App → Client
200
The miss path is also the stampede path. Jitter, single-flight locks, and soft TTL are not optional extras. Partial-object SET on write races — prefer DEL and reload a coherent snapshot, or version the blob.
Write-through — warm on the write path
Write-through means: the cache is updated as part of the write, so a read immediately after should hit.
The safe order is the same as cache-aside's durability order:
- Write the primary,
COMMIT SETthe cache (with a jittered TTL)- Ack the client
A tempting inversion is SET then DB. That gives read-your-writes even if the DB is slow — and it also lets other pods read a value that never committed. Prefer commit-then-SET. If you need read-your-writes on the same request, return the value you just wrote from memory; do not make Redis the proof.
When it wins
- A brand-new session / API token must be readable on the next request to any pod. A cache-aside miss that races a slow DB looks like "logged in but 401."
- Tiny, hot objects where almost every write is read back immediately and write QPS is modest.
- Config blobs that you would rather pay memory for than pay a miss after every deploy-time write.
When it loses
- High write QPS on keys nobody reads (analytics counters, last-seen timestamps). You are buying Redis memory and write latency for cold data.
- Wide rows / huge blobs. You amplify write amplification into the cache cluster.
- Multi-region.
SETin region A does not magicallySETregion B. You still have an invalidation or replication story. Write-through only syncs the local cache+DB pair — not linearizability across regions. Near-cache / distributed caching.
Write-through does not retire TTL. A later writer can skip the cache. A replica can be restored from backup. A bug can SET the wrong JSON. TTL is the staleness bound when the write path is incomplete — same as cache-aside. Depth: TTL + jitter and invalidation.
Write-through does not retire stampede control. Keys still expire. The fill path on miss is still the single-flight lesson.
Deep dive · Write-through through a cache proxy
Some stores (write-through Redis + a DB connector, or an ORM interceptor) hide the two writes behind one API. The failure mode does not disappear — it moves. You still need to know which write is the commit, what happens if SET fails after COMMIT (you have a valid DB and a stale or empty cache: that is cache-aside's world, retry SET or DEL), and what happens if COMMIT fails after SET (you must roll the cache back or DEL). An interceptor that "just writes both" without that undo is an outage generator.
Write-behind — ack from cache, flush later
Write-behind (write-back) means: the client is told success when the cache and a queue accept the write. A worker later persists to the primary.
SETcache- Enqueue
(key, value)(or a mutation event) - Ack
- Async flush: pop queue, write DB
This is how you absorb a write spike that the primary cannot take synchronously: likes, last-seen, impression counts, ephemeral presence. The product question is "is a lost update a bug or a shrug?"
The only honest durability story is: the queue is a WAL. Kafka / Pulsar / Redis Streams with a consumer group and replication, or a Postgres outbox table. An in-process deque is a teaching model, not a store. If the pod dies, that deque is gone.
If Redis holds the only copy and Redis dies before flush, the write is gone. Replicated Redis is still not a ledger: eviction, FLUSHALL, and failovers with data-loss are real. See eviction policies. A write-behind buffer in Redis usually wants AOF (or a Redis Stream / external WAL). RDB-only snapshots can lose the recent buffer.
Read-your-writes is easy in write-behind: the cache has the new value. That is why the happy-path demo looks great. After a cache restart, readers fall through to a DB that never saw the write. You will debug "it worked in staging" for a day.
Picking a pattern
| Workload | Pick | Why |
|---|---|---|
| User profile, product catalog | Cache-aside + DEL | Read-heavy; writes are rare; stale-until-TTL is acceptable if invalidate fails |
| New session / CSRF token | Write-through | Next request on any pod must hit a warm key |
| Payments, inventory decrement | Primary first (cache-aside or WT). Never WB on an in-memory queue | Lost write is money or oversell |
| Likes, last-seen, impressions | Write-behind onto a durable stream | Lossy or delayed is OK; primary cannot take the QPS |
| Multi-region profile | Cache-aside + local invalidate; bounded staleness | WT does not sync regions |
Do not mix patterns on the same key without writing it down. One service write-throughs user:42 while a batch job updates Postgres and forgets Redis. The cache is now a liar with a long TTL.
Failure modes worth drawing
Decisions
- 1
Write returns 200
- nextWhere did we ack?
- ?
Where did we ack?
- after COMMITPrimary has the row
- after cache plus queueCache has it
- 3
Primary has the row
- nextDEL or SET issued?
- ?
DEL or SET issued?
- DEL lostStale cache until TTL
- SET doneWarm and durable
- 5
Stale cache until TTL
- 6
Warm and durable
- 7
Cache has it
- nextCrash before flush?
- ?
Crash before flush?
- queue was in-memoryLost write
- queue is a WALWorker replays
- 9
Lost write
- 10
Worker replays
Other races that show up in interviews:
- SET before COMMIT — other pods read uncommitted / rolled-back state.
- SET-on-write with two writers — older blob wins. Prefer
DELor CAS with a version. - Partial success — DB updated, cache neither SET nor DEL. Retry the cache op; TTL is the backstop.
- Worker double-flush — flush must be idempotent. Last-write-wins on a full blob, or version/compare on a patch.
- Bypass writers — anyone who can
UPDATEthe primary must invalidate or you do not have a cache, you have a rumor.
Redis down is still an explicit policy from the hub: fail-open (hit DB, risk overload) vs fail-closed (503 / local stale). Write-behind cannot fail-open the flush; the write already returned 200.
Working sketches
Production uses a real client. The sandbox below is the protocol: a dict for the DB, a dict for Redis, a deque for the flush queue.
// Sketch — same three verbs. Not runnable here.
type Store = {
db: Map<string, string>;
cache: Map<string, string>;
queue: Array<[string, string]>;
};
export function writeThrough(s: Store, key: string, value: string) {
s.db.set(key, value); // commit source of truth first
s.cache.set(key, value); // then warm
}
export function writeBehind(s: Store, key: string, value: string) {
s.cache.set(key, value);
s.queue.push([key, value]);
}
export function flushBehind(s: Store, n = 100) {
let flushed = 0;
while (s.queue.length && flushed < n) {
const pair = s.queue.shift();
if (!pair) break;
s.db.set(pair[0], pair[1]);
flushed++;
}
return flushed;
}In-memory store (run this)
Dict-backed DB + cache + queue. Crash is "drop the process-local queue." No network, no redis-py.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Change payment:9 to a like-counter and argue whether "cache only" is acceptable. Then give the queue a fake disk (flushed_log = list(s.queue) before crash) and replay. That is the difference between a demo and a WAL.
Interview Q&A
Cache-aside vs read-through?
Answer
Same miss pattern. Cache-aside puts load logic in the app; read-through centralizes it in a library or proxy. Consistency races and stampede issues remain — the library does not delete them. Whoever fills the cache owns serialization and TTL policy.
Default for user profile reads — which write path?
Answer
Cache-aside + invalidate-on-write. Profiles are read-heavy, writes are rare, and memory should follow the hot set. Write-through wastes cache on users nobody is looking at. Write-behind adds a durability story you do not need.
When does write-through win?
Answer
When almost every write is read back immediately and the DB can take the extra write latency. Classic: a session just created must be readable on the next request to any pod. Write volume must stay manageable — you pay Redis memory and a SET on every mutation.
What is the real danger of write-behind?
Answer
Process death or queue loss before flush. The client already got 200. If the queue was in-memory, the write is gone. If Redis was the only copy and Redis dies, same outcome. Only a replicated WAL/stream makes this honest. Payments do not belong here.
Why not update the cache before the DB commit?
Answer
Other readers (and other pods) may see uncommitted or rolled-back data. Prefer COMMIT then SET or DEL. Same-request read-your-writes can return the in-memory value you just wrote without publishing it early.
Does write-through remove the need for TTL and stampede control?
Answer
No. Writers will skip the cache. Restores, bugs, and replica lag still exist. TTL is the staleness bound when invalidation/SET fails. Expiry still stampedes — combine with jitter and single-flight on miss.
How do you get read-your-writes with cache-aside?
Answer
Return the value from the request that just committed. Optionally SET after COMMIT if the next hop is another pod. Do not make "it is in Redis" the definition of success.
Multi-region — does write-through save you?
Answer
No. A SET in one region is not a SET in another. Invalidate locally, accept bounded staleness, or replicate invalidation events. Instant global cache consistency is not a Redis feature you get for free. Write-through is not linearizability across regions: near-cache.
Redis persistence for a write-behind buffer?
Answer
Often AOF or a Redis Stream / external queue. RDB-only snapshots can lose the recent buffer. Eviction of an unflushed key is a lost write. Eviction, persistence, and types.
What do you monitor on write-behind?
Answer
Queue depth, oldest-unflushed age, flush failures, and the gap between ack and durable write. Alert on lag before customers notice missing rows. Cache hit rate will look healthy while the DB is lying.
Can you mix patterns in one service?
Answer
Yes, per key class, if you document it. Sessions write-through, profiles cache-aside, impression counts write-behind to a stream. Mixing on the same key — or letting a batch job bypass the cache — is how stale data becomes an incident.
Write-behind flush failed; what should the worker do?
Answer
Retry with backoff on the same durable offset. The flush must be idempotent (full blob last-write-wins, or a version). Do not ack the queue item until the primary accepts it. Dropping the event to "unblock" the queue is a silent lost write.
Pitfalls
Pick one: profile bio, new session, payment capture, like-counter. For each, say (1) where ack happens, (2) who has the value if the pod dies now, (3) who has it if Redis dies now, (4) what the next GET returns. Then run the playground, comment out flush_behind, and re-read "who still has it?" for payment:9.
If you would still write-behind the capture, name the WAL (topic, replication, consumer group) and the lag SLO. If you cannot, you wanted cache-aside.
Go Deeper
Official docs
- Redis cache-aside overview
- Redis — Caching patterns
- Microsoft Learn — Cache-Aside pattern
- AWS whitepaper — caching patterns
- AWS ElastiCache — Caching strategies
In this cluster