Caching
Part 6 of 8 · Redis cacheRedis Eviction Policies & Memory Management
maxmemory plus a policy. Caches usually allkeys-lru or allkeys-lfu. volatile-* protects keys without TTL. noeviction makes writes fail. Pick string/hash/list/set/zset/stream from the access pattern; RDB vs AOF vs none from durability. Eviction is not capacity planning.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
TTL (previous lessons) is proactive: you decided the key should die. Eviction is reactive: the process is at the ceiling and something has to give. Mixing those up is how sessions disappear while a million cold product cards remain — or how writes start failing at peak.
By the end of this lesson you should be able to:
- Set
maxmemoryon purpose and name the policy - Contrast
allkeys-*vsvolatile-*vsnoeviction - Explain that Redis LRU is sampled, not a perfect linked list
- Say when LFU beats LRU
- Map blob / fields / queue / unique / ranked / log to Redis types
- Contrast RDB snapshots vs AOF
everysecvs persistence off for a disposable cache - List the metrics that show an eviction storm (and why hit rate drops)
The instance is full. What should Redis do?
Prefer
Pure cache: allkeys-lru or allkeys-lfu, sized on purpose
Every key is disposable. Evict the cold (LRU) or the rarely used (LFU). Keep maxmemory under the host so the kernel OOM killer is not your policy.
- Writes keep succeeding; the working set stays hot.
- TTL still exists as a staleness bound, not as the only way keys leave.
- You watch evicted_keys and hit rate — a storm is an incident, not 'Redis working as designed' without a page.
Alternative
noeviction on the cache, or allkeys-lru on mixed durable keys
noeviction: SET returns an error at the worst time. allkeys-* on an instance that also holds sessions without TTL: LRU will drop a session to make room for a one-hit cache fill.
- noeviction is correct for a primary-like Redis, wrong for a cache tier.
- volatile-lru plus TTL on cache keys only is the mixed-workload move.
- Leaving maxmemory unset means the host runs out of RAM instead.
A SET when the process is at the ceiling
Vertical path. Policy is the branch that decides whether the write lands.
- 1
SET / write arrives
New key, or a bigger value on an existing key. used_memory would exceed maxmemory. - 2
Under maxmemory? store it
Happy path. No eviction. - 3
noeviction → error
OOM command not allowed. The client must handle a failed SET. Cache fills start failing closed. - 4
allkeys-lru → sample and evict cold
Any key can go, TTL or not. Sessions without expire are eligible. - 5
volatile-lru → evict among keys with TTL
Durable keys without expire stay. If nothing is volatile, the write still fails.
maxmemory is a ceiling, not a plan
maxmemory is the hard cap on the dataset (with overheads — see used_memory vs used_memory_rss). When a command would pass the cap, Redis runs maxmemory-policy.
If you do not set it, Redis grows until the host is unhappy. The Linux OOM killer is an eviction policy with worse telemetry. Always set maxmemory on a cache, leaving headroom for the OS, replication buffers, and fragmentation.
Eviction is not capacity planning:
- If the working set does not fit, you will evict hot keys, miss, reload, evict, miss — a eviction storm
- TTL hygiene still matters; eviction does not fix unbounded key growth from a bug
- Hot-key CPU is not a memory problem; split keys or add L1
- Sizing still starts from "what must stay resident" plus overhead, then you pick a policy for the overflow
Decisions
- 1
Step 1 Write arrives
- nextStep 2 used_memory over maxmemory?
- ?
Step 2 used_memory over maxmemory?
- noStep 3 Store the value
- yesStep 4 maxmemory-policy
- 3
Step 3 Store the value
- ?
Step 4 maxmemory-policy
- noeviction - the defaultStep 5a Reject the write with OOM; reads still work
- allkeys-lru or allkeys-lfuStep 5b Sample keys, evict the coldest
- volatile-lru, volatile-lfu, volatile-ttlStep 5c Any keys with a TTL?
- 5
Step 5a Reject the write with OOM; reads still work
- 6
Step 5b Sample keys, evict the coldest
- nextStep 3 Store the value
- ?
Step 5c Any keys with a TTL?
- yesStep 6 Evict only among keys that have a TTL
- noStep 5a Reject the write with OOM; reads still work
- 8
Step 6 Evict only among keys that have a TTL
- nextStep 3 Store the value
Lesson map
Redis Eviction Policies & Memory Management
maxmemory plus a policy. Caches usually allkeys-lru or allkeys-lfu. volatile-* protects keys without TTL. noeviction makes writes fail. Pick string/hash/list/set/zset/stream from the access pattern; RDB vs AOF vs none from durability. Eviction is not capacity planning.
Architecture. Step 1 Write arrives Ready. Step 2 used_memory over maxmemory? Ready. Step 3 Store the value Ready. Step 4 maxmemory-policy Ready. Step 5a Reject the write with OOM; reads still work Ready. Step 5b Sample keys, evict the coldest Ready. Step 5c Any keys with a TTL? Ready. Step 6 Evict only among keys that have a TTL Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB W["Step 1 Write arrives Ready"] M["Step 2 used_memory over maxmemory? Ready"] OK["Step 3 Store the value Ready"] P["Step 4 maxmemory-policy Ready"] E["Step 5a Reject the write with OOM reads still work Ready"] L["Step 5b Sample keys, evict the coldest Ready"] V["Step 5c Any keys with a TTL? Ready"] X["Step 6 Evict only among keys that have a TTL Ready"] W -->|continues| M M -->|no| OK M -->|yes| P P -->|noeviction - the default| E P -->|allkeys-lru or allkeys-lfu| L P -->|volatile-lru, volatile-lfu, volatile-ttl| V V -->|yes| X V -->|no| E L -->|continues| OK X -->|continues| OK
The policies
Redis ships several. You will be asked to pick, not to recite the man page.
| Policy | Evicts from | Order | Typical use |
|---|---|---|---|
noeviction | nobody | writes fail | Redis as a durable store; never a cache tier |
allkeys-lru | every key | approximate least-recently-used | Default cache |
allkeys-lfu | every key | approximate least-frequently-used | Cache where recency ≠ importance |
allkeys-random | every key | random | Rarely; uniform junk |
volatile-lru | keys with an expire | sampled LRU | Mixed: cache keys have TTL, sessions do not |
volatile-lfu | keys with expire | sampled LFU | Mixed, frequency-shaped |
volatile-ttl | keys with expire | nearest to expire | Prefer dropping what would die soon anyway |
volatile-random | keys with expire | random | Mixed, cheap, rough |
allkeys vs volatile is the interview distinction. The Redis default is noeviction. volatile-* ignores keys with no TTL. volatile only considers keys with an expire set; if no key has a TTL, volatile-* behaves like noeviction and writes fail. That protects a session or a lock you stored without EX — and it means a cache that forgot to set TTL will not evict those keys. You can fill the instance with immortal junk and then fail writes. If every cache key has a TTL, volatile-lru and allkeys-lru behave similarly until someone SETs without EX.
Sampled LRU and LFU
Redis does not keep a perfect LRU list of every key (that costs memory you wanted for data). It samples a few keys (maxmemory-samples, default 5) and evicts the best victim in the sample. Higher samples → closer to true LRU, more CPU per eviction.
That approximation is why a "I just used that key" story can still lose it under pressure — unlikely if it is truly hot (it keeps getting resampled as young), possible if it is warm and unlucky. The playground uses the same idea with an item-count ceiling.
LRU (recency): a weekly report that was scanned an hour ago looks as hot as a session hit every second, until time passes. A one-off bulk read can protect a huge blob and evict a user key.
LFU (frequency): better when some keys are hit steadily and others are scanned once. Counter saturates and decays; you are not expected to derive the decay in an interview, only to say "recurring hot keys beat a recent scan."
volatile-ttl: among keys that will expire, drop the one closest to death. Good when TTLs are honest; bad when a 24h key is cold and a 30s key is hot — you might evict the hot one because it is "about to die," then immediately miss-reload it.
Operations
CONFIG SET maxmemory 2gb
CONFIG SET maxmemory-policy allkeys-lru
INFO memory
INFO statsWhat you actually page on:
| Signal | Reading |
|---|---|
used_memory / maxmemory | How close to the ceiling |
evicted_keys (rate, not total) | Pressure. A spike with a hit-rate drop is a storm |
keyspace_hits / misses | Eviction turning into DB load |
mem_fragmentation_ratio | RSS vs used. High → allocator holes; very low → swapping |
expired_keys vs evicted_keys | TTL doing its job vs pressure |
CONFIG SET is the live lever; persist it in redis.conf / the operator's CRD or you will lose the policy on restart.
Fragmentation: Redis uses jemalloc. After a huge delete wave, used_memory falls and RSS may not. That is not "eviction failed." Restart or a compacting story is an ops choice; do not "fix" it by switching to noeviction.
Replication and persistence eat memory too (client-output-buffer, AOF rewrite). maxmemory does not mean "the rest of the box is free."
Deep dive · Eviction storms
The working set is 1.4× maxmemory. Every fill evicts something still needed. The next request misses, reloads, fills, evicts a neighbor. CPU sits in eviction + Lua + SQL. Latency explodes while evicted_keys and misses climb together. Fixes: bigger instance, smaller values (IDs not blobs), shorter TTL on the cold class, split hot keys, stop writing keys you never read. Changing LRU to random does not add RAM.
Tie-back to the cluster
- Cache-aside assumes a miss can rebuild. Evicting a hot key is just a miss — unless you stampede. Keep single-flight.
- Write-behind that treated Redis as the only copy of an unflushed write: eviction is a lost write. Do not do that.
- TTL jitter still applies. Eviction is not a substitute for expiry cliffs.
- Locks (
SET NX PX) have a TTL; they are volatile. Avolatile-*policy may evict a lock under pressure — which looks like a released mutex. Keep lock keys small and lock TTL short; do not store 2 MB values next to them on a tinymaxmemory.
Data structures — pick from the access pattern
Redis is an in-memory data structure server, not only a string blob store. Wrong type is a hot-key tax.
- String — blobs, counters (
INCR), bitmaps / HyperLogLog overlays. - Hash — object fields (
HSET user:1 name Ajay); partial reads. - List — queues / stacks (
LPUSH/BRPOP). - Set — unique membership, intersections.
- Sorted set — leaderboards, time indexes (
ZADDscore). - Stream — durable-ish log for consumers / CDC-like patterns.
Flow
- 1
Cache blob
- nextString
- 2
String
- 3
Object fields
- nextHash
- 4
Hash
- 5
Job queue
- nextList
- 6
List
- 7
Unique tags
- nextSet
- 8
Set
- 9
Leaderboard
- nextZSet
- 10
ZSet
- 11
Event log
- nextStream
- 12
Stream
Hash vs JSON string: hash allows partial field updates/reads; JSON string is simpler but rewrite-whole — races with invalidation. Streams vs lists for jobs: streams give consumer groups / ACK; lists are simpler blind queues.
Do not cache huge blobs or unbounded lists. Paginate, or cache IDs and hydrate. Hot keys: local near-cache, hashing/splitting, read replicas carefully (lag).
Persistence: RDB vs AOF vs none
| Knob | Option A | Option B | Pick when |
|---|---|---|---|
| Persistence | RDB | AOF (or both) | RDB = faster restart/snapshots; AOF = less data loss |
| AOF fsync | everysec | always | everysec common; always = latency hit |
| Durability need | Cache OK to empty | Session / queue | Cache can disable persistence; queues often need AOF/replicas |
RDB: fork + point-in-time snapshot. Compact; faster restores; can lose minutes of data since last save.
AOF: append every write (or batched). appendfsync everysec is the usual compromise — caps loss at about 1s without syncing every write. Rewrite compacts AOF over time.
Both: common production — AOF for durability, RDB for backups.
None: pure disposable cache with a known cold-start plan and stampede defenses. RDB/AOF would only slow failover for data you can rebuild.
Treat Redis as a durable DB with no AOF/replicas → node death loses "acknowledged" writes (especially write-behind buffers). Write-behind on RDB-only Redis can lose the recent buffer — prefer AOF or an external WAL.
Sequence
- 1
Client
Step1 write
- 2
Client → Redis
SET k v
- 3
Redis
Step2 buffer then fsync ~1s
- 4
Redis → Disk
AOF append
- 5
Redis
Step3 periodic snapshot
- 6
Redis → Disk
RDB dump
Async replicas for read scaling / HA (Sentinel or managed ElastiCache). Replication lag ⇒ stale reads from replicas — same class of problem as near-cache staleness. Do not read replicas for a compare-and-set that assumed the primary.
Working sketches
# Sketch against a real client — not runnable here.
def memory_report(r) -> dict:
info = r.info("memory")
stats = r.info("stats")
conf = r.config_get("maxmemory*")
return {
"used_memory_human": info.get("used_memory_human"),
"maxmemory_policy": conf.get("maxmemory-policy"),
"evicted_keys": stats.get("evicted_keys"),
}// INFO memory + INFO stats + CONFIG GET maxmemory-policy
// Alert on evicted_keys rate, not on the counter's lifetime totalTiny sampled-LRU (run this)
Item count stands in for maxmemory. Contrast noeviction (write error), allkeys-lru (cold keys leave, even durable ones), volatile-lru (no-TTL session survives).
Press Run. Snippets must be self-contained — no network, files, or native modules.
noeviction keeps a,b,c and errors on d. volatile-lru keeps session:1 because it has no TTL. allkeys-lru may keep the session because we touched it — recency — but it could have evicted it; that is the mixed-workload argument.
Mini structures (run this)
Strings with a tiny LRU ceiling, plus hash and zset. Educational only — not Redis protocol.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
allkeys-lru vs volatile-lru?
Answer
allkeys may evict any key. volatile only considers keys with an expire set; if no key has a TTL, volatile-* behaves like noeviction and writes fail. The Redis default is noeviction. Use allkeys on a pure cache. Use volatile when the same process holds durable keys you must not drop — and put TTL on every disposable key, or eviction will find nothing and writes fail.
When do you pick LFU over LRU?
Answer
When recency is a bad proxy for importance: a nightly scan of old reports should not protect those keys over a user session that is hit steadily. LFU tracks frequency (with decay). LRU is the usual default and is easier to reason about.
Is noeviction OK for a cache?
Answer
Usually wrong. When full, SET/cache fills error. Clients fail closed or stampede the DB. Prefer eviction plus sizing. noeviction belongs on Redis-as-a-store where dropping a key is data loss.
Eviction vs TTL?
Answer
TTL is proactive and is your staleness bound. Eviction is what happens under memory pressure. You want both: TTL for correctness windows, eviction so a burst of keys cannot take the host down. Eviction is not a substitute for sizing or for jitter.
Will eviction save me from a hot key?
Answer
Hot keys stay resident under LRU/LFU; cold neighbors leave. You still need to split extreme hot keys for CPU and bandwidth. Memory policy does not fix one key taking a core.
What do you monitor?
Answer
used_memory vs maxmemory, evicted_keys rate, hit/miss, expired_keys, mem_fragmentation_ratio. Alert when eviction rate jumps and hit rate drops together — that is a working-set problem, not 'Redis doing LRU.'
Why is Redis LRU approximate?
Answer
A perfect LRU list per key costs RAM. Redis samples a handful of keys and evicts the coldest in the sample. maxmemory-samples trades CPU for accuracy. Interview answer: sampled LRU, not a linked-list LRU.
What if volatile-lru has no keys with TTL?
Answer
Nothing to evict. The write fails like noeviction. A cache that 'forgot EX' plus volatile-* is a trap. Either give cache keys TTL or use allkeys-*.
Should maxmemory be 100% of the VM?
Answer
No. Leave room for the OS, buffers, AOF rewrite, fragmentation. Host OOM is worse than a few evictions. Set maxmemory explicitly; do not wait for the kernel.
Can I use eviction instead of picking TTLs?
Answer
No. Eviction order is recency/frequency/random, not 'this price must not be stale more than 30s.' TTL is a product bound. Eviction is a safety valve. See the jitter lesson for cliffs; this lesson for the ceiling.
RDB vs AOF?
Answer
RDB = snapshots, faster restore, more potential loss. AOF = write log, less loss, rewrite cost. appendfsync everysec caps loss at about 1s. Production often runs both. Pure cache: persistence off, plus a cold-start + stampede plan.
Hash vs JSON string?
Answer
Hash allows partial field updates/reads. JSON string is simpler but you rewrite the whole blob — races with invalidation.
Streams vs lists for jobs?
Answer
Streams give consumer groups / ACK. Lists are simpler blind queues (BRPOP). Pick from delivery semantics, not fashion.
Pitfalls
Run the playground. Then: (1) stop touching session:1 before the fourth SET under allkeys-lru and see if the session is the victim; (2) under volatile-lru, SET four keys all without TTL and catch the OOM — nothing volatile to evict; (3) sketch the INFO fields you would graph during a load test that overfills the cache.
On a design doc: pick policy + maxmemory for (a) a 4 GB dedicated cache, (b) a Redis that also stores pub/sub backlog, (c) a primary-like Redis for jobs. If (a) and (c) get the same answer, redo (c).
Then for (1) product catalog cache, (2) session store, (3) write-behind flush buffer: name type, persistence, eviction policy, and TTL. Say what a node reboot does to each.
Go Deeper
Official docs
- Redis eviction policies
- Redis memory optimization
- Redis CONFIG SET
- Redis data types
- Redis persistence
In this cluster