Caching
Part 3 of 8 · Redis cacheCache TTL Design & Jitter
TTL is a staleness bound. Identical TTLs create expiry cliffs. Add jitter; use soft TTL / XFetch on hot keys; size TTL from business tolerance, not habit.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
The cache-aside hub told you to put a TTL on every fill. This lesson is how to pick that number, why a constant is a footgun, and what to do when one key is too hot for expiry.
By the end of this lesson you should be able to:
- Say out loud: TTL is a staleness bound, not a freshness guarantee
- Draw an expiry histogram with identical TTLs vs jittered TTLs
- Pick base TTL from "how wrong can this be, for how long" plus reload cost
- Explain jitter vs single-flight vs XFetch / stale-while-revalidate — they are layers, not rivals
- Keep negative-cache TTLs much shorter than positive ones
How hot keys should expire
Prefer
Jittered TTL, sized from business staleness
base + U(0, J) so a batch-warmed slab does not vanish together. Ultra-hot keys also get XFetch or a soft TTL, plus a mutex on true miss.
- Cheap: one extra random int on SET. No new Redis feature.
- Cliff width becomes J seconds instead of one second.
- Invalidation-on-write stays the primary freshness mechanism.
Alternative
Everyone gets EX 300 because that is the house default
A deploy, a cache flush, or a cron that warms 50k product cards will expire them in lockstep five minutes later.
- The DB QPS chart grows a comb: a spike every TTL.
- Jitter wider than the base makes TTL meaningless.
- A viral key still stampedes — jitter spreads many keys, not one key.
Fill, then expire without a cliff
Phone-friendly vertical path. The mermaid below is the same story.
- 1
Pick a base from staleness tolerance
Config can be hours. A feed card might be 30 seconds. Negatives are shorter still. - 2
SET with EX base plus jitter
ttl = base + rand(0, J). J is a slice of the base, not 10x the base. - 3
Many keys now expire at different seconds
The histogram spreads. DB QPS stops looking like a square wave. - 4
Hot key still needs more
Jitter does not split one key. Single-flight the miss; optionally XFetch before expiry. - 5
On write: DEL, do not wait for TTL
TTL is the backstop when invalidate is lost — not the happy path.
TTL is a staleness bound
Redis EXPIRE / SET … EX makes the key vanish. The next GET is a miss. Until that moment, you might be serving a value that the primary has already replaced — if some writer forgot to DEL.
That is the contract:
- Invalidate-on-write (or write-through SET) is how you try to be fresh.
- TTL is how long you are willing to be wrong when that path fails.
- No TTL is not
EX 0. SET with EX 0 is rejected (invalid expire time) and EXPIRE key 0 deletes the key at once. A key with no TTL at all (TTL returns -1) never expires - memory risk; use an eviction policy consciously. See eviction.
Too-short TTL: constant reloads, Redis becomes a noisy proxy, the DB stays hot, you paid for a cache and kept the original bottleneck.
Too-long TTL: a bug or a missed invalidate becomes a prolonged incident. "We shipped the price change at 10:00 and customers saw it at noon" is a TTL number you cannot defend.
Expiry cliffs
Warm a catalog at deploy: 50,000 keys, all EX 300, all SET within a few seconds. Five minutes later they expire together. Every pod misses. Even with single-flight you now have 50,000 distinct keys to reload — the mutex helps per key, not across the keyspace.
Identical TTLs are a synchronized herd. Jitter turns one spike into a ridge.
Decisions
- 1
Step 1 SET key EX base plus random 0 to J
- nextStep 2 Expiries spread over the J window
- same TTL on a batch warmFailure path - synchronized cliff, DB QPS spike
- 2
Step 2 Expiries spread over the J window
- nextStep 3 GET - hit or miss?
- ?
Step 3 GET - hit or miss?
- hitStep 4 Inside the soft-TTL or XFetch window?
- missStep 6 Single-flight load, SET with a new jittered TTL
- ?
Step 4 Inside the soft-TTL or XFetch window?
- noStep 5a Return the value
- yesStep 5b Return the value and refresh in the background
- 5
Step 5a Return the value
- 6
Step 5b Return the value and refresh in the background
- 7
Step 6 Single-flight load, SET with a new jittered TTL
- 8
Failure path - synchronized cliff, DB QPS spike
Lesson map
Cache TTL Design & Jitter
TTL is a staleness bound. Identical TTLs create expiry cliffs. Add jitter; use soft TTL / XFetch on hot keys; size TTL from business tolerance, not habit.
Architecture. Step 1 SET key EX base plus random 0 to J Ready. Step 2 Expiries spread over the J window Ready. Step 3 GET - hit or miss? Ready. Step 4 Inside the soft-TTL or XFetch window? Ready. Step 5a Return the value Ready. Step 5b Return the value and refresh in the background Ready. Step 6 Single-flight load, SET with a new jittered TTL Ready. Failure path - synchronized cliff, DB QPS spike Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB S["Step 1 SET key EX base plus random 0 to J Ready"] E["Step 2 Expiries spread over the J window Ready"] G["Step 3 GET - hit or miss? Ready"] H["Step 4 Inside the soft-TTL or XFetch window? Ready"] R["Step 5a Return the value Ready"] B["Step 5b Return the value and refresh in the background Ready"] L["Step 6 Single-flight load, SET with a new jittered TTL Ready"] F["Failure path - synchronized cliff, DB QPS spike Ready"] S -->|continues| E E -->|continues| G G -->|hit| H H -->|no| R H -->|yes| B G -->|miss| L S -->|same TTL on a batch warm| F
Jitter recipe
ttl_seconds = base + uniform(0, J)Typical J is 10–20% of base (30s on a 300s base). Wider than the base and you no longer have a bound — some keys live 2 * base, which you did not agree to with the product.
Apply jitter on every fill, including the SET after a single-flight load. Forgetting it there re-creates the cliff the mutex just survived.
Soft TTL, hard TTL, XFetch
Jitter spreads many keys. One viral key still expires at one instant. Layer more:
| Technique | What it does | Enough alone? |
|---|---|---|
| Jitter | Desynchronizes many keys' hard expiry | Medium keys, yes. Viral key, no |
| Single-flight | One loader per missed key; waiters poll | Stops N loads of the same key. Does not stop 50k distinct misses |
| Soft TTL + background refresh | Serve slightly stale while one worker rebuilds | Best UX under load; needs a stored expiry |
| XFetch | Probabilistic early refresh on hit, before hard miss | Smooths hot keys without a global lock |
Hard TTL is Redis EXPIRE. The key is gone. Next read is a miss.
Soft TTL is an application timestamp you store next to the value (or a separate logical expiry). Past soft TTL you still return the blob and trigger a refresh. Past hard TTL Redis dropped it and you must load. This is stale-while-revalidate for application caches.
XFetch (Vattani, Grosvenor, Rodriguez, VLDB 2015) — Optimal Probabilistic Cache Stampede Prevention. Refresh probabilistically before expiry; the chance grows with recompute time (delta) and a tuning factor β. Store (value, delta, logical_expiry) with a physical Redis TTL slightly longer than logical expiry. On each hit, refresh early if:
now - delta * beta * ln(random()) >= expirydelta= measured recompute costbeta≈ 1.0 (higher = earlier refresh)random()in (0, 1) solnis negative: that product is a random lead time- Hot keys are sampled more often, so they refresh earlier — no global lock
- Occasional redundant refreshes are OK
- Combine with a mutex for hard misses (value already gone)
Deep dive · Why the log of a random number?
random() is in (0, 1), so ln(random()) is negative and - delta * beta * ln(random()) is a random lead time. Far from expiry the inequality almost never fires; near expiry it fires with rising probability. A key that is hit 10k times per second will almost certainly refresh early. A key hit once a minute probably will not — and it does not need to. Pair it with a mutex so the rare true miss still single-flights.
Clock skew: prefer Redis's own TTL (PTTL, EX) over app clocks when you can. If you store an absolute expires_at, use Redis TIME (or the same clock the SET used), not each pod's wall clock. A pod 2 seconds fast will XFetch late or early as a herd.
Per-key classes
Do not run one TTL for the whole cache.
| Class | Base | Jitter | Notes |
|---|---|---|---|
| Feature flags / config | 5–60 min | small | Invalidate on change; long bound is OK |
| User profiles | 2–10 min | ~10–20% | Invalidate on write |
| Feed cards / rankings | 15–60 s | comparable slice | Stale ranking is a product decision |
| Session (write-through) | session lifetime | small | Still expire; do not live forever |
| Negative cache | seconds to ~1 min | yes | Far below the positive TTL |
Reload cost matters. A 200ms join that is acceptable at 10 QPS is not acceptable at 10k QPS when a cliff hits. Expensive keys get longer TTL or XFetch, not "make TTL 5 seconds so it feels fresh." Freshness comes from DEL on write.
What to measure
Hit ratio alone hides cliffs. Watch:
- Miss rate and miss latency over time (look for a comb at
baseseconds) - DB QPS correlated with TTL boundaries
- Lock acquire rate and lock wait timeouts (stampede the mutex did not fully hide)
- Rebuild duration
delta(feeds XFetch) - Stale-serve count (soft TTL is working)
- Evicted keys vs expired keys (eviction is not TTL; see the last lesson)
Alert when miss rate or DB QPS spikes at TTL multiples after a warm or a deploy.
Decisions
- 1
GET hit
- nextXFetch predicate?
- ?
XFetch predicate?
- noReturn value
- yesOne worker refreshes
- 3
Return value
- 4
One worker refreshes
- nextReturn value
- 5
GET miss
- nextSET NX PX single-flight
- 6
SET NX PX single-flight
- nextLoad primary
- 7
Load primary
- nextSET EX base plus jitter
- 8
SET EX base plus jitter
Working sketches
# Sketch of the helpers. Playground below runs them on 200 keys.
import random, time
def ttl_with_jitter(base_seconds: int, jitter: int = 30) -> int:
if base_seconds < 1:
raise ValueError("base_seconds must be >= 1")
return base_seconds + random.randint(0, max(0, jitter))
def should_soft_refresh(expires_at: float, now: float | None = None, window: float = 30.0) -> bool:
now = time.time() if now is None else now
return expires_at - now <= windowexport function ttlWithJitter(baseSeconds: number, jitter = 30): number {
if (baseSeconds < 1) throw new Error("baseSeconds must be >= 1");
return baseSeconds + Math.floor(Math.random() * (Math.max(0, jitter) + 1));
}
export function shouldSoftRefresh(expiresAt: number, now = Date.now() / 1000, window = 30): boolean {
return expiresAt - now <= window;
}In-memory histogram (run this)
200 keys, identical TTL vs jittered TTL. Print the expiry-second histogram. Then sample XFetch near vs far from expiry. No Redis, no network.
Press Run. Snippets must be self-contained — no network, files, or native modules.
You should see one bucket of 200 at t+300 without jitter, and ~31 buckets with a max around 10 when J = 30. XFetch fires often at now = 290 and almost never at now = 10 for expiry = 300.
Soft TTL plus try-lock (run this)
Store {value, soft_exp, hard_exp}. After soft expiry, losers serve stale while one winner reloads. A simplified early-refresh coin-flip (β / (ttl_remaining + β)) is a teaching model — production XFetch uses the log formula above.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Why add jitter to TTL?
Answer
Identical TTLs on keys warmed together expire together — a cliff of misses and a DB spike. base + U(0, J) spreads that work across J seconds. It is cheap and should be on every SET, including post-lock fills.
Is jitter enough for a viral key?
Answer
No. Jitter desynchronizes many keys. One hot key still expires at one instant. Add single-flight so only one loader runs, and optionally XFetch or soft TTL so you refresh before the hard miss.
How do you pick the base TTL?
Answer
Maximum acceptable staleness if invalidate fails, considering reload cost. Config can be long; feed cards short; negatives shorter. Do not copy "300 seconds" from another service. Invalidation-on-write remains the primary freshness mechanism.
Soft TTL vs hard TTL?
Answer
Hard TTL is Redis EXPIRE — the key is gone. Soft TTL is an app timestamp: serve the stale blob and refresh in the background. Soft TTL keeps UX stable under load; hard TTL is the memory and last-resort bound.
What about clock skew?
Answer
Prefer Redis TTL rather than each pod's wall clock. If you store absolute expiry, use Redis TIME (or the clock at SET) consistently. Skewed pods will XFetch as a herd or never XFetch.
What does a TTL of zero / no EXPIRE mean?
Answer
SET with EX 0 is rejected (invalid expire time) and EXPIRE key 0 deletes the key at once. A key with no TTL at all (TTL returns -1) never expires - memory risk; use an eviction policy consciously. Use a missing TTL only for durable keys you intend to manage.
Can jitter be wider than the base?
Answer
You can mechanically, and you should not. The bound you agreed with the product was base. J ≈ 10–20% of base is the usual slice. A jitter of 300 on a base of 30 means some keys live 11x longer than the others.
Where does XFetch sit relative to the mutex?
Answer
XFetch reduces how often you ever miss on a hot key. The mutex still protects the true miss. Use both: probabilistic early refresh on hit, SET NX PX on miss. XFetch without a mutex still stampedes when the physical TTL elapses.
Why still invalidate-on-write if TTL exists?
Answer
TTL is the backstop. Waiting for expiry after a password change or a price cut is a product bug. DEL after commit; TTL limits damage when DEL is lost.
What metric catches an expiry cliff?
Answer
Miss rate and DB QPS vs time, aligned with TTL multiples after a warm or deploy. Hit ratio can stay 'fine' as an average while a 5-second spike pages the database.
Pitfalls
Run the playground. Write down max-bucket size with J=0 vs J=30 vs J=300. Then pick TTLs for (1) a feature flag, (2) a profile, (3) a home-feed card, (4) a missing user id. Defend each with a staleness sentence and a reload-cost sentence.
Mentally load-test: 50k keys warmed at t=0 with EX 300, 200 pods, no jitter. How many distinct DB loads in one second at t=300 even with a perfect per-key mutex? (Up to 50k.) Then add J=30.
Go Deeper
Official docs and the paper
- Redis EXPIRE
- Redis SET
- Vattani et al., Optimal Probabilistic Cache Stampede Prevention (VLDB 2015)
- Preventing cache stampede with Redis and XFetch — Jim Nelson
In this cluster