Caching
Part 5 of 8 · Redis cacheNegative Caching
Cache the fact that a key does not exist with a distinct sentinel and a short TTL. Invalidate on create. Miss storms on absent keys are as dangerous as hot hits.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Cache-aside on a missing id is an infinite miss: every request hits the primary, gets zero rows, returns 404. Bots, typos, deleted accounts, and scrapers love this shape. Negative caching stores the absence so the next nine lookups never reach SQL.
By the end of this lesson you should be able to:
- Distinguish Redis miss (no key) from cached absence (sentinel)
- Pick a short negative TTL with jitter, far below the positive TTL
DELthe negative key in the create path (and undelete / restore)- Refuse to cache transient DB errors
- Apply single-flight to the first miss of an absent key
Ten lookups of a user that does not exist
Prefer
Typed sentinel, short TTL, DEL on create
First lookup hits the DB, SETs a miss marker, returns None. The next nine are cache hits. When the user is created, DEL so the next read can see them.
- Distinct marker: never {} or empty string if those are real payloads.
- TTL in seconds to a minute, plus jitter — not the profile TTL.
- Same code path and similar timing for cached miss vs cached user.
Alternative
Do not cache misses / cache the 500 / keep the sentinel forever
Uncached absence is a DB DoS. Cached 5xx turns an outage into a long false-empty window. A long negative TTL hides a user who just signed up.
- Scanners iterate ids and hold a connection per miss.
- A colliding sentinel makes a real user look deleted.
- Create without DEL looks like 'signup succeeded, GET 404'.
Miss, remember the miss, forget it on create
Vertical path. Create without step 5 is the classic bug.
- 1
GET cache:user:999
Redis miss is not the same as 'user does not exist'. It only means we do not know yet. - 2
Load the primary
Zero rows is a definite absence. A timeout is not. Single-flight this load. - 3
SET a typed sentinel with short jittered TTL
Something that cannot collide with a serialized user. Seconds to ~1 minute. - 4
Later GETs return None without SQL
Ten lookups, one db hit. This is the whole point. - 5
On create: write DB, then DEL the cache key
Drops the negative (or a stale positive). Next read fills the real row.
Absence is a value
A Redis GET that returns nil means unknown. Your options:
- Load the DB every time — correct, and expensive when the id is hot-missing
- Cache the row when found, do nothing when not — the miss storm
- Cache a negative — remember that you looked and there was nothing
Option 3 is how HTTP CDNs cache 404s for a short Cache-Control. Application caches need the same idea with an explicit marker, because Redis has no status code — only bytes.
Decisions
- 1
Step 1 GET user:999
- nextStep 2 Cached?
- ?
Step 2 Cached?
- sentinel __MISSING__Step 3a Return not found, no DB hit
- valueStep 3b Return the user
- missStep 4 Single-flight load from the DB
- 3
Step 3a Return not found, no DB hit
- 4
Step 3b Return the user
- 5
Step 4 Single-flight load from the DB
- nextStep 5 Row exists?
- DB timeout or 5xxFailure path - do not cache errors as absent
- ?
Step 5 Row exists?
- yesStep 6a SET value EX 300 plus jitter
- 0 rowsStep 6b SET sentinel EX 60 plus jitter
- 7
Step 6a SET value EX 300 plus jitter
- 8
Step 6b SET sentinel EX 60 plus jitter
- 9
Step 7 User 999 created - DEL user:999
- DEL skippedFailure path - 404 for up to the negative TTL
- 10
Failure path - 404 for up to the negative TTL
- 11
Failure path - do not cache errors as absent
Lesson map
Negative Caching
Cache the fact that a key does not exist with a distinct sentinel and a short TTL. Invalidate on create. Miss storms on absent keys are as dangerous as hot hits.
Architecture. Step 1 GET user:999 Ready. Step 2 Cached? Ready. Step 3a Return not found, no DB hit Ready. Step 3b Return the user Ready. Step 4 Single-flight load from the DB Ready. Step 5 Row exists? Ready. Step 6a SET value EX 300 plus jitter Ready. Step 6b SET sentinel EX 60 plus jitter Ready. Step 7 User 999 created - DEL user:999 Ready. Failure path - 404 for up to the negative TTL Ready. Failure path - do not cache errors as absent Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB G["Step 1 GET user:999 Ready"] M["Step 2 Cached? Ready"] N["Step 3a Return not found, no DB hit Ready"] V["Step 3b Return the user Ready"] D["Step 4 Single-flight load from the DB Ready"] Z["Step 5 Row exists? Ready"] P["Step 6a SET value EX 300 plus jitter Ready"] S["Step 6b SET sentinel EX 60 plus jitter Ready"] C["Step 7 User 999 created - DEL user:999 Ready"] F["Failure path - 404 for up to the negative TTL Ready"] T["Failure path - do not cache errors as absent Ready"] G -->|continues| M M -->|sentinel __MISSING__| N M -->|value| V M -->|miss| D D -->|continues| Z Z -->|yes| P Z -->|0 rows| S C -->|DEL skipped| F D -->|DB timeout or 5xx| T
The third branch is the one people skip. Caching a timeout as "user does not exist" turns a database blip into a 60-second outage for that id — or a signup that cannot be read back.
Sentinels that cannot collide
Pick a marker the real payload will never equal.
| Store | Pros | Watch |
|---|---|---|
Dedicated string __MISSING__ | Simple; cheap | If the value type is "JSON user", compare before json.loads |
Envelope with kind=neg vs kind=user | Explicit; can add checked_at | Every reader must understand the envelope |
Separate key neg:user:<id> | Cannot collide with cache:user:<id> | Two-key protocol; invalidate both on create |
| Bloom filter of absent ids | Tiny memory for huge id spaces | False positives: "probably missing" must still allow a DB check on the create/read-your-writes path |
Never use an empty object, empty string, JSON null, or a zero-id stub if any of those can be a real record or an ORM default. The playground forces a collision on purpose so you see the warning.
TTL: short, jittered, invalidated
Positive profile TTL might be 5 minutes. Negative TTL should be much shorter — seconds to a couple of minutes.
Why short:
- A user creates an account at that id (or reclaims a deleted handle). Until you
DEL, they look missing. - Wrong negatives are worse than wrong positives: a stale profile is old data; a stale negative is "this person does not exist."
- Memory: miss storms often cover a huge id space. Short TTL lets eviction and expiry keep the junk moving.
Still add jitter. A botnet that warmed 100k negative keys with the same EX 60 will cliff those misses into the DB a minute later — the same comb you already fixed for positive keys.
Invalidate on create is mandatory even with a short TTL:
- Insert the row,
COMMIT DEL cache:user:<id>(negative or stale positive — you do not care which)- Optionally write-through SET the new blob if the next GET must hit
Restore-from-delete, admin undelete, and "claim this username" are create paths too. If they write SQL and skip Redis, the sentinel wins until TTL.
Stampede, errors, privacy
The first 200 requests for a deleted celebrity are a stampede of empty SELECTs. Negative caching without single-flight still lets those 200 through on the first expiry. Same mutex, same waiter poll. After the sentinel lands, waiters GET the miss marker and return.
Do not cache 5xx / timeouts / connection errors. Those are not absences. A brief circuit breaker is fine; a negative SET is not. Some teams cache a definite application 404 (user deleted) and never cache "Postgres said 57P01."
Privacy / enumeration. If a cached hit returns in 2 ms and a cached miss returns in 2 ms, you did not add a timing oracle on top of the one you already had. If misses skip the cache and hits do not, you leak existence. Return the same 404 shape; rate-limit lookups by ip / account; do not give a 400 vs 404 that means "this email is registered."
HTTP analogy (MDN caching): shared caches may store 404 with a short lifetime. They also must not store 502 as if it were the document. Your Redis sentinel is that 404, not that 502.
Deep dive · Bloom filters as a first layer
A bloom filter of "ids we know are absent" (or "ids that exist," inverted) sits in front of Redis. False positive on an existence filter means you might skip Redis and hit the DB — safe, extra load. False positive on an absence filter means you might return 404 for a user who exists — unsafe unless you treat bloom as advisory and still check on authenticated / create paths. Use blooms to cheaply drop scrapers; never as the only negative cache for accounts.
Architecture notes
GET /users/999
→ GET cache:user:999
hit user → 200
hit NEG → 404
miss → SET NX lock → SELECT
row → SET json EX pos+jitter → 200
empty → SET NEG EX neg+jitter → 404
error → no SET → 503 / fail-open policy
POST /users
→ INSERT COMMIT → DEL cache:user:{id}Writers that bypass the API (imports, consoles) must still DEL. This is the same "all write paths" rule as the hub, with a worse user-visible symptom.
Metrics: negative hit rate, negative SET rate, DB QPS for empty-result queries, create-then-GET 404s (invalidate missed), sentinel decode errors (schema collision).
Working sketches
# Sketch of the protocol. Playground below uses a dict instead of Redis.
NEG = b"__MISSING__"
def get_user(r, user_id, load, pos_ttl=300, neg_ttl=60):
key = f"user:{user_id}"
cached = r.get(key)
if cached == NEG:
return None
if cached is not None:
return json.loads(cached)
row = load(user_id)
if row is None:
r.set(key, NEG, ex=neg_ttl + random.randint(0, 15))
return None
r.set(key, json.dumps(row), ex=pos_ttl + random.randint(0, 30))
return row
def on_user_created(r, user_id):
r.delete(f"user:{user_id}") # drop negative or stale positiveconst NEG = "__MISSING__";
// GET key; if NEG → null; if JSON → user; if nil → load
// empty load → SET NEG EX short+jitter
// create path → DEL keyIn-memory Redis (run this)
Ten lookups of missing 99 must produce one DB hit. Create deletes the sentinel. A colliding payload is printed as a warning.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comment out r.delete in on_create and run again. found stays None and db_hits stays 1 — signup succeeded in the DB, GET still 404. That is the invalidate bug.
Interview Q&A
Why negative cache at all?
Answer
Absent keys are often hotter than present ones: scrapers, deleted celebrities, typo ids. Without a sentinel, every lookup is a DB miss. Miss storms are as dangerous as hot hits.
How long should the negative TTL be?
Answer
Short — seconds to about a minute, plus jitter — and always shorter than the positive TTL. Invalidate on create so a new row is not hidden for the rest of the window.
Do you cache database errors?
Answer
Generally no for 5xx, timeouts, or connection failures — those are not absences. A definite application 404 (row deleted) is fair to cache briefly. Caching a timeout as NEG turns an outage into 'not found.'
What about bloom filters?
Answer
Useful as a compact first layer over a huge negative space (scraper ids). Treat false positives carefully: an absence bloom that can hide a real user is unsafe. Fall back to Redis/DB on the create and authenticated-read paths.
Security / existence leaks?
Answer
Keep timing and status codes similar for cached hit and cached miss. Rate-limit id and email lookups. Do not return 400 vs 404 in a way that enumerates accounts. Negative caching should not make enumeration easier than the uncached API.
Does stampede control apply to negatives?
Answer
Yes. The first wave after expiry is N empty SELECTs. SET NX PX on the fill; waiters poll until the sentinel (or the user) appears.
What sentinel do you store?
Answer
A typed marker that cannot equal a real payload: a dedicated string compared before JSON parse, an envelope with kind=neg, or a separate key. Never an empty object or empty string if those are valid documents.
What happens on create if you forget DEL?
Answer
The user exists in the primary and looks missing in the API until negative TTL elapses. Tests that only hit the DB will pass. Include a create-then-GET case.
Positive TTL 300, negative TTL 300 — OK?
Answer
No. Wrong negatives last as long as wrong positives, but the product cost is higher. Keep negatives much shorter; rely on DEL for correctness, TTL for the hole if DEL is lost.
Two-key design vs one key with a sentinel?
Answer
Separate neg: keys cannot collide with user JSON. One-key sentinels are simpler and one DEL covers both. Either works; colliding sentinels do not. Pick one and document it.
Pitfalls
Run the playground, then delete the on_create DEL and run again. Write the two printouts down. Add a third case: load_user raises, and get_user must not SET NEG (catch, re-raise, assert 99 not in r.kv).
On the whiteboard: bot walks ids 1..1_000_000. With no negatives, that is 1e6 SELECTs per pass. With negatives and TTL 60, it is 1e6 / 60 per second once warm — still a reason to rate-limit, but the cache is doing its job.
Go Deeper
Official docs
- Redis cache-aside overview
- MDN — HTTP caching (404s vs errors in shared caches)
- AWS ElastiCache best practices
In this cluster