Networking
Part 3 of 6 · CDN & edge cacheSWR & stale-if-error
max-age is hard freshness; stale-while-revalidate hides revalidation; stale-if-error keeps serving through origin 5xx. Together they are a latency and availability tool, not a correctness tool.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Traditional caching is binary: fresh or miss. After max-age, every request waits on origin. A popular expiry looks like a thundering herd; an origin blip looks like a site outage.
RFC 5861 adds two Cache-Control extensions that RFC 9111 caches still speak:
- stale-while-revalidate — serve the stale body immediately; one refresh runs in the background.
- stale-if-error — if that refresh (or a synchronous fetch) fails with 5xx / timeout / unreachable, keep serving stale for a longer window.
Together they are a latency and availability tool. They do not make yesterday's HTML correct. For publish-time correctness use purge and generation tokens.
Why SWR plus SIE wins
| Concern | max-age only | SWR only | SWR + SIE (winner for HTML/API) |
|---|---|---|---|
| First request after expiry | Wait for origin | Serve stale; refresh async | Same |
| Origin 5xx / timeout | User sees 502 | User sees 502 | Serve stale while the error window is open |
| Burst at TTL | N origin fetches unless you collapse | One background fetch; waiters get stale | Same, plus outage cover |
| Storage | One expiry | Hard + soft timestamps | Hard + soft + error-fallback |
| Correctness | Stale never served | Stale served on purpose | Stale served on purpose — not a purge |
- 1
popular HTML → origin stampede
Hard miss after max-age
Every edge that expires together issues a fetch. Without a shield and collapsing, a 60s TTL on a viral page is a metronome DDoS.
- 2
Winner: short hard TTL, long SWR, longer SIE
Browsers can stay honest with max-age=60. The CDN holds s-maxage plus stale-while-revalidate so users never wait on origin for that page. stale-if-error rides out origin deploys and 5xx.
- ?
When you must not use SWR
Price, inventory, authz, GDPR takedown — anything where serving yesterday is a bug. Soft-purge into a tiny SWR window, or bump a generation token. SWR hides latency; it does not publish.
Header syntax and windows
Cache-Control: public, max-age=60, s-maxage=3600, stale-while-revalidate=86400, stale-if-error=604800| Directive | Who | Meaning |
|---|---|---|
max-age | Browsers (and shared caches if s-maxage is absent) | Hard freshness in seconds |
s-maxage | Shared caches / CDNs | Overrides max-age at the edge |
stale-while-revalidate | Caches that implement SWR | After hard TTL, serve stale this long while one refresh runs |
stale-if-error | Caches that implement SIE | Serve stale on origin error for this many seconds past freshness (vendor clocks differ — read the docs) |
Age | All caches | Seconds since the origin generated the object — remaining freshness is TTL minus Age |
Hard TTL = max-age / s-maxage. Inside it, serve without talking to origin.
Soft TTL = SWR window. Inside it, serve stale and revalidate.
Error window = SIE. Inside it, origin failure does not become a user-facing 5xx.
A useful mental model for one stored object:
| Clock | State | Behavior |
|---|---|---|
now < stored + hard | Fresh | HIT, no origin |
stored + hard ≤ now < stored + hard + swr | SWR | HIT stale + single background fetch |
Revalidate returns 5xx, now still in SIE | SIE | HIT stale, do not propagate 502 |
| Past every window | Dead | Synchronous fetch; failure is a real error |
Vendors disagree on whether SIE is "hard + swr + sie" stacked or "sie from stored time." Interviewers care that you name the three clocks, then say you would read Fastly/Cloudflare stale docs for the arithmetic.
Revalidation triggers
- Time-based — hard TTL elapsed; SWR fires one refresh.
- Event-based — purge, soft purge, surrogate-key. Soft purge is "mark stale and use SWR."
- Conditional GET — send
If-None-Match/If-Modified-Sinceso a refresh can be 304. SWR still hid the RTT from the user who got stale.
SWR without collapsing is still a herd: every waiter might start a refresh. The cache must single-flight the revalidation. The playground below uses a revalidating flag for that.
Sequence
- 1
Step1 Client → Step2 Edge cache
GET /resource
- 2
Step2 Edge cache → Step1 Client
200 fresh
- 3
Step2 Edge cache → Step1 Client
200 stale
- 4
Step2 Edge cache → Step3 Origin
Step4 async revalidate once
- 5
Step3 Origin → Step2 Edge cache
200 new body
- 6
Step2 Edge cache → Step2 Edge cache
Step5 replace entry, reset stored
- 7
Step2 Edge cache → Step3 Origin
revalidate
- 8
Step3 Origin → Step2 Edge cache
503
- 9
Step2 Edge cache → Step1 Client
200 stale fallback
- 10
Step2 Edge cache → Step3 Origin
synchronous fetch
- 11
Step3 Origin → Step2 Edge cache
503
- 12
Step2 Edge cache → Step1 Client
502
Lesson map
SWR & stale-if-error
max-age is hard freshness; stale-while-revalidate hides revalidation; stale-if-error keeps serving through origin 5xx. Together they are a latency and availability tool, not a correctness tool.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["Step1 Client"] e["Step2 Edge cache"] o["Step3 Origin"] c -->|GET /resource| e e -->|200 fresh| c e -->|200 stale| c e -->|Step4 async| o o -->|200 new body| e e -->|revalidate| o
Age across tiers
The hierarchy stores the same bytes at shield then edge. Age is how you stop a shield-stale object from being re-cached as brand new at the edge.
If the shield generated the object 50 minutes ago and s-maxage is 3600, remaining freshness is 10 minutes if the edge honors Age. If a fill ignores Age, the edge treats shield-old bytes as Age: 0. That bug shows up as "we waited an hour and users still see old HTML."
SIE at the shield is especially powerful: one stale HIT there feeds many edges during an origin outage. Pair it with collapsing so failover-to-origin does not stampede.
TTL jitter. If every edge stores the object with the same s-maxage and the same Age origin, they expire together. SWR hides that metronome from users; collapsing hides it from origin. Jitter (a few percent of TTL) still helps the refresh wave. Do not rely on jitter alone.
must-revalidate and 304. A shared cache that must revalidate can still be cheap: conditional GET, 304, no body. SWR is "do that in the background." If the vendor treats s-maxage as "never serve stale," you need Surrogate-Control (Fastly) or an explicit stale setting, not a hope that RFC 5861 directives override RFC 9111 shared-cache language.
Cache-Control: public, max-age=60
Surrogate-Control: max-age=3600, stale-while-revalidate=86400, stale-if-error=604800The browser stays honest; the CDN is allowed to lie for a day of SWR and a week of SIE. Strip Surrogate-Control before the response hits the client.
Deep dive · Browsers vs CDN vs Surrogate-Control
Browsers honor max-age and, in the Cache API / some browsers, SWR. They do not honor s-maxage. HTML that must be "eventually fresh" in the tab uses a short max-age; the CDN uses a longer shared TTL. Fastly prefers Surrogate-Control for the shared-cache policy and strips it before the browser. If you only set Cache-Control: s-maxage=3600 and skip browser max-age, some clients will cache for an hour too — or none will, depending on other directives. Set both deliberately.
Playground: fake-clock state machine
No fetch. Origin is a function you flip to 503. Advance now and watch state: fresh → swr → sie → dead.
Press Run. Snippets must be self-contained — no network, files, or native modules.
This model completes SWR refresh on the triggering request so the log is easy to read. Production serves this client stale and swaps the entry when the background fetch finishes — same states, different overlap.
Interview Q&A
Difference between max-age and stale-while-revalidate?
Answer
max-age (or s-maxage at the CDN) is the hard freshness period. After it, the object is stale. stale-while-revalidate allows serving that stale body while one revalidation runs. Users see TTFB of a HIT; origin sees a trickle, not a herd, if you also collapse.
When do you use stale-if-error?
Answer
When origin 5xx, timeouts, or shield failover would otherwise 502 the user, and serving the last-known-good body is acceptable. Dashboards, marketing HTML, product descriptions: yes. Account balances, authz decisions, legally taken-down pages: no. Size the window to cover deploys and the longest origin incident you will not page freshness for.
How does SWR prevent a thundering herd?
Answer
It does not, by itself. SWR means waiters can take stale instead of joining a synchronous miss. You still need a single in-flight refresh per key (collapsing) and a shield so 200 POPs do not each refresh. Interview answer: SWR + single-flight + shield.
Soft TTL vs hard TTL?
Answer
Hard TTL is max-age / s-maxage: the object is fresh. Soft TTL is the SWR window: the object is stale but servable while a refresh runs. SIE is a third clock for errors. Quote all three. Do not call SWR "a longer max-age" — caches and browsers treat must-revalidate / s-maxage differently.
Is SWR a correctness tool?
Answer
No. It is a latency tool. Clients can observe stale bytes for the whole SWR window. Publish and takedown need purge, soft purge, or a generation token that changes the key. If you cannot tolerate a 120s lie, do not set SWR to 120.
What does Age do with shielding?
Answer
Age is seconds since origin (or the generating cache) produced the response. Remaining freshness is TTL minus Age. If the edge fills from a shield and zeroes Age, shield-stale content becomes edge-fresh. Honor Age on every fill; it is how multi-tier TTL stays honest.
Who honors s-maxage vs max-age?
Answer
s-maxage is for shared caches (CDN). max-age is the browser knob (and the shared default if s-maxage is missing). Typical HTML: short max-age, longer s-maxage, long SWR/SIE on the CDN. Hashed assets: long both plus immutable.
What happens when SWR refresh returns 304?
Answer
The body did not change. Reset freshness (stored / Age accounting) so the object is fresh again without downloading bytes. Still collapse the conditional GET. Users who arrived during the refresh already got stale — that is the point.
Should you SWR 5xx or 404?
Answer
Do not long-cache 5xx. Negative-cache 404 briefly so scanners do not melt origin, but a newly published URL must not be stuck missing. SIE applies to successful cached objects when a later origin fetch errors — it is not "cache the 503."
How do you design origin-outage behavior with SWR/SIE?
Answer
Long stale-if-error at edge and shield, last-known-good, do not disable collapsing, accept that the freshness SLO slips while availability holds. Say that tradeoff out loud. Pair with TLS / Anycast failover so a bad POP withdraws instead of serving errors with an empty cache.
Event-based revalidation vs the SWR clock?
Answer
A soft purge marks the object stale now, even if max-age has not elapsed. SWR then applies: serve stale, one refresh. Hard purge deletes the object, so there is nothing to SWR until origin fills again. Do not wait for the hard TTL when a CMS publish must be visible inside the freshness SLO — that is what purge is for.
Pitfalls
| Pitfall | What actually happens |
|---|---|
| Treating SWR as purge | Users keep stale through a legal takedown or price change |
| SWR without collapsing | Every waiter starts a refresh; origin still melts |
Ignoring s-maxage must-revalidate | You thought the CDN would SWR; RFC 9111 shared-cache semantics differ — verify vendor |
Zeroing Age on shield fill | Stale object reborn as fresh at the edge |
| SIE window shorter than a deploy | Origin 5xx during rollout, cache already dead, users see 502 |
| Immutable unversioned HTML | SWR never saves you; the URL must change or you must purge |
| Caching 5xx with a long TTL | You "helped" origin by locking the site into error |
On paper: max-age=60, SWR=120, SIE=300, stored at t=0. Mark t=30, t=70, t=200 with origin down, t=500 with origin down. Write what the user body is and whether origin is called. Then change the story to a CMS publish at t=10 that must be visible by t=20 — and show why SWR alone fails (you need purge / generation).
Go Deeper
- RFC 5861 — stale-while-revalidate / stale-if-error
- RFC 9111 — HTTP Caching
- Fastly — Serving stale content
- MDN — Cache-Control
Sibling pages: CDN hierarchy · Cache keys and Vary · Purge and generation tokens · TLS / Anycast · Request collapsing