Retry Storms, Backoff & Jitter
Naive retries amplify outages. Combine exponential backoff, full jitter, Retry-After, circuit breakers, client budgets, and idempotency keys.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why full jitter wins when 50 clients wake up together
Prefer
Exponential cap plus full jitter
Each client sleeps Uniform(0, min(base * 2^attempt, max)). The herd smears across the window instead of colliding.
- AWS measured this: full jitter spreads load better than decorrelated or equal jitter for the same cap.
- Works with Retry-After: delay = max(server_hint, jittered_backoff).
- Cheap: one random number per attempt. No coordination.
- Pairs with a circuit breaker so you stop sending when the dependency is dark.
Alternative
Retry immediately, or exponential with no jitter / equal jitter
No jitter: every client waits 100ms, 200ms, 400ms and hits at the same tick. Equal jitter (half deterministic + half random) still clumps.
- A 5xx at T=0 becomes 50 requests at T=100ms, 50 at T=200ms — worse than the original spike.
- Equal jitter: delay = cap/2 + Uniform(0, cap/2) still peaks around the same center.
- Immediate retry (no sleep) is how a blip becomes a mesh-wide outage.
- Backoff without a budget or breaker still runs until MAX_ATTEMPTS even when the circuit should be open.
One call, many guards
Budget, then breaker, then send, then jittered retry. Vertical path for phones.
- 1
Acquire a budget token
Token bucket or leaky bucket of retries per minute for this process. Denied: fail fast (load shed). Do not retry a retry. - 2
Check the circuit
Closed: send. Open: fail fast until open_timeout, then half-open with a probe. Do not hammer a dark dependency. - 3
Send with Idempotency-Key
Same key on every attempt of this logical intent. Timeouts are at-least-once on the wire. - 4
On 5xx / timeout
Parse Retry-After if present. delay = max(hint, Uniform(0, min(base*2^n, cap))). Record a failure on the breaker. - 5
On 429 / 503 + Retry-After
The server owns the cadence. Sleep at least that long. Do not apply a shorter client backoff. - 6
On most 4xx
Do not retry. 422 fingerprint mismatch and 401/403 will not heal. 409 in-flight waits Retry-After then retries the same key.
Overview
When a downstream service goes dark, naive clients retry immediately. If 200 pods each had 50 in-flight calls, you now have 10_000 retries in the same millisecond the process is trying to recover. That is a retry storm (thundering herd). It:
- Amplifies a small outage into a full one (availability).
- Burns CPU, sockets, and cloud spend (cost).
- Duplicates writes unless the operation is idempotent (integrity).
- Floods logs and pages the wrong symptom (observability).
The fix is a stack, not a sleep:
- Exponential backoff (the envelope).
- Full jitter (break synchrony).
- Honor
Retry-After. - Circuit breaker (stop sending).
- Client budget (one noisy tenant cannot flood).
- Idempotency keys (retries must be safe).
- Load shedding at the edge when you are the overloaded service.
Core knobs
| Concept | Definition | Typical numbers | Interacts with |
|---|---|---|---|
| Exponential backoff | sleep = min(base * 2^attempt, max_delay) | base 50–100 ms, cap 20–32 s | Sets the ceiling jitter samples from |
| Full jitter | delay = Uniform(0, sleep) | one random() per attempt | Destroys lockstep |
| Equal jitter | delay = sleep/2 + Uniform(0, sleep/2) | same cap | Still peaks; worse than full in AWS’s sim |
| Retry-After | Header: integer seconds or HTTP-date | overrides if larger | 429 and 503 |
| Circuit breaker | Closed → Open → Half-Open | threshold 5, open 10–30 s | Stops traffic while the dep recovers |
| Client budget | Token bucket of retry attempts / window | e.g. 100 / 60 s | Caps one process |
| Idempotency key | Client UUID for the logical intent | 255 chars, stored | Makes HTTP retries effectively-once |
| Load shed | Drop when local CPU / queue too high | queue > N | Protects this process too |
Exponential envelope
Attempt 0 failed. You are choosing how long to wait before attempt 1.
cap(attempt) = min(base * 2^attempt, max_delay)
full jitter = random(0, cap(attempt))If max_delay is tiny (250 ms) and the dependency needs 8 s to restart, you still flood. If it is 5 minutes, a transient DNS blip becomes an awful p99. Derive the cap from how long recovery actually takes (health-check interval, k8s probe, downstream SLA), not from a copied constant.
Full jitter vs equal vs none
AWS Architecture’s classic post (2015, still the interview answer):
- None: all clients wait
cap, collide atnow+cap. - Equal: wait
cap/2 + U(0, cap/2)— the mass sits in the top half of the window. - Full: wait
U(0, cap)— the mass is flat across the window. Fewest collisions for the same cap. - Decorrelated jitter:
delay = U(0, 3 * prev_delay)clamped — also good; full jitter is simpler to explain.
The playground below counts collisions in time buckets. Full jitter wins.
Retry-After
MDN Retry-After: either a delay in seconds (Retry-After: 120) or an HTTP-date. On 429 and 503, delay = max(parsed_header, jittered_backoff). Ignoring a 30 s hint to retry in 100 ms is how you get banned.
Circuit breaker
Three states:
- Closed — traffic flows; failures increment a counter; success resets.
- Open — fail fast (or queue locally, usually fail). After
open_timeout, go half-open. - Half-open — allow a probe. Success closes; failure re-opens.
Without a breaker, backoff still sends MAX_ATTEMPTS from every pod. With a breaker, the herd stops.
Client budget
A token bucket of retry attempts (not original requests) per process per minute. When empty, shed: better a fast error than a self-inflicted DDoS. Tune per client; public SDKs need this more than a single internal job.
Idempotency
Retries turn one POST /charges into N. That is at-least-once. Attach the same Idempotency-Key every time. Without it, backoff is how you double-bill politely.
Decisions
- 1
Call
- nextBudget token?
- ?
Budget token?
- noShed BudgetExceeded
- yesCircuit closed or half-open?
- 3
Shed BudgetExceeded
- ?
Circuit closed or half-open?
- openFail CircuitOpen
- okSend plus Idempotency-Key
- 5
Fail CircuitOpen
- 6
Send plus Idempotency-Key
- nextResult
- ?
Result
- 2xxSuccess reset breaker
- 429 or 503max Retry-After, jitter
- 5xx timeoutfull jitter backoff
- other 4xxDo not retry
- 8
Success reset breaker
- 9
max Retry-After, jitter
- nextAttempts left and budget?
- 10
full jitter backoff
- nextAttempts left and budget?
- 11
Do not retry
- ?
Attempts left and budget?
- yesSend plus Idempotency-Key
- noGive up
- 13
Give up
Lesson map
Retry Storms, Backoff & Jitter
Naive retries amplify outages. Combine exponential backoff, full jitter, Retry-After, circuit breakers, client budgets, and idempotency keys.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Call"] b["Budget token?"] z["Shed BudgetExceeded"] c["Circuit closed or half-open?"] a -->|Call to Budget token?| b b -->|no| z b -->|yes| c
Architecture choices
| Choice | Pros | Cons | Pick when |
|---|---|---|---|
| Client exponential + jitter | No server change; works on gRPC too | Still sends traffic | Default mesh RPC |
Server Retry-After | Service owns cadence | Clients must parse it | Public APIs, 429 |
| Central retry-policy service | Uniform policy | Latency, SPOF | Huge orgs only |
| Per-downstream breaker | Isolates a dark dep | State per dest; distributed sync is hard | High QPS |
| Idempotency key on mutations | Safe at-least-once | Downstream must implement keys | Payments, orders |
| Token-bucket budget | Caps a noisy process | Needs tuning | SDKs, multi-tenant |
| Shed before retry | Protects yourself | Higher user-visible errors | Gateways |
Cap vs latency. Too low: flood during recovery. Too high: a blip becomes a 30 s click. Sweet spot ≈ downstream recovery SLA (often 10–30 s) with 3–7 attempts, never unbounded.
Do not retry POST without a key. See HTTP methods.
Collision sim: 50 clients, one down service
No network. Fifty clients fail together at t=0. First we look at one retry wave with a 2 s cap (the AWS picture): no jitter is a single spike, equal jitter occupies only the top half of the window, full jitter smears [0, cap]. Then we stack five exponential attempts. A collision is an extra retry sharing a 100 ms bucket; peak occupancy is what the recovering service feels.
Press Run. Snippets must be self-contained — no network, files, or native modules.
None should print peak=50 on the one-wave line (one bucket). Equal occupies roughly the top half of [0, 2000] so the peak stays high. Full fills more buckets and drops the peak — that is the recovered service’s chance to live.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Client policy (in memory, no HTTP)
A tiny state machine you can quote. Retry-After: 5 beats a 0.3 s jitter roll.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
What is a retry storm and why does exponential backoff alone not stop it?
Answer
Many clients fail at once and retry on the same schedule. Exponential backoff without jitter keeps them synchronized: everyone waits 100 ms, then 200 ms. You still get spikes; they are just spaced. Jitter desynchronizes. Breakers stop the spikes entirely while the dependency is down.
Why full jitter over equal jitter?
Answer
Equal jitter samples the top half of [0, cap], so the histogram still has a mode. Full jitter is uniform on [0, cap], which minimizes collisions for a given cap. That is the AWS Architecture result. Decorrelated jitter is also fine; full jitter is the one to derive on a whiteboard.
How do you parse Retry-After?
Answer
If the value is a number, treat it as delta-seconds. If it is a date, subtract now (clamp at 0). Then sleep = max(that, your jittered backoff). Never retry sooner than the server asked on 429/503.
Should you retry a 400? A 409? A 429? A 500?
Answer
400/401/403/404/422: no (fix the request). 409 in-flight on an idempotency key: wait Retry-After, same key. 429: yes, honor Retry-After. 5xx and timeouts: yes, with jitter, budget, breaker, and a key on mutations.
Where does the circuit breaker sit relative to backoff?
Answer
In front. If the circuit is open you do not compute a backoff and send anyway. Open means fail fast; half-open means one probe. Backoff is for Closed (and the probe).
What is a client retry budget?
Answer
A cap on how many retries this process may perform in a window (token bucket). Original traffic can still flow until other limits hit. The budget exists so a retry loop cannot outrun the original QPS. When empty, shed.
Why must retries carry an Idempotency-Key?
Answer
Because a timeout does not mean the server did not commit. Retry is at-least-once. The key plus server claim makes it effectively-once. Backoff without a key is a polite double charge. See idempotency keys.
How do you pick base, cap, and max attempts?
Answer
Base ≈ one RTT (50–100 ms). Cap ≈ downstream recovery time (often 10–32 s). Attempts 3–7 so the last wait still sits under the cap. Unbounded retries are a bug. Tie the cap to an SLA, not to a blog snippet.
Load shedding vs retrying — who sheds?
Answer
The overloaded service sheds (429/503 + Retry-After, or drop). The client budgets and breaks so it does not become the overload. Both layers. Client-only retry is how you DDoS a friend.
Do safe GET retries need jitter too?
Answer
Yes. GET is safe, so duplicates are not money bugs, but synchronized GET storms still knock over read replicas. Jitter and breakers apply to reads. Keys are optional on GET.
Pitfalls
On paper, mark T=0 as the 503. Draw 50 arrows all landing at 100 ms, 200 ms, 400 ms (no jitter). Then draw 50 arrows uniformly in [0, 100], [0, 200], [0, 400] (full jitter). Circle the max arrows in any 50 ms bucket. That picture is the interview.