System design
Part 5 of 6 · Rate limitingHTTP 429, RateLimit Headers & Retry-After
Rejecting a request is half the job; the response has to tell the client what to do next. Reply 429 Too Many Requests (RFC 6585) with a Retry-After header (RFC 9110) that says when capacity will exist, and publish the budget on every response so good clients slow down before they hit the wall: the legacy X-RateLimit-Limit, Remaining and Reset triple, or the IETF RateLimit-Policy and RateLimit fields. Clients must honor Retry-After, add random jitter, and use exponential backoff when no header is present; otherwise every throttled client comes back in the same second and the 429s arrive in waves.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Rejecting a request is half the job. The response has to tell the client what to do next. Reply 429 with a Retry-After that says when capacity will exist, and publish the budget on every response so good clients slow down before they hit the wall.
By the end you should be able to:
- Pick 429, 503, 401, or 403 for the rejection in front of you
- Send
Retry-Afteras delay-seconds, and a truthful wait - Map the legacy triple and the IETF draft fields onto the same budget
- Add jitter so a herd does not return in one slot
- Recognize a
lock_timeout429 as contention, not volume
Why it matters
Interview signal. Name the right status code for each kind of rejection, explain Retry-After and the rate limit headers, and show why jitter matters for retry storms. Bonus: some 429s are lock contention, not traffic volume.
Production signal. Wrong codes and missing headers cause real incidents. Clients treat 429 as an auth failure and refresh tokens in a loop. SDKs retry immediately and amplify an overload. A daily quota reset brings every client back at midnight.
Core concepts (deep)
What 429 means. RFC 6585 defines 429 as "the user has sent too many requests in a given amount of time". It is a client-side, per-identity signal: this caller should slow down. It is not an auth failure (401), not a permission failure (403), and not a server-wide outage (503). The response may include a Retry-After header, and a cache must not store it.
Retry-After. Defined in RFC 9110, section 10.2.3. The value is either delay-seconds (Retry-After: 30) or an HTTP-date. Prefer delay-seconds: it does not depend on client clock skew. Send a truthful value, at least the time until the bucket has a token again. A value that is too short just schedules the next 429.
Rate limit headers. Many APIs (GitHub, for example) send X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset, but the meaning of Reset varies: seconds to wait, a Unix timestamp, or milliseconds. The IETF httpapi draft (draft-ietf-httpapi-ratelimit-headers-11, May 2026, still a draft) replaces the triple with two structured fields: RateLimit-Policy: "default";q=100;w=60 (quota q per window w seconds) and RateLimit: "default";r=42;t=30 (r units left, t seconds of effective window). The draft says Retry-After takes precedence when both are present, and that remaining quota is a hint, not a guarantee. Diagram 1 still labels the legacy triple. The legacy fields and the current draft fields carry the same information.
Client behavior. On 429 or 503 with Retry-After: wait at least that long, then add random jitter. Without a header: exponential backoff with full jitter (random between 0 and base x 2^attempt, capped) and a maximum attempt count. Retry only idempotent requests, or requests carrying an idempotency key. Pace proactively: when RateLimit says r=0, wait t seconds before sending instead of sending and failing.
Retry storms and jitter. If 1000 clients are all told Retry-After: 10 and all wait exactly 10 s, they all return in the same instant and trip the limit again. Jitter spreads the return over a window. Servers can also add jitter to the advertised value. The same applies to quota resets: a daily reset at midnight UTC brings everyone back at 00:00:00.
Which code when. 429 when one identity exceeds its quota. 503 with Retry-After when the server sheds load for everyone. 401 when credentials are missing or invalid. 403 when the caller is authenticated but not allowed. Stripe adds a nuance: some of its 429s are lock_timeout errors from concurrent writes to the same object, and slowing down globally does not help. Serialize writes per object instead.
Wrong pick, both ways. Returning 401 or 403 for rate limits makes clients refresh tokens or give up instead of slowing down. Returning 500 or 503 for a per-client limit makes the client think the server is broken, triggers retries from generic HTTP libraries, and pages on-call for a client problem.
This caller is too fast
Prefer
429 plus Retry-After plus the rate limit fields
The origin is healthy. This identity is over quota. Clients wait, add jitter, and pace from the published budget.
- RFC 6585 exists so you do not overload 503.
- Prefer delay-seconds. A cache must not store the 429.
- Retry-After takes precedence over the draft RateLimit hint.
Alternative
401, 403, 500, or a bare 429
Clients refresh tokens, give up, or treat a quota deny as a server fault. A shared Retry-After with no jitter comes back as one wave.
- Generic HTTP libraries retry 500 and 503.
- Reset parsed as the wrong unit waits 0 seconds or 50 years.
- A Stripe lock_timeout is object contention, not a slower global rate.
Happy path: 429, Retry-After, jittered retry
The six steps match Diagram 1. The failure path is a client that ignores the header.
- 1
Client sends the request
The request carries an API key or token so the limiter can name the identity. - 2
Server limiter denies it
The gateway or middleware denies the call because the key has no budget left in the current window. - 3
Reply 429 with the budget
Send Retry-After plus the legacy X-RateLimit-Limit, Remaining, and Reset, or the draft RateLimit-Policy and RateLimit fields. They carry the same information. - 4
Client reads Retry-After
Accept delay-seconds or an HTTP-date. Ignore a malformed value. - 5
Wait that long plus jitter
The floor is Retry-After. Random jitter keeps throttled clients from returning in the same instant. - 6
Retry once the budget refills
The next call succeeds. A client that ignores the header and retries at once joins a synchronized retry storm.
Diagrams - step by step
Three small diagrams for 429 and rate limit headers. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: 429, Retry-After, jittered retry
Flow
- 1
Step 1 Client sends a request
- nextStep 2 Server limiter denies it
- 2
Step 2 Server limiter denies it
- nextStep 3 Reply 429 with Retry-After plus RateLimit-Limit, Remaining, Reset
- 3
Step 3 Reply 429 with Retry-After plus RateLimit-Limit, Remaining, Reset
- nextStep 4 Client reads Retry-After
- 4
Step 4 Client reads Retry-After
- nextStep 5 Wait that long plus random jitter
- no header, retry at onceFailure path - synchronized retry storm
- 5
Step 5 Wait that long plus random jitter
- nextStep 6 Retry succeeds once the budget refills
- 6
Step 6 Retry succeeds once the budget refills
- 7
Failure path - synchronized retry storm
429 means slow down. Retry-After tells the client when to come back, and the rate limit fields let well-behaved clients pace themselves before they hit the limit.
Lesson map
HTTP 429, RateLimit Headers & Retry-After
Diagram 1 walks 6 steps from Step 1 Client sends a request through Step 6 Retry succeeds once the budget refills.
Architecture. Step 1 Client sends a request Ready. Step 2 Server limiter denies it Ready. Step 3 Reply 429 with Retry-After plus RateLimit-Limit, Remaining, Reset Ready. Step 4 Client reads Retry-After Ready. Step 5 Wait that long plus random jitter Ready. Step 6 Retry succeeds once the budget refills Ready. Failure path - synchronized retry storm Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 Client sends a request Ready"] B["Step 2 Server limiter denies it Ready"] C["Step 3 Reply 429 with Retry-After plus RateLimit-Limit, Remaining, Reset Ready"] D["Step 4 Client reads Retry-After Ready"] E["Step 5 Wait that long plus random jitter Ready"] F["Step 6 Retry succeeds once the budget refills Ready"] X["Failure path - synchronized retry storm Ready"] A -->|continues| B B -->|continues| C C -->|continues| D D -->|continues| E E -->|continues| F D -->|no header, retry at once| X
Diagram 2 - Failure path: every client retries at the same instant
Sequence
- 1
1000 clients → API
Step 1 burst hits the limit
- 2
API → 1000 clients
Step 2 429 Retry-After 10 to every client
- 3
1000 clients → 1000 clients
Step 3 every client sleeps exactly 10 s
- 4
1000 clients → API
Step 4 all 1000 retry in the same second
- 5
API → 1000 clients
Step 5 limit hit again, another wave of 429s
- 6
1000 clients
Fix - clients add random jitter and exponential backoff, servers report a truthful reset
Identical waits line clients up into waves. Jitter spreads retries out. A truthful reset time keeps clients from coming back before capacity exists.
Diagram 3 - Decision: which status code to return
Decisions
- ?
Step 1 Too many requests from this identity?
- yes429 plus Retry-After
- noStep 2 Not authenticated?
- 2
429 plus Retry-After
- nextStep 4 Stripe-style lock_timeout?
- ?
Step 2 Not authenticated?
- yes401
- noStep 3 Authenticated but not allowed?
- 4
401
- Wrong pick for rate limitsClients refresh tokens instead of slowing down
- ?
Step 3 Authenticated but not allowed?
- yes403
- no503 plus Retry-After
- 6
403
- 7
503 plus Retry-After
- ?
Step 4 Stripe-style lock_timeout?
- yesSerialize writes per object
- noClient backs off with jitter
- 9
Serialize writes per object
- 10
Client backs off with jitter
- 11
Clients refresh tokens instead of slowing down
Status codes drive client behavior. 429 is not an auth failure and not a server error. Some Stripe 429s are lock contention on one object, which needs serialization rather than slower traffic. Using 401 for a rate limit makes clients refresh tokens instead of slowing down.
Working Python
Run python3 http429.py (stdlib only, about 3 seconds; starts a server on 127.0.0.1). Header names follow draft-ietf-httpapi-ratelimit-headers-11 (RateLimit-Policy, RateLimit) plus the legacy X-RateLimit-* triple. Run this file locally. It binds a port.
"""HTTP 429 end to end: server sends Retry-After + RateLimit fields, client waits with jitter.
Run: python3 http429.py (stdlib only, about 3 seconds; starts a server on 127.0.0.1)
Header names follow draft-ietf-httpapi-ratelimit-headers-11 (RateLimit-Policy, RateLimit)
plus the legacy X-RateLimit-* triple that many public APIs still send.
"""
import email.utils
import math
import random
import threading
import time
import urllib.error
import urllib.request
from collections import Counter
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
LIMIT, WINDOW = 5, 2 # 5 requests per 2 s per client, fixed window to keep the math visible
class Limiter:
def __init__(self) -> None:
self.lock = threading.Lock()
self.windows: dict = {} # client -> (window_start, count)
def check(self, client: str):
now = time.time()
with self.lock:
start, count = self.windows.get(client, (now, 0))
if now - start >= WINDOW:
start, count = now, 0
reset = max(1, math.ceil(start + WINDOW - now)) # whole seconds, like delay-seconds
if count >= LIMIT:
return False, 0, reset
self.windows[client] = (start, count + 1)
return True, LIMIT - count - 1, reset
LIMITER = Limiter()
class Handler(BaseHTTPRequestHandler):
def do_GET(self) -> None:
client = self.headers.get("X-Client-Id", "anon") # Step 1: identify the caller
ok, remaining, reset = LIMITER.check(client) # Step 2: limiter decision
self.send_response(200 if ok else 429)
self.send_header("RateLimit-Policy", f'"default";q={LIMIT};w={WINDOW}')
self.send_header("RateLimit", f'"default";r={remaining};t={reset}')
self.send_header("X-RateLimit-Limit", str(LIMIT))
self.send_header("X-RateLimit-Remaining", str(remaining))
self.send_header("X-RateLimit-Reset", str(reset))
if not ok:
self.send_header("Retry-After", str(reset)) # Step 3: truthful wait, in seconds
self.send_header("Content-Type", "application/problem+json")
self.end_headers()
self.wfile.write(b"ok" if ok else b'{"title":"Too Many Requests","status":429}')
def log_message(self, *args) -> None: # keep the demo output quiet
pass
def parse_retry_after(value, now=None):
"""Retry-After is either delay-seconds ("7") or an HTTP-date. Returns seconds or None."""
if value is None:
return None
value = value.strip()
if value.isdigit():
return float(value)
try:
when = email.utils.parsedate_to_datetime(value).timestamp()
except (TypeError, ValueError):
return None # malformed: ignore it, fall back to backoff
return max(0.0, when - (now if now is not None else time.time()))
def backoff_delay(attempt: int, retry_after, base=0.2, cap=30.0, rng=random) -> float:
"""Step 4-5: honor Retry-After as a floor, then add jitter. No header: full-jitter backoff."""
if retry_after is not None:
return min(cap, retry_after + rng.uniform(0, max(1.0, 0.5 * retry_after)))
return rng.uniform(0, min(cap, base * 2 ** attempt))
def get_with_retry(url: str, client: str, max_attempts: int = 5):
for attempt in range(max_attempts):
req = urllib.request.Request(url, headers={"X-Client-Id": client})
try:
with urllib.request.urlopen(req) as resp:
return resp.status, attempt, resp.headers["RateLimit"]
except urllib.error.HTTPError as e:
if e.code not in (429, 503): # only throttling and overload are retryable here
raise
wait = backoff_delay(attempt, parse_retry_after(e.headers.get("Retry-After")))
time.sleep(wait) # Step 6: retry after the budget refills
return 429, max_attempts, None
def herd(clients=1000, retry_after=10, jitter=True, seed=1):
"""Diagram 2: when do 1000 throttled clients come back? Busiest 100 ms slot and slot count."""
rng = random.Random(seed)
arrivals = Counter(
int(10 * (backoff_delay(0, retry_after, rng=rng) if jitter else retry_after))
for _ in range(clients)
)
return max(arrivals.values()), len(arrivals)
if __name__ == "__main__":
srv = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
threading.Thread(target=srv.serve_forever, daemon=True).start()
url = f"http://127.0.0.1:{srv.server_address[1]}/items"
print("8 calls from one client, limit 5 per 2 s:")
for i in range(8):
status, retries, rl = get_with_retry(url, "c1")
print(f" call {i + 1}: status={status} retries={retries} RateLimit={rl}")
srv.shutdown()
date = email.utils.formatdate(time.time() + 3, usegmt=True)
print(f"Retry-After as HTTP-date 3 s ahead -> wait {parse_retry_after(date):.1f} s (1 s resolution)")
for j in (False, True):
peak, slots = herd(jitter=j)
print(f"herd jitter={j!s:5s}: busiest 100 ms slot has {peak:4d} of 1000 retries, spread over {slots} slots")Sample run on the box: 8 calls with limit 5 per 2 s. Calls 1-5 returned 200 with RateLimit r=4..0. Call 6 got one 429, waited Retry-After plus jitter, then 200. Retry-After as an HTTP-date 3 s ahead gave a 2.9 s wait. Herd of 1000 clients with Retry-After 10: without jitter all 1000 land in one 100 ms slot; with jitter the busiest slot has 29, spread over 50 slots.
Working TypeScript
Run npx tsx http429.ts (Node 18+, node:http and global fetch, about 3 seconds). The client honors Retry-After plus jitter and paces when RateLimit says r=0.
/**
* 429 client handling in TypeScript: honor Retry-After, add jitter, and pace on the RateLimit field.
* Run: npx tsx http429.ts (Node 18+, uses node:http and global fetch, about 3 s)
*/
import { createServer } from "node:http";
import type { AddressInfo } from "node:net";
const LIMIT = 5, WINDOW_S = 2;
let windowStart = Date.now(), used = 0;
// Server: fixed window per process, Step 2-3 of Diagram 1.
const server = createServer((_req, res) => {
const now = Date.now();
if (now - windowStart >= WINDOW_S * 1000) { windowStart = now; used = 0; }
const reset = Math.max(1, Math.ceil((windowStart + WINDOW_S * 1000 - now) / 1000));
const ok = used < LIMIT;
if (ok) used++;
res.setHeader("RateLimit-Policy", `"default";q=${LIMIT};w=${WINDOW_S}`);
res.setHeader("RateLimit", `"default";r=${LIMIT - used};t=${reset}`);
if (!ok) res.setHeader("Retry-After", String(reset)); // seconds, never earlier than the window end
res.statusCode = ok ? 200 : 429;
res.end(ok ? "ok" : "slow down");
});
/** Retry-After: delay-seconds or HTTP-date. Returns milliseconds, or null if absent or malformed. */
export function parseRetryAfter(v: string | null, nowMs = Date.now()): number | null {
if (!v) return null;
if (/^\d+$/.test(v.trim())) return Number(v) * 1000;
const t = Date.parse(v);
return Number.isNaN(t) ? null : Math.max(0, t - nowMs);
}
/** RateLimit: "default";r=0;t=2 -> { r: 0, t: 2 }. Malformed fields must be ignored. */
export function parseRateLimit(v: string | null): { r: number; t: number } | null {
const m = v?.match(/;r=(\d+)(?:;t=(\d+))?/);
return m ? { r: Number(m[1]), t: Number(m[2] ?? 0) } : null;
}
const sleep = (ms: number) => new Promise<void>((r) => setTimeout(r, ms));
/** Step 4-6: wait Retry-After plus jitter; with no header, exponential backoff with full jitter. */
function delayMs(attempt: number, retryAfterMs: number | null): number {
if (retryAfterMs !== null) return retryAfterMs + Math.random() * Math.max(1000, retryAfterMs / 2);
return Math.random() * Math.min(30_000, 200 * 2 ** attempt);
}
let pauseUntil = 0; // proactive pacing: shared by all calls of this client
async function call(url: string, maxAttempts = 5): Promise<{ status: number; retries: number }> {
for (let attempt = 0; attempt < maxAttempts; attempt++) {
const wait = pauseUntil - Date.now();
if (wait > 0) await sleep(wait); // budget is known to be 0: do not even send
const res = await fetch(url);
await res.text();
const rl = parseRateLimit(res.headers.get("RateLimit"));
if (rl && rl.r === 0) pauseUntil = Date.now() + rl.t * 1000; // pace before hitting the wall
if (res.status !== 429 && res.status !== 503) return { status: res.status, retries: attempt };
await sleep(delayMs(attempt, parseRetryAfter(res.headers.get("Retry-After"))));
}
return { status: 429, retries: maxAttempts };
}
server.listen(0, "127.0.0.1", async () => {
const url = `http://127.0.0.1:${(server.address() as AddressInfo).port}/items`;
const t0 = Date.now();
for (let i = 1; i <= 8; i++) {
const { status, retries } = await call(url);
console.log(`call ${i}: status=${status} retries=${retries} t=${((Date.now() - t0) / 1000).toFixed(1)}s`);
}
console.log("parseRetryAfter('7') =", parseRetryAfter("7"), "ms; bad value ->", parseRetryAfter("soon"));
server.close();
});Sample run on the box: calls 1-5 at t=0.1s. The 5th returned RateLimit r=0;t=2, so the client paused and calls 6-8 succeeded at t=2.1s with zero 429s. parseRetryAfter("7") = 7000 ms. "soon" returns null. tsc --strict: OK.
Interview Q&A
When do you return 429 and when 503?
Answer
429 when one caller exceeded its own quota: the fix is on the client side. 503, ideally with Retry-After, when the service is overloaded or shedding load for everyone. Mixing them up sends the wrong signal to clients and to your own alerting.
What should a 429 response contain?
Answer
Retry-After with a truthful delay in seconds, the rate limit fields so the client can see the policy and what is left, and a short machine-readable body (problem+json is a good fit). Caches must not store it.
Why do clients need jitter if the server already sends Retry-After?
Answer
Every client throttled at the same time gets the same value. If they all wait exactly that long, they come back in one wave and hit the limit again. Random jitter spreads them over a window, which turns a spike into a steady stream the limiter can admit.
What is the difference between X-RateLimit-Reset and the IETF RateLimit field?
Answer
X-RateLimit-Reset is unstandardized: some APIs send seconds to wait, others a Unix timestamp. The IETF draft uses delay-seconds only (the t parameter), structured field syntax, and a separate RateLimit-Policy field for the quota and window. Clients should parse defensively and ignore malformed values.
A client gets 429 on a POST. Should the SDK retry it automatically?
Answer
Only if the request is idempotent or carries an idempotency key, so a retry cannot create a second charge or order. For a 429 the request was rejected before processing, but the SDK cannot always be sure, especially behind proxies, so the idempotency key is the safe rule.
Stripe returns 429 lock_timeout. Does backing off fix it?
Answer
Not by itself. That error means concurrent requests are updating the same object, so the fix is to serialize writes per object (a queue or lock per customer id). General backoff only reduces the collision rate.
How would you design an SDK retry policy?
Answer
Retry on 429 and 503 (and connection errors) only. Honor Retry-After as a floor and add jitter. Otherwise use capped exponential backoff with full jitter. Limit total attempts and total time. Pace on the RateLimit fields. Expose metrics so callers can see throttling instead of silent latency.
Pros and cons
| Approach | Pros | Cons |
|---|---|---|
| 429 + Retry-After + rate limit fields | Clear contract. Clients can pace themselves and recover without support tickets. Standard status code. | Headers leak some capacity information. Reset semantics differ between APIs. Server must compute a truthful wait. |
| Bare 429, no headers | Trivial to implement. | Clients guess the wait, usually retry too soon. Retry storms. More 429s overall. |
| Wrong status (401, 403, 500, 503) for per-client limits | None worth having. | Clients refresh tokens, give up, or treat it as a server fault. Generic retries amplify load. Wrong alerts. |
Pitfalls
From a fixed window of 5 per 2 seconds that is already full, write status, Retry-After, RateLimit-Policy, RateLimit, and the legacy triple. Then place 1000 clients that all received Retry-After 10. Say where they land with and without jitter. If the status is 401, start over.
Go Deeper
- RFC 6585, section 4: 429 Too Many Requests
- RFC 9110: Retry-After
- IETF draft: RateLimit header fields for HTTP
- AWS Architecture Blog: Exponential backoff and jitter
- MDN: 429 Too Many Requests
- GitHub REST API: rate limits and headers
- Stripe docs: rate limits and lock_timeout
Related
- Rate Limiting: Token Bucket, Leaky Bucket & Sliding Window (
rate-limiting) - Token Bucket vs Leaky Bucket vs Sliding Window (
token-leaky-sliding-window) - Redis + Lua Atomic Rate Limiters (
redis-lua-atomic-rate-limiters) - Distributed Rate Limits Across Gateways (
distributed-rate-limits-gateways) - Fairness, Quotas & Noisy Neighbors (
fairness-quotas-noisy-neighbors)
Prev / Next
- Prev: Distributed Rate Limits Across Gateways (
distributed-rate-limits-gateways) - Next: Fairness, Quotas & Noisy Neighbors (
fairness-quotas-noisy-neighbors) - Hub: Rate Limiting: Token Bucket, Leaky Bucket & Sliding Window (
rate-limiting)