Rate limiting
Studies in this cluster, in series order. Each one keeps its own URL.
System design
Capacity, trade-offs, and request paths you can defend on a whiteboard.
Rate limiting
6 studies- 1.Rate Limiting: Token Bucket, Leaky Bucket & Sliding WindowFixed-window 2× burst; sliding O(1) counter; token vs leaky bucket; Redis+Lua; 429/Retry-After.
- 2.Token Bucket vs Leaky Bucket vs Sliding WindowThree classic limiters: token bucket (burst + sustained rate), leaky bucket (smooth drain), sliding window (fairer than fixed windows). Interviews want tradeoffs, not just names.
- 3.Redis + Lua Atomic Rate LimitersDistributed limiters need atomic read-modify-write. Redis + Lua (EVAL/EVALSHA) runs check+debit in one script so concurrent replicas cannot both undercount. Prefer hash tags for Cluster slot affinity; keep scripts short.
- 4.Distributed Rate Limits Across GatewaysA limit of 100 req/s per key means nothing if each of N gateway pods enforces it on its own: the real limit becomes N x 100 and changes every time the fleet autoscales. Accurate limits need one shared counter (Redis with an atomic Lua script); fast limits need a local check that costs no network hop. Production systems layer them: the edge drops obvious abuse, a local per-pod bucket rejects clear overage for free, and a global Redis bucket enforces the real quota. When Redis is slow or down, degrade to the local share and alert, instead of either blocking everything or admitting everything.
- 5.HTTP 429, RateLimit Headers & Retry-AfterRejecting a request is half the job; the response has to tell the client what to do next. Reply 429 Too Many Requests (RFC 6585) with a Retry-After header (RFC 9110) that says when capacity will exist, and publish the budget on every response so good clients slow down before they hit the wall: the legacy X-RateLimit-Limit, Remaining and Reset triple, or the IETF RateLimit-Policy and RateLimit fields. Clients must honor Retry-After, add random jitter, and use exponential backoff when no header is present; otherwise every throttled client comes back in the same second and the 429s arrive in waves.
- 6.Fairness, Quotas & Noisy NeighborsA single global RPS limit protects the servers but not the tenants: whoever sends the most wins the budget, so one batch job can push every other customer into 429s. Fairness needs two layers. Admission control gives each tenant its own quota on each scarce dimension (rate, concurrency, burst, usage per billing period), so a 429 hits only the tenant over its quota. Scheduling then shares the workers among admitted requests with weighted fair queuing (in practice deficit round robin), with a strict priority lane for critical traffic, so a deep queue from one tenant cannot starve the rest.