Observability
Part 4 of 6 · Resilience PatternsLoad Shedding & Admission Control — Drop Early, Protect the Core
Load shedding deliberately refuses or degrades work when you are the bottleneck—before queues explode and every request times out. Admission control at the edge/gateway drops low-priority or expensive work first, returns 503 + Retry-After, and protects critical paths (checkout over recommendations). Unlike rate limiting (fairness/quotas), shedding is survival under overload.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Traffic 10× — what do you drop first?
Prefer
Priority admission: shed recs, keep checkout
Reserve slots for critical routes. Non-critical hits a softer queue limit and gets 503 + Retry-After. Degrade to cached catalog when product allows.
- Autoscale is slow; shedding is immediate survival.
- Early 503 frees workers for requests that can still succeed.
- shed_total sits on the same dashboard as SLO burn.
Alternative
Scale only, or drop randomly, or treat it like a quota
Scale still needed for sustained load — it does not buy the next 30 seconds. Random shedding may drop checkout. Token-bucket rate limits are fairness under normal load, not overload survival.
- Breaker protects you from a remote sick dep; shed protects others from you.
- No Retry-After → client hammer loops. See [Retry Storms](/studies/retry-storms-backoff-jitter).
- Shedding without metrics is flying blind.
Admit critical, shed best-effort
Gateway measures queue and priority before the service spends a worker.
- 1
Measure
Queue depth, criticality, maybe p99 wait vs budget. - 2
POST /recs
Low priority + queue high → shed. - 3
503 Retry-After
Hint clients to back off harder as you saturate. - 4
POST /checkout
Critical class uses reserved slots. - 5
Forward with deadline
Remaining budget rides with the request. Depth: timeouts lesson.
Overview
Autoscale is slow; GC pauses and hot partitions appear suddenly. Senior interviews ask: "Traffic 10×—what do you drop first?"
Production: good shedding keeps golden signals green for money paths while secondary features return errors or stale cache. Point to rate-limit lessons for token buckets — here we focus on overload survival.
Shedding strategies
| Strategy | Signal | Drop preference | Notes |
|---|---|---|---|
| Queue-length | Depth greater than N | Newest or lowest priority | Simple; reactive |
| Latency-based | p99 or queue wait greater than budget | Expensive handlers first | Ties to SLOs |
| Priority classes | Request header / route | Recommendations before checkout | Product must define priority |
| Cost-based | CPU/IO estimate | Heavy reports / fan-out | Needs cost model |
| Coarse edge | Gateway concurrency | Whole routes | Protects origins |
Sequence
- 1
Gateway
Step1 Measure queue and priority
- 2
Client → Gateway
POST /recs
- 3
Gateway → Gateway
Step2 Admit? high queue, low pri
- 4
Gateway → Client
Step3 503 Retry-After 2
- 5
Client → Gateway
POST /checkout
- 6
Gateway → Gateway
Step4 Admit critical path
- 7
Gateway → Service
Step5 Forward with deadline
- 8
Service → Gateway
200
- 9
Gateway → Client
200
Lesson map
Load Shedding & Admission Control — Drop Early, Protect the Core
Load shedding deliberately refuses or degrades work when you are the bottleneck—before queues explode and every request times out. Admission control at the edge/gateway drops low-priority or expensive work first, returns 503 + Retry-After, and protects critical paths (checkout over recommendations). Unlike rate limiting (fairness/quotas), shedding is survival under overload.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["Client"] g["Gateway"] s["Service"] c -->|POST /recs| g g -->|Step3 503| c c -->|POST /checkout| g g -->|Step5 Forward| s s -->|200| g g -->|200| c
Keep participant names and notes short so sequence lifelines stay inside the diagram card.
Shed vs rate limit
| Rate limiting | Load shedding | |
|---|---|---|
| Goal | Fairness, abuse, quotas | Survive overload; protect SLO |
| Typical signal | Tokens / fixed window | Queue depth, latency, CPU |
| Healthy steady state | Often always on | Mostly idle until stress |
| Response | 429 often | 503 + Retry-After common |
Do not re-teach token buckets here — link rate limiting, token / leaky / sliding window, and fairness / quotas.
Protect the core
- Classify routes: critical (auth, checkout, pay) vs best-effort (recs, personalization, analytics ingest).
- Shed best-effort first; never starve critical class until last resort.
- Prefer degrade (cached catalog) over hard 503 when product allows.
- Emit
shed_totalby route and reason, and dashboard next to SLO burn.
Sandbox: priority-aware admission (Python)
Reserve slots for critical work. Non-critical sheds first.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same admission idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comparative teaching
- Vs scaling out only: Shed buys time; scale still needed for sustained load.
- Vs breaker: Breaker protects you from a remote sick dep; shed protects others from you when you are sick/overloaded.
- Vs dropping randomly: Priority shedding preserves revenue/UX; random shedding is fairer but may drop checkout.
Pitfalls
10× traffic, checkout + recs + analytics ingest + a heavy report endpoint. Who do you shed first, what status and header, and which metric sits next to SLO burn? Where does the gateway stop vs the service?
Interview Q&A
What status code for shed?
Answer
Often 503 with Retry-After; 429 if the policy is quota-like. Be consistent and metric it. Quota-like 429 details live on HTTP 429 headers.
Where should admission live?
Answer
Edge/gateway for coarse protection; service entry for fine priority; both in layered defense.
Why drop early rather than queue?
Answer
Queued work often finishes after the client's deadline — wasted capacity. Early 503 frees workers for requests that can still succeed. Pair with timeouts / budgets.
How does this relate to SLOs?
Answer
Shed to keep critical SLO burn rates acceptable; track shed rate as a user-visible reliability signal too. Depth: SLIs / SLOs.
Anti-pattern?
Answer
Shedding without metrics; shedding critical and non-critical equally; no Retry-After → client hammer loops.
Shed vs breaker?
Answer
Breaker: dependency is sick. Shed: you are overloaded. Mixing them up means you either hammer a dark dep or drop traffic while the dep is fine.
Why not only autoscale?
Answer
Autoscale has lag. Shedding is immediate survival so checkout stays within SLO while capacity catches up. Scale is still required for sustained load.
Random shed vs priority shed?
Answer
Random is fairer across routes; priority preserves revenue/UX. Interviews want checkout protected over recommendations.
Degrade vs hard 503?
Answer
Prefer cached/partial responses when product allows. Hard shed when there is nothing safe to serve or even degrade is too expensive.
What belongs in shed_total labels?
Answer
Route and reason (queue, latency, CPU, priority). High-cardinality user ids explode metrics — cardinality lives on the triad / Prometheus labels lesson; do not recap it here.
Go Deeper
- Google SRE book — Handling Overload
- AWS Well-Architected — Reliability pillar
- Envoy overload manager / circuit breakers (edge admission)
- Microsoft Learn — Throttling pattern (related; distinct from pure shed)
- Next: Timeouts, Budgets & Deadline Propagation