Observability
Part 1 of 6 · Resilience PatternsResilience Patterns — Circuit Breakers, Bulkheads & Load Shedding
When systems overload, naive retries amplify traffic, connection pools saturate, and one sick dependency can take the whole site down. Resilience patterns—circuit breakers, bulkheads, load shedding, and deadline budgets—give you deliberate ways to stop calling sick dependencies, isolate blast radius, drop work early, and bound wait time. This hub maps the cluster and shows when to use each pattern versus the alternatives.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Payment service is timing out — what do you do?
Prefer
Interrupt the cascade at the right layer
Shed if you are the bottleneck. Break if the dependency is sick. Isolate pools so recs cannot starve checkout. Bound every hop with remaining budget.
- Admission control drops work you cannot finish (503 + Retry-After).
- A breaker fail-fasts instead of hammering a dark dependency.
- Bulkheads cap how much capacity one dep may consume.
- Retries belong elsewhere — jitter, keys, and budgets — not unbounded.
Alternative
Retry harder, share one pool, wait forever
Naive retries amplify traffic. A shared thread pool lets one slow client take the site down. Unbounded waits fill queues past the client's SLO.
- No jitter + no breaker = retry storm. Depth: Retry Storms — do not recap backoff here.
- Shared pools turn a recs blip into checkout timeouts.
- Rate limiting is not shedding: quotas vs survival.
- Hedging races parallel requests; it does not replace an open breaker.
Where each pattern interrupts the cascade
Interview order: protect yourself, then the dependency, then bound wait time.
- 1
Request arrives
Classify critical vs best-effort before you spend a worker. - 2
Admission control
If we are overloaded, shed non-critical work. Return 503 + Retry-After. - 3
Deadline budget
If remaining budget cannot finish the hop, fail fast. Do not start ghost work. - 4
Circuit for dep X
Open → fail-fast / fallback. Closed or half-open → try to acquire a bulkhead permit. - 5
Bulkhead permit
Pool full → queue with a bound, or reject locally. Acquired → call with remaining timeout. - 6
Record outcome
Success releases the permit. Fail/slow records toward a trip. Export the metrics.
Overview
Senior interviews love failure modes: "Your payment service is timing out—what do you do?" Production on-calls live this weekly.
Observability (metrics, traces, SLOs) tells you that you are overloaded. These patterns are what you do about it. The Observability Triad, SLOs, and alerting lessons cover detecting and paging. API Retry Storms and Idempotency cover safe retries. This cluster is the next layer — stopping amplification and protecting the core path.
You should be able to:
- Draw the overload cascade and mark where each pattern cuts it.
- Answer breaker vs bulkhead in one sentence each.
- Refuse to conflate load shedding with rate-limit fairness.
The overload cascade
- Dependency slows or errors.
- Clients retry without jitter → traffic multiplies. Pair retries with budgets and breakers — see Retry Storms; do not recap jitter here.
- Shared thread/connection pools fill → unrelated endpoints lag.
- Queues grow → latency exceeds SLOs → more retries.
- Cascading failure across the fleet.
Resilience patterns interrupt that cascade at different points.
Comparative overview
| Pattern | Primary question | Stops | Does NOT do |
|---|---|---|---|
| Circuit breaker | Is this dependency sick? | Calling a failing dependency | Isolate capacity between deps |
| Bulkhead | Can one dep starve others? | Blast radius via separate pools | Decide if dep is healthy |
| Load shedding | Are we the bottleneck? | Accepting work we cannot finish | Fix a remote failure |
| Timeouts / budgets | How long may we wait? | Unbounded hangs | Choose retry vs give up |
| Retry (elsewhere) | Was it transient? | Giving up too early | Prevent storms without jitter/breaker |
Vs alternatives: Rate limiting is fairness/quotas under normal load; load shedding is survival under overload — do not conflate, and do not re-teach token buckets here. Hedging races parallel requests; breakers stop all calls when open. Feature-flag degradation returns cached/partial responses — complementary to shedding.
Architecture (request path)
Single-column path from admission to outcome. Each diamond is a later lesson.
Decisions
- 1
1 Request arrives
- next2 Admission control
- ?
2 Admission control
- shed503 + Retry-After
- admit3 Deadline budget OK?
- 3
503 + Retry-After
- ?
3 Deadline budget OK?
- noFail fast: budget exhausted
- yes4 Circuit for dep X
- 5
Fail fast: budget exhausted
- ?
4 Circuit for dep X
- openFail fast / fallback
- closed or half-open5 Acquire bulkhead permit
- 7
Fail fast / fallback
- ?
5 Acquire bulkhead permit
- pool fullQueue or reject locally
- acquired6 Call with remaining timeout
- 9
Queue or reject locally
- 10
6 Call with remaining timeout
- next7 Outcome
- ?
7 Outcome
- successRelease + record success
- fail or slowRecord failure, maybe trip
- 12
Release + record success
- 13
Record failure, maybe trip
Lesson map
Resilience Patterns — Circuit Breakers, Bulkheads & Load Shedding
When systems overload, naive retries amplify traffic, connection pools saturate, and one sick dependency can take the whole site down. Resilience patterns—circuit breakers, bulkheads, load shedding, and deadline budgets—give you deliberate ways to stop calling sick dependencies, isolate blast radius, drop work early, and bound wait time. This hub maps the cluster and shows when to use each pattern versus the alternatives.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Request arrives"] b["2 Admission control"] d["3 Deadline budget OK?"] f["4 Circuit for dep X"] a -->|1 Request arrives| b b -->|admit| d d -->|yes| f
Hub map — read order
- This hub — when overload + retries amplify; pattern roles.
- Circuit Breaker Mechanics — Closed → Open → Half-Open; thresholds; probes.
- Bulkheads — thread/connection/semaphore pools; failure domains.
- Load Shedding & Admission Control — drop early; protect checkout over recommendations.
- Timeouts, Budgets & Deadline Propagation — per-hop vs remaining budget.
- Retry vs Break vs Shed — decision matrix and anti-patterns.
Sandbox: decide break vs shed vs call (Python)
Conceptual orchestration — not a library API. Protect self first, then stop hammering a sick dependency.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same decision matrix (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Payment is timing out. Draw the five-step overload cascade. Mark where you shed, where you break, where you split pools, and where remaining budget fails the hop. What do you not retry without an Idempotency-Key?
Interview Q&A
Circuit breaker vs bulkhead — what is the difference?
Why do retries make outages worse?
Answer
Without jitter, breaker, and idempotency, every timeout spawns N more requests. Amplification turns a partial outage into a full one. Pair retries with budgets and breakers — see Retry Storms. Do not recap full jitter here.
When is load shedding better than scaling?
Answer
Autoscale has lag; shedding is immediate survival. Shed expensive/low-priority work so critical paths (checkout, auth) stay within SLO while capacity catches up. Depth: load shedding.
Fail-open vs fail-closed for a breaker?
Answer
Fail-closed (open = reject) protects the dependency and your pools. Fail-open (open = still try) risks amplifying a sick dep — use only when availability of a soft feature matters more than dependency health.
How do observability and resilience connect?
Answer
Breaker state, shed count, pool saturation, and budget exhaustion must be metrics/traces. Without them you cannot tune thresholds or prove the pattern worked. Correlate with RED / exemplars and SLO burn.
Shed vs rate limit?
Answer
Rate limiting is fairness and quotas under normal load. Load shedding is survival under overload. Same 429/503 family of responses can appear; the signal differs (tokens vs queue depth / latency / CPU). Point at rate limiting — do not re-teach token bucket.
Where does hedging sit?
Answer
Hedging races parallel requests to hide tail latency. Breakers stop all calls when Open. Hedging still needs the same deadline and idempotency. It is not a substitute for an open circuit.
What does a timeout / budget stop?
Answer
Unbounded hangs. It does not choose retry vs give up — that is the decision matrix. Depth: timeouts and budgets.
What belongs on the whiteboard first?
Answer
The cascade, then which lever matches the signal: local saturation → shed; dep sick → break; wait unbounded → budget; one dep starving others → bulkhead.
Feature-flag degradation vs shedding?
Answer
Degradation returns cached or partial responses. Complementary to shedding: prefer degrade over a hard 503 when the product allows. Shedding still protects the core when there is nothing safe to serve.
Go Deeper
- resilience4j CircuitBreaker — circuit breaker and bulkhead concepts
- Microsoft Learn — Circuit Breaker pattern
- AWS Well-Architected — Reliability pillar
- Google SRE book — Handling Overload
- Envoy outlier detection — data-plane cousin of app-level breakers
- Next: Circuit Breaker Mechanics — Closed, Open, Half-Open