Observability
Part 6 of 6 · Resilience PatternsRetry vs Break vs Shed — Decision Matrix & Anti-Patterns
Choose the lever that matches the failure: retry transient, idempotent blips; break when a dependency is sick; shed when you are overloaded; degrade when a partial/cached answer beats an error. Mixing them wrong creates retry storms, stuck-open breakers, or silent UX death. This lesson is the cluster's decision matrix—not a re-teach of retry jitter or token buckets.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Errors spiked — retry or back off?
Prefer
Signal first, then one primary lever
Local queue/CPU/pool saturated → shed. Budget gone → fail fast. Dep sick → break + fallback. Only then retry if idempotent and transient.
- Protect yourself before you protect the dependency.
- Non-idempotent POST with unknown outcome: do not blind retry — use keys or break.
- Emit the decision as a metric; page on prolonged shed/break for critical classes.
Alternative
Retry everything, or mix the levers
Wrong mix: retry storms, stuck-open breakers, or silent UX death. Token-bucket rate limits are not this matrix.
- Retry storms without jitter + breaker — see [Retry Storms](/studies/retry-storms-backoff-jitter).
- Breaker that never half-opens — permanent Outage Mode.
- Shared pool + aggressive retries — one dep melts the site.
Playbook snapshot
Same order as the sandbox: self, budget, breaker, then maybe retry.
- 1
Local saturation
Shed or degrade non-critical. Scale out async — do not accept more work. - 2
Budget
Fail fast if remaining time cannot finish the hop. - 3
Breaker / dep health
Open or error rate high → break + fallback. Degrade if the feature is optional. - 4
Transient + idempotent
Retry with jitter inside budget. Keys: [Idempotency-Key](/studies/api-idempotency-keys). - 5
Emit and page
Decision metrics. Page on prolonged shed/break for critical classes.
Overview
On-call and interviews both ask: "Errors spiked—retry or back off?" The answer depends on signals: dependency error rate vs your queue depth vs remaining budget vs idempotency. Wire those signals into playbooks and code defaults.
This page is the cluster's decision matrix — not a re-teach of retry jitter or token buckets.
Decision matrix
| Situation | Primary action | Also consider | Avoid |
|---|---|---|---|
| Single 503/timeout, idempotent, budget left | Retry with jitter | Hedging if safe | Immediate stampedes |
| Dependency error rate / slow-call high | Break (open circuit) | Fallback / degrade | More retries |
| Local queue / CPU / pool saturated | Shed non-critical | Scale out async | Accepting more work |
| Feature optional | Degrade (cache, flag off) | Shed if no cache | Blocking critical path |
| Non-idempotent POST unknown outcome | Do not blind retry | Idempotency keys (other lesson) | Duplicate charges |
| Budget less than hop needs | Fail fast / shed | — | Start call anyway |
Decisions
- 1
1 Signals: health, load, budget
- next2 Local overload?
- ?
2 Local overload?
- yesShed or degrade non-critical
- no3 Budget enough?
- 3
Shed or degrade non-critical
- ?
3 Budget enough?
- noFail fast
- yes4 Dep breaker open?
- 5
Fail fast
- ?
4 Dep breaker open?
- yesBreak + fallback
- no5 Transient and idempotent?
- 7
Break + fallback
- ?
5 Transient and idempotent?
- yesRetry with jitter in budget
- noSingle attempt or degrade
- 9
Retry with jitter in budget
- 10
Single attempt or degrade
Lesson map
Retry vs Break vs Shed — Decision Matrix & Anti-Patterns
Choose the lever that matches the failure: retry transient, idempotent blips; break when a dependency is sick; shed when you are overloaded; degrade when a partial/cached answer beats an error. Mixing them wrong creates retry storms, stuck-open breakers, or silent UX death. This lesson is the cluster's decision matrix—not a re-teach of retry jitter or token buckets.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s["1 Signals: health, load, budget"] q1["2 Local overload?"] q2["3 Budget enough?"] q3["4 Dep breaker open?"] s -->|1 Signals: health, load, budget| q1 q1 -->|no| q2 q2 -->|yes| q3
Comparative pros/cons
| Lever | Pros | Cons |
|---|---|---|
| Retry | Hides blips; improves availability | Amplifies outages without jitter/breaker/budget |
| Break | Protects pools and dependency | Wrong thresholds → flapping or forever Open |
| Shed | Protects SLOs under overload | User-visible errors; needs priority design |
| Degrade | Better UX than hard fail | Stale/wrong data risk; product buy-in |
Anti-patterns
- Retry storms without jitter + breaker — synchronized retries hammer the dep (see Retry Storms).
- Breaker that never half-opens — permanent Outage Mode; missing reset or probes always fail due to bad timeouts. Depth: mechanics.
- Shed without metrics — you drop traffic blindly; cannot tune or prove SLO save. Depth: shedding.
- Timeout greater than client wait — ghost work after abandon (budgets).
- Shared pool + aggressive retries — one dep melts the site (bulkheads).
- Retry non-idempotent writes — duplicate side effects; use keys / safe retries or break instead.
Sandbox: unified policy hook (Python)
Single policy function combining cluster lessons.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same policy (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Playbook snapshot
- Check local saturation → shed/degrade.
- Check budget → fail fast if insufficient.
- Check breaker / dep health → break + fallback.
- Else if idempotent + transient → retry with jitter inside budget.
- Emit decision metrics; page on prolonged shed/break for critical classes.
Pitfalls
(1) Dep 40% errors on checkout. (2) Your p99 climbs, dep RED looks healthy. (3) Unknown outcome on a charge POST. Name the primary lever, the metric, and what you refuse to do.
Interview Q&A
Dependency 40% errors — retry or break?
Answer
Approaching break territory; limited retries with jitter maybe, but trip breaker if sustained. Prefer degrade for non-critical. Depth: breaker mechanics.
Your p99 climbs, dep looks healthy — what?
Answer
You may be the bottleneck → shed and scale; breakers will not help if the dep is fine.
How do idempotency keys change the matrix?
Answer
Safe retries for writes become possible; without keys, prefer break/degrade over retry for POST/PUT side effects. Depth: Idempotency-Key and safe retries — do not re-teach key storage here.
Name three anti-patterns.
Answer
Retry storms; breaker stuck Open; shed without metrics (also: timeout greater than client wait).
What dashboards complete the story?
Answer
Breaker state, shed rate by priority, pool saturation, deadline exceeded, retry attempts — alongside SLOs from Observability Triad / SLOs lessons.
Fail-open breaker ever OK?
Answer
Rarely — for soft UX where a stale call is better than empty, and load is known safe. Default fail-fast when Open.
Where does rate limiting sit on this matrix?
Answer
Fairness and quotas under normal load — not the overload survival lever. Point at rate limiting; do not recap token bucket.
Budget less than hop needs?
Answer
Fail fast or shed. Do not start the call, and do not retry. Depth: timeouts.
Optional feature, dep flaky?
Answer
Degrade (cache, flag off). Do not block the critical path. Shed if there is no cache.
What is the first question on-call should ask?
Answer
Are we saturated or is the dependency sick? Local load → shed. Dep health → break. Only then consider a jittered retry inside budget.
How do bulkheads change retries?
Answer
A shared pool plus aggressive retries lets one dep melt the site. Cap capacity first (bulkheads), then decide retry vs break.
What must never be re-taught on this page?
Answer
Jitter/backoff formulas, token-bucket math, Idempotency-Key internals, and OAuth. Cross-link those live lessons.