Resilience Patterns
Studies in this cluster, in series order. Each one keeps its own URL.
Observability
SLIs, traces, and the dashboard you would actually page on.
Resilience Patterns
6 studies- 1.Resilience Patterns — Circuit Breakers, Bulkheads & Load SheddingWhen systems overload, naive retries amplify traffic, connection pools saturate, and one sick dependency can take the whole site down. Resilience patterns—circuit breakers, bulkheads, load shedding, and deadline budgets—give you deliberate ways to stop calling sick dependencies, isolate blast radius, drop work early, and bound wait time. This hub maps the cluster and shows when to use each pattern versus the alternatives.
- 2.Circuit Breaker Mechanics — Closed, Open, Half-Open, Thresholds & ProbesA circuit breaker wraps calls to a dependency and tracks recent failures (and often slow calls). In Closed it passes traffic; after a threshold it Opens and fail-fasts; after a reset timeout it Half-Opens with a limited probe. Tuned well, it stops retry storms against a sick service; tuned poorly, it flaps or never recovers.
- 3.Bulkheads — Thread Pools, Connection Pools, Queues & Failure DomainsBulkheads isolate failure domains so a slow or chatty dependency cannot exhaust shared threads, connections, or CPU and starve unrelated work. Separate pools, queue depth limits, and coarse Kubernetes resource limits are all bulkheads at different layers. Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
- 4.Load Shedding & Admission Control — Drop Early, Protect the CoreLoad shedding deliberately refuses or degrades work when you are the bottleneck—before queues explode and every request times out. Admission control at the edge/gateway drops low-priority or expensive work first, returns 503 + Retry-After, and protects critical paths (checkout over recommendations). Unlike rate limiting (fairness/quotas), shedding is survival under overload.
- 5.Timeouts, Budgets & Deadline Propagation — End-to-End Latency CapsTimeouts bound how long one hop may wait; deadlines/budgets bound the whole user request across nested calls. Each hop must shrink the remaining budget (gRPC deadlines, context.WithDeadline, AbortSignal). Anti-pattern: server timeout longer than the client's wait—work continues after the client is gone. Pair with breakers and shedding so you fail fast when budget is nearly exhausted.
- 6.Retry vs Break vs Shed — Decision Matrix & Anti-PatternsChoose the lever that matches the failure: retry transient, idempotent blips; break when a dependency is sick; shed when you are overloaded; degrade when a partial/cached answer beats an error. Mixing them wrong creates retry storms, stuck-open breakers, or silent UX death. This lesson is the cluster's decision matrix—not a re-teach of retry jitter or token buckets.