Observability
Part 2 of 6 · Resilience PatternsCircuit Breaker Mechanics — Closed, Open, Half-Open, Thresholds & Probes
A circuit breaker wraps calls to a dependency and tracks recent failures (and often slow calls). In Closed it passes traffic; after a threshold it Opens and fail-fasts; after a reset timeout it Half-Opens with a limited probe. Tuned well, it stops retry storms against a sick service; tuned poorly, it flaps or never recovers.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How the breaker trips — and how it recovers
Prefer
Rate window + min calls + one probe
Trip on failure (or slow-call) rate once you have enough samples. Stay Open for a reset timeout. Probe with concurrency=1. Success closes; failure re-opens.
- Handles partial outages a consecutive-N counter misses.
- min_calls avoids flapping at 1 QPS.
- One probe prevents a half-open stampede.
Alternative
Retries only, or a breaker that never probes
No breaker + retries only amplifies outages. A breaker that never half-opens is permanent Outage Mode. Fail-open under overload hammers the sick dep.
- Consecutive-only misses a steady 10% error rate that still burns capacity.
- Reset timeout misconfigured, or probes always fail from a too-tight timeout.
- Tune breakers together with [budgets](/studies/timeouts-budgets-deadline-propagation).
Closed, Open, Half-Open
Success stays Closed. The self-loop is omitted on purpose so the diagram does not stack labels.
- 1
Closed
Calls go through. Record successes and failures (and slow calls if configured). - 2
Threshold hit
Failure rate or consecutive count exceeds the trip line, with enough samples. - 3
Open
Fail-fast or fallback. No calls to the dependency until reset timeout. - 4
Half-Open
Allow a limited probe (prefer concurrency=1). - 5
Probe outcome
Success → Closed (clear the window). Failure → Open again.
Overview
Interviewers ask you to draw the state machine and pick thresholds under SLO pressure. In production, breakers prevent your fleet from DDoSing a struggling dependency — and free bulkhead slots for healthy paths.
Pair with Observability Triad metrics (error rate, latency percentiles). Avoid re-teaching retry storms here.
States
| State | Behavior | Transitions |
|---|---|---|
| Closed | Calls go through; record successes/failures | → Open when failure (or slow-call) threshold hit |
| Open | Fail fast (or fallback); no calls to dep | → Half-Open after reset timeout |
| Half-Open | Allow limited probes (often concurrency=1) | Success → Closed; failure → Open |
Flow
- 1
1 Closed: pass traffic
- fail rate or count2 Open: fail-fast
- 2
2 Open: fail-fast
- reset timeout3 Half-Open: one probe
- 3
3 Half-Open: one probe
- probe success1 Closed: pass traffic
- probe failure2 Open: fail-fast
Lesson map
Circuit Breaker Mechanics — Closed, Open, Half-Open, Thresholds & Probes
A circuit breaker wraps calls to a dependency and tracks recent failures (and often slow calls). In Closed it passes traffic; after a threshold it Opens and fail-fasts; after a reset timeout it Half-Opens with a limited probe. Tuned well, it stops retry storms against a sick service; tuned poorly, it flaps or never recovers.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["1 Closed: pass traffic"] o["2 Open: fail-fast"] h["3 Half-Open: one probe"] c -->|fail rate or| o o -->|reset timeout| h h -->|probe success| c h -->|probe failure| o
Stay-Closed on success is the default while below threshold — not a second labeled edge from Closed.
Thresholds and windows
- Consecutive failures: Simple; trips on N fails in a row. Good for hard errors; blind to intermittent 5% error rates.
- Failure rate in sliding window: e.g. 50% of last 100 calls. Better for partial outages.
- Slow-call threshold: Count calls slower than X ms as failures (dependency "alive" but saturating you).
- Minimum throughput: Do not trip on 1/1 failures at low QPS — require N samples in the window.
- Reset / wait duration: How long to stay Open before probing.
- Half-open probe concurrency: Prefer 1 (or small N) so a surge of probes does not re-kill the dep.
Pros of rate-based: Handles partial failures. Cons: Needs enough volume; more config. Pros of consecutive: Simple. Cons: Misses flaky 10% error rates that still burn capacity.
Fail-fast vs fail-open
- Fail-fast (fail-closed): When Open, reject immediately → protect pools and dependency. Default for critical deps.
- Fail-open: When Open, still attempt calls → last-resort for soft features where any data beats none. Dangerous under overload.
Libraries (concept level): resilience4j (Java), opossum (Node), Polly (.NET), older Hystrix ideas. Envoy outlier detection is the data-plane cousin — eject hosts from a cluster on consecutive 5xx/timeouts without re-teaching service mesh here.
Sandbox: sliding-window breaker (Python)
Educational failure-rate window. Probe concurrency = 1.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comparative notes
| Approach | When better | Tradeoff |
|---|---|---|
| App-level breaker | Per-RPC semantics, fallbacks in code | Every service implements / configures |
| Envoy outlier detection | Host ejection without app changes | Coarser; less business fallback |
| No breaker + retries only | Never — amplifies outages | Classic anti-pattern |
Pitfalls
Dependency SLO is 99.9% availability, typical 200 QPS, p99 80 ms. Sketch failure-rate window size, min_calls, slow-call threshold, reset timeout, and half-open concurrency. What metric proves a trip was correct?
Interview Q&A
Draw Closed → Open → Half-Open and name the transitions.
Answer
Closed→Open on threshold; Open→Half-Open after reset timeout; Half-Open→Closed on probe success; Half-Open→Open on probe failure.
Why require a minimum number of calls before tripping?
Answer
At 1 QPS, one failure is 100% — you would flap open constantly. min_calls (or volume threshold) avoids false trips.
Why limit half-open concurrency to 1?
Answer
When the reset timer fires, a traffic spike could send hundreds of probes and re-crush the recovering dependency.
Slow-call vs error-based trip?
Answer
Errors miss "stuck but 200 OK" dependencies that hold threads. Slow-call thresholds treat latency as a failure signal. Correlate with triad latency percentiles.
How do you observe breakers?
Answer
Export state gauges, trip counters, probe success/fail. Alert on prolonged Open for critical deps; correlate with dependency RED metrics and SLO burn.
Breaker never half-opens — what is wrong?
Answer
Reset timeout misconfigured, clock skew, or a code path that never transitions; or probes always fail due to a too-aggressive timeout — tune together with budgets.
Fail-fast vs fail-open?
Answer
Fail-closed (Open = reject) protects pools and the dependency — default for critical deps. Fail-open still attempts calls: last-resort for soft features, dangerous under overload.
App-level breaker vs Envoy outlier detection?
Answer
App-level: per-RPC semantics and in-process fallbacks. Outlier detection: host ejection without app changes, coarser, less business fallback. Point, do not re-teach mesh.
Why is no-breaker plus retries an anti-pattern?
Answer
Retries amplify a sick dependency. The breaker stops sending. Jitter details live on Retry Storms.
How do breakers free bulkhead slots?
Answer
Fail-fast when Open so workers are not stuck waiting on a dark dep. Isolation still needs per-dep pools — a breaker without a bulkhead still shares threads.
Go Deeper
- resilience4j CircuitBreaker — sliding window, slow call
- Microsoft Learn — Circuit Breaker pattern
- Netflix Hystrix wiki — historical; concepts still interview-relevant
- Envoy outlier detection
- AWS Well-Architected — Reliability
- Google SRE — Handling Overload
- Next: Bulkheads — Thread Pools, Connection Pools, Queues