Observability
Topic, then cluster, then study. Recently added is the short list at the top.
Recently added
Show more- 3.Bulkheads — Thread Pools, Connection Pools, Queues & Failure DomainsBulkheads isolate failure domains so a slow or chatty dependency cannot exhaust shared threads, connections, or CPU and starve unrelated work. Separate pools, queue depth limits, and coarse Kubernetes resource limits are all bulkheads at different layers. Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
- 2.Circuit Breaker Mechanics — Closed, Open, Half-Open, Thresholds & ProbesA circuit breaker wraps calls to a dependency and tracks recent failures (and often slow calls). In Closed it passes traffic; after a threshold it Opens and fail-fasts; after a reset timeout it Half-Opens with a limited probe. Tuned well, it stops retry storms against a sick service; tuned poorly, it flaps or never recovers.
- 4.Load Shedding & Admission Control — Drop Early, Protect the CoreLoad shedding deliberately refuses or degrades work when you are the bottleneck—before queues explode and every request times out. Admission control at the edge/gateway drops low-priority or expensive work first, returns 503 + Retry-After, and protects critical paths (checkout over recommendations). Unlike rate limiting (fairness/quotas), shedding is survival under overload.
- 1.Resilience Patterns — Circuit Breakers, Bulkheads & Load SheddingWhen systems overload, naive retries amplify traffic, connection pools saturate, and one sick dependency can take the whole site down. Resilience patterns—circuit breakers, bulkheads, load shedding, and deadline budgets—give you deliberate ways to stop calling sick dependencies, isolate blast radius, drop work early, and bound wait time. This hub maps the cluster and shows when to use each pattern versus the alternatives.
- 6.Retry vs Break vs Shed — Decision Matrix & Anti-PatternsChoose the lever that matches the failure: retry transient, idempotent blips; break when a dependency is sick; shed when you are overloaded; degrade when a partial/cached answer beats an error. Mixing them wrong creates retry storms, stuck-open breakers, or silent UX death. This lesson is the cluster's decision matrix—not a re-teach of retry jitter or token buckets.
- 5.Timeouts, Budgets & Deadline Propagation — End-to-End Latency CapsTimeouts bound how long one hop may wait; deadlines/budgets bound the whole user request across nested calls. Each hop must shrink the remaining budget (gRPC deadlines, context.WithDeadline, AbortSignal). Anti-pattern: server timeout longer than the client's wait—work continues after the client is gone. Pair with breakers and shedding so you fail fast when budget is nearly exhausted.
Observability
SLIs, traces, and the dashboard you would actually page on.
Observability triad
5 studies- 1.Observability Triad — Metrics, Logs & Distributed TracingMetrics, logs, and traces answer different questions on one incident; correlate with exemplars and trace_id.
- 2.SLIs, SLOs & Error BudgetsSLI is the experience ratio; SLO is the target; error budget is 1 minus SLO. Choose nines and journey vs request SLIs deliberately.
- 3.Metric Cardinality & Prometheus Label DesignA series is metric times labels. Unbounded labels explode memory — bound labels; use histograms and recording rules.
- 4.Trace Context Propagation (W3C) & Sampling StrategiesW3C traceparent and tracestate keep traces intact. Head vs tail sampling trades cost against capturing rare failures.
- 5.Alerting — Multi-Window Burn Rates vs Static ThresholdsStatic thresholds flap. Multi-window multi-burn-rate pages on fast burn and tickets on slow burn.
Resilience Patterns
6 studies- 1.Resilience Patterns — Circuit Breakers, Bulkheads & Load SheddingWhen systems overload, naive retries amplify traffic, connection pools saturate, and one sick dependency can take the whole site down. Resilience patterns—circuit breakers, bulkheads, load shedding, and deadline budgets—give you deliberate ways to stop calling sick dependencies, isolate blast radius, drop work early, and bound wait time. This hub maps the cluster and shows when to use each pattern versus the alternatives.
- 2.Circuit Breaker Mechanics — Closed, Open, Half-Open, Thresholds & ProbesA circuit breaker wraps calls to a dependency and tracks recent failures (and often slow calls). In Closed it passes traffic; after a threshold it Opens and fail-fasts; after a reset timeout it Half-Opens with a limited probe. Tuned well, it stops retry storms against a sick service; tuned poorly, it flaps or never recovers.
- 3.Bulkheads — Thread Pools, Connection Pools, Queues & Failure DomainsBulkheads isolate failure domains so a slow or chatty dependency cannot exhaust shared threads, connections, or CPU and starve unrelated work. Separate pools, queue depth limits, and coarse Kubernetes resource limits are all bulkheads at different layers. Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
- 4.Load Shedding & Admission Control — Drop Early, Protect the CoreLoad shedding deliberately refuses or degrades work when you are the bottleneck—before queues explode and every request times out. Admission control at the edge/gateway drops low-priority or expensive work first, returns 503 + Retry-After, and protects critical paths (checkout over recommendations). Unlike rate limiting (fairness/quotas), shedding is survival under overload.
- 5.Timeouts, Budgets & Deadline Propagation — End-to-End Latency CapsTimeouts bound how long one hop may wait; deadlines/budgets bound the whole user request across nested calls. Each hop must shrink the remaining budget (gRPC deadlines, context.WithDeadline, AbortSignal). Anti-pattern: server timeout longer than the client's wait—work continues after the client is gone. Pair with breakers and shedding so you fail fast when budget is nearly exhausted.
- 6.Retry vs Break vs Shed — Decision Matrix & Anti-PatternsChoose the lever that matches the failure: retry transient, idempotent blips; break when a dependency is sick; shed when you are overloaded; degrade when a partial/cached answer beats an error. Mixing them wrong creates retry storms, stuck-open breakers, or silent UX death. This lesson is the cluster's decision matrix—not a re-teach of retry jitter or token buckets.
SLOs & tracing
6 studies- 1.SLOs, Error Budgets & Distributed TracingSLI/SLO/SLA; error budget & burn rates; multi-window burn alerts; OTel spans + W3C Trace Context; histograms not averages.
- 2.SLI Design PatternsAn SLI is good/total of something users feel. Availability, latency (threshold ratio), freshness, quality, and journey SLIs each fail differently if you count retries or averages.
- 3.Multi-Window Burn AlertsPage on burn rate times window, not error rate ≥ SLO for 10 minutes. Multi-window alerts catch fast burns and slow leaks without paging on a single failed request.
- 4.Histogram vs Average Latency SLIsAverages hide the tail. Latency SLIs are histogram threshold ratios (good = faster than T); exemplars attach traces to the buckets users actually felt.
- 5.Trace Sampling StrategiesHead sampling decides at the root; tail sampling waits until the trace is complete. Mix probabilistic, rate-limiting, and error-biased policies or you will miss the incident or drown storage.
- 6.Exemplars, RED & USE MethodsExemplars pin a trace to a metric bucket. RED (rate, errors, duration) for request services; USE (utilization, saturation, errors) for resources — use both, not one dashboard religion.