TopicsObservability
Observability
SLIs, traces, and the dashboard you would actually page on.
Common tags: metrics, tracing, sli
- Observability
Timeouts, Budgets & Deadline Propagation — End-to-End Latency Caps
Cluster · Resilience Patterns
Timeouts bound how long one hop may wait; deadlines/budgets bound the whole user request across nested calls. Each hop must shrink the remaining budget (gRPC deadlines, context.WithDeadline, AbortSignal). Anti-pattern: server timeout longer than the client's wait—work continues after the client is gone. Pair with breakers and shedding so you fail fast when budget is nearly exhausted.
Open study →- observability
- sre
- timeouts
- deadlines
- resilience
- overload
- interview
- Observability
Retry vs Break vs Shed — Decision Matrix & Anti-Patterns
Cluster · Resilience Patterns
Choose the lever that matches the failure: retry transient, idempotent blips; break when a dependency is sick; shed when you are overloaded; degrade when a partial/cached answer beats an error. Mixing them wrong creates retry storms, stuck-open breakers, or silent UX death. This lesson is the cluster's decision matrix—not a re-teach of retry jitter or token buckets.
Open study →- observability
- sre
- resilience
- circuit-breaker
- load-shedding
- overload
- interview
- Observability
Resilience Patterns — Circuit Breakers, Bulkheads & Load Shedding
Cluster · Resilience Patterns
When systems overload, naive retries amplify traffic, connection pools saturate, and one sick dependency can take the whole site down. Resilience patterns—circuit breakers, bulkheads, load shedding, and deadline budgets—give you deliberate ways to stop calling sick dependencies, isolate blast radius, drop work early, and bound wait time. This hub maps the cluster and shows when to use each pattern versus the alternatives.
Open study →- observability
- sre
- resilience
- circuit-breaker
- bulkhead
- load-shedding
- overload
- interview
- Observability
Load Shedding & Admission Control — Drop Early, Protect the Core
Cluster · Resilience Patterns
Load shedding deliberately refuses or degrades work when you are the bottleneck—before queues explode and every request times out. Admission control at the edge/gateway drops low-priority or expensive work first, returns 503 + Retry-After, and protects critical paths (checkout over recommendations). Unlike rate limiting (fairness/quotas), shedding is survival under overload.
Open study →- observability
- sre
- load-shedding
- admission-control
- resilience
- overload
- interview
- Observability
Circuit Breaker Mechanics — Closed, Open, Half-Open, Thresholds & Probes
Cluster · Resilience Patterns
A circuit breaker wraps calls to a dependency and tracks recent failures (and often slow calls). In Closed it passes traffic; after a threshold it Opens and fail-fasts; after a reset timeout it Half-Opens with a limited probe. Tuned well, it stops retry storms against a sick service; tuned poorly, it flaps or never recovers.
Open study →- observability
- sre
- circuit-breaker
- resilience
- fail-fast
- overload
- interview
- Observability
Bulkheads — Thread Pools, Connection Pools, Queues & Failure Domains
Cluster · Resilience Patterns
Bulkheads isolate failure domains so a slow or chatty dependency cannot exhaust shared threads, connections, or CPU and starve unrelated work. Separate pools, queue depth limits, and coarse Kubernetes resource limits are all bulkheads at different layers. Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
Open study →- observability
- sre
- bulkhead
- resilience
- overload
- interview
- Observability
Trace Sampling Strategies
Cluster · SLOs & tracing
Head sampling decides at the root; tail sampling waits until the trace is complete. Mix probabilistic, rate-limiting, and error-biased policies or you will miss the incident or drown storage.
Open study →- observability
- sre
- tracing
- Observability
Trace Context Propagation (W3C) & Sampling Strategies
Cluster · Observability triad
W3C traceparent and tracestate keep traces intact. Head vs tail sampling trades cost against capturing rare failures.
Open study →- tracing
- w3c
- traceparent
- sampling
- otel
- Observability
SLIs, SLOs & Error Budgets
Cluster · Observability triad
SLI is the experience ratio; SLO is the target; error budget is 1 minus SLO. Choose nines and journey vs request SLIs deliberately.
Open study →- sli
- slo
- error-budget
- sre
- sla
- Observability
SLI Design Patterns
Cluster · SLOs & tracing
An SLI is good/total of something users feel. Availability, latency (threshold ratio), freshness, quality, and journey SLIs each fail differently if you count retries or averages.
Open study →- observability
- sre
- tracing
- Observability
Observability Triad — Metrics, Logs & Distributed Tracing
Cluster · Observability triad
Metrics, logs, and traces answer different questions on one incident; correlate with exemplars and trace_id.
Open study →- observability
- metrics
- logs
- traces
- exemplars
- otel
- sre
- Observability
Multi-Window Burn Alerts
Cluster · SLOs & tracing
Page on burn rate times window, not error rate ≥ SLO for 10 minutes. Multi-window alerts catch fast burns and slow leaks without paging on a single failed request.
Open study →- observability
- sre
- tracing
- Observability
Metric Cardinality & Prometheus Label Design
Cluster · Observability triad
A series is metric times labels. Unbounded labels explode memory — bound labels; use histograms and recording rules.
Open study →- prometheus
- cardinality
- labels
- histograms
- Observability
Histogram vs Average Latency SLIs
Cluster · SLOs & tracing
Averages hide the tail. Latency SLIs are histogram threshold ratios (good = faster than T); exemplars attach traces to the buckets users actually felt.
Open study →- observability
- sre
- tracing
- Observability
Exemplars, RED & USE Methods
Cluster · SLOs & tracing
Exemplars pin a trace to a metric bucket. RED (rate, errors, duration) for request services; USE (utilization, saturation, errors) for resources — use both, not one dashboard religion.
Open study →- observability
- sre
- tracing
- Observability
Alerting — Multi-Window Burn Rates vs Static Thresholds
Cluster · Observability triad
Static thresholds flap. Multi-window multi-burn-rate pages on fast burn and tickets on slow burn.
Open study →- alerting
- burn-rate
- slo
- on-call
- Observability
SLOs, Error Budgets & Distributed Tracing
Cluster · SLOs & tracing
SLI/SLO/SLA; error budget & burn rates; multi-window burn alerts; OTel spans + W3C Trace Context; histograms not averages.
Open study →- observability
- sre
- tracing