SLOs & tracing
Studies in this cluster, in series order. Each one keeps its own URL.
Observability
SLIs, traces, and the dashboard you would actually page on.
SLOs & tracing
6 studies- 1.SLOs, Error Budgets & Distributed TracingSLI/SLO/SLA; error budget & burn rates; multi-window burn alerts; OTel spans + W3C Trace Context; histograms not averages.
- 2.SLI Design PatternsAn SLI is good/total of something users feel. Availability, latency (threshold ratio), freshness, quality, and journey SLIs each fail differently if you count retries or averages.
- 3.Multi-Window Burn AlertsPage on burn rate times window, not error rate ≥ SLO for 10 minutes. Multi-window alerts catch fast burns and slow leaks without paging on a single failed request.
- 4.Histogram vs Average Latency SLIsAverages hide the tail. Latency SLIs are histogram threshold ratios (good = faster than T); exemplars attach traces to the buckets users actually felt.
- 5.Trace Sampling StrategiesHead sampling decides at the root; tail sampling waits until the trace is complete. Mix probabilistic, rate-limiting, and error-biased policies or you will miss the incident or drown storage.
- 6.Exemplars, RED & USE MethodsExemplars pin a trace to a metric bucket. RED (rate, errors, duration) for request services; USE (utilization, saturation, errors) for resources — use both, not one dashboard religion.