Observability triad
Studies in this cluster, in series order. Each one keeps its own URL.
Observability
SLIs, traces, and the dashboard you would actually page on.
Observability triad
5 studies- 1.Observability Triad — Metrics, Logs & Distributed TracingMetrics, logs, and traces answer different questions on one incident; correlate with exemplars and trace_id.
- 2.SLIs, SLOs & Error BudgetsSLI is the experience ratio; SLO is the target; error budget is 1 minus SLO. Choose nines and journey vs request SLIs deliberately.
- 3.Metric Cardinality & Prometheus Label DesignA series is metric times labels. Unbounded labels explode memory — bound labels; use histograms and recording rules.
- 4.Trace Context Propagation (W3C) & Sampling StrategiesW3C traceparent and tracestate keep traces intact. Head vs tail sampling trades cost against capturing rare failures.
- 5.Alerting — Multi-Window Burn Rates vs Static ThresholdsStatic thresholds flap. Multi-window multi-burn-rate pages on fast burn and tickets on slow burn.