Observability
Part 1 of 6 · SLOs & tracingSLOs, Error Budgets & Distributed Tracing
SLI/SLO/SLA; error budget & burn rates; multi-window burn alerts; OTel spans + W3C Trace Context; histograms not averages.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
- 1
naive alert → fatigue or missed incidents
Page on error rate ≥ SLO for 10 minutes
A quiet hour with one 500 looks like 100% burn. A slow leak never trips 10 minutes. This is the alert that trains people to ignore pages.
- 2
Winner: multi-window burn rate on a user SLI
Budget = 1 − SLO. Burn = error_rate / budget. Page on fast×short AND slower×longer (classic 14.4× 1h/5m). Ticket on ~1× long window. Full recipes live on the burn-alerts page.
- ?
When the page fires, jump via exemplars and traces
Histograms, not averages. Exemplars pin a trace to the bucket. Head vs tail sampling decides whether you even have that trace. Those are the later lessons in this cluster.
Overview
Senior interviews probe whether you can turn "we monitor stuff" into user-centric reliability contracts. Production systems fail; the question is whether you fail within a budget the business accepted, and whether you can find the cause across services in minutes rather than hours.
Three skills this lesson drills:
- SLIs / SLOs / SLAs — measure what users feel, set a target below 100%, and know a contract from an internal objective.
- Error budgets — translate SLO misses into burn rate, freeze/feature tradeoffs, and multi-window burn alerts (SRE Workbook style).
- Distributed tracing — spans, W3C Trace Context, sampling, and the OpenTelemetry path from app → collector → backend, correlated with metrics and logs.
Prior lessons (consistent hashing, MVCC, Redis stampede prevention, structured outputs, idempotency) all create failure modes. Observability is how you detect, budget, and debug them.
You should be able to:
- Separate SLI, SLO, and SLA, and set the SLO stricter than the SLA.
- Compute a 99.9% window budget (~43 minutes per 30 days) and a burn rate.
- Design multi-window page/ticket alerts instead of "error rate ≥ SLO for 10 minutes."
- Propagate W3C Trace Context through OpenTelemetry; name head vs tail sampling.
- Defend histograms + threshold-ratio SLIs against average latency.
Flow
- 1
Client
- W3C traceparentEdge / LB
- 2
Edge / LB
- nextAPI
- relatedSLI: good / total
- 3
API
- nextPayments
- relatedOne trace of nested spans
- 4
Payments
- relatedOne trace of nested spans
- 5
SLI: good / total
- nextSLO target
- 6
SLO target
- nextError budget
- 7
Error budget
- nextMulti-window burn alerts
- 8
Multi-window burn alerts
- 9
One trace of nested spans
Lesson map
SLOs, Error Budgets & Distributed Tracing
SLI/SLO/SLA; error budget & burn rates; multi-window burn alerts; OTel spans + W3C Trace Context; histograms not averages.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB client["Client"] edge["Edge / LB"] api["API"] pay["Payments"] client -->|W3C traceparent| edge edge -->|Edge / LB to API| api api -->|API to Payments| pay
SLI vs SLO vs SLA
| Term | What it is | Who owns it | Example |
|---|---|---|---|
| SLI | Quantitative measure of service level (prefer good / total) | Eng / SRE | Successful HTTP requests / total HTTP requests |
| SLO | Internal target for an SLI over a window | Eng + product | 99.9% of API requests succeed over 28 days |
| SLA | External business contract with consequences (credits, penalties) | Legal / sales | 99.5% monthly uptime or 10% fee credit |
An SLA is not an SLO. SLAs are legal/commercial; SLOs are engineering decision tools. Set SLOs stricter than SLAs so you have room before customers get credits.
Prefer SLIs of the form good events / total events (0–100%). Then:
error_budget = 1 − SLOA ~4-week rolling window (integer weeks so weekend count stays stable) is the usual starting point, plus weekly ops summaries and quarterly planning rollups.
100% SLO is wrong: change causes outages; the internet between you and users is not 100%; a 100% target leaves only reactive work. Reliability is a tradeoff against velocity; the budget makes that tradeoff explicit.
Worked numbers
99.9% availability over 30 days → allowed downtime ≈ 43.2 minutes (0.1% of the window). Over 28 days that is ≈ 40.3 minutes. If the SLI is event-based, the budget is 0.1% of requests, not wall-clock.
If you serve 3,000,000 requests in the window with a 99.9% success SLO:
- Error budget =
0.001 × 3_000_000 = 3_000failed requests - One outage that causes 1,500 failures burns 50% of the budget
Quantify incidents by % of budget burned, not by "severity vibes." A bad release that burns 13% of quarterly budget twice a year often beats a rare disk failure that burns 65% once every five years — prioritize the higher expected budget cost.
Choosing good SLIs
Request-driven services (most APIs)
Use RED (Tom Wilkie): Rate, Errors, Duration.
Complement with USE (Brendan Gregg) on resources: Utilization, Saturation, Errors (CPU, memory, disk, network, queues). USE is for capacity; RED is closer to what users feel.
Recommended SLI shapes (SRE Workbook):
- Availability: successful responses / total (define "success" carefully — often non-5xx, or business-OK codes).
- Latency: fraction of requests faster than a threshold (for example, 99% under 300 ms). Prefer threshold-ratio SLIs for budgets over nested percentile targets.
- Quality: undegraded responses / total when you have graceful degradation (stale cache, generic avatars, and so on).
Pipelines and storage
- Freshness: records (or user-facing reads) updated within threshold / total.
- Correctness: verified-correct outputs / total (synthetic golden data or independent checker).
- Coverage: successfully processed records / expected records.
- Durability: readable records / written records (watch: users may care about today's data, not the 10-year archive).
Where you measure (quality vs cost)
| Placement | Upside | Gap |
|---|---|---|
| App logs / metrics | Cheap, detailed | Misses requests that never reach the app |
| Load balancer | Closer to the user | May miss client-side failures |
| Black-box probes | Catches total outages | Sparse coverage of paths |
| Client instrumentation | Closest to UX | Needs client code + a reliable telemetry pipeline |
Start cheap (LB/logs), iterate toward the user. Separate SLI specification ("homepage loads under 100 ms") from SLI implementation (how you measure it). Document the gap — server-side misses DNS/CDN failures.
Bucket request classes
Don't invent unique SLOs per endpoint. Bucket into classes:
- CRITICAL — login, checkout (tightest)
- HIGH_FAST — interactive UI paths
- HIGH_SLOW — reports / exports
- LOW / NO_SLO — polls, dark launches
Error budgets: formula, burn rate, freeze vs fix
SLI = good / total
SLO = target # e.g. 0.999
Error budget = 1 − SLO # as a fraction
Burn rate = actual_error_rate / (1 − SLO)Burn rate 1 means you will exhaust the budget exactly at the end of the window if the rate continues.
Burn rate 14.4 over 1 hour ≈ 2% of a 30-day budget consumed in that hour.
Error budget policy (the "teeth")
Without a written policy approved by product + eng + SRE, an SLO is just a dashboard KPI. Typical actions when budget is exhausted (or critically burned):
- Feature freeze / change freeze for risky deploys until budget recovers.
- Reliability focus — all (or most) eng capacity on reliability bugs until back in SLO.
- Escalation path when stakeholders disagree on whether to freeze.
When in budget with low toil and happy users: you can increase velocity or reduce SRE engagement. When met but users unhappy: tighten the SLO. When missed but users happy: consider loosening.
Multi-window burn alerts
Alert on significant budget consumption, not on "error rate ≥ SLO for 10 minutes" (that pages constantly while barely spending budget). Full recipes, PromQL sketches, and severity mapping: Multi-window burn alerts. How to pick the SLI that feeds the burn: SLI design patterns.
Recommended starting config (~99.9% SLO, ~30-day window). Page when long AND short windows both exceed the burn threshold (AND-gate reduces flapping). Short window ≈ 1/12 of long window so the alert resets quickly when burning stops.
| Severity | Long window | Short window | Burn | ~Budget consumed |
|---|---|---|---|---|
| Page | 1 hour | 5 minutes | 14.4× | ~2% |
| Page | 6 hours | 30 minutes | 6× | ~5% |
| Ticket | 3 days | 6 hours | 1× | ~10% |
Conceptual Prometheus-style rule (availability SLI as error ratio) — sketch only, not runnable here:
# Page: 14.4x burn over 1h AND still burning over 5m
expr: |
(
job:slo_errors_per_request:ratio_rate1h > (14.4 * 0.001)
and
job:slo_errors_per_request:ratio_rate5m > (14.4 * 0.001)
)
or
(
job:slo_errors_per_request:ratio_rate6h > (6 * 0.001)
and
job:slo_errors_per_request:ratio_rate30m > (6 * 0.001)
)Low-traffic caveats. One failed request in a quiet hour can look like a 1000× burn. Mitigations: synthetic traffic, combine related services, client retries with backoff, lower SLO / longer window, or minimum-request guards.
Extreme availability. 99.999% monthly → ~26 seconds of 100% outage exhausts the budget. You cannot page your way out; you need canaries, progressive delivery, and blast-radius limits.
Burn-rate math (run this)
Self-contained: 99.9% over 30 days, 14.4× / 6× / 1× windows, and the 1,500-of-3,000,000 outage.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Distributed tracing: spans, context, sampling
Core model
- Span: a timed operation (name, start/end, attributes, events, status, links).
- Trace: a tree/DAG of spans sharing one trace ID.
- Span context: trace ID + span ID + flags (+ tracestate).
- Parent/child: causal nesting (HTTP handler → DB query → cache get).
Logs are per-process narratives. Traces stitch causality across services with a shared trace ID. You see which dependency, queue, or serialization step ate the latency budget for one request.
W3C Trace Context
Propagate across process boundaries via HTTP headers:
traceparent:version-trace_id-parent_id-flags- Example:
00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01 - Flags bit
01= sampled
- Example:
tracestate: vendor-specific key/value list (OpenTelemetry may carry sampling thresholdth, randomnessrv, and so on)
OpenTelemetry's default propagator uses W3C Trace Context. Do not invent custom header formats between services you control — you will break cross-vendor correlation.
Security notes:
- Sanitize/ignore untrusted inbound
traceparentif abuse is a concern (forged IDs / header injection). - Avoid leaking internal baggage to public clients.
Sampling
Head vs tail, rate limits, and error-biased keep rules are a full lesson: Trace sampling strategies. Pair kept traces with metric buckets via exemplars / RED / USE. | --- | --- | --- | | Head-based (root decides) | Simple, low overhead | May drop the interesting rare error | | Parent-based | Honor upstream sampled flag | Consistent trees | | Probability / TraceIDRatio | Steady % of traffic | Need enough volume for rare paths | | Tail-based (collector) | Keep errors / slow / interesting | Higher buffer cost; best for high-cardinality debug |
Production pattern: low head sample rate (for example 1–10%) + tail sampling that always keeps errors and high-latency traces + exemplars linking histogram buckets to example trace IDs.
Too much sampling: miss the one bad customer shard. Too little: melt storage and money. Always-on 100% trace export rarely scales.
Architecture: apps → OTel SDK → collector → backend
Flow
- 1
App + OTel SDK
- OTLPOTel Collector
- 2
OTel Collector
- batch / sample / scrubTempo / Jaeger / Honeycomb
- nextPrometheus / Mimir
- nextLoki / ELK + trace_id
- 3
Tempo / Jaeger / Honeycomb
- 4
Prometheus / Mimir
- exemplarsTempo / Jaeger / Honeycomb
- 5
Loki / ELK + trace_id
- nextTempo / Jaeger / Honeycomb
Why a collector?
- Centralize exporters, auth, sampling, PII scrubbing, and fan-out.
- Keep app SDKs thin and restart-safe (batch + retry in the collector).
- One place to switch backends (Jaeger → Tempo → vendor) without redeploying every service.
Correlation triad:
- Metrics — cheap, aggregatable, great for SLIs and burn alerts.
- Logs — high-cardinality narrative; inject
trace_id/span_id. - Traces — causal path across services for why.
Exemplars: attach a trace ID to a Prometheus histogram observation so Grafana can jump from a p99 spike to a representative slow trace. Metrics/logs/traces as silos (no trace_id in logs, no exemplars) is how incidents stretch from minutes to hours.
Histograms, not averages
A 100 ms average can hide 1% of 5 s requests. Fan-out makes backend p99 become frontend median. Use histograms, threshold-ratio SLIs, and percentiles — never averages for latency SLOs. You also cannot aggregate means of means meaningfully. Worked numbers and heatmap pitfalls: Histogram vs average latency SLIs.
Put an exact bucket on the SLO threshold. If the latency SLI is "under 300 ms," the histogram needs le="0.3" (or equivalent). Missing that bucket biases the ratio SLI.
Example PromQL ideas (sketch):
# Latency SLI (good if under 300ms):
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/ sum(rate(http_request_duration_seconds_count[5m]))
# Availability SLI:
sum(rate(http_requests_total{code!~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))Deep dive · Why threshold-ratio beats nested percentiles
A budget wants a single good / total number. "p99 under 300 ms" nested inside "99.9% of windows" is hard to burn-alert. The threshold-ratio SLI ("fraction of requests faster than 300 ms ≥ 99%") is the same shape as availability, so the 14.4× / 6× / 1× windows apply unchanged.
OTel and Prometheus sketches (not runnable here)
These need SDKs, a collector, and /metrics. Read them as wiring diagrams; run the burn-rate sandboxes above.
Python span + W3C extract/inject:
# pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.propagate import inject, extract
from opentelemetry.trace import Status, StatusCode
resource = Resource.create({"service.name": "checkout-api", "service.version": "1.4.0"})
provider = TracerProvider(resource=resource)
provider.add_span_processor(BatchSpanProcessor(
OTLPSpanExporter(endpoint="localhost:4317", insecure=True)
))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("checkout-api")
def handle_checkout(headers: dict, cart_id: str) -> dict:
ctx = extract(headers) # W3C traceparent from inbound HTTP
with tracer.start_as_current_span("checkout.handle", context=ctx) as span:
span.set_attribute("cart.id", cart_id)
span.set_attribute("http.route", "/v1/checkout")
try:
charge = charge_card(cart_id)
span.set_attribute("payment.status", charge["status"])
return {"ok": True, "charge_id": charge["id"]}
except Exception as exc:
span.record_exception(exc)
span.set_status(Status(StatusCode.ERROR, str(exc)))
raise
def charge_card(cart_id: str) -> dict:
with tracer.start_as_current_span("payment.charge") as span:
span.set_attribute("cart.id", cart_id)
outbound_headers: dict[str, str] = {}
inject(outbound_headers) # sets traceparent / tracestate
return {"id": "ch_123", "status": "succeeded"}Python histogram with an exact 300 ms bucket:
# pip install prometheus-client flask
from prometheus_client import Counter, Histogram, generate_latest, CONTENT_TYPE_LATEST
from flask import Flask, Response, request
import time
REQUESTS = Counter("http_requests_total", "HTTP requests", ["method", "route", "code"])
LATENCY = Histogram(
"http_request_duration_seconds",
"Request latency",
["method", "route"],
buckets=(0.05, 0.1, 0.2, 0.3, 0.5, 1.0, 2.5, 5.0, float("inf")),
)
app = Flask(__name__)
@app.before_request
def _start_timer():
request._start = time.perf_counter()
@app.after_request
def _record(resp):
route = request.url_rule.rule if request.url_rule else "unknown"
elapsed = time.perf_counter() - getattr(request, "_start", time.perf_counter())
REQUESTS.labels(request.method, route, str(resp.status_code)).inc()
LATENCY.labels(request.method, route).observe(elapsed)
return resp
@app.get("/metrics")
def metrics():
return Response(generate_latest(), mimetype=CONTENT_TYPE_LATEST)TypeScript OpenTelemetry Node SDK sketch:
// npm i @opentelemetry/sdk-node @opentelemetry/auto-instrumentations-node
// @opentelemetry/exporter-trace-otlp-grpc @opentelemetry/api express
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-grpc";
import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";
import { Resource } from "@opentelemetry/resources";
import { ATTR_SERVICE_NAME } from "@opentelemetry/semantic-conventions";
import { trace, SpanStatusCode, context, propagation } from "@opentelemetry/api";
import express from "express";
const sdk = new NodeSDK({
resource: new Resource({ [ATTR_SERVICE_NAME]: "checkout-api" }),
traceExporter: new OTLPTraceExporter({ url: "http://localhost:4317" }),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();
const tracer = trace.getTracer("checkout-api");
const app = express();
app.post("/v1/checkout", async (req, res) => {
return tracer.startActiveSpan("checkout.handle", async (span) => {
span.setAttribute("cart.id", String(req.body?.cartId ?? ""));
span.setAttribute("http.route", "/v1/checkout");
try {
const charge = await chargeCard(String(req.body?.cartId));
span.setAttribute("payment.status", charge.status);
res.json({ ok: true, chargeId: charge.id });
} catch (err) {
span.recordException(err as Error);
span.setStatus({ code: SpanStatusCode.ERROR, message: String(err) });
res.status(502).json({ ok: false });
} finally {
span.end();
}
});
});
async function chargeCard(cartId: string) {
return tracer.startActiveSpan("payment.charge", async (span) => {
span.setAttribute("cart.id", cartId);
const headers: Record<string, string> = {};
propagation.inject(context.active(), headers); // W3C traceparent
span.end();
return { id: "ch_123", status: "succeeded" };
});
}TypeScript prom-client histogram sketch (same PromQL ideas as above):
// npm i prom-client express
import express from "express";
import client from "prom-client";
const register = new client.Registry();
client.collectDefaultMetrics({ register });
const httpRequests = new client.Counter({
name: "http_requests_total",
help: "HTTP requests",
labelNames: ["method", "route", "code"] as const,
registers: [register],
});
const httpDuration = new client.Histogram({
name: "http_request_duration_seconds",
help: "Request latency",
labelNames: ["method", "route"] as const,
buckets: [0.05, 0.1, 0.2, 0.3, 0.5, 1, 2.5, 5],
registers: [register],
});Interview Q&A
What's the difference between SLI, SLO, and SLA?
Answer
SLI is the measurement (good / total). SLO is the internal target over a window. SLA is the external contract with remedies. Engineers manage to SLOs; legal/sales own SLAs. SLOs should be tighter than SLAs.
Why not set a 100% availability SLO?
Answer
Impossible in practice (change, dependencies, client networks). It removes the error budget, forces pure firefighting, and blocks improvement. Reliability is a tradeoff against velocity; the budget makes that tradeoff explicit.
How do you alert on SLOs without alert fatigue?
Answer
Multi-window, multi-burn-rate alerts: page on fast burns (14.4× over 1h and 5m), ticket on slow burns (1× over 3d and 6h). Alert on budget consumption, not raw threshold flaps. Suppress overlapping severities. AND-gate long + short windows so a blip does not page.
Average latency looks fine — users still complain. Why?
Answer
Averages hide the tail. A 100 ms average can hide 1% of 5 s requests; fan-out makes backend p99 become frontend median. Use histograms, threshold-ratio SLIs, and percentiles — never averages for latency SLOs. Put an exact histogram bucket on the threshold (300 ms → le="0.3").
How does distributed tracing help beyond logs?
Answer
Logs are per-process narratives. Traces stitch causality across services with shared trace IDs via W3C traceparent. You see which dependency, queue, or serialization step ate the latency budget for one request.
Head sampling vs tail sampling?
Answer
Head decides at the root (cheap, may drop rare failures). Tail buffers and keeps errors/slow traces (better recall for incidents, more collector cost). Many orgs combine low head rate + tail keep-rules + exemplars.
Where should the SLO be measured — service or edge?
Answer
Prefer the closest practical point to the user for product SLOs (edge / LB / client). Use component SLOs for dependency contracts. Document the implementation gap (server-side misses DNS/CDN failures).
Error budget exhausted mid-quarter — ship the launch?
Answer
Per policy: usually no for risky changes; freeze or require reliability work until budget recovers. Exceptions go through the documented escalation path with product ownership — not silent heroics.
What is an error budget, operationally?
Answer
1 − SLO over the window, in events (or time, if you must). A 99.9% monthly availability SLO on 10M requests is ~10k allowed failures. The budget is the shared currency between product velocity and reliability work. When it is gone, the default is to stop risky ships — that is the point of having one.
How do you define a good vs valid event for an availability SLI?
Answer
Valid = requests that should count (successful or failed user-facing calls; exclude probes, bots, and bad-client 4xx if that is policy). Good = the subset that met the promise (2xx/3xx, or latency ≤ threshold). SLI = good / valid. Mixing health-check 200s into valid is how dashboards stay green while checkout is dead.
What are exemplars and why pair them with traces?
Answer
An exemplar is a sample trace id attached to a metric bucket (for example a histogram _bucket that just observed a 2s checkout). From a red p99 chart you jump straight into the trace that caused it. Without exemplars you grep logs; with them, metrics and traces are one workflow.
Pitfalls
- Vanity metrics — CPU "looks fine," uptime ping to
/healthzwhile checkout is broken. Measure user journeys. - Averaging latency — destroys tail signal; cannot aggregate means of means for SLOs.
- 100% SLO / no error budget policy — SLO without teeth; reliability work loses to features forever (until a big outage).
- Alerting on raw error rate ≥
(1 − SLO)— floods pages while barely spending budget; poor precision. - Too much sampling — miss the one bad customer shard. Too little — melt storage and money.
- Inconsistent context propagation — custom headers, missing middleware on one language → broken traces at service boundaries.
- Histogram buckets missing the SLO threshold — ratio SLI biased; put an exact
lebucket on the threshold (for example 0.3 s). - Low-traffic paging — one failure at night pages on-call; add synthetics / aggregation / floors.
- Metrics/logs/traces silos — no
trace_idin logs, no exemplars → slow incident response. - Treating dependency outages as "not our burn" — users don't care whose ticket queue failed; policy should address user impact.
Pick one user journey (checkout). Define: the SLI event and "good," the window, the SLO (not 100%), the 30-day budget in minutes and in failed requests at 3e6 traffic, one 14.4× page and one 1× ticket, where traceparent is injected, and the histogram bucket that matches the latency threshold. If any of those is hand-wavy, you are not done.
Go Deeper
- Google SRE Workbook — Implementing SLOs
- Google SRE Workbook — Alerting on SLOs (multi-window burn rates)
- OpenTelemetry — Context propagation (W3C Trace Context)
- W3C Trace Context
- Prometheus — Histograms and summaries
- Grafana — The RED Method (Tom Wilkie)
- Honeycomb — SLOs get budget rate alerts
- YouTube: Observability: the present and future (Charity Majors)
- YouTube: Actionable SLOs Based on What Matters Most
- OpenTelemetry Python docs
Cheat sheet
SLI = good / total (user-centric)
SLO = target over window (e.g. 99.9% / 28d or 30d)
Budget = 1 − SLO
Burn = error_rate / (1 − SLO)
Page when (long burn AND short burn) at 14.4× (1h/5m) or 6× (6h/30m)
Ticket at ~1× (3d/6h)
Trace = tree of spans; propagate W3C traceparent
OTel: SDK → Collector → Tempo/Jaeger/Honeycomb + Prom + logs(trace_id)
Latency SLIs: histogram + threshold ratio, not averages