Observability
Part 4 of 6 · SLOs & tracingHistogram vs Average Latency SLIs
Averages hide the tail. Latency SLIs are histogram threshold ratios (good = faster than T); exemplars attach traces to the buckets users actually felt.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Latency is a distribution. Users live in the tail. A service with ninety-nine 100 ms calls and one 5 s call has a 149 ms average that looks “fine” in a weekly review while 1% of users waited five seconds. Fan-out makes it worse: if a page issues 20 parallel backend calls, a backend p99 becomes something like the frontend median.
The SLI that belongs in a budget is not “mean < 200 ms” and not a nested “p99 < 300 ms in 99.9% of 5-minute windows.” It is a threshold ratio:
good = requests with duration ≤ T
SLI = good / validThat ratio is the same shape as availability, so multi-window burn alerts apply unchanged. Histograms are how you implement the count of duration ≤ T without storing every sample. Exemplars are how you jump from a hot bucket to a trace — full pairing on exemplars / RED / USE.
Why histograms win over averages
| Aspect | Average-only | Histogram threshold SLI (winner) |
|---|---|---|
| What you see | One mean per interval | Distribution, quantiles, heatmaps |
| Alert fidelity | Fires always or never | Matches the user promise “under T” |
| Aggregation | Mean of means is not the mean | Buckets add; then quantile or ratio |
| Root cause | Grep logs | Exemplar → trace id |
| Storage | Tiny | Buckets × labels (moderate) |
| OTel / Prom | DIY timer | Native histogram instruments |
- 1
mean looks green → tail users rage
SLO on average latency
Ninety-nine requests at 100 ms plus one at 5 s average 149 ms. Capacity planning using the mean provisions for nobody’s p99. You also cannot add means across shards and get a global mean without weights.
- 2
Winner: histogram + threshold ratio + exact bucket on T
Count observations ≤ T. SLI = that count / valid. Put le on T. Burn-alert the miss ratio like availability. Heatmaps show when the tail moved.
- ?
Exemplars and sampling decide if you can explain the bucket
A red p99 without a trace is a heatmap you stare at. Head vs tail sampling is the next lesson; exemplars attach the kept trace to the bucket.
Why the mean lies
Three failure modes show up in interviews:
- Tail hiding. Outliers pull the mean a little and the users a lot. Medians hide the tail even more (the 5 s call may not move p50 at all).
- Bursty correlation. GC pauses, stop-the-world, downstream timeouts cluster. A mean over 5 minutes dilutes a 20-second outage into “+3 ms.” A heatmap (histogram per time column) shows the spike as a bright band.
- Fan-out. One slow dependency among many parallel calls dominates the parent span. Backend averages stay pretty; the user’s waterfall does not.
Percentiles (p50 / p95 / p99) are the right dashboard language. They are still awkward as a budget: “p99 under 300 ms” is not one good/valid number. Convert to a threshold ratio (“99% of requests under 300 ms”) and you can spend error budget.
Histograms, buckets, quantiles
A histogram is a set of cumulative buckets: each le (less-or-equal) bound counts observations at or below that bound, plus +Inf. Prometheus estimates quantiles by interpolating in the bucket where the cumulative count crosses the percentile. Too few buckets around T makes p99 and the threshold ratio both mushy. Too many plus high-cardinality labels explodes series.
Exponential buckets (0.001, 0.002, 0.004, … or OTel’s exponential histograms) cover microseconds to seconds with a handful of bounds — the usual latency choice. Linear buckets (0–10 ms, 10–20 ms, …) give uniform resolution and explode if the range is large.
Put T on a bound. If the SLI is “under 300 ms,” you want a bucket at 0.3 s. If the nearest bounds are 0.25 and 0.5, the ratio le="0.25" is too strict and le="0.5" is too loose; interpolation for quantiles will not save a threshold-ratio SLI that reads a single le.
Sketch (not scraped here):
# good if under 300ms
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(http_request_duration_seconds_count[5m]))OpenTelemetry histograms record into SDK buckets (classic explicit, or exponential) and exporters map them to Prometheus _bucket series. Units are seconds unless you set otherwise — recording milliseconds into a 0–10 s explicit histogram parks everything in the first bucket.
A heatmap is that histogram over time (for example 1-minute columns). Use it to see when the tail moved (deploys, GC, traffic shape). A single quantile line cannot show a bimodal split (fast cache hits vs slow origin).
Flow
- 1
Request
- nextTimer start
- 2
Timer start
- nextHandler / span
- 3
Handler / span
- nextTimer stop
- 4
Timer stop
- nextHistogram observe duration
- 5
Histogram observe duration
- nextAdd to bucket le ≥ duration
- 6
Add to bucket le ≥ duration
- nextOptional exemplar: trace id
- 7
Optional exemplar: trace id
- nextExporter OTLP / Prom
- 8
Exporter OTLP / Prom
- nextRatio SLI + heatmap + quantile
- 9
Ratio SLI + heatmap + quantile
- nextBurn alerts on 1 − ratio
- 10
Burn alerts on 1 − ratio
Lesson map
Histogram vs Average Latency SLIs
Averages hide the tail. Latency SLIs are histogram threshold ratios (good = faster than T); exemplars attach traces to the buckets users actually felt.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Request"] b["Timer start"] c["Handler / span"] d["Timer stop"] a -->|Request to Timer start| b b -->|Timer start to Handler / span| c c -->|Handler / span to Timer stop| d
Step notes: 1 time the user-visible work, 2 observe once per valid event, 3 land in a bucket, 4 optionally attach an exemplar, 5 export, 6 SLI is the threshold ratio, not the mean.
Deep dive · Quantile alerts vs threshold-ratio burns
histogram_quantile(0.99, …) > 0.3 for 5 minutes is a fine symptom alert. It is still not a budget: a p99 that flaps around 300 ms may spend little error budget, and a 1% mass at 2 s can keep p99 “OK” depending on interpolation. The threshold ratio counts every request slower than T as a bad event, which is what users and error budgets understand. Use quantiles on dashboards; use the ratio on burn alerts.
Exemplars, labels, cold start
An exemplar is a sample attached to a bucket — typically a trace id (and span id) for one observation that landed there. When the 1–2.5 s bucket spikes, Grafana can jump to Tempo/Jaeger on that id. Without exemplars you correlate by timestamp and hope. Sampling still has to keep that trace; see trace sampling.
Label discipline: route, method, status (or class). Not user_id. Each unique label set multiplies every _bucket.
Cold start / JIT / first-connection TLS can fill the tail for a minute after deploy. Exclude a warmup window from the SLO, or use canaries, rather than loosening T globally.
Worked comparison (in-memory)
I/O: a list of latencies in, printed mean vs threshold ratio vs bucket counts. No /metrics.
Press Run. Snippets must be self-contained — no network, files, or native modules.
You should see: mean ≈ 149 ms, threshold ratio 0.990 at 300 ms, the 300 ms bucket matches good/total, unweighted mean-of-means 240 vs weighted ~83, and a missing T bucket either under- or over-counts good.
_sum / _count is the average. Ignore it for SLOs.
Prometheus histograms also export _sum and _count. Their ratio is the mean. It is useful for capacity (“we did 2,000 CPU-seconds of handler time”) and useless as a latency SLO. Interview trap: “we already have a histogram, we SLO on _sum/_count.” That is the average with extra steps.
Summaries compute quantiles in-process. You cannot merge summaries across replicas and get a global p99. Histograms merge by adding buckets. That is why latency SLIs at more than one replica must be histograms (or OTel exponential histograms exported in a mergeable form), not summaries.
histogram_quantile interpolation. Prometheus assumes a uniform distribution inside the bucket where the cumulative count crosses the percentile. If 90% of requests are under 200 ms and the next bound is 1 s, a p95 can be interpolated somewhere in 200 ms–1 s and look “about 400 ms” while the real tail is all at 900 ms. Tighter bounds in the SLO region reduce that lie. The threshold ratio that reads a single le does not interpolate — which is another reason it is the better budget SLI.
+Inf bucket. Every observation is ≤ inf. The count there is _count. If a large mass sits only in +Inf (no bound above your worst real latency), quantiles in the tail are fiction. Include a bound above T and above the timeouts you actually return (gateway 15 s, handler 2.5 s, …).
Native / exponential histograms (Prometheus native, OTel exponential) spend fewer series on a wide range. You still want to know T and confirm the schema has resolution around T. “Exponential therefore we skipped the 300 ms bound” is not a defense.
Fan-out numbers worth memorizing
If each of N independent backends has probability p of exceeding T, the parent that waits for all of them exceeds T with probability 1 − (1−p)^N. For N=20 and p=0.01 (backend p99 = T), parent P(slow) ≈ 18%. The frontend median can sit near a backend p99. This is why a pretty average on each microservice is compatible with a miserable user, and why a journey latency SLI (parent span, or client RUM) beats averaging child means.
Heatmap reading: a bright band at 2–5 s after a deploy, with the 0–200 ms band still populated, is bimodal (cache hit vs origin, or new slow path) — a quantile line might barely move. The threshold ratio at 300 ms will drop as soon as the slow mode has mass. That drop is what you burn-alert.
Timeouts: a 5 s handler timeout is a bad latency event and usually a bad availability event. Count it once in each SLI (failed, and not ≤ T) or you will double-burn; most teams accept both because the user felt both. Say the choice.
Designing buckets around T
Start from the SLO, not from a copied default.
- Place an exact bound on T (0.3 s if T is 300 ms).
- Put two or three bounds below T for heatmaps of the happy path (0.05, 0.1, 0.2).
- Put bounds above T at likely timeouts and “user gave up” (0.5, 1, 2.5, 5, 10).
- End with
+Inf. - Keep the list identical across languages so a Python and a Node service merge.
Default Prometheus buckets (0.005 … 10) may skip your T. OTel exponential histograms usually cover the range but you must still verify resolution near T. Changing buckets later resets historical quantile compatibility — treat the list as an API.
Heatmap pitfalls: log-scale axes hide a 300 ms SLO line; always overlay T. Mixing seconds and milliseconds on one heatmap is how a 250 ms cluster looks like 250 s. Exclude +Inf from color if a few timeouts would paint the whole column.
Client RUM histograms (navigation timing) are the implementation closest to the spec “page usable under 2 s.” Server histograms miss DNS/TLS/CDN. Document that gap the same way you do for availability.
The _sum series is still worth scraping: it is the work integral for capacity, not the SLO. Use it on a USE-ish “seconds spent” panel; do not burn-alert it.
Cold start again: a canary with 50 rps will fill the first histogram minutes with JIT/TLS. Either exclude job=canary from the product SLI or wait until the instance is in the serving set. Mixing canary and stable in one histogram without a label is how a good deploy looks like a tail regression.
If you need a single spoken example in an interview: “99 requests at 100 ms, one at 5 s: mean 149 ms, fraction under 300 ms is 99%. I SLO the fraction, I chart the heatmap, I put a bucket on 300 ms, I attach an exemplar to the 5 s call.”
Recording 250 “milliseconds” into a seconds histogram lands 250 in +Inf (or an enormous exponential bucket). The playground uses milliseconds in memory and compares to T=300 in the same unit on purpose — convert to seconds at the instrument, not in the SLO sentence.
Interview Q&A
Why is mean latency a poor SLI?
Answer
The mean is pulled a little by outliers while users in the tail have a terrible time. A 149 ms mean can hide 1% of 5 s calls. Medians hide them even more. SLOs need a user promise: fraction under T, implemented with a histogram.
OpenTelemetry histogram vs Prometheus histogram?
Answer
OTel is the SDK instrument (explicit or exponential buckets, optional exemplars). Prometheus is a storage/exposition format (_bucket, _count, _sum). Exporters translate. Do not record milliseconds into a seconds histogram.
How does an exemplar help?
Answer
It stores a raw sample — usually a trace id — on a bucket. When that bucket or p99 spikes, you open the trace instead of grepping logs by timestamp. Useless if sampling dropped the trace or you have no trace backend.
Why not alert on p99 > 300 ms for 5 minutes as the SLO?
Answer
Fine as a symptom. As a budget it is not a good/valid ratio, interpolation can hide mass above T, and 14.4× burn math wants one error rate. Prefer fraction of requests slower than 300 ms, with an exact le bucket on 0.3 s.
Exponential vs linear buckets?
Answer
Exponential covers a wide dynamic range with few bounds — default for latency. Linear gives uniform resolution and explodes series if the range is seconds and the step is milliseconds. Always include a bound at the SLO threshold T.
Heatmap vs histogram?
Answer
A histogram is the distribution now (or over one range). A heatmap is histograms over time columns. Use heatmaps to see bursty tails and deploys; use the threshold ratio for the SLO.
Fan-out: why does backend p99 become frontend median?
Answer
If the browser or BFF waits on N parallel calls, the parent finishes when the slowest child finishes. The distribution of the max of N draws sits far to the right of one draw. Quote that when someone shows you a pretty backend average.
What labels belong on a latency histogram?
Answer
Low cardinality: route template, method, status or class, maybe region. Never user id, raw URL, or trace id as a label (trace id belongs on the exemplar). Each extra distinct combination multiplies every bucket series.
Units?
Answer
OTel histograms default to seconds. Prometheus examples use http_request_duration_seconds. Recording 250 ms as 250 instead of 0.25 overflows buckets and makes quantiles nonsense. Standardize; lint it.
How does this SLI feed a burn alert?
Answer
error_rate = 1 − (count ≤ T / valid). Then burn = error_rate / (1 − SLO) with the same 14.4× 1 h/5 m AND-gate as availability. Nested percentile targets do not drop into that formula cleanly.
Why is _sum / _count not a latency SLO?
Answer
That ratio is the mean. Capacity likes it; users do not. Summaries cannot merge across replicas; histograms add buckets. SLO the threshold ratio, chart quantiles, ignore the mean for paging.
Pitfalls
- Average on the SLO — hides tails; cannot merge shards.
- No bucket at T — ratio biased to the nearest bound.
- Too few buckets in the tail — p99 and heatmaps lie.
- High-cardinality labels —
user_id× buckets melts Prometheus. - Storing every raw latency — histograms exist so you do not.
- Missing exemplars — no jump to a trace; configure sampler + exemplar on the histogram.
- Cold-start bias — first N seconds after deploy look like an SLO miss.
- Millisecond vs second mix — everything lands in +Inf or in the first bucket.
- Quantile-only alerts — flap without spending budget; prefer threshold ratio for burns.
- Canary mixed into product histogram — warmup looks like a tail regression; label or exclude it.
- Timeouts only in +Inf — without a bound near the gateway timeout, heatmaps and p99 become fiction.
Write 99 numbers at 100 ms and one at 5 s. Compute mean, p50, and fraction under 300 ms. Then merge a 1000-qps 80 ms shard with a 10-qps 400 ms shard using means-of-means vs adding counts. If those two exercises surprise you, you are not ready to defend a latency SLO.
Go Deeper
- Prometheus — Histograms and summaries
- Prometheus — histogram_quantile
- OpenTelemetry — Histogram data model
- Grafana — Heatmap
- Google SRE Book — Embracing Risk
- Next: Trace sampling strategies
Cheat sheet
Do not SLO on average latency
SLI = count(duration ≤ T) / valid
Histogram implements the count; put le on T
Buckets add across shards; means of means do not
Heatmap = histogram over time
Exemplar = trace id on a bucket
Burn the miss ratio with 14.4× windows