Observability
Part 6 of 6 · SLOs & tracingExemplars, RED & USE Methods
Exemplars pin a trace to a metric bucket. RED (rate, errors, duration) for request services; USE (utilization, saturation, errors) for resources — use both, not one dashboard religion.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
When a 14.4× burn page fires, the next minute should not be “grep production.” You need a metric that already carries a trace id for a representative bad request, and you need to know whether you are looking at a service problem or a resource problem.
- Exemplars pin a trace to a metric point (almost always a latency histogram bucket). Grafana jumps to Tempo/Jaeger.
- RED is the request-centric trio that lines up with SLI design: how much traffic, how many errors, how long.
- USE is the resource-centric trio from Brendan Gregg’s USE method: how busy, how queued, how many device errors.
They are complementary. A checkout handler with a bad deploy is a RED incident on a quiet CPU. A disk at 100% utilization with a rising wait queue is a USE incident that will become RED in a few minutes. One dashboard religion leaves half of those dark.
Why “both, plus exemplars” wins
| Aspect | Exemplars | RED | USE |
|---|---|---|---|
| Scope | One observation ↔ one trace | Service / RPC | CPU, disk, nic, pool, queue |
| Primary job | Fast jump from spike to cause | SLO-shaped rate / error / duration | Catch saturation before user errors |
| Data | Histogram + exemplar labels | Counters + histogram | Gauges + counters |
| Alert | Not by itself | Error ratio, duration SLI burns | Saturation, resource errors |
| When | Tracing already exists | HTTP/gRPC/queue consumers | Every resource you can run out of |
- 1
metrics religion or traces religion → slow MTTR
One giant dashboard of everything
RED without USE misses the full disk. USE without RED misses the 500s on idle hosts. Traces without exemplars means you already know the minute, not the request. Metrics without a sampled tail means you never open a waterfall.
- 2
Winner: RED for services, USE for resources, exemplars on the tail
Product SLIs and burns sit on RED (and threshold-ratio duration). Capacity sits on USE. Histogram observe() attaches a trace id on slow/error samples. Tail sampling keeps those traces.
- ?
Do not attach an exemplar to every request
That is another 100% export. Sample the interesting buckets. Without a trace backend, exemplars are dead strings. Sampling policy: previous lesson.
Exemplars
A Prometheus exemplar is a sampled trace attached to a high-resolution point — for histograms, to the observation that just incremented a bucket. Fields: trace id, often span id, optional attributes (component, region). Grafana’s exemplar dots on a heatmap or p99 panel are those ids.
Enable them on the histogram instrument (exemplar / exemplarLabels in client libraries), pass the id at observe(), scrape with a Prometheus that understands exemplars, and configure the datasource to Tempo/Jaeger. Missing backend = pretty dots that 404.
Attach exemplars on high-latency and error observations, not on every 20 ms 200. Combined with tail sampling, the id you stored is an id you actually kept.
Decisions
- 1
Request
- nextRED counters
- nextDuration histogram
- 2
RED counters
- nextPrometheus
- 3
Duration histogram
- nextobserve + exemplar trace id
- 4
observe + exemplar trace id
- nextPrometheus
- 5
Prometheus
- nextQuery
- ?
Query
- REDrate / error% / duration
- USEutil / sat / resource errors
- Exemplartrace id
- 7
rate / error% / duration
- nextBurn alert
- 8
util / sat / resource errors
- nextCapacity alert
- 9
trace id
- nextTempo / Jaeger waterfall
- 10
Tempo / Jaeger waterfall
- 11
Burn alert
- nextTempo / Jaeger waterfall
- 12
Capacity alert
Lesson map
Exemplars, RED & USE Methods
Exemplars pin a trace to a metric bucket. RED (rate, errors, duration) for request services; USE (utilization, saturation, errors) for resources — use both, not one dashboard religion.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Request"] b["RED counters"] c["Duration histogram"] d["observe + exemplar trace id"] a -->|Request to RED counters| b a -->|Request to Duration histogram| c c -->|Duration histogram| d
Step notes: 1 count rate and errors, 2 observe duration, 3 attach id on the tail, 4 scrape, 5 RED/USE panels, 6 click through to the trace that sampling kept.
RED (services)
Tom Wilkie’s RED method, aimed at request-driven services:
- Rate — requests per second (traffic).
rate(http_requests_total[5m]). - Errors — failed requests per second, or error ratio (better for SLIs). Failed / total. Define fail the same way as the availability SLI (usually 5xx + timeouts, not all 4xx).
- Duration — latency distribution: histogram, threshold ratio, p95/p99 on the dashboard. Not an average.
RED is the request SLO surface. Availability SLI ≈ 1 − error ratio. Latency SLI ≈ duration ≤ T. Burns consume those ratios. A “RED dashboard” that uses mean latency has failed the histogram lesson.
Categorize errors: 5xx vs 429 vs client 4xx or you will page on bad input. Separate CRITICAL routes from scrapes.
USE (resources)
USE is a checklist for every resource: CPU, memory, disk, network, thread pools, connection pools, locks, queues.
- Utilization — busy time / total time, or used / capacity (percent busy). A CPU at 80% busy is a fact, not yet a user outage.
- Saturation — extra work queued that cannot be served now: run queue, disk wait, pool waiters, load average as a rough CPU saturation signal. Saturation is not utilization. Mixing them is the classic false alert.
- Errors — retries, checksum fails, dropped packets, 5xx from the device, allocator failures.
Walk the checklist per resource during an incident: for each, name util, sat, err. That is the method — not a single “USE panel” with 200 lines.
USE tells you capacity and approaching cliffs. It does not replace product SLIs. A saturated thread pool will show up in RED duration next; catching sat first is the point.
Deep dive · Golden signals vs RED vs USE
Google’s four golden signals are latency, traffic, errors, saturation. RED is golden signals without saturation, scoped to a service. USE is saturation/util/errors scoped to a resource. You need the union. “We monitor golden signals” without saying where (service vs disk) is hand-waving. Put RED on the request path that feeds SLOs; put USE on every limiter (CPU, DB connections, queue depth).
Scoring both (in-memory)
I/O: snapshots of a service and a resource in, printed RED and USE scores. Thresholds are illustrative, not a vendor product. No scrape.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The first pair is why “CPU looks fine” is a failed interview. The second is why a green request dashboard is not permission to ignore saturation. Exemplars attach to the two slow/error samples, not to every 40 ms call.
Walk one incident with all three
Checkout 14.4× page, 02:14 UTC.
- RED. Error ratio on
class=CRITICALjumped from 0.04% to 8%. Rate (QPS) is flat — not a traffic spike. Duration p95 moved 120 ms → 1.8 s; the threshold-ratio at 300 ms is the number that actually burned. - USE. API CPU util 22%, run queue empty, disk wait 0, connection pool 15% used. Resources are fine. Do not scale the cluster. Look at the service.
- Heatmap + exemplar. The 1–2.5 s bucket is hot. An exemplar id opens a waterfall:
checkout.handle→inventory.reserve1.7 s, status ERROR. Tail sampling kept it because duration ≥ 500 ms (sampling). - Logs. Same
trace_idin the inventory log line: timeout talking to a new replica. Fix is a dependency, not a CPU limit.
Flip the story: RED still 99.95% good, p95 140 ms, but disk util 98%, sat waiters 40, device errors climbing. USE pages before RED. That is capacity. Exemplars can still show the occasional slow write; they are not the first panel.
Logs need trace_id / span_id injected or the triad is two silos plus a third. Exemplars without log injection still get you the waterfall; the log line is how you see the message the span did not store.
PromQL sketches (not a scrape)
# RED — error ratio (availability SLI miss)
sum(rate(http_requests_total{code=~"5..",class="CRITICAL"}[5m]))
/
sum(rate(http_requests_total{class="CRITICAL"}[5m]))
# RED — duration threshold ratio (good)
sum(rate(http_request_duration_seconds_bucket{le="0.3",class="CRITICAL"}[5m]))
/
sum(rate(http_request_duration_seconds_count{class="CRITICAL"}[5m]))
# USE — CPU util (busy / capacity); saturation is a *queue* metric, not this
avg(cpu_busy_ratio{job="api"})
# e.g. node_runq or pool waiters for sat; device error counters for errGrafana maps exemplars on the histogram query to Tempo. If the query is a recording-rule ratio, you may lose exemplar attachment — keep a raw histogram panel next to the burn.
Golden signals (latency, traffic, errors, saturation) = RED + the saturation leg of USE. Saying “we do golden signals” is naming the union; drawing where each signal lives (service vs disk) is the senior answer.
What to put on which panel
Keep the walls separate so a new hire can find them at 3am:
| Panel | Metrics | Click-through |
|---|---|---|
| Service health (RED) | QPS, error ratio, duration heatmap + threshold ratio | Exemplar → Tempo |
| Resource health (USE) | Per CPU / disk / nic / pool: util, queue, device errors | Host → logs; not traces first |
| SLO burn | 1 h / 5 m / 6 h / 3 d burn for each SLI class | Same heatmap + playbook |
| Trace search | Tail-kept errors and slow traces | Only after RED/USE named a system |
Instrument RED on the request histogram and counters you already need for SLIs — do not create a second parallel “RED-only” schema. Instrument USE from node exporters, DBMS views, and pool gauges. Exemplars live only on histograms/summaries that represent user-visible duration (or a rare custom bucket). CPU gauges do not get exemplars; there is no trace of a stolen millisecond of steal time.
Storage: exemplars are sampled and cheap compared with 100% traces. RED counters are cheap. Histograms are moderate. USE gauges are cheap. The bill you will actually get is trace ingest. Sampling policy is the cost control; exemplars are the index.
If you only remember one sentence for this last cluster page: RED tells you the user is hurting, USE tells you which resource is full, exemplars tell you which request to open.
Interview Q&A
What are Prometheus exemplars and why use them?
Answer
Sampled trace ids attached to a metric point (histogram bucket). They let you jump from a p99 spike or a burn page to a real waterfall. They are not labels. They need a tracing backend and a sampler that kept the id.
Explain RED.
Answer
Rate = traffic, Errors = failed requests (prefer a ratio), Duration = latency distribution. For request services this is the SLO surface. Duration is a histogram / threshold ratio, not a mean. Tom Wilkie / Grafana popularized the trio for microservices.
When do you prefer USE over RED?
Answer
You do not prefer one. USE is for resources (CPU, disk, pools, nics) where capacity and saturation matter before requests fail. RED is for the request path users feel. A database host gets USE; the query API gets RED; exemplars join a slow query bucket to a trace.
How do you enable exemplars in client libraries?
Answer
Use a histogram that supports them, declare exemplar label names (trace_id), and pass the map on observe(). Scrape with a Prometheus that stores exemplars; point Grafana at Tempo. Sampling still has to keep the span.
Saturation vs utilization?
Answer
Utilization is how busy the resource is (busy / capacity). Saturation is queued excess (run queue, waiters). A CPU can be 90% utilized with no queue, or 60% with a pile of waiters depending on thread model. Alerting on util alone misses queues; alerting on util as if it were sat causes false pages. Gregg’s USE method keeps them separate.
What alerts sit on RED vs USE?
Answer
RED: error ratio and latency-threshold burns (14.4× windows), not raw p99 flaps. USE: saturation above a wait SLO, resource error spikes, disk full. Do not put both on one panel; link them with the exemplar drill-down.
Can exemplars replace tail sampling?
Answer
No. An exemplar that points at a dropped trace is a broken link. Keep errors and slow traces in the sampler; attach exemplars on those same observations. Metrics stay complete; traces stay sampled.
Why not one dashboard with RED, USE, and every exemplar?
Answer
Visual noise. Service health panel (RED), resource panel (USE), heatmap with exemplar dots that open traces. Three surfaces, one incident workflow. Staff interviews want that split, not a 40-row Grafana screenshot.
RED without error categorization?
Answer
Lumping all non-2xx hides client 4xx vs your 5xx vs shedding 429. The availability SLI already had this fight. Reuse that predicate in the Errors leg or the burn will disagree with the RED panel.
Walk the cluster in one minute.
Answer
SLI = good/valid (design). Latency good = histogram ≤ T. Burn = error_rate / (1 − SLO) with 14.4× 1 h AND 5 m. When it pages: RED tells you the service, USE tells you the resource, an exemplar opens a trace that tail sampling kept.
Pitfalls
- Over-sampling exemplars — every request’s id is a second trace export; sample the tail.
- No trace backend — dots that go nowhere.
- trace_id as a metric label — cardinality bomb; use exemplar fields.
- RED with average duration — tail is gone again.
- USE saturation = utilization — false pages and missed queues.
- Dashboard overload — one panel for all three methods.
- RED-only shop — miss the disk that is about to take RED down.
- USE-only shop — miss the 500s on idle boxes (bad deploy, deadlock, dependency).
- Exemplars on a 100% dropped sampler — ids do not exist in Tempo.
Take a 14.4× availability page on checkout. Name: (1) the RED numbers you open first, (2) one USE check that could still be green, (3) the histogram bucket you click, (4) why that trace was kept. If hop 3 has no id, your exemplars are off. If hop 4 was head-sampled at 1% with no tail keep, you got unlucky — fix the sampler, not the dashboard.
Go Deeper
- Prometheus — Histogram / exemplars
- Grafana — Exemplars
- The RED method (Tom Wilkie / Grafana)
- USE method (Brendan Gregg)
- Google SRE Book — Monitoring Distributed Systems
- Cluster hub: SLOs, error budgets & tracing
Cheat sheet
Exemplar = trace id on a histogram observation (not a Prom label)
RED = Rate, Errors, Duration -> request services, SLIs, burns
USE = Utilization, Saturation, Errors -> resources (sat ≠ util)
Use both. Sample exemplars on errors/slow. Keep those traces.
SLI good/valid → burn 14.4× → heatmap → exemplar → tail-kept trace