Observability
Part 1 of 5 · Observability triadObservability Triad — Metrics, Logs & Distributed Tracing
Metrics, logs, and traces answer different questions on one incident; correlate with exemplars and trace_id.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Same checkout incident: which pillar first?
Prefer
Triad + exemplars (production default)
Detect on metrics, scope on a sampled trace, detail on structured logs for that trace_id. Best MTTR if labels stay bounded and sampling keeps errors.
- Burn / capacity / deploy p99 → metrics (histograms + exemplars).
- Critical path across services → traces (span waterfall).
- Forensics for one user or audit → structured logs with the same trace_id.
Alternative
One pillar as a religion
Metrics-only knows that error rate spiked, not why. Logs-only drowns in volume and cannot see fan-out. Traces-only without sampling or metrics explodes cost and still lacks SLO burn.
- Unstructured log.info(body) with PII is a DDoS on yourself.
- http_requests_total with user_id explodes cardinality.
- Sample 100% of prod, or 0.1% with no error bias — both fail.
Overview
The observability triad is metrics (aggregates over time), logs (discrete events), and traces (causal request trees across services). Events and exemplars glue them together. Senior engineers pick the right pillar first, then correlate with shared identifiers — not dump everything into one unstructured firehose.
The same production incident answers different questions depending on which pillar you open:
| Question | Metrics win | Logs win | Traces win |
|---|---|---|---|
| Are we burning the error budget right now? | Burn rate / RED | Noisy | Expensive |
| Why did this checkout fail for user X? | Averages hide | Structured error + attrs | If sampled |
| Which downstream hop added 800 ms? | Blames the service, not the span | Hard to order | Span waterfall |
| Do we need more pods? | Saturation / QPS | No | No |
| Did a deploy change p99? | Histograms + exemplars | Grep deploy id | Compare traces |
You should be able to:
- Name which pillar wins for budget burn, per-request forensics, and critical-path latency.
- Sketch detect → scope → detail → decide (metrics → trace → logs → page / ticket / freeze).
- Define an exemplar and why every log line carries
trace_id. - Contrast RED vs USE and say when each is the first dashboard.
- List the three classic anti-patterns interviewers love.
Decisions
- 1
1 Detect Metrics RED
- exemplar2 Scope Trace waterfall
- 2
2 Scope Trace waterfall
- trace_id3 Detail Logs by trace_id
- 3
3 Detail Logs by trace_id
- next4 Decide
- ?
4 Decide
- pageRemediate
- ticketBacklog
- budgetFreeze
- 5
Remediate
- 6
Backlog
- 7
Freeze
Lesson map
Observability Triad — Metrics, Logs & Distributed Tracing
Metrics, logs, and traces answer different questions on one incident; correlate with exemplars and trace_id.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB app["App"] prom["Prometheus"] tempo["Traces"] loki["Logs"] app -->|histogram plus| prom app -->|spans same| tempo app -->|structured log| loki prom -->|open exemplar| tempo tempo -->|filter by| loki
- 1
error_rate / latency / saturation → page or heatmap
Detect on metrics
Cheap if labels are bounded. Perfect for SLOs, alerting, and capacity. You know that something burned — not which request.
- 2
exemplar / alert link → span waterfall
Scope on a trace
The histogram bucket points at a trace id. Open the waterfall. That is the hop that added 800 ms — not "payments is slow on average."
- ?
same trace_id → page / ticket / freeze
Detail on structured logs
Filter logs by that id. Then decide: remediate now, file debt, or freeze features because the budget is gone.
Pillar deep-dives
Metrics — when they beat the alternatives
Pros: cheap cardinality if labels are bounded; perfect for SLOs, alerting, capacity; histograms plus exemplars link to traces.
Cons: lose the per-request story; high-cardinality labels explode memory; averages lie — use histograms.
Anti-pattern: http_requests_total{user_id=...} → metric explosion. Design labels on metric cardinality. Latency SLIs are threshold ratios, not means — that argument lives on histogram vs average.
Use metrics first for: error-budget burn, "do we need more pods?", "did this deploy move p99?", saturation. Those questions are SLI / SLO / budget and burn-rate alerts.
Logs — when they beat the alternatives
Pros: rich context, cheap to add fields; forensic detail; audit trails.
Cons: unstructured text is unqueryable at scale; logging every debug line in prod is a DDoS on yourself; without trace_id you cannot join to traces.
Anti-pattern: log.info(f"got {body}") with PII and no schema.
Use logs first when you already have a request id / user id / audit obligation. Do not start a Sev-1 by grepping "error" across 40 services.
Traces — when they beat the alternatives
Pros: causal graph across services; critical path; tail-latency debugging.
Cons: storage cost; head-based sampling misses rare failures unless combined with tail sampling; broken context propagation = orphan spans.
Anti-pattern: sample 100% of prod traffic with no budget; or sample 0.1% with no error bias.
W3C traceparent, parent-based flags, and head vs tail are a full lesson: trace context & sampling. The September 15 cluster's trace sampling strategies is the keep/drop policy deep dive — use both pages, do not paste them together.
Exemplars and structured correlation
An exemplar is a metric sample that points at a trace id (Prometheus exemplars / OpenTelemetry). Alert fires on p99 → click exemplar → open that slow trace → jump to logs with the same trace_id.
Without exemplars you correlate by timestamp across clock-skewed hosts. That is swivel-chair observability.
Sequence
- 1
App → Prometheus
histogram plus exemplar
- 2
App → Traces
spans same trace_id
- 3
App → Logs
structured log plus trace_id
- 4
Prometheus → Traces
open exemplar
- 5
Traces → Logs
filter by trace_id
The pairing of buckets ↔ traces ↔ RED/USE dashboards is exemplars, RED & USE. This hub's job is the workflow: detect, scope, detail.
RED vs USE
| Method | Letters | Use on | First question |
|---|---|---|---|
| RED (Tom Wilkie) | Rate, Errors, Duration | Request-driven services | Is the user SLI burning? |
| USE (Brendan Gregg) | Utilization, Saturation, Errors | CPU, disk, nic, queues, pools | Is a resource about to shed? |
Checkout 500s on idle CPUs is a RED incident. Disk at 100% with a rising wait queue is a USE incident that becomes RED in minutes. Alert on user SLIs; keep USE on investigation dashboards, not as the primary page. Full pairing: exemplars / RED / USE.
Correlation demo (in-memory)
I/O: fake checkout hops in → counters, JSON logs, and a span tree out. No OpenTelemetry SDK — this is the data model.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
You should see 200 and 502 counters, JSON lines that share a trace_id with the span tree, and an exemplar pointing at the last request's id. That is the join key — not the wall clock.
Pros / cons summary
| Approach | Pros | Cons | Prefer when |
|---|---|---|---|
| Metrics-first | Cheap alerts, SLOs | No causal path | Page-worthy symptoms |
| Logs-first | Detail | Cost + noise | Known request id / audit |
| Traces-first | Causal | Sampling bias / cost | Latency in a mesh |
| Triad + exemplars | Best MTTR | Needs discipline | Production default |
Interview Q&A
Metrics vs logs vs traces — one sentence each?
Answer
Metrics are aggregates for health and SLOs. Logs are timestamped events with context. Traces are distributed causal request trees (spans sharing one trace id).
Checkout p99 just jumped. Which pillar do you open first?
Answer
Metrics: confirm the burn / heatmap / which route. Then an exemplar into a slow trace for the hop. Then logs for that trace_id if you need the payload-level reason (card declined, timeout class). Opening logs first is how you drown.
Why not just log everything?
Answer
Cardinality of text, cost, PII, and you still lack cheap burn-rate alerts and critical-path views. Logs without a schema and trace_id cannot join to the other pillars.
What is an exemplar?
Answer
A link from a metric data point (usually a histogram bucket) to an example trace id so you jump from aggregate to instance. Grafana/Prometheus exemplars and OTel exemplars are the same idea.
Why put trace_id on every log line?
Answer
Joins pillars without brittle timestamp correlation across clock-skewed hosts. Filter logs by the id you got from the trace or the exemplar.
When do metrics beat traces for debugging?
Answer
Capacity, saturation, global error-budget burn, and "is this deploy worse than yesterday?" trend questions. Traces answer one request; metrics answer the fleet.
RED vs USE?
Answer
RED (Rate, Errors, Duration) for request-driven services and user SLIs. USE (Utilization, Saturation, Errors) for resources (CPU, queues, disks). Use both: a bad deploy can be RED-red / USE-green; a full disk is USE-red first.
Anti-patterns interviewers love?
Answer
Unstructured logs; unbounded metric labels (user_id, raw URL); 100% tracing in prod; 0.1% random sample with no error bias; alerts on CPU only while user SLIs burn unnoticed; three tools with no shared id.
What fails if you choose the wrong pillar?
Answer
Metrics-only: you know error rate spiked, not why. Logs-only: volume and no fan-out. Traces-only with no sampling budget: cost explodes and you still lack SLO burn signals.
How does this relate to SLOs and burn alerts?
Answer
The triad tells you where to look. SLIs/SLOs tell you whether the user promise is held. Multi-window burn alerts tell you when to page. Open SLIs, SLOs & error budgets next, then alerting.
Pitfalls
- Alerting on infrastructure metrics while user SLIs are green or red unnoticed.
- Trace every request with no sampling budget.
- Metric labels from unconstrained user input.
- Logs without schema /
trace_id/ severity. - Three tools, zero correlation — swivel-chair observability.
- Starting a Sev-1 in the log search box.
- Averages as a latency SLO (use histograms; see the SLOs cluster).
- Mesh-only traces with no app spans — network hops without business logic.
Write a 60-second incident script: checkout 502s after a payments deploy, p99 +800 ms, one VIP tweet. For each of detect / scope / detail, name the pillar, the identifier you carry, and the decision (page, ticket, freeze). If you cannot name the join key, you do not have a triad.
Go Deeper
- OpenTelemetry — Signals
- Google SRE Book — Monitoring Distributed Systems
- Prometheus — Histograms and summaries
- Grafana — Exemplars
- Grafana — The RED method
- Brendan Gregg — USE method
- YouTube: Observability: the present and future (Charity Majors)
- Companion cluster: SLOs, error budgets & tracing
Cheat sheet
Metrics = aggregates (SLOs, burn, capacity)
Logs = events with context (forensics, audit)
Traces = causal span trees (critical path)
Detect on metrics → exemplar → trace → logs(trace_id) → decide
RED = Rate, Errors, Duration (services)
USE = Utilization, Saturation, Errors (resources)
Anti-patterns: unstructured logs, user_id labels, 100% tracing