Observability
Part 2 of 6 · SLOs & tracingSLI Design Patterns
An SLI is good/total of something users feel. Availability, latency (threshold ratio), freshness, quality, and journey SLIs each fail differently if you count retries or averages.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
An SLI (Service Level Indicator) is a quantitative measure of reliability from the consumer’s perspective. The SRE shape is a ratio over a window:
SLI = good events / valid events # 0–1, report as a percentThe SLO is the target for that ratio (99.9% over 28 or 30 days). The SLA is the external contract. This lesson is the indicator: what you count, what you drop, and which pattern matches the failure users feel. The hub covers the contract; burn alerts consume this ratio.
If you cannot name the event and the predicate for good, you are not ready to pick 99.9%. Interviewers stop you there.
Why the ratio wins over a vibe metric
| Pattern | What it measures | When it wins | How it lies |
|---|---|---|---|
| good / valid (winner) | Fraction of user-meaningful attempts that met the promise | Budgets, burn alerts, freeze policy | Only if valid is polluted (probes, retries, wrong 4xx) |
Uptime ping to /healthz | Process is up | Almost never for product SLOs | Checkout can be 100% 5xx while health is green |
| Average latency | Mean of a distribution | Capacity trivia | Hides the tail; see histograms |
| Raw error count | Absolute fails | Debugging a spike | 10 fails at noon ≠ 10 fails at 3am; no budget |
- 1
vanity metric → green board, dead checkout
Ping /healthz and call it availability
Probes are not users. Mixing them into valid events inflates the ratio. CPU "looks fine" is the same class of mistake.
- 2
Winner: user-centric good / valid, one class at a time
Name the event (checkout HTTP, homepage paint, pipeline record). Drop invalid. Count good. Bucket CRITICAL vs HIGH_SLOW. Latency uses a threshold ratio, not an average.
- ?
Then attach a window, SLO, and burn policy
Same ratio feeds 14.4× pages and 1× tickets. If the SLI is wrong, every alert downstream is theater. Journey SLIs need traces; freshness needs timestamps — pick the pattern that matches the pain.
Good events vs valid events
Two counters, always:
- Valid — attempts that should count toward the promise. Successful or failed user-facing calls. This is the denominator.
- Good — the subset that met the promise (2xx/3xx, or latency under the threshold, or data younger than T). This is the numerator.
Request success ratio = good / valid over the measurement window (fixed 5-minute scrape for alerting; rolling ~4-week window for the SLO).
Invalid is not “failed.” Invalid is “this row should not be in the fraction.” Typical drops:
- Load-balancer health checks and synthetics you do not want in the product SLO (or keep them in a separate probe SLI).
- Auth bot 401/403 if product policy says those are client mistakes, not your unavailability.
- Internal scrapes, admin, and dark-launch traffic.
- Duplicate retries of the same user attempt — see below.
Decisions
- 1
Incoming request
- nextValid user attempt?
- ?
Valid user attempt?
- NoDrop from denominator
- YesStart timer
- 3
Drop from denominator
- nextExport counters + histogram
- 4
Start timer
- nextExecute business logic
- 5
Execute business logic
- nextGood by the SLI predicate?
- ?
Good by the SLI predicate?
- YesNumerator +1
- NoBad: burns budget
- 7
Numerator +1
- nextEmit latency histogram too
- 8
Bad: burns budget
- nextEmit latency histogram too
- 9
Emit latency histogram too
- nextExport counters + histogram
- 10
Export counters + histogram
Lesson map
SLI Design Patterns
An SLI is good/total of something users feel. Availability, latency (threshold ratio), freshness, quality, and journey SLIs each fail differently if you count retries or averages.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Incoming request"] b["Valid user attempt?"] d["Drop from denominator"] c["Start timer"] a -->|Incoming request to Valid user attempt?| b b -->|No| d b -->|Yes| c
Step notes: 1 validate (denominator), 2 time the work, 3 apply the success predicate, 4 observe latency whether good or bad, 5 export. The histogram is not the availability SLI; it feeds a latency-threshold SLI.
Availability, latency, freshness, quality
Availability (request-driven APIs)
Good = responses that are successful by your contract — often non-5xx, sometimes “business-OK” (HTTP 200 with paid: false may still be a failed charge). Valid = user-facing attempts that reached a complete response or timed out (timeouts are bad, not invalid).
Do not invent a unique SLO per endpoint. Bucket:
- CRITICAL — login, checkout, pay (tightest).
- HIGH_FAST — interactive reads.
- HIGH_SLOW — exports, reports (looser latency, still high availability).
- LOW / NO_SLO — polls, debug, dark launches.
Latency (threshold ratio, not a percentile nested in a percentile)
“p99 under 300 ms” is a useful dashboard. It is a clumsy budget. The SLI that composes with availability is:
good = requests with duration ≤ T
SLI = good / valid
SLO = that ratio ≥ 99% (example)Put an exact histogram bucket on T. Missing le="0.3" biases the ratio. Averages hide the tail — full worked numbers on histogram vs average.
Freshness (pipelines, caches, replicas)
Users care that the data they read is young enough, not that the job “ran.” Emit age at the point of consumption (now − record_ts). Good = age under threshold (for example 5 s for a session cache, 15 min for a warehouse tile). Valid = reads (or records) that should have been fresh.
You need trustworthy timestamps: NTP or the producer’s event time, not a clock that drifted 30 s on one AZ. OpenTelemetry trace IDs do not replace event time; they correlate the read to the writer.
Quality / correctness
Graceful degradation is still a worse user experience: stale catalog, generic avatar, approximate search. Quality SLI = undegraded responses / valid. Correctness = schema or business-rule pass rate (golden synthetic, independent checker). Expensive to compute; use on money paths and ML scores, not on every static asset.
User-journey (composite) SLIs
Checkout is not “payments 99.9%” and “inventory 99.9%” multiplied in your head. A journey SLI is one end-to-end event: the user started checkout and received a paid confirmation within T, with no manual retry. That needs a shared trace id (W3C traceparent) so you do not double-count hops. Per-request SLIs stay for component contracts with dependencies.
| Pattern | Scope | Typical metric | Cost | Use when |
|---|---|---|---|---|
| Per-request | One endpoint | Success ratio, threshold latency | Low | Component SLOs, one service you own |
| User journey | Several services | Composite success + overall latency | Tracing | Checkout, login, “publish post” |
| Freshness | Data at read time | Age ≤ T ratio | Event timestamps | CDN, cache, replication, ETL |
| Correctness | Business predicate | Checker pass rate | High | Ledger, pricing, model output |
| Hybrid | Availability AND latency | Both must be good to increment numerator | Slightly pickier | SLAs that cap both errors and delay |
A hybrid “good” means succeeded and fast. That is stricter than two separate SLOs; say so, or you will double-punish the budget. Many teams keep two SLIs (availability, latency) on the same valid stream and two budgets.
Where you measure
| Placement | Upside | Gap |
|---|---|---|
| App logs / metrics | Cheap, detailed codes | Misses requests that never reach the process |
| Load balancer / edge | Closer to the user | May miss client-side failures and DNS |
| Black-box probes | Catches total outages | Sparse paths; easy to overfit to the probe URL |
| Client instrumentation | Closest to UX | Needs client code and a reliable pipeline |
Separate SLI specification (“homepage becomes interactive under 100 ms”) from SLI implementation (edge http_request_duration under 100 ms). Document the gap: server-side misses DNS and CDN failures. Start at the LB, iterate toward the user.
Measurement windows: short (1–5 min) for burn alerts; long (~28–30 days, integer weeks so weekend count stays stable) for the SLO. Do not report a 30-day SLO off a 5-minute scrape without recording rules.
Deep dive · Specification vs implementation, and time skew
The workbook’s point: product language is the spec; Prometheus/OTel is an implementation. If the spec is “checkout completes,” and the implementation is “HTTP 200 from checkout-api,” you will miss payments that 200 with a failed capture, and you will miss edge 502s that never hit the app. Write both sentences.
Freshness SLIs die on clock skew. If producer and consumer disagree by more than the freshness threshold, you will page on a healthy pipeline or miss a stalled one. Prefer event time in the payload, plus a bounded allowed skew. Do not use span duration as freshness — duration is latency, not age of data.
Worked ratios (in-memory)
No scrape, no HTTP. A tiny request log in memory: each row is one attempt with flags for retry, probe, latency, payload age, and degradation. The functions below are the I/O contract: in → classified events, out → printed SLI ratios.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Read the output: the probe and the successful retry make a naive success ratio look healthier than the user journey. The 2.5 s call fails the latency SLI while still “available.” The stale degraded 200 fails freshness and quality. That is the whole point of several SLIs on the same valid stream — one number cannot name every failure mode.
Instrument the same split in production with two counters (requests_valid, requests_good) plus a histogram whose buckets include the latency threshold. OTel and Prometheus are exporters; they do not choose the predicate for you.
Pipelines, coverage, durability
Request APIs are not the only consumers. The workbook’s other SLI families show up in interviews as soon as you mention Kafka, a warehouse, or object storage.
Coverage. Expected records vs successfully processed records. A consumer that “is up” at 99.9% while dropping 8% of messages is a coverage miss, not an availability miss. Valid = records the contract said you would handle (offset high-water, files in the landing bucket). Good = those that made it to the sink with a checksum.
Durability. Readable records / written records. Users often care about today’s objects, not the ten-year archive: say which population you count. A 99.999% durability SLO on “all objects ever” can hide a bad day of new writes.
Correctness vs freshness on the same pipeline. A job can be on time and wrong (schema drift) or late and right (catch-up). Two SLIs, two budgets. Do not OR them into one “pipeline healthy” boolean — that boolean cannot feed a burn rate.
Worked availability still applies: 99.9% over 30 days is 43.2 minutes of equivalent full outage, or 0.1% of events. If checkout sees 3,000,000 attempts in the window, the budget is 3,000 failed attempts. Pick the unit that matches the SLI (time for “are we up,” events for “did this request work”). The hub computes both; this page decides which events exist.
Instrumentation without a scrape loop
A wrapper around a handler is enough conceptually:
- Decide valid (user attempt, not probe, not extra retry).
- Start a monotonic timer.
- Run the work.
- Apply the good predicate (status contract, and/or duration ≤ T, and/or age ≤ T).
- Observe the histogram always (failed calls have latency too).
- Export via OTel or Prometheus — the backend is not the SLI.
Reuse the wrapper across endpoints; put class (CRITICAL / HIGH_SLOW) on labels, not a unique SLO name per URL. Externalize T and the class map so a latency SLO change is config, not a binary. High-frequency debug endpoints can sample debug attributes; the SLI counters themselves stay on 100% of valid events — sampling traces is a different knob (trace sampling).
Client-side retries plus idempotency keys: the user journey is still one attempt. Count the journey once. The retry storm belongs on a dependency RED panel, not as extra good/valid pairs.
Interview Q&A
What is the difference between an SLI and an SLO?
Answer
An SLI is the measurement: good / valid over a window. An SLO is the internal target for that measurement (99.9% over 30 days). An SLA is the external contract with money attached. Engineers manage to SLOs; legal owns SLAs. Set the SLO stricter than the SLA.
Why measure good events vs valid events separately?
Answer
Valid is the denominator (attempts that should count). Good is the numerator (attempts that met the promise). If you only count successes, you hide outages. If you stuff probes and retries into valid, you inflate the ratio. The split is how you keep the fraction honest.
How would you design an SLI for data freshness?
Answer
Measure age at consumption: now minus the record’s event time. Good = age under a threshold users would accept. Valid = reads (or records) that were supposed to be fresh. Ratio = good / valid. Clock skew larger than the threshold will page you; propagate event time, do not use span duration as age.
Why is a latency percentile a poor budget SLI by itself?
Answer
Nested “p99 under 300 ms in 99.9% of windows” does not yield one good / total you can burn-alert. A threshold-ratio SLI (fraction of requests faster than 300 ms) is the same shape as availability, so 14.4× / 6× / 1× windows apply unchanged. Put an exact histogram bucket on 300 ms.
Per-request SLI vs user-journey SLI?
Answer
Per-request is cheap and blind to downstream: payments can be 99.9% while checkout fails on inventory. Journey SLIs capture end-to-end success and need a trace id so hops are one attempt. Use per-request for component contracts; use a journey SLI for the product promise.
How does an error budget use this SLI?
Answer
Budget = 1 − SLO, in events or in time. Burn rate = actual error rate / budget. A wrong SLI (healthz, retries as good) feeds a wrong budget, so freeze/launch policy becomes theater. Details and the 14.4× AND-gate: multi-window burn alerts.
Should 4xx count as availability failures?
Answer
Usually no for bad-client 400/401/404 if those are the client’s mistake. Yes for 429 if you are shedding load the user felt, and for 409/422 if they mean “we could not complete the business action.” Write the rule. Mixing health-check 200s into valid is worse than a slightly wrong 4xx policy.
Where should a product SLO be measured?
Answer
Closest practical point to the user: client or edge. Component SLOs live at the service. Document the implementation gap (server-side misses DNS/CDN). Start cheap at the load balancer, then move out.
Why bucket request classes instead of one SLO per endpoint?
Answer
Cardinality and politics. Fifty SLOs will not be maintained; checkout and a metrics scrape should not share a 300 ms budget. CRITICAL / HIGH_FAST / HIGH_SLOW / LOW is enough for most APIs.
What goes wrong if you count every HTTP attempt including retries?
Answer
You either mint extra good events when a retry finally works, or you inflate valid with failures the user already experienced as one wait. Count user attempts. Tag retries separately (retries_total) for dependency health, not for the product SLI numerator.
Pitfalls
- Counting retries as good — one painful checkout becomes a success after attempt 4.
- Health-check inflation —
/healthz200s in valid while checkout is dead. - Average latency as an SLI — tail disappears; cannot aggregate means of means.
- Window mismatch — alerting on 30 days of scrape noise, or reporting SLO off 5 minutes.
- Clock skew on freshness — threshold of 5 s with 8 s NTP drift.
- Hard-coded thresholds in binaries — T and class maps should be config, not a redeploy.
- Over-instrumenting every URL — high-cardinality labels (user id) explode series; sample hot debug paths, keep SLIs on low-cardinality class/route.
- Mixed units — milliseconds in one service, seconds in another; standardize on seconds for OTel histograms.
- One SLO for every endpoint — unmaintainable; bucket classes.
- Treating a dependency outage as “not our SLI” — users do not care whose queue failed; product SLIs stay on user impact.
Pick one journey (checkout). Write four sentences: the event; what is valid; what is good; what you drop (retries, probes, 4xx policy). Then name T for latency and T for freshness if they apply. If any sentence needs a dashboard to exist first, you started at the wrong end.
Go Deeper
- Google SRE Book — Monitoring Distributed Systems
- Google SRE Workbook — Implementing SLOs
- Google SRE Book — Data Integrity (freshness / correctness cousins)
- OpenTelemetry — Metrics
- Prometheus — Histograms and summaries
- Next in cluster: Multi-window burn alerts
Cheat sheet
SLI = good / valid (user attempt, not probe, not extra retry)
Availability: success by contract / valid
Latency: duration ≤ T / valid (histogram bucket on T)
Freshness: age ≤ T / valid (event time, not span duration)
Quality: undegraded / valid
Journey: end-to-end good / valid (needs trace id)
Then: SLO target, budget = 1 − SLO, multi-window burn