Observability
Part 2 of 5 · Observability triadSLIs, SLOs & Error Budgets
SLI is the experience ratio; SLO is the target; error budget is 1 minus SLO. Choose nines and journey vs request SLIs deliberately.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
What do you put the nines on?
Prefer
User-felt good / valid, SLO stricter than SLA
Pick the event first. Availability or latency-threshold ratios feed a budget. 99.9% / 28d is a starting contract, not a slogan.
- Journey SLIs for revenue paths; request SLIs for paging.
- One nine tighter shrinks budget about 10× — pay for it or don't set it.
- Burn rate drives freeze vs ship. Static 1% error alerts flap.
Alternative
CPU SLO, 100% target, or average latency
Infrastructure green while checkout dies. Zero budget means only reactive work. Averages hide the tail the user felt.
- SLA 99.5% copied as the internal SLO leaves no slack before credits.
- Four nines without capacity is a promise you will miss.
- Synthetics-only while real users fail is a vanity contract.
Overview
An SLI measures what users experience (usually good / valid events). An SLO is the target for that SLI over a window (for example 99.9% over 28 days). An error budget is 1 − SLO — the allowed unreliability you can spend on change, incidents, and risk. SLAs are external contracts with money attached; set SLOs stricter than SLAs.
Choosing 99.9% vs 99.99%, request-count vs user-journey SLIs, and how fast you burn the budget are product decisions. The math is there so the argument is not a vibe.
The triad hub told you which pillar to open. This lesson is the contract those metrics feed.
You should be able to:
- Separate SLI / SLO / SLA and put the SLO inside the SLA.
- Compute remaining event budget and burn rate from good/valid counts.
- Say what moving 99.9% → 99.99% costs (~10× less room to ship).
- Contrast journey vs request SLIs without pretending one replaces the other.
- Preview why paired long+short burn windows beat a static threshold.
Architecture
1 Define SLI
- 1
User journey
- nextgood
- nextvalid
- 2
good
- nextSLI = good/valid
- 3
valid
- 4
SLI = good/valid
- nextSLO 99.9% / 28d
2 Set SLO
- 5
SLO 99.9% / 28d
- nextBudget = 1 - SLO
- 6
Budget = 1 - SLO
- nextBudget left?
3 Spend / burn
- 7
Budget left?
- YesShip
- LowFreeze
- GonePage + review
- 8
Ship
- 9
Freeze
- 10
Page + review
Lesson map
SLIs, SLOs & Error Budgets
SLI is the experience ratio; SLO is the target; error budget is 1 minus SLO. Choose nines and journey vs request SLIs deliberately.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB biz["Business"] eng["Engineering"] cust["Customer"] biz -->|SLA 99.5% or| cust eng -->|burn report /| biz
SLI types and targets
| Choice | Pros | Cons | Fails when |
|---|---|---|---|
| Availability (success ratio) | Simple, alertable | Hides slow success | Timeouts return 200 with empty body |
| Latency threshold ratio (for example under 300 ms) | User-felt speed | Needs histograms | Using averages instead of percentiles |
| Freshness (data age) | Critical for feeds / ML | Harder instrumentation | Only monitoring API 5xx |
| Correctness | Catches silent bugs | Expensive validators | Trusting 2xx alone |
| Request-count SLI | High volume, stable | Bots skew | Ignoring multi-step journeys |
| User-journey SLI | Matches revenue path | Complex composition | Counting only edge QPS |
| 99.9% / 28d | ~40.3 min downtime budget | May feel "loose" | Treating it as the customer SLA |
| 99.99% / 28d | ~4.3 min budget | Freezes change; costly | Promising without capacity |
99.9% vs 99.99%: moving one nine shrinks budget about 10×. If you cannot afford the engineering cost of four nines, do not set it as an internal SLO.
Valid vs good, retries, probes, and endpoint classes are drilled on SLI design patterns. Latency as good = duration ≤ T is histogram vs average. This page keeps the target and the budget.
- 1
vanity → green board, dead checkout
SLO the CPU and the /healthz ping
Users did not buy your idle percentage. Availability that counts probes is a fiction. Average latency is the same class of lie.
- 2
Winner: good/valid on a user event, then nines you can pay for
Request SLIs page. Journey SLIs match login → search → checkout. 99.9% leaves ~40 minutes per 28 days. Write the freeze policy before the dashboard.
- ?
Spend the budget on purpose
Incidents and risky deploys consume it. Burn rate 1× is on pace. Fast burn pages — recipes on the alerting page, not a static 1% line.
Worked numbers
- 99.9% over 28 days → budget ≈ 40.3 minutes of downtime, or 0.1% of events.
- 99.9% over 30 days → ≈ 43.2 minutes. Memorize both; interviews mix windows.
- 3,000,000 requests, 99.9% success SLO → budget = 3,000 failed requests.
- One bad deploy causing 1,500 failures burns 50% of the monthly event budget in hours → freeze.
Burn rate 14.4× on a 30-day 99.9% budget exhausts the month in about 2 days if sustained (30 / 14.4 ≈ 2.08). That constant is derived on multi-window burn alerts; this page needs you fluent in the ratio.
SLI vs SLO vs SLA
Sequence
- 1
Business → Customer
SLA 99.5% or credit
- 2
Engineering → Engineering
SLO 99.9% internal
- 3
Engineering → Engineering
SLI = good/valid
- 4
Engineering → Business
burn report / freeze
Interview trap: reciting "SLA is the SLO we tell customers." No. SLA is contractual. SLO is the engineering target with slack. SLI feeds both.
Never 100%. Change causes outages; the dependency stack is not 100%; zero budget means only reactive work.
Event-based vs time-based availability. Event-based (good/valid requests) matches user attempts. Time-based (uptime minutes) can hide partial failures — half the fleet 500ing while "the service is up."
Error budget calculator (run this)
I/O: SLO + good/valid counts in → remaining failures, burn multiple, and a paired-window page boolean out.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Expect: budget fraction 0.001, remaining 700 after 300 failures in 1e6, burn 2× at 0.2% errors, page? True when both windows sit at 2% errors (20×), remaining 1500 after the 1500-of-3e6 outage, and four nines = 0.0001 budget.
Preview: multi-window burn alerts
Static error_rate > 1% flaps. The SRE workbook uses paired windows: page on 14.4× burn over 1 h and 5 m; ticket on ~1× over 3 d. Full comparison of static vs burn, AND-gates, and page-vs-ticket routing: alerting — multi-window burn rates. The 14.4 derivation and Prom recording-rule sketch: multi-window burn alerts in the older cluster.
Interview Q&A
SLI vs SLO vs SLA?
Answer
SLI is the measure (good / valid). SLO is the internal target over a window. SLA is the external contract with consequences. Set the SLO stricter than the SLA so you have slack before credits.
Why never a 100% SLO?
Answer
Change causes outages; the internet and dependency stack are not 100%; zero budget leaves only reactive work. Reliability trades against velocity — the budget makes that trade explicit.
Event-based vs time-based availability?
Answer
Event-based (good / valid requests) matches user attempts. Time-based (uptime minutes) can hide partial failures: the process is "up" while half of checkouts 500.
User journey vs request SLI?
Answer
Journey matches revenue (login → search → checkout) and is harder to compose. Request SLIs are stabler for paging. Use request SLIs to wake humans; use journey SLIs to argue about product reliability.
What is burn rate 14.4× on a 30-day 99.9% budget?
Answer
It exhausts the monthly budget in about 2 days if sustained (30 / 14.4 ≈ 2.08). Equivalently it spends ~2% of the 30-day budget in one hour. Page only when a short window agrees.
Pros of a looser SLO?
Answer
More feature velocity. Cons: more user pain and SLA credit risk if you loosen past the contract. If users are unhappy while you meet the SLO, the SLO is wrong — tighten it.
3,000,000 requests, 99.9% SLO, 1,500 failures. What now?
Answer
Budget is 3,000 events. You spent 50%. Policy should freeze risky deploys or shift capacity to reliability until the budget recovers — not "ship anyway, we still have half."
Why not SLO on CPU or disk?
Answer
Those are USE resource signals. Users feel request success and speed. Page on user SLIs; keep saturation on investigation dashboards. See the triad hub for RED vs USE.
99.9% vs 99.99% — say the cost?
Answer
About 10× less error budget (43 minutes → 4.3 minutes per 30 days, or 0.1% → 0.01% of events). Four nines freeze change unless you already paid for the architecture.
How do multi-window alerts beat a static 1% line?
Answer
Static ignores remaining budget and traffic shape. Paired long AND short burn windows page on fast exhaustion and ticket on slow leaks without flapping on a 2-minute blip. Next lesson in this cluster.
Pitfalls
- SLOs on CPU / disk instead of user journeys.
- Averaging latency as an SLI.
- Setting four nines without budget for the cost.
- No policy when budget hits 0% (ship anyway).
- Counting synthetic probes only while real users fail.
- Copying the SLA number as the SLO (no slack).
- One SLO per URL instead of a few request classes.
- Treating a dependency outage as "not our burn" while users fail.
Pick checkout. Name valid vs good, 28-day vs 30-day window, 99.9% vs 99.99% with the minute budget, request vs journey, and what happens at 50% burn. If the freeze rule is "we'll see," you do not have an SLO.
Go Deeper
- Google SRE Book — Service Level Objectives
- Google SRE Workbook — Implementing SLOs
- Google SRE Workbook — Alerting on SLOs
- OpenTelemetry — Metrics
- Companion: SLOs, error budgets & tracing · SLI design patterns
Cheat sheet
SLI = good / valid
SLO = internal target (stricter than SLA)
SLA = contract with money
Budget = 1 − SLO
Burn = error_rate / budget
99.9% / 28d ≈ 40.3 min 99.9% / 30d ≈ 43.2 min
99.99% / 28d ≈ 4.3 min one nine ≈ 10× less budget
Request SLI pages; journey SLI matches revenue
Never 100%. Never CPU as the user SLO.