Observability
Part 3 of 6 · SLOs & tracingMulti-Window Burn Alerts
Page on burn rate times window, not error rate ≥ SLO for 10 minutes. Multi-window alerts catch fast burns and slow leaks without paging on a single failed request.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
A multi-window burn alert watches how fast the error budget is being spent over several overlapping windows (classically 1 hour, 6 hours, 3 days). It fires only when the burn in a long window and a short window both exceed a threshold. Severity scales with how quickly the 30-day (or 28-day) budget would vanish.
This is the SRE workbook recipe from Alerting on SLOs. It is not a PromQL hobby: it is how you get paged for a fast outage and ticketed for a slow leak, without training the team to ignore 10-minute threshold flaps.
The SLI that feeds the burn must already be honest — SLI design patterns. Latency burns use a threshold-ratio, not an average — histograms.
Why multi-window AND-gates win
| Situation | “error rate ≥ SLO for 10 min” | Multi-window burn (winner) |
|---|---|---|
| One 500 in a quiet hour | Can look like 100% errors → page | Short window recovers; min-request guard |
| 30 s full outage | Often pages, barely spends budget | 1 h burn stays under 14.4× → no page |
| 5 min full outage at 99.9% | Maybe, depending on scrape | 14.4× 1 h and 5 m both trip → page |
| 0.15% errors for a week | Never trips 10 min vs 0.1% SLO | 1× 3 d / 6 h ticket on the leak |
| Graded response | Every alert feels like P1 | Page vs ticket from window × burn |
- 1
naive threshold → fatigue or missed leaks
Page when error rate ≥ SLO for 10 minutes
A quiet hour with one failure looks like a 1000× burn. A week-long 0.15% leak never trips 10 minutes. This is the alert people learn to ignore.
- 2
Winner: burn rate × window, long AND short
burn = error_rate / (1 − SLO). Classic page: 14.4× on 1 h AND 5 m (~2% of a 30-day budget). Second page: 6× on 6 h AND 30 m (~5%). Ticket: 1× on 3 d AND 6 h (~10%).
- ?
When it fires, jump with traces — do not raise a third threshold
Exemplars pin a trace to the histogram bucket. Sampling decides whether that trace exists. Those are later lessons. Here, get the page right first.
Error budget and burn rate
SLI = good / valid
SLO = target # 0.999 for 99.9%
error_budget = 1 − SLO # 0.001
error_rate = 1 − SLI # bad / valid
burn_rate = error_rate / error_budgetBurn 1× means: if this error rate continues, you exhaust the budget at the end of the SLO window. Burn 14.4× means you would exhaust it 14.4 times faster — in window / 14.4.
For 99.9% over 30 days:
- Window =
30 × 24 × 60 = 43,200minutes - Budget time =
0.001 × 43,200 = 43.2minutes of equivalent full outage - If the SLI is event-based, budget =
0.1%of requests, not wall clock. Same math: replace minutes with events.
Quantify incidents as percent of budget burned, not as “sev-2 vibes.”
The classic workbook table (99.9%, ~30-day window)
Page when both windows exceed the burn. Short window ≈ 1/12 of the long window so the alert resets quickly after burning stops (the 1 h condition would otherwise keep paging for an hour).
| Severity | Long window | Short window | Burn | ~Budget consumed if it stays true |
|---|---|---|---|---|
| Page | 1 hour | 5 minutes | 14.4× | ~2% |
| Page | 6 hours | 30 minutes | 6× | ~5% |
| Ticket | 3 days | 6 hours | 1× | ~10% |
Worked companions:
- 6× for 6 h: 5% of 43.2 min = 2.16 min equivalent.
2.16 / 360 = 0.006error rate →0.006 / 0.001 = 6. - 1× for 3 d: 3/30 of the window at budget pace → 10% of monthly budget.
For 99.95% or 99.99%, recompute; do not copy 14.4 blindly. Extreme 99.999% monthly ≈ 26 seconds of full outage — you cannot page your way out; you need canaries and blast-radius limits.
AND-gate, not three independent sirens
Decisions
- 1
SLI: good / valid
- nexterror_rate
- 2
error_rate
- nextburn = error_rate / (1 − SLO)
- 3
burn = error_rate / (1 − SLO)
- next1h burn
- next5m burn
- next6h burn
- next30m burn
- next3d burn
- next6h burn
- 4
1h burn
- next1h ≥ 14.4× AND 5m ≥ 14.4×
- 5
5m burn
- 6
6h burn
- next6h ≥ 6× AND 30m ≥ 6×
- 7
30m burn
- 8
3d burn
- next3d ≥ 1× AND 6h ≥ 1×
- 9
6h burn
- ?
1h ≥ 14.4× AND 5m ≥ 14.4×
- yesPage ~2% hour
- noStay quiet
- ?
6h ≥ 6× AND 30m ≥ 6×
- yesPage ~5% / 6h
- ?
3d ≥ 1× AND 6h ≥ 1×
- yesTicket leak
- 13
Page ~2% hour
- 14
Page ~5% / 6h
- 15
Ticket leak
- 16
Stay quiet
Lesson map
Multi-Window Burn Alerts
Page on burn rate times window, not error rate ≥ SLO for 10 minutes. Multi-window alerts catch fast burns and slow leaks without paging on a single failed request.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB d1["1h burn"] g["1h ≥ 14.4× AND 5m ≥ 14.4×"] f1["3d burn"] i["3d ≥ 1× AND 6h ≥ 1×"] d1 -->|1h burn to 1h ≥ 14.4× AND 5m ≥ 14.4×| g f1 -->|3d burn to 3d ≥ 1× AND 6h ≥ 1×| i
Why AND. A 30-second full outage:
- 5 m window:
30/300 = 0.10error rate → burn 100× (trips short). - 1 h window:
30/3600 ≈ 0.0083→ burn 8.3× (under 14.4). - AND fails → no page. Correct: 30 s is ~1.2% of a 43.2 min budget, not a 2% hour-burn.
A 5-minute full outage: 1 h error rate 5/60 ≈ 0.083 → burn 83×; 5 m burn is huge. AND fires. You just spent 5 / 43.2 ≈ 12% of monthly budget — worth a page.
After the incident stops, the 5 m window heals in minutes; the 1 h window would keep you up. The short leg is the reset.
Suppress overlapping severities: if 14.4× is paging, do not also open the 1× ticket for the same burn. Route page → human; ticket → next business day.
Conceptual rule (sketch only — not scraped here):
page_fast =
rate_1h > 14.4 * 0.001
AND
rate_5m > 14.4 * 0.001
page_mid =
rate_6h > 6 * 0.001
AND
rate_30m > 6 * 0.001
ticket =
rate_3d > 1 * 0.001
AND
rate_6h > 1 * 0.001Recording rules pre-aggregate ratio_rate1h and friends so Alertmanager is not computing 3-day rates on every eval. On-the-fly queries are fresher and more expensive at high cardinality.
Low traffic and other teeth
One failed request in a 10-request hour → error rate 0.1 → burn 100×. The 14.4× page will fire even though you spent almost no monthly budget. Mitigations: synthetic traffic so the denominator is never tiny; combine related services; a minimum-request guard; or a lower SLO / longer window for that class.
Without a written error-budget policy (freeze, reliability focus, escalation), the page is just noise. When budget is exhausted: stop risky ships. When in budget and users are happy: ship. When met but users hate you: tighten the SLI, not the alert math.
Deep dive · Burn vs PromQL, and why not to query Prometheus from the app
Some sketches divide “sum of rates” by budget × window_seconds twice or call Prometheus over HTTP from the app. The definition that matches the workbook is simply error_rate / (1 − SLO) on each window’s trailing ratio. You do not need a sidecar that GETs the query API. Recording rules plus Alertmanager are enough.
Honeycomb-style “budget rate” products are the same idea with a nicer UI: alert on how fast the remaining budget is disappearing, not on a static error-rate line.
Worked numbers (in-memory)
I/O: constants and synthetic error rates in, printed burns and AND-gate decisions out. No Prometheus.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Expected shape: budget minutes 43.2; 14.4× hour consumes 2%; 6× / 6 h consumes 5%; 1× / 3 d consumes 10%; 30 s blip does not page; 5 min outage does; the 0.15% leak tickets but does not 14.4× page; 1/10 requests looks like a 100× fire.
Windows, 28 days, and other SLOs
Integer weeks (28 days) keep weekend count stable in a rolling window. 28-day budget at 99.9% is 0.001 × 28 × 24 × 60 ≈ 40.3 minutes. Recompute 14.4: 2% of 40.3 min in one hour is 0.02 × (28 × 24) = 13.44×, not 14.4. Teams that say “14.4” on a 28-day SLO are approximating; say the approximation out loud or derive the constant from target_fraction × window_hours.
Other useful derivations (always: burn = (budget_fraction_you_want_to_spend) × (window / alert_long_window)):
| SLO | Window | Budget | 2% of budget in 1 h |
|---|---|---|---|
| 99.9% | 30 d | 43.2 min | 14.4× |
| 99.9% | 28 d | 40.3 min | 13.44× |
| 99.95% | 30 d | 21.6 min | 28.8× |
| 99.99% | 30 d | 4.32 min | 144× |
99.99% paging at 144× means a ~25-second full outage in an hour trips the “2% of budget” page. That is how tight 99.99% is — not a PromQL trick.
Inhibit / suppress. Alertmanager should inhibit the 1× ticket while a 14.4× or 6× page is firing for the same SLI, and inhibit 6× while 14.4× is already paging if they represent the same incident. Otherwise on-call gets three issues for one outage.
Recording rules (sketch): emit slo:burn_rate:ratio_rate1h, …5m, …6h, …30m, …3d as error_ratio / budget. Alert on those time series. Evaluating a 3-day rate() on every rule group without recording is how you learn about Prometheus CPU.
Latency-threshold SLIs use the same table: error_rate is the fraction slower than T (or failed). Do not invent a second 10-minute p99 alert and call it the SLO page.
Policy teeth. A page without freeze rules is a dashboard with sound. When 14.4× fires, the playbook is: halt risky deploys, find the burn with exemplars, restore the SLI, then spend remaining budget on purpose. When the 1× ticket fires, you still have ~90% of the window — schedule reliability work this week, do not wake someone at 3am.
Interview Q&A
What is an error budget, and why alert on it?
Answer
The allowed SLO miss over the window: 1 − SLO. For 99.9% / 30 days that is 43.2 minutes or 0.1% of events. Alerting on budget consumption ties pages to user-facing reliability, not to an arbitrary error-rate line.
Write the burn-rate formula.
Answer
burn_rate = error_rate / (1 − SLO). Burn 1× spends the budget exactly over the SLO window. Burn 14.4× spends it 14.4 times faster. Do not invent a second division by window seconds unless you are converting units — the ratio of rates is already dimensionless.
Why multiple windows?
Answer
A short window catches spikes; a medium window catches sustained moderate pain; a long window catches slow leaks. Together they grade severity (page vs ticket) and avoid both missed incidents and alert fatigue.
Derive 14.4× for 99.9% / 30 days.
Answer
Page at ~2% of budget in one hour. Budget = 43.2 min, 2% = 0.864 min, 0.864/60 = 0.0144 error rate, 0.0144/0.001 = 14.4. Equivalently 0.02 × 720 hours = 14.4. AND with 5 minutes (1/12 of 1 h) so the page clears when burning stops.
30-second full outage — do you page the 14.4× rule?
Answer
No. Five-minute burn is huge, one-hour burn is ~8.3×, under 14.4. The AND-gate is the feature: you do not page on a blip that spent ~1% of monthly budget. A five-minute full outage does page.
How is this different from latency > 500 ms for 5 minutes?
Answer
That alert ignores the SLO window and the remaining budget. A brief p99 spike may spend almost no budget; a 0.2% error leak may never trip 500 ms. Burn alerts are budget-centric and reuse the same windows for availability and latency-threshold SLIs.
1 fail in 10 quiet requests — what happens?
Answer
Error rate 10%, burn 100×, you page. Mitigate with synthetics, aggregation across small services, or a minimum-request guard. Low traffic is the standard caveat on this recipe.
Recording rules vs computing burn on the fly?
Answer
Recording rules make 1 h / 6 h / 3 d ratios cheap and stable for Alertmanager; they add storage and one eval delay. On-the-fly queries are fresher and can melt Prometheus at high cardinality. Production: record the ratios, alert on the records.
Error budget gone mid-quarter — ship the launch?
Answer
Per policy, usually no for risky changes. Freeze or spend eng time on reliability until the budget recovers. Exceptions go through a written escalation with product, not silent heroics. Alerts without a freeze policy have no teeth.
What severity for 1 h at 3× vs 3 d at 0.8×?
Answer
3× on one hour is under 14.4× — no fast page (watch the 6× / 1× rows). 0.8× on three days is under the 1× ticket. Neither fires. If 3 d is 1.2× and 6 h is 1.2×, you ticket a leak. Map severity from the table, not from gut P1.
Why short window ≈ 1/12 of long?
Answer
So the compound alert resets about 1/12 of the long window after the burn stops. Without the short leg, a 1 h condition keeps paging long after the incident is over. 1 h → 5 m, 6 h → 30 m, 3 d → 6 h.
Pitfalls
- Alerting on error rate ≥ (1 − SLO) for 10 minutes — floods pages, misses leaks.
- No AND-gate — every 5-minute spike pages; on-call dies.
- Copied 14.4× onto a 99.99% SLO — the constant is for ~2% of a 30-day 99.9% budget; recompute.
- Low-traffic paging — one fail at night looks like 1000×; add floors or synthetics.
- Overlapping page + ticket — suppress the weaker severity for the same burn.
- Thresholds too aggressive — paging on 2× / 5 m trains people to ignore the app.
- Wrong SLI feeding the burn — retries and
/healthzin valid; fix the indicator first. - No budget policy — you paged, then shipped the launch anyway.
- 99.999% heroics — 26 seconds of outage exhaust the month; need progressive delivery, not a tighter PromQL.
For 99.9% / 30 d, write budget minutes (43.2), 14.4× error rate (1.44%), 2% of budget in minutes (0.864), and AND a 5-minute window. Then pick a 45-second outage and a 2-hour 0.2% leak and say which row fires. If you need a cluster to do the arithmetic, you do not own the alert yet.
Go Deeper
- Google SRE Workbook — Alerting on SLOs (the 14.4× recipe)
- Google SRE Workbook — Implementing SLOs
- Honeycomb — SLOs get budget rate alerts
- Grafana Cloud — SLO and burn-rate alerting
- Previous: SLI design patterns · Next: Histogram vs average latency
Cheat sheet
budget = 1 − SLO
burn = error_rate / budget
99.9% / 30d budget ≈ 43.2 min
Page 14.4× for 1h AND 5m (~2% budget)
Page 6× for 6h AND 30m (~5% budget)
Ticket 1× for 3d AND 6h (~10% budget)
Do not page on “error rate ≥ SLO for 10 minutes”