Observability
Part 5 of 5 · Observability triadAlerting — Multi-Window Burn Rates vs Static Thresholds
Static thresholds flap. Multi-window multi-burn-rate pages on fast burn and tickets on slow burn.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
What wakes a human?
Prefer
Multi-window multi-burn-rate on a user SLI
Budget-aware. Short window catches meltdown; long window catches leaks. AND-gate kills flaps. Page vs ticket by burn multiple.
- Tied to the SLO you actually sold internally.
- Recording rules keep 3-day rates cheap.
- Inhibit page > ticket for the same SLI during one outage.
Alternative
Static error_rate > 1% for 5 minutes
Ignores remaining budget and traffic shape. Pages on canary blips. Misses a week-long 0.15% leak. One size of P1.
- Quiet hour + one failure looks like 50% errors.
- Absolute RPS thresholds ignore diurnal patterns.
- Raw CPU spike is not an SLI — do not page it.
Overview
Static threshold alerts (error_rate > 1% for 5 minutes) are easy and often wrong: they flap on traffic dips, ignore your error budget, and page for symptoms that are still on pace. Multi-window multi-burn-rate alerts (Google SRE Workbook) detect both fast burns (page now) and slow burns (ticket), using paired long/short windows so a brief blip alone does not page.
Severity routing and alert fatigue control are as important as the math. The SLI / SLO page defined the budget; this lesson decides who wakes up.
You should be able to:
- Contrast static vs burn on flap, fast outage, and slow leak.
- Write the AND-gate in one sentence.
- Map 14.4× / 6× / 1× to page vs ticket for 99.9% / 30d.
- Route symptomatic (user SLI) vs causal (CPU) signals.
- Name inhibit rules, recording rules, and "alert without a runbook."
| Dimension | Static threshold | Multi-window burn rate |
|---|---|---|
| Tied to SLO? | Usually no | Yes — budget-aware |
| Fast outage (minutes) | May miss if threshold high | Short window catches |
| Slow leak (days) | May never fire | Long window catches |
| Flapping / noise | High | Lower (AND of windows) |
| Severity | One size | Page vs ticket by burn |
| Complexity | Low | Medium (recording rules) |
What fails if you choose static only: on-call pages on every deploy blip; or never pages while budget silently drains over a week.
Architecture
1 Inputs
- 1
SLI good/valid
- nextBurn = err / (1-SLO)
- 2
Burn = err / (1-SLO)
- next1h AND 5m
- next6h AND 30m
- next3d AND 6h
2 Windows
- 3
1h AND 5m
- 14.4x bothPage P1
- 4
6h AND 30m
- 6x bothPage P2
- 5
3d AND 6h
- 1x bothTicket
3 Decisions
- 6
Page P1
- 7
Page P2
- 8
Ticket
Lesson map
Alerting — Multi-Window Burn Rates vs Static Thresholds
Static thresholds flap. Multi-window multi-burn-rate pages on fast burn and tickets on slow burn.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB prom["Prometheus"] am["Alertmanager"] oc["On-call"] tick["Tickets"] prom -->|Burn14_4 firing| am am -->|page| oc am -->|ticket slow burn| tick
- 1
easy PromQL → fatigue or silent drain
Static 1% for 5 minutes
Canary 2% for three minutes pages. A 0.15% leak for a week never pages. Traffic dip + one 500 looks like a meltdown.
- 2
Winner: paired windows on burn multiples
Page 14.4× (1 h AND 5 m) and 6× (6 h AND 30 m). Ticket 1× (3 d AND 6 h). Short ≈ 1/12 of long so the page clears when burning stops.
- ?
Route and inhibit
Fast burn → page. Slow burn → ticket. Disk 90% → ticket. CPU spike → dashboard. Inhibit weaker alerts for the same SLI.
Classic Google parameters (30d, 99.9% SLO)
For a 30-day window, budget ≈ 0.1%. Approximate pairs:
| Severity | Long window | Short window | Burn multiple | Budget consumed if sustained |
|---|---|---|---|---|
| Page | 1 h | 5 m | 14.4× | ~2% of 30d budget in 1 h |
| Page | 6 h | 30 m | 6× | ~5% in 6 h |
| Ticket | 3 d | 6 h | 1× | on-pace consumption |
Exact workbook tables vary by SLO and window; the pattern is: AND a long and short window so you don't page on a 2-minute spike alone, and don't wait forever for a slow burn.
Do not copy 14.4× onto a 99.99% SLO. Recompute. The derivation (2% of 43.2 minutes in one hour) is multi-window burn alerts.
Sequence
- 1
Prometheus → Prometheus
recording rules per window
- 2
Prometheus → Alertmanager
Burn14_4 firing
- 3
Alertmanager → On-call
page
- 4
Alertmanager → Tickets
ticket slow burn
Why static flaps
- Traffic drops → fewer requests → one failure looks like 50% error rate.
- Canary deploy → brief 2% errors → page even if budget impact is tiny.
- Absolute RPS thresholds ignore diurnal patterns.
Burn-rate alerts scale with traffic because they use ratios against the budget. Still: one fail in ten quiet requests is a 100× burn — add synthetics or a minimum-request guard. That caveat is shared with the older burn page.
Recording rules: precompute job:slo_burn:1h / :5m / :3d so alert eval is cheap and dashboards share definitions. Alerting on raw high-card series is how you learn cardinality the hard way.
When the page fires, jump with exemplars and traces from the triad hub — do not raise a third static threshold.
Burn windows (run this)
I/O: SLO + good/valid per window → PAGE_FAST / PAGE_SLOWISH / TICKET / OK.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Expect spike_only → OK (5 m is on fire, 1 h is not), fast_burn → PAGE_FAST, and a ~1× 3 d / 6 h case → TICKET. That AND-gate is the whole trick.
Alert fatigue and severity routing
| Signal | Route | Why |
|---|---|---|
| Fast burn 14.4× | Page | User-visible meltdown |
| Medium 6× | Page / high Pri | Serious but not instant death |
| Slow 1× | Ticket | Trend / reliability debt |
| Disk 90% foresight | Ticket | Not yet user-facing |
| Raw CPU spike | Often none | Not an SLI |
Symptomatic vs causal: page on user SLIs; use causal metrics (CPU, saturation) as investigation dashboards, not primary pages. That is the RED vs USE split from the triad hub.
Reduce fatigue: fewer pages, SLO-based, inhibit related alerts, require runbooks, delete alerts nobody has acted on in a quarter.
Interview Q&A
Why AND two windows?
Answer
Short alone flaps; long alone is slow. Together you page on sustained badness. The short leg also clears the alert about 1/12 of the long window after burning stops.
What is burn rate 1×?
Answer
Consuming budget exactly on pace for the SLO window. A 1× ticket on 3 d / 6 h means "this leak will exhaust the month if we ignore it" — not "wake someone at 3am."
Problem with a static 1% error alert?
Answer
Ignores budget and traffic shape. Pages when remaining budget is fine (canary blip, quiet-hour single fail) and misses a slow leak under 1%.
Page vs ticket?
Answer
Page when the budget will exhaust soon without human action. Ticket for slow leaks. Map from the window × multiple table, not from gut P1.
Recording rules?
Answer
Precompute burn rates so alert eval is cheap and dashboards share definitions. Three-day rate() on every eval cycle without recording is a Prometheus incident.
How do you reduce fatigue?
Answer
Fewer pages, SLO-based, inhibit related alerts, require runbooks, delete never-acted alerts. Do not page on CPU.
5-minute spike, healthy 1-hour window — page?
Answer
No. The AND-gate fails. That is the point of the playground spike_only case. A five-minute full outage at 99.9% usually does page because both windows trip 14.4×.
Copy 14.4× onto 99.99% / 30d?
Answer
No. 14.4× is ~2% of a 0.1% budget in one hour. Four nines have a 0.01% budget; the equivalent multiple is ~144×. Recompute or you will under-page.
What belongs in Alertmanager inhibit?
Answer
Inhibit the 1× ticket while 14.4× or 6× is already paging for the same SLI. Inhibit 6× while 14.4× is firing if they are the same incident. One outage, one wake-up.
Where does this sit versus the September 15 burn page?
Answer
That page derives 14.4× and walks PromQL. This page is why static dies, how you route page vs ticket, and the AND decision. Link them; do not merge them into one mega-lesson.
Pitfalls
- Paging on every 5xx without validity filters (client disconnects).
- No inhibit rules → storm during total outage.
- Copying 14.4× without matching your SLO window.
- Alert without owner / runbook.
- Static thresholds on raw RPS.
- Paging on CPU / disk while the user SLI is the actual contract.
- No min-request guard at night.
- Firing page and ticket and Slack and email for one burn.
(1) 45-second full outage. (2) 5-minute full outage. (3) 0.15% errors for eight days. For 99.9% / 30d, say which row fires (none / 14.4× / 6× / 1×) and whether you page or ticket. Then write the inhibit rule so (2) does not also open a ticket.
Go Deeper
- Google SRE Workbook — Alerting on SLOs
- Prometheus — Alerting rules
- Alertmanager — Configuration
- Honeycomb — SLOs get budget rate alerts
- Companion: multi-window burn alerts · SLIs, SLOs & error budgets
Cheat sheet
burn = error_rate / (1 − SLO)
AND long window with short window (short ≈ 1/12 long)
99.9% / 30d:
Page 14.4× 1h AND 5m (~2% budget / hour)
Page 6× 6h AND 30m
Ticket 1× 3d AND 6h
Static 1% for 5m flaps or misses leaks.
Page user SLIs; ticket slow burns; don't page CPU.
Inhibit overlapping severities. Recording rules for windows.