Observability
Part 3 of 5 · Observability triadMetric Cardinality & Prometheus Label Design
A series is metric times labels. Unbounded labels explode memory — bound labels; use histograms and recording rules.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Where does identity go?
Prefer
Bounded labels + histograms + recording rules
method, status, route template. Buckets for mergeable p99. Precompute burns. Put user_id on structured logs and traces.
- series = Cartesian product of label values.
- Relabel drop is the prod fix; raising limits is last resort.
- Exemplars link a bucket to one trace without a series per request.
Alternative
user_id / raw URL / summary quantiles as the SLO
Each user is a series forever (until retention). Scrapes die. Quantiles computed per pod cannot be averaged into a fleet p99.
- Pod on every app metric multiplies by replica count.
- Unlimited custom buckets per route is another explosion.
- Alerting on raw high-card series melts rule eval.
Overview
In Prometheus, a time series is a metric name plus a unique set of label key/value pairs. Cardinality is the count of those series. Bounded labels (method, status, route template) are powerful. Unbounded labels (user_id, request_id, raw URLs) cause cardinality explosion — OOM on Prometheus, slow queries, and broken recording rules.
Design labels for aggregation. Put high-cardinality identity in logs and traces. Use histograms (not summaries) when you need mergeable quantiles across instances. Use exemplars to link buckets back to traces.
The triad hub said metrics are the cheap detect signal. This lesson is how they stay cheap.
You should be able to:
- Write
series = metric × distinct label setsand multiply an example. - Name five forbidden label values and where they belong instead.
- Explain why summaries cannot be the fleet p99 SLI.
- Say what recording rules and relabel
labeldropare for. - Attach an exemplar instead of a
trace_idlabel.
Architecture
1 Emit
- 1
App
- labelsTime series set
- 2
Time series set
- nextGET /metrics
2 Scrape
- 3
GET /metrics
- nextPrometheus TSDB
- 4
Prometheus TSDB
- nextCardinality OK?
3 Explode?
- 5
Cardinality OK?
- user_idMemory / scrape death
- boundedRules + alerts
- 6
Memory / scrape death
- nextdrop/keep relabel
- nextMove id to logs
- 7
Rules + alerts
- nextExemplars to traces
4 Fix
- 8
drop/keep relabel
- 9
Move id to logs
- 10
Exemplars to traces
Lesson map
Metric Cardinality & Prometheus Label Design
A series is metric times labels. Unbounded labels explode memory — bound labels; use histograms and recording rules.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB app["App"] prom["Prometheus TSDB"] alert["Alertmanager"] app -->|duration_bucket| prom prom -->|burn or latency| alert
Label design
| Label choice | Series growth | Query power | What fails |
|---|---|---|---|
code, method | Low | High for RED | — |
route = template /users/:id | Low–med | High | Using raw path /users/42 |
user_id | Unbounded | "Filter one user" | Prometheus dies; use logs |
pod on every app metric | × replicas | Debug instance | Global SLO queries heavy |
| Histogram buckets | × buckets | Mergeable p99 | Too many custom buckets |
| Summary quantiles | Per-instance | Local p99 only | Cannot aggregate across pods |
Mental model:
http_requests_total{method="GET",code="200",route="/checkout"}
→ 1 series
http_requests_total{method="GET",code="200",route="/checkout",user="u1"}
→ a new series per userAggregation sum by (route) (rate(http_requests_total[5m])) only works well if route cardinality is bounded. Templates, not interpolated ids.
- 1
debug one customer → TSDB OOM
Ship user_id on the counter
Each user is a series until retention. Scrapes slow, then die. Recording rules that grouped by user never finish.
- 2
Winner: bounded RED labels, identity in logs/traces
method / code / route template. Histograms with a sane bucket set. Exemplar carries the one trace you needed.
- ?
If it already exploded
Relabel labeldrop, stop emitting, tombstone the metric. Raise limits last. Then move the id off the metric.
Histograms vs summaries
| Histogram | Summary | |
|---|---|---|
| Client | Counts per bucket | Computes quantiles client-side |
| Aggregate across pods | histogram_quantile | Quantiles are not mergeable |
| Cost | × buckets | × quantile streams |
| Exemplars | Common | Rare |
Prefer histograms for SLIs you will aggregate fleet-wide. The SLO-shaped argument (threshold ratio, exact le bucket, never _sum/_count as the SLO) is histogram vs average latency. Do not thin-duplicate that page here — this table is why summaries are a cardinality and math trap.
le="+Inf", _count, and _sum still exist. Forgetting +Inf breaks histogram_quantile. Using _sum/_count as a latency SLO is the average with extra steps.
Sequence
- 1
App → Prometheus
duration_bucket plus exemplar
- 2
Prometheus → Prometheus
recording rule drops labels
- 3
Prometheus → Alertmanager
burn or latency on the record
Recording rules: precompute expensive queries (burn rates, fleet p99) into new metrics with fewer labels. Alert on the record, not on the raw high-card series.
Exemplars: attach trace_id to a bucket sample so a latency alert opens a real slow request. That is not a label on every point — it is a sampled link. See exemplars / RED / USE.
Cardinality simulator (run this)
I/O: label value sets in → series count out. Then a naive histogram bucket pick.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Expect 18 bounded series, 180000 with user_id, 18 after drop, 144 with 8 histogram buckets; TypeScript 12, 10000, le="0.3" for 0.22 s, +Inf for 9 s.
Interview Q&A
What is Prometheus cardinality?
Answer
The number of unique time series: metric name × distinct label sets. Each unique combination of label values is stored, indexed, and scraped separately.
Why is user_id as a metric label bad?
Answer
Each user creates a series (until retention). Scrapes and the TSDB choke. You wanted "this one customer" — that query belongs in logs or traces with a bounded metric for the fleet.
Histogram vs summary for p99 across 100 pods?
Answer
Histogram. Quantiles from summaries cannot be averaged or merged correctly. Add buckets, then histogram_quantile (or, better for SLOs, a threshold ratio on an exact le).
How do you fix an explosion already in prod?
Answer
Relabel labeldrop, stop emitting the label, tombstone the high-card metric, raise limits only as last resort. Move identity to logs and traces. Do not "add more RAM" as the design.
What are recording rules for?
Answer
Pre-aggregate and reduce label sets for alerts and dashboards. Burn-rate rules especially should run on low-card records, not raw request series.
What are exemplars?
Answer
A sampled trace id attached to a bucket observation so a latency alert opens a real slow request. They are not a trace_id label on every series.
Raw URL vs route template?
Answer
/users/42 is unbounded. /users/:id is one series. Path templates at the HTTP framework / middleware are the usual fix; doing it in Prom relabel with regex is a backup.
Does putting pod on every metric help debugging?
Answer
Sometimes, at the cost of × replica cardinality on every query. Prefer pod on resource USE metrics; keep user SLIs aggregated. Debug instances from traces.
Unlimited custom histogram buckets per route?
Answer
Each extra bound multiplies series. Align buckets on the SLO threshold T, plus a few around it and +Inf. Native/exponential histograms help range, not a free pass on labels.
Where does this sit in the triad cluster?
Answer
Metrics stay the detect pillar only if cardinality is bounded. Next: trace context so the exemplar you attached actually has a tree. Alerts must run on recorded low-card burns: alerting.
Pitfalls
- Raw URL / email / session id as labels.
- Aggregating summaries across instances and calling it fleet p99.
- Alerting on raw high-cardinality metrics.
- Unlimited custom histogram buckets per route.
- Forgetting
le="+Inf"and_count/_sumin histogram queries. pod+user_id+ 20 buckets on a RED counter.- Raising Prometheus flags instead of dropping the label.
- Putting
trace_idon the metric instead of an exemplar.
For method (4) × code (5) × route (30) × pod (40) × 10 buckets, write the series count. Then drop pod and user_id (10k) that a teammate wanted. If you need a calculator for the second number, you already know the answer is no.
Go Deeper
- Prometheus — Metric and label naming
- Prometheus — Histograms and summaries
- Grafana Mimir — Cardinality management
- OpenTelemetry — Metrics
- Companion: histogram vs average latency · exemplars / RED / USE
Cheat sheet
series = metric_name × distinct label sets
OK: method, code, route_template
Never: user_id, request_id, email, /users/42
Histograms merge; summaries do not
Recording rules: fewer labels for alerts
Exemplars: trace_id on a sample, not a label
Fix: labeldrop → stop emit → tombstone → limits last