Observability
Part 5 of 6 · SLOs & tracingTrace Sampling Strategies
Head sampling decides at the root; tail sampling waits until the trace is complete. Mix probabilistic, rate-limiting, and error-biased policies or you will miss the incident or drown storage.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Overview
Every span you keep costs CPU, network, and storage. Every span you drop might have been the one checkout that burned 12% of the monthly budget. Sampling is the policy that chooses keep vs drop. It is not “set sample_rate=0.1 and go home.”
Two decision times matter:
- Head-based — at the start of the request (root span). The decision propagates downstream. Low overhead, no buffer. Blind to the outcome: the 500 and the 5 s tail are unknown yet.
- Tail-based — after the trace is complete. A collector (or batch processor) buffers spans, then keeps errors, high latency, interesting attributes, or a sample of the rest. Best recall for incidents; memory and delay cost.
Production pattern that interviews want: low head rate (1–10%) so you always have a baseline, plus tail rules that always keep errors and slow traces, plus exemplars so metric buckets point at a kept id. Always-on 100% export rarely survives contact with traffic.
Why a mixed policy wins
| Strategy | When it decides | Wins | Loses |
|---|---|---|---|
| Head / root | First span | Simple, cheap | Drops late failures |
| Parent-based | Any child | Whole trees, W3C flag | Propagates DROP too |
| Probabilistic / TraceIDRatio | Anywhere | Predictable volume | No bias to errors |
| Rate-limiting | Usually edge | Caps bursts | May drop a failure storm |
| Tail (collector) | After complete | Keep errors / slow | Buffer RAM, complexity |
| Mix (winner) | Head + tail keep-rules | Baseline + incident recall | Two systems to operate |
- 1
no policy → bill shock or missed 500s
100% export or head-only 0.1%
100% melts storage. Head-only p=0.1 keeps a random 10% and will often drop the one poisoned shard. Rate-limit alone drops the incident burst you most need.
- 2
Winner: low head probability + parent-based + tail keep errors/slow
Roots sample a few percent for baseline. Children honor the flag. The collector always retains error and high-latency traces (and a sample of OK). Rate-limit the OK rest so a traffic spike cannot fork the bill.
- ?
If you cannot buffer, bias the head sampler
X-Ray-style rules and Jaeger per-operation rates approximate “keep more of /checkout.” Still blind to unknown errors. Prefer tail when the collector can assemble traces.
Head, parent, and W3C flags
W3C traceparent looks like 00-trace_id-parent_id-flags. The flags byte bit 01 means sampled. OpenTelemetry’s default propagator carries that. Do not invent custom headers between services you own.
Parent-based sampler: if a parent exists, inherit its decision; otherwise use a root sampler (often TraceIdRatioBased(p)). That keeps traces as trees instead of Swiss cheese. It also means an upstream DROP hides downstream errors — the child never records. If a mesh sidecar samples 1% and your app samples 100% without honoring parent, you get partial traces and angry humans.
Head-based is the right default at a latency-sensitive edge where you refuse to buffer. It is a poor only-line of defense for rare, late failures (the 500 happens in payments after the root already chose DROP).
Decisions
- 1
Start request
- nextHead sampler
- ?
Head sampler
- DROPNo spans recorded
- KEEPRoot span + sampled flag
- 3
No spans recorded
- 4
Root span + sampled flag
- nextPropagate traceparent
- 5
Propagate traceparent
- nextParent-based child?
- ?
Parent-based child?
- InheritSame KEEP or DROP
- OverrideLocal sampler - usually a mistake
- 7
Same KEEP or DROP
- nextOptional rate limit on KEEP
- 8
Local sampler - usually a mistake
- 9
Optional rate limit on KEEP
- nextEnd request
- 10
End request
- nextTail sampler at collector
- ?
Tail sampler at collector
- Error or slowForce KEEP
- OK sampleKeep with p or rate
- elseDROP buffered spans
- 12
Force KEEP
- nextExport
- 13
Keep with p or rate
- nextExport
- 14
DROP buffered spans
- 15
Export
Lesson map
Trace Sampling Strategies
Head sampling decides at the root; tail sampling waits until the trace is complete. Mix probabilistic, rate-limiting, and error-biased policies or you will miss the incident or drown storage.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Start request"] b["Head sampler"] d["No spans recorded"] c["Root span + sampled flag"] a -->|Start request to Head sampler| b b -->|DROP| d b -->|KEEP| c
Step notes: 1 head at ingress, 2 propagate the flag, 3 children inherit, 4 tail re-judges complete traces, 5 export only what the mix kept.
Tail sampling, probability, rate limits
Tail requires the collector to assemble spans by trace id until the trace ends (or a timeout). Then filters: status=ERROR, duration > 500 ms, attribute user.plan=enterprise, HTTP 5xx, etc. Cost: RAM proportional to in-flight traces × span size. Unbounded buffers OOM. Bound by count, age, and attributes you actually need.
Probabilistic / TraceIDRatio uses bits of the trace id so every service that sees the same id makes the same decision — unlike Math.random() at each hop. Predictable reduction; no incident bias.
Rate-limiting (token bucket of N traces/sec) protects the pipeline when traffic jumps 10×. Reset the window every second; a stuck counter throttles forever. Combined with probability: probability sets the long-run fraction, rate limit caps the spike. Do not rate-limit errors the same as OK traffic if the goal is incident recall.
OpenTelemetry Sampler.ShouldSample is the plug-in: parent-based wrapping ratio, plus a collector-side tail processor. X-Ray and Jaeger expose similar knobs (reservoir + fixed rate, per-operation).
Deep dive · Why head cannot see errors, and why 100% parent-based still explodes
Head decides before status codes exist. Error-biased head hacks (sample 100% of /checkout only) help known hot paths and miss unknown ones. Tail sees the real status and duration.
A 100% parent-based sampler with an aggressive upstream is a cost bomb: every kept root fans out to every descendant. A 100% parent-based sampler with a DROP-happy upstream is a blindness bomb: payments can 500 all day inside unsampled trees. Set the root policy explicitly; do not assume “parent-based true” is a strategy.
Consistency, PII, and SLOs
Mismatched libraries (Python ratio 0.1, Node always-on, a batch worker that starts new roots) produce partial traces. One sampler config story; collector as the place that can still tail-filter.
Sampling is not a substitute for PII scrubbing. Dropping 90% of traces still stores the 10% of payloads you kept. Scrub in the collector.
Traces are for why. The SLO still comes from metrics (SLI design, histograms). If you sample 1% of requests, you cannot compute a 99.9% availability SLI from traces. Metrics stay 100%; traces stay sampled; exemplars join them.
Keep vs drop (in-memory)
I/O: synthetic traces in, printed keep/drop counts for head vs tail vs mix. Deterministic “random” from trace id so reruns are stable. No network.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Head p=0.1 will miss most errors (they were not special at start). Tail and mix keep all errors and slow traces, plus a 10% baseline of boring OK traffic. That is the interview diagram.
A rate limit is a counter per second on KEEP; if you forget to reset it, you drop everything after the first N traces for the process lifetime — a classic bug.
Remote sampling, reservoirs, and tracestate
Jaeger can push sampling strategies from a remote service: default probabilistic plus per-operation rates (higher on POST /checkout, lower on health). That is still head sampling — the operation name is known at start, the status code is not. Use it to bias known money paths; do not confuse it with tail keep-on-error.
X-Ray sampling rules use a reservoir (guaranteed N traces per second) plus a fixed rate for the rest. Reservoir is how you still see something at 3am when QPS is tiny (the low-traffic cousin of the burn-alert request floor). The remaining fraction is probabilistic. Rules can match service, URL, method. Outcome-blind unless you add tail processing downstream.
W3C tracestate. OpenTelemetry may carry a sampling threshold (th) and randomness (rv) so consistent probability can be adjusted along the path without breaking the id. You do not need to recite the spec; you do need to propagate tracestate with traceparent or vendor-specific consistent sampling falls apart.
Collector architecture. Apps export via OTLP to a collector; the collector batches, scrubs PII, tail-samples, and fans out to Tempo/Jaeger. Head sampling in the SDK reduces app overhead and export volume. Tail sampling in the collector can only see traces that were recorded. If the SDK DROPped at head, there is nothing to tail-keep — unless the SDK records without sampling (parent-based record-only) and the collector makes the export decision. Two different “drop” times: not-recorded vs recorded-but-not-exported. Interviews like that distinction.
Adaptive sampling (raise p when error rate climbs) is a collector policy, not magic ML. It still needs bounds (rate limit) or a bad deploy exports 100% and you pay for the outage twice.
Partial traces from async. A queue consumer that starts a new root instead of extracting traceparent from the message breaks the journey SLI and the waterfall. Sampling cannot stitch what you never linked. Inject/extract on every hop you own, including brokers.
Head vs tail keep/drop in one sentence: head is a coin flip on identity; tail is a predicate on the finished tree. The playground is that sentence with numbers.
Sampler decisions, RECORD vs DROP
OpenTelemetry’s ShouldSample returns roughly:
- DROP — do not record, do not export. Children that honor parent also drop.
- RECORD_ONLY — keep locally (and maybe tail-export later) but the sampled flag is off for other vendors.
- RECORD_AND_SAMPLE — record and set the W3C sampled bit so the rest of the graph keeps.
Head-only shops use DROP vs RECORD_AND_SAMPLE. Tail-friendly shops often record at the SDK (or record a subset of attributes) and let the collector DROP export. That is how tail can see errors that a 1% head coin flip would have destroyed — at the cost of app CPU and collector RAM. If you cannot afford to record 100%, record 100% of errors you already know (span status set before end) plus 5–10% of OK; that is a hybrid, still not as complete as collector tail on fully recorded traces.
Composite order matters. “First KEEP wins” vs “any DROP wins” vs “AND of rate-limit and probability” are different policies. A useful mix:
- If parent says DROP, DROP (consistency) — unless you are the collector tail re-judging recorded spans.
- If error or duration ≥ T, KEEP (incident recall).
- Else KEEP with TraceIdRatio p.
- Else apply rate limit on the OK remainder.
Do step 4 after step 2 so an error storm is not the thing you throttle. A naive “first KEEP wins” mix with a rate limiter before the error rule is how you drop the outage.
Fire-and-forget and span links: a parent that returns before the child finishes cannot tail-sample the child in the same trace without links. Either wait, or link the async trace and accept two ids. Journey SLIs should use the user-visible parent, not the batch worker’s root.
Always-on debug for one trace_id (forced KEEP via a header you only honor internally) is an incident tool. Do not leave that header working on the public internet.
Cost math worth saying: 5k QPS × 20 spans × 100% × 1 kB/span is ~100 MB/s ingest. At 5% head plus tail-keep of 5% errors you might be under 10 MB/s and still have every error. The playground’s mix policy is that trade in miniature: keep the boring 10% for baseline, keep 100% of errors/slow, drop the rest.
If a vendor “adaptive sampler” cannot explain keep/drop on one trace, treat it as a black box and add an explicit tail rule for status=ERROR anyway.
Jaeger remote sampling, X-Ray rules, and OTel ParentBased(TraceIdRatioBased(p)) are all head tools. Name them that way. Tail is a collector processor that sees finished trees. If the interview only allows one sentence: low head rate, parent-based children, collector keeps errors and slow, rate-limit the OK rest.
Interview Q&A
Head-based vs tail-based sampling?
Answer
Head decides at the first span and propagates the sampled flag; cheap, may drop rare late failures. Tail buffers until the trace completes and can filter on status, latency, and attributes; better incident recall, higher collector memory. Many shops do both.
Why combine probabilistic and rate-limiting?
Answer
Probability sets a predictable long-run fraction. Rate-limiting caps a traffic spike so ingestion cost is bounded. Together: steady bill, no 10× surprise. Do not rate-limit away the error burst if tail-keep-errors is the point.
What does the OpenTelemetry ParentBased sampler do?
Answer
If a parent span context exists, inherit KEEP or DROP. If this is a root, use the configured root sampler (often TraceIdRatioBased). That keeps traces intact across microservices when everyone honors W3C flags.
How does tail sampling affect memory?
Answer
The collector must hold spans until the trace ends or times out. High concurrency × verbose attributes = RAM. Bound buffer size and age; drop incomplete traces rather than OOM. Selective attribute capture beats “keep every header.”
When is head-only acceptable?
Answer
Edge/front-end where buffering would add user latency, or fire-and-forget pipelines that never finish a trace in one collector. Compensate with higher sample on known critical routes, metrics+exemplars, and accepting that some errors will be missing from storage.
Trade-offs of 100% parent-based?
Answer
Full trees when upstream keeps; cost explosion if upstream keeps too much; total blindness if upstream drops. Parent-based is a consistency tool, not a volume policy. Set the root rate (and tail keep-rules) explicitly.
Can I compute a 99.9% SLO from 1% sampled traces?
Answer
No. Sampling biases what you stored. SLIs come from metrics on 100% of events (or unbiased stats). Traces explain a kept example. Exemplars join a metric bucket to one trace id.
TraceIdRatio vs Math.random at each service?
Answer
Ratio on the trace id is stable across hops: every service makes the same decision. Independent random at each hop shreds trees (some children keep, parents drop). Always hash the id, do not flip a coin per process.
What goes wrong if samplers disagree across languages?
Answer
Partial traces: Python kept the root, Node dropped the child, the waterfall stops at the boundary. One collector pipeline and parent-based honor everywhere. Treat custom headers as an incident.
Rate-limit reset bug?
Answer
If the per-second counter never resets, you keep the first N traces after process start and drop forever. Token buckets must refill. Test it with a clock you control — the playground’s mix policy is the logic; the counter is the ops bug.
Does tail sampling work if the SDK already DROPped at head?
Answer
No. Tail can only judge recorded spans. Either record at the SDK and export-filter in the collector, or accept that head DROP is gone forever. That is why “1% head plus tail keep errors” only works if errors were still recorded.
Pitfalls
- Inconsistent decisions across services — mismatched configs, custom headers, ignoring
traceparentflags. - Unbounded tail buffers — OOM under load; bound by count and age.
- Head-only on rare errors — the incident is the thing you dropped.
- Sampling drift in short windows — ratio is exact in expectation, noisy in a 5-minute debug session.
- Parent DROP hiding downstream 500s — inherit is consistency, not omniscience.
- Rate-limit without refill — permanent throttle.
- 100% export “for a week” — the week becomes the new bill.
- Using traces as the SLI — biased sample; keep metrics at 100%.
- PII in the 10% you kept — sampling is not redaction; scrub in the collector.
- New roots on queue consumers — lost journey traces; extract
traceparentfrom the message. - Rate-limiting errors with the OK traffic — the outage is the burst you dropped.
Take 200 traces, 5% errors. Head-sample 10% using id % 10 === 0. Count missed errors. Then keep if error OR slow OR id % 10 === 0. The second policy is what you describe in the interview before you name vendors.
Go Deeper
- OpenTelemetry — Sampling
- Jaeger — Sampling
- AWS X-Ray — Sampling rules
- Google Cloud — Distributed tracing at scale
- Token bucket (rate-limit refill)
- Next: Exemplars, RED & USE
Cheat sheet
Head = decide at root, propagate sampled flag (cheap, outcome-blind)
Tail = decide after complete (keep errors/slow; needs a buffer)
Parent-based = inherit; do not start a new coin flip
Ratio on trace id, not Math.random per hop
Mix: low head % + tail keep interesting + rate-limit OK
SLI from metrics at 100%; traces sampled; exemplars join them