Observability
Part 4 of 5 · Observability triadTrace Context Propagation (W3C) & Sampling Strategies
W3C traceparent and tracestate keep traces intact. Head vs tail sampling trades cost against capturing rare failures.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How do you keep the tree — and the bill?
Prefer
W3C + parent-based + mixed sampling
Inject traceparent everywhere. Children inherit the sampled bit. Low head rate plus tail keep errors/slow. App spans for business logic.
- Same trace_id across hops; new span_id per operation.
- Always-on errors still need a baseline sample of OK traffic.
- Queue metadata gets the same carrier as HTTP.
Alternative
Custom headers, new ids per service, 100% or 0.1% blind
Orphan spans. Head-only 1% drops the poisoned request. Tail-only without a cap OOMs the collector. Mesh-only traces miss the handler bug.
- Truncating or lowercasing hex incorrectly breaks vendors.
- 100% in prod 'temporarily' becomes the bill.
- Baggage with user email leaks PII to every hop.
Overview
Distributed tracing only works if every hop shares the same trace context. W3C Trace Context standardizes traceparent and tracestate HTTP headers (plus optional baggage for cross-cutting key/values).
Sampling decides which traces you keep:
- Head-based — decide at the root. Cheap. May miss rare failures.
- Tail-based — decide after completion. Can keep errors and slow traces. Needs a buffer.
- Parent-based — children inherit so trees stay intact.
- Rate-limited — cap N/sec to protect the backend.
- Always-on errors — high signal; still need a base sample of successes.
Broken propagation → orphan spans. 100% sample in prod → cost crisis. 0.1% blind random → invisible outages.
The triad hub uses traces to answer "which hop?" This lesson is how the tree exists at all.
You should be able to:
- Decode a
traceparentline on a whiteboard. - Contrast head / parent / tail / rate-limit / error-keep.
- Name baggage risks (PII, size, accidental coupling).
- Say why mesh-only instrumentation is not enough.
- Propagate context on async queues.
Sequence
- 1
Client → Gateway
GET /orders plus traceparent
- 2
Gateway → Orders
same trace_id new span_id
- 3
Orders → Payments
traceparent
- 4
Payments → Orders
500
- 5
Orders → Gateway
error spans exported
Lesson map
Trace Context Propagation (W3C) & Sampling Strategies
W3C traceparent and tracestate keep traces intact. Head vs tail sampling trades cost against capturing rare failures.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB client["Client"] gw["Gateway"] ord["Orders"] pay["Payments"] client -->|GET /orders plus| gw gw -->|same trace_id| ord ord -->|traceparent| pay pay -->|500| ord ord -->|error spans| gw
Sampling strategies
| Strategy | When decided | Pros | Cons | Prefer |
|---|---|---|---|---|
| Head-based probabilistic | Root span start | Cheap, simple | Misses rare errors | Steady high QPS |
| Parent-based | Inherit upstream | Intact traces | Upstream drop = blind | Meshes / multi-service |
| Tail-based | After trace ends | Keep errors / slow | Buffer memory / cost | Critical paths |
| Rate-limited | Cap N/sec | Predictable cost | Drops bursts | Protect backend |
| Always-on errors | Status-based | High signal | Still need base sample | Combined policies |
What fails if you choose the other: head-only at 1% → the one poisoned request may never be stored. Tail-only without capacity → collector OOM. No parent-based → half-traces that confuse on-call.
Production mix interviews want: ParentBased(root = TraceIdRatioBased(0.05)) plus tail keep errors / latency outliers. The policy catalog with worked keep/drop tables is trace sampling strategies — complementary, not a stub of this page.
- 1
HTTP request → root or parse
Ingress
No header → create trace_id / span_id. Header present → parse W3C; do not mint a new trace id.
- 2
head or parent → record bit
Sample
New root: probabilistic head decision. Child: honor parent sampled flag. Tail may still drop/keep after status is known.
- ?
Propagate and export
Inject new span_id, same trace_id. Batch processor. Tail policy: keep error/slow; drop normal traffic over budget.
Architecture
1 Ingress
- 1
HTTP request
- nextHas traceparent?
- 2
Has traceparent?
- NoCreate root ids
- YesParse W3C
- 3
Create root ids
- nextHead sampler
- 4
Parse W3C
- nextParentBased
2 Sample
- 5
Head sampler
- nextRecord?
- 6
ParentBased
- nextRecord?
- 7
Record?
- YesInject child header
- NoPropagate unsampled
3 Export
- 8
Inject child header
- nextTail policy?
- 9
Propagate unsampled
- 10
Tail policy?
- error/slowExport
- over budgetDrop
- 11
Export
- 12
Drop
W3C traceparent anatomy
traceparent: 00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01
ver-trace_id (32 hex)-----------span_id (16 hex)-flagsFlags: sampled bit 00000001 → record this trace (01 vs 00).
tracestate: vendor-specific list (rojo=...,congo=...) for multi-vendor hops.- Baggage: separate header for app-level key/values. Careful: PII, size limits, accidental coupling via undocumented keys.
- Service mesh vs app instrumentation: mesh can inject headers at the sidecar, but app spans are still needed for business logic. Prefer OTel SDK in the app plus optional mesh for network spans.
- Security: sanitize untrusted inbound headers if abuse is a concern; do not leak internal baggage to public clients.
Parse, inject, parent-based sample (run this)
I/O: header string in → (trace_id, span_id, sampled) out; child header preserves trace_id.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Expect: child same trace_id, 01 flags when parent sampled, tail keep on 500 even if head said drop, None/null on a short id. TypeScript also shows an unsampled parent stays dropped under parent-based (shouldSample returns false).
Interview Q&A
What is in traceparent?
Answer
Version (00), 16-byte trace id (32 hex), 8-byte span id (16 hex), flags (sampled bit). Example: 00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01.
Head vs tail sampling?
Answer
Head decides at start (cheap, biased against rare failures). Tail decides after (can keep errors/slow, needs buffering). Combine: low head rate + tail keep-rules.
Why parent-based?
Answer
Keeps a single decision for the whole tree so you do not get partial traces. Children honor the upstream sampled flag in W3C flags.
Baggage risks?
Answer
PII leakage, header size limits, accidental coupling of services via undocumented keys. Baggage is not a substitute for trace ids or for a real API.
Is mesh-only tracing enough?
Answer
Network spans help, but business-logic spans need app instrumentation. Sidecar injects headers; your handler still has to create the span that says payments.charge.
How do you combine policies?
Answer
ParentBased(root = TraceIdRatioBased(0.05)) plus tail keep errors and latency outliers. Rate-limit the exporter so a retry storm cannot melt storage.
Four services, payments returns 500, head p=0.01 with no tail. What happens?
Answer
Most error traces are dropped. Metrics still show the burn (triad detect). You cannot open a waterfall. That is why tail keep-errors exists, and why exemplars must land on kept traces.
How do queues break traces?
Answer
HTTP middleware injects headers; the worker that reads Kafka does not. Put the W3C carrier in message metadata. Minting a new root on the consumer orphans the producer span.
100% sampling 'just until we find the bug'?
Answer
Becomes the default. Cost crisis. Use tail keep on the erroring route / attribute instead of global 100%.
tracestate vs baggage?
Answer
tracestate is vendor tracing state (sampling threshold, vendor ids). Baggage is application key/values. Do not stuff user email into either if it will hop the public internet.
Pitfalls
- Truncating or lowercasing hex incorrectly.
- Creating a new
trace_idat each service (breaks the tree). - Sampling 100% in prod "temporarily."
- Tail sampling without memory limits.
- Forgetting to propagate over async queues (carrier in message metadata).
- Mesh-only spans with no handler names.
- Trusting inbound
traceparentfrom the public internet without a policy. - Head-only 0.1% with no error bias — the incident is invisible.
Gateway → orders → inventory → payments. Payments 500s, inventory is fine. Write the headers at each hop (same trace_id, new span_id). Then say which traces exist under (a) head p=0.01, (b) parent-based drop from gateway, (c) tail keep 5xx. If (a) and (c) look the same to you, redo it.
Go Deeper
- W3C Trace Context
- W3C Baggage
- OpenTelemetry — Sampling
- OpenTelemetry — Context propagation
- Companion: trace sampling strategies · exemplars / RED / USE
Cheat sheet
traceparent = 00-{32 hex trace}-{16 hex span}-{flags}
flags 01 = sampled
Head: decide at start (cheap, miss rares)
Tail: decide at end (keep errors/slow, buffer)
ParentBased: inherit flag (whole trees)
Baggage ≠ context. PII / size / coupling.
Mesh ≠ app spans. Queues need a carrier.
Never mint a new trace_id at the next service.