Approximate Aggregations
Studies in this cluster, in series order. Each one keeps its own URL.
Data engineering
Pipelines, sketches, approximate aggregations, object storage, and stream processing with event time, windows, and exactly-once sinks.
Approximate Aggregations
6 studies- 1.Approximate Aggregations — Sketches for Quantiles, Cardinality & Merge PipelinesExact p99 over billions of events does not fit in memory, and averaging shard percentiles is wrong. Interviewers expect mergeable sketches: fixed-size summaries you update online, serialize, and combine leaves to regions to global.
- 2.KLL Quantile Sketches — Error Bounds, k Parameter & Merge SemanticsKLL (Karnin–Lang–Liberty) is a mergeable quantile sketch with a provable rank-error bound. Interviewers want the k parameter, what ε means, and how merges preserve guarantees — not vendor trivia.
- 3.T-Digest — Centroid Compression, Tail Accuracy & Heuristic LimitsT-Digest stores mergeable centroids with a compression parameter that spends accuracy on the tails. Interviewers want how centroids work, why p99 looks good, and where the heuristic breaks versus KLL’s theorems.
- 4.Sketch Merge Pipelines — Incremental, Hierarchical & Cross-Shard AggregationSketches only pay off when the pipeline merges them correctly: incrementally on a stream, hierarchically across regions, and across shards without double-counting. Interviewers want leaf-to-region-to-global and the retry failure mode.
- 5.KLL vs T-Digest vs HDRHistogram vs Exact — When to Choose WhatSenior interviews are decision matrices, not brand loyalty. Pick the structure that matches error model, merge needs, value domain, and ops cost — KLL, T-Digest, HDR/Prom hist, or exact offline.
- 6.Production Sketch Ops — Serialization, Idempotent Merges, Bias & MonitoringSketches fail in production through bytes, retries, and silent bias — not through forgetting the paper. Version the payload, apply once by sketch_id, and watch coverage plus dual-read error.