Recall, Latency, and Memory — Tuning and Evaluation
Every ANN index trades recall, latency, and memory, with build time as the fourth wheel. The first artifact is an evaluation harness: frozen real queries, exact ground truth, recall at k plus tail recall, p50 and p99, and bytes per replica. Sweep efSearch, nprobe, quantization, and re-rank depth, then pick the cheapest point that meets the target.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is k-recall at k?
Answer
The overlap of the returned top k with the true top k, divided by k. It is an index-quality metric against exact search, not a relevance grade.
L2
How do you build ground truth?
Answer
A frozen sample of real queries, exact brute force, the same metric as production. Do not take the queries from the corpus, or each query's nearest neighbor is itself.
L3
Why report tail recall?
Answer
Mean recall 0.95 can hide 6 percent of queries under 0.8. If those users matter, the target is the mean and the tail, which pushes the operating point to a wider beam or more probes.
L4
Recall target 0.95 and p99 under 50 ms. How do you get there?
Answer
Sweep efSearch or nprobe at production concurrency, plot recall against p99, and pick the cheapest point that meets both. If none does, change M, efConstruction, nlist, quantization, or shard layout, then sweep again.
L5
Why might p99 be 5 times p50?
Answer
Dense regions need more hops. Skewed IVF lists are larger. Filters change the walk. Shards and segments add stragglers. GC, cold pages, and merges sit on top. Profile visited counts, not only wall time.
L6
You doubled efSearch, latency doubled, recall barely moved. What now?
Answer
You are past the knee, or the graph is the bottleneck. Drop back to the knee. If recall is still short, rebuild with higher M or efConstruction, check normalization and metric, or add a re-rank.
L7
What does quantization cost beyond the codes?
Answer
Most systems keep the original vectors somewhere for re-scoring, in RAM, on disk, or memory-mapped. Budget for both, plus codebooks or per-segment parameters.
Failure modes
Ground truth from the ANN index
You measured the index against itself. Recall looks perfect and means nothing.
Queries sampled from the corpus
The nearest neighbor is the query vector. Recall is inflated before you tune anything.
Mean recall only
A cluster of queries at 0.3 disappears into a 0.95 average.
Warm benchmark, cold production
DiskANN and memory-mapped indexes look fast until the pages are cold.
Misconceptions
Index recall and search relevance are the same number.
An index at 0.99 recall cannot fix a bad embedding. A great model cannot rescue an index at 0.7. Measure them separately. nDCG and MRR need labels.
The highest recall you can buy is the right operating point.
Pick the cheapest p99 that meets the target. Past the knee you pay latency for overlap you do not use.
Latency on an idle laptop is the SLO.
Measure at production concurrency, with the filters you actually send.
Interviewer traps
Quoting ANN-Benchmarks QPS as your capacity plan.
Use those plots to see the shape of a library. Size your corpus with your dimension, your replica count, and your headroom.
Tuning efSearch and calling the embedding model done.
If the exact top-k is already irrelevant, the harness will not save the product. Say that, then go back to the index knob.
Design scenario
Same prompt for every reader.
Requirements
Mean recall at 10 at least 0.95, fewer than 1 percent of queries under 0.8 if the tail matters, and p99 inside the interactive budget at the real QPS.
Traffic / scale
Online queries with filters, plus a weekly job that recomputes exact top-k for a frozen query sample.
Latency
Report p50 and p99 at production concurrency. A single-threaded idle number is not the SLO.
Consistency
Ground truth uses the production metric and the production model id. A changed model invalidates the oracle.
Availability
The harness runs offline. It must not share the serving heap with queries.
Failure assumptions
- Someone computes ground truth with the ANN index.
- The benchmark is warm and production is cold.
- Quantization saved RAM and the re-rank vectors were forgotten.
Constraints
- State the target as mean and tail, not mean alone.
- Pick the cheapest point that meets it.
Prompt
Stand up a recall harness before you pick efSearch or nprobe for a production corpus.
API
What does the weekly job compare, and who consumes recall at 10 versus 1-recall at 1?
Data
Where do the frozen queries, the exact top-k, and the byte budget live?
Architecture
Which knob moves first for low recall, and which one requires a rebuild?
The evaluation loop
Ground truth is exact. The pick is the cheapest point that meets the target, not the best recall you can buy.
- 1
Freeze real queries
Hundreds to a few thousand queries, held out. Do not sample them from the vectors you will search. - 2
Compute exact top-k
Brute force, same metric as production. This is the oracle. The ANN index is not allowed to grade itself. - 3
Sweep one query knob
efSearch, nprobe, or re-rank depth. Record mean recall, the share of queries under 0.8, p50, p99, and bytes. - 4
Pick or rebuild
Take the cheapest p99 that meets the target at production QPS. If none does, change the build, the codes, or the index type, and sweep again.
Overview
Every ANN index is a three-way trade between recall, latency, and memory, with build time and cost as the fourth wheel. You cannot tune what you do not measure. The first artifact of a vector search project is an evaluation harness: a fixed set of real queries, exact ground truth, recall at k plus tail recall, p50 and p99, and memory per replica. Then you sweep the knobs and pick the cheapest point that meets the target.
Two different questions get mixed in interviews. Index recall asks whether ANN returned what exact search would return. Search relevance asks whether those exact neighbors are even good. An index at 0.99 recall cannot fix a bad embedding model. A great model cannot rescue an index at 0.7. Measure them separately. A semantic cache that reuses embeddings is a different system: Caching for Agents — Prompt Cache, KV Reuse, Semantic Cache & Tool Result Cache.
Metrics that matter
| Metric | Definition | Use it for |
|---|---|---|
| k-recall at k | Overlap of the returned top k with the true top k, divided by k | Index quality against exact search. The ANN-Benchmarks standard |
| 1-recall at 1 | Did the single best result equal the true best | Dedup, nearest-match lookups |
| Tail recall | Share of queries with recall under a threshold | Clusters where the index fails |
| nDCG, MRR, precision at k | Relevance against human or LLM labels | End-to-end search quality, including the embedding model |
| p50, p95, p99 | Latency percentiles at the target QPS | SLOs. ANN cost varies by query region |
| QPS per core or per dollar | Throughput at a fixed recall | Capacity and cost |
| Bytes per vector | Vectors plus graph or codes, plus ids and metadata | RAM sizing |
| Build and merge time | Full build, segment merges | Reindex windows, freshness |
The loop
Decisions
- 1
1. Hold out real queries
- next2. Exact top k is truth
- 2
2. Exact top k is truth
- next3. Build index configs
- 3
3. Build index configs
- next4. Sweep ef, nprobe, re-rank
- 4
4. Sweep ef, nprobe, re-rank
- next5. Recall, tail, p99, RAM
- 5
5. Recall, tail, p99, RAM
- next6. Meets recall target?
- ?
6. Meets recall target?
- yes7. Cheapest p99 at target
- no7b. Rebuild or change index
- 7
7. Cheapest p99 at target
- next8. Freeze set, alert weekly
- 8
7b. Rebuild or change index
- next3. Build index configs
- 9
8. Freeze set, alert weekly
Lesson map
Recall, Latency, and Memory — Tuning and Evaluation
Every ANN index trades recall, latency, and memory, with build time as the fourth wheel. The first artifact is an evaluation harness: frozen real queries, exact ground truth, recall at k plus tail recall, p50 and p99, and bytes per replica. Sweep efSearch, nprobe, quantization, and re-rank depth, then pick the cheapest point that meets the target.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Hold out real queries"] b["2. Exact top k is truth"] c["3. Build index configs"] d["4. Sweep ef, nprobe, re-rank"] a -->|1. Hold out real queries| b b -->|2. Exact top k is truth| c c -->|3. Build index configs| d
Classic mistakes
- Ground truth computed with the ANN index itself, or with a different metric than production.
- Queries sampled from the corpus, so each query's nearest neighbor is itself, which inflates recall.
- Latency measured single-threaded on an idle box, instead of at production concurrency with filters.
- Mean recall only, hiding a cluster of queries at 0.3.
- Warm caches in the benchmark, cold disks in production. This shows up on DiskANN and on memory-mapped indexes.
- Comparing configs at different k, or at different candidate counts.
Sandbox: a harness and a knee
Ground truth comes from exact search. Mean recall hides tail failures, so the harness also prints the share of queries under 0.8. The pick is the cheapest nprobe whose mean meets 0.95, and the printout shows why the tail may push you further.
Press Run. Snippets must be self-contained — no network, files, or native modules.
On the seeded run, nprobe 5 reaches mean recall about 0.956, and about 6 percent of queries still sit under 0.8. If those users matter, the target is "mean at least 0.95 and under 1 percent of queries below 0.8," which pushes the pick to nprobe 8. p99 moves faster than p50 because some queries hit the oversized cells. 1-recall at 1 and k-recall at k disagree at low nprobe. One metric does not fit every product. Millisecond columns move with the machine. The recall columns should not.
Symptom to knob
| Symptom | Likely cause | Turn this |
|---|---|---|
| Recall too low, latency fine | Beam or probes too narrow | Raise efSearch, num_candidates, or nprobe |
| Recall plateaus as efSearch rises | Graph quality | Raise M or efConstruction and rebuild. Check metric and normalization |
| Recall low only for some queries | Skew, hot clusters, filters | Tail analysis, more lists, the filter strategy |
| p99 high, p50 fine | Hot IVF lists, GC, segment count, cold cache | Rebalance lists, fewer segments, warm-up, more replicas |
| Memory over budget | fp32 vectors plus a graph | int8 or binary quantization with re-score, PQ, DiskANN, fewer RAM replicas |
| Recall fell after quantization | Codes too aggressive | Oversample and re-score at higher precision. Larger m |
| Latency rises with shard count | Fan-out and per-shard k | Fewer, larger shards. The production page owns this |
| Build too slow | efConstruction, single-threaded build | Parallel build, a lower efConstruction, bulk load then a force-merge |
Memory math
Per vector:
- fp32 is d times 4. fp16 is d times 2. int8 is d. Binary is d / 8. PQ is m bytes for 8-bit codes.
- HNSW links are about M times 2 times 4 bytes on layer 0, plus about 10 percent for upper layers. M = 16 is about 141 bytes.
- Ids and metadata are 8 bytes and often more inside the database.
- Multiply by replicas, then leave about 30 to 40 percent headroom for the page cache, merges, and query buffers.
Worked example, 100 million vectors, d = 768:
- fp32 HNSW: 100 million times (3072 + 141) is about 321 GB per copy, about 642 GB with one replica.
- int8 HNSW: 100 million times (768 + 141) is about 91 GB per copy.
- Binary plus rescore: 100 million times (96 + 141) is about 24 GB in RAM, plus full vectors on disk for re-scoring.
- IVF-PQ with m = 96: about 10 GB of codes plus ids.
Sandbox: a capacity planner
Index choice is usually decided by memory before latency. The planner prices each option, then applies replicas and a usable fraction of node RAM. Little's law at the end turns QPS and p99 into in-flight queries and a core estimate.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Ten million by 768, two replicas, 64 GB nodes at 60 percent usable: flat and fp32 HNSW need 2 nodes, int8 HNSW fits on 1, and IVF-PQ is about 1 GB per copy. Five hundred million by 1024 on 256 GB nodes: fp32 HNSW is on the order of 28 nodes, int8 about 8, binary plus rescore about 2, IVF-PQ about 1. Three thousand QPS at 25 ms p99 is about 75 queries in flight. At 4 ms of CPU per query that is about 12 busy cores per replica set, before headroom.
Dimensionality is a knob
Higher-dimensional embeddings cost linearly in memory and bandwidth. Matryoshka-trained models let you truncate, for example 3072 down to 1024 or 256, with a modest quality loss. A common pattern is a short-vector first stage and a full-vector re-score. Measure relevance, nDCG, after truncation, not only index recall.
Interview Q&A
How do you measure recall in production without labels?
Answer
Keep a frozen sample of real queries. Compute the exact top k offline with brute force, on a GPU or a batch job, and compare it to what the ANN index returns. Run that on a schedule and after each reindex. Relevance quality needs labels, human or model-judged, and it is a separate number.
The recall target is 0.95 and p99 must stay under 50 ms. How do you get there?
Answer
Build the harness. Sweep efSearch or nprobe at production concurrency. Plot recall against p99. Pick the cheapest point that meets both. If none does, change build parameters, quantization, or the shard layout, then sweep again.
Why might p99 be five times p50 for an ANN index?
Answer
Query cost varies. Dense regions need more hops. Skewed IVF lists are larger. Filters change the traversal. Segments and shards add fan-out and stragglers. GC pauses, cold pages, and merges sit on top. Profile per-query visited counts, not only wall time.
What is the memory cost of quantization beyond the codes?
Answer
Most systems keep the original vectors somewhere for re-scoring, in RAM, on disk, or memory-mapped. Some keep both the quantized form and the raw form in the same index. Budget for both, plus codebooks or per-segment quantization parameters.
1-recall at 1, or 10-recall at 10?
Answer
Whichever the product consumes. A RAG re-ranker over the top 50 cares about recall at 50. A duplicate detector cares about 1-recall at 1. Report that one, plus tail recall.
You doubled efSearch. Latency doubled and recall barely moved. What now?
Answer
You are past the knee, or the graph is the bottleneck. Lower efSearch back to the knee. If recall is still short, rebuild with a higher M or efConstruction, check normalization and the metric, or add a re-rank stage.
Why keep index recall separate from nDCG?
Answer
Index recall answers "did ANN match exact search?" nDCG answers "were those neighbors useful?" A perfect index over a bad embedding still fails the product. A perfect embedding over a sloppy index fails it too. The harness that mixes them cannot tell you which knob to turn.
What headroom do you leave on a vector node?
Answer
About 30 to 40 percent beyond the index bytes, for the page cache, segment merges, and query buffers. A plan that fills the DIMM with vectors has no room for the merge that compaction requires.
Pitfalls
Sketch recall against p99 for seven nprobe values. Circle the cheapest point at mean 0.95. Then move the circle if 6 percent of queries are still under 0.8. Write the byte line for 100 million by 768 next to the circle.