Hybrid Retrieval — BM25 + Vectors, Reranking & Metadata Filters
Production RAG rarely ships vectors alone. Hybrid retrieval fuses BM25 with dense similarity, reranks a shortlist, and applies metadata filters for tenant, ACL, and time before the model ever sees a chunk.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The user pastes error code ERR-42 and also paraphrases the symptom
Prefer
Filter, then BM25 plus vectors, then rerank
ACL and tenant first. Lexical search catches the code. Dense search catches the paraphrase. Reciprocal rank fusion merges the lists. A cross-encoder keeps a handful of chunks for the prompt.
- The two retrievers fail on different queries, so the union recalls more.
- RRF does not require the scores to share a unit.
- Forbidden chunks never enter the shortlist.
Alternative
Top-k cosine on the whole corpus
The error code is a rare token the embedder treats as noise. A neighboring tenant’s similar paragraph can outrank the real runbook.
- Exact ids and new jargon fall out of the dense top-k.
- A post-filter after generation is a leak, not a control.
- Reranking a million hits is not a plan.
Overview
Lexical rankers and dense rankers are wrong in different ways. BM25 rewards term frequency and rarity, so an error code, SKU, or exact identifier floats up. It misses synonyms. A dense index does the opposite: paraphrases work, exact rare tokens often do not, especially if they were scarce in the embedder’s training mix.
Hybrid retrieval runs both, fuses the lists, then reranks a shortlist. Metadata filters (tenant, ACL, time, product) are part of the same path. They are not a cleanup step after the model has already read the text.
The ANN structure under the dense branch is Vector Indexes. This page is the query path on top of it.
Who wins which query
| Signal | Wins on | Fails on |
|---|---|---|
| BM25 / sparse | Exact ids, error codes, rare tokens | Synonyms and paraphrases |
| Dense vectors | Paraphrase and fuzzy intent | Exact SKUs and jargon absent from the embedding space |
| Cross-encoder rerank | Precision on a top-20 to top-50 | Cost. It cannot scan millions |
| Metadata filters | Security and freshness | Over-filtering yields an empty set |
Query path
Flow
- 1
1. User query
- next2. Apply ACL and tenant filters
- 2
2. Apply ACL and tenant filters
- next3. BM25 on the filtered set
- next4. Dense ANN on the same set
- 3
3. BM25 on the filtered set
- next5. Fuse with RRF or weights
- 4
4. Dense ANN on the same set
- next5. Fuse with RRF or weights
- 5
5. Fuse with RRF or weights
- next6. Cross-encoder rerank
- 6
6. Cross-encoder rerank
- next7. Pack for the generator
- 7
7. Pack for the generator
Lesson map
Hybrid Retrieval — BM25 + Vectors, Reranking & Metadata Filters
Production RAG rarely ships vectors alone. Hybrid retrieval fuses BM25 with dense similarity, reranks a shortlist, and applies metadata filters for tenant, ACL, and time before the model ever sees a chunk.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB q["Query"] h["Hybrid"] r["Rerank"] l["LLM"] q -->|ACL filter then| h h -->|Fused shortlist| r r -->|Top chunks| l l -->|Answer plus| q
Run the sparse and dense branches in parallel after the filter. Cap each list. If the reranker times out, fall back to the fused list and record the miss. Do not block the user path on an unbounded third model.
Sequence
- 1
Query → Hybrid
ACL filter then BM25 and ANN
- 2
Hybrid → Rerank
Fused shortlist
- 3
Rerank → LLM
Top chunks
- 4
LLM → Query
Answer plus cites
Fusion
| Method | Pros | Cons |
|---|---|---|
| Reciprocal rank fusion | Scale-free and simple | Ignores how large the score gap was |
| Weighted sum | Tunable | Needs a shared score scale |
| Cascade (sparse then dense, or the reverse) | Cheap first stage | The first stage’s misses never return |
| Learned fusion | Can fit your query mix | Needs training labels |
RRF adds 1 / (k + rank) across lists. The constant k (often 60) keeps the top of one list from erasing the other. It is the default when BM25 scores and cosine scores are not comparable. A weighted sum is reasonable after you calibrate both sides on a held-out set. A cascade is a latency optimization you must audit for the queries the first stage drops.
RRF (run this)
Two ranked id lists. A document that appears near the top of both lists beats a document that won only one. Scores are not consulted, only ranks.
ProblemFuse a BM25 ranking and a dense ranking.
Expectederr-42 and hnsw-overview tie at the top. timeout-guide, present on only one list, ranks last.
Edge cases
- An id that appears in only one list still gets a score.
- k dampens the contribution of rank 1.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Filter, then keep (run this)
This is the predicate, not the index. The production rule is the same shape: tenant and a freshness bound are inputs to retrieval, not a filter you apply after the model has quoted the text.
ProblemKeep rows for one tenant at or after a minimum timestamp.
ExpectedOnly the acme row that is new enough survives.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Where the reranker sits
- Bi-encoder / ANN: millions of chunks down to about 50 to 200.
- Cross-encoder or late interaction: that shortlist down to about 5 to 20 chunks for the prompt.
- LLM-as-reranker: a last resort on a tiny list. It is expensive and the ranking jitters between calls.
A cross-encoder reads the query and the passage together. That is why it is more precise than a bi-encoder and why it cannot be the first stage.
Empty filters are a product bug, not a model bug. If the ACL predicate removes every hit, refuse or relax a non-security constraint (time window, product area) with an explicit message. Do not drop the ACL predicate to “get some context.”
ACL and tenants
- Partition by tenant at index time when isolation and QPS demand it. More indexes to operate. Stronger boundary.
- Filter
tenant_idandallowed_groupson every query when one index must serve many tenants. You must test it. A selective filter can also hurt HNSW recall. That interaction is in the index lesson. - Never retrieve a forbidden chunk and rely on the application to hide it after the LLM has seen the text.
RAG is a data path. It is not your authorization system. The gateway still authenticates the caller. The retrieval predicate must match that decision. A missing predicate is a cross-tenant leak, which the failure-modes lesson treats as a severe incident class.
Document changes can arrive as events. Reliable delivery of those events is a messaging concern. The score you compute after the chunk is indexed is this lesson.
Latency budget
A practical shape: parallel sparse and dense, fuse, rerank at most about 50 passages, then the generator. Put a timeout on the reranker. If it fires, serve the fused shortlist and mark the request degraded. Unbounded “rerank until it returns” becomes the p99.
Packing that shortlist into a token budget, and the citation contract, is RAG Context Assembly.
Interview Q&A
Why hybrid instead of vectors only?
Answer
Lexical errors and semantic errors are complementary. Exact codes fail in the embedding space. Paraphrases fail in BM25. Fusion lifts recall at k on real enterprise queries that contain both.
What is reciprocal rank fusion?
Answer
Each list contributes 1 / (k + rank) for every document. You sum those contributions and sort. You never have to put BM25 and cosine on one numeric scale.
Where do ACL filters go?
Answer
Before retrieval, or as an index-time partition, so a forbidden chunk is not a candidate. Filtering after the model has read the text is a leak. Over-tight filters return nothing. Treat emptiness as a branch, not as a cue to drop the predicate.
Why rerank if you already have a vector score?
Answer
Bi-encoders are fast and shallow. A cross-encoder on the top 20 to 50 chunks is slower per row and much better at “does this passage answer this question.” You cannot afford that model on the full corpus.
Is BM25 alone ever enough?
Answer
Sometimes, for code-search-like queries dominated by exact tokens. It is usually not enough for a natural-language support bot. Say which query mix you measured.
How do you budget the latency?
Answer
Filter first. Run sparse and dense in parallel. Fuse. Rerank a capped shortlist. Generate. If the reranker times out, degrade to the fused list. Do not let one stage be unbounded.
RRF or a weighted sum?
Answer
RRF when you have not calibrated the two score distributions. A weighted sum after you fit weights on labeled queries. Learned fusion only if you will keep the labels fresh.
What is wrong with an LLM reranker over the top 200?
Answer
Cost and variance. You pay a full generation-class model to sort passages, and the order moves between calls. Keep that pattern for a tiny shortlist or skip it.
Partition or filter for multi-tenant?
Answer
Partition when a tenant is hot or the isolation bar is high. Metadata filters when you need one index and you have tests that user A never receives tenant B. Both are valid. “Search globally and hide in the UI” is not.
What do you do with an empty retrieval?
Answer
Do not invent context by dropping the ACL filter. Refuse, or relax a non-security constraint and say so. The assembly lesson turns that empty set into a prompt contract. The failure lesson turns a cross-tenant hit into an incident.
Pitfalls
Half your golden queries are pasted error codes. Half are “the job dies after the retry.” Vectors alone miss the codes. BM25 alone misses the paraphrases. Sketch the two lists, the fusion, and the one number you watch so a weight change does not silently drop the codes.