Query DSL, Relevance Scoring (TF-IDF/BM25) & Filters vs Queries
Query DSL composes bool, match, term, and range clauses. Relevance defaults to BM25. Filters answer yes or no and cache well. Queries compute the score that ranks hits.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a filter return that a query also returns?
Answer
Membership. The filter stops there. The query also returns a score.
L2
Which clauses are the usual filters?
Answer
term, terms, range, and exists on exact fields such as stock, price, tenant id, and status.
L3
What do must, should, filter, and must_not each mean?
Answer
must is required and scored. should adds optional score, and it can be required when no must or filter is present. filter and must_not are not scored.
L4
How does BM25 treat a repeated term versus classic TF-IDF?
Answer
Classic TF often keeps growing with raw or log frequency. BM25 approaches a ceiling controlled by k1, so stuffing a word helps less after the first occurrences.
L5
What does b do?
Answer
It sets how strongly document length changes the score. b near 1 penalizes long documents more. b at 0 ignores length. The default in Elasticsearch is 0.75, with k1 at 1.2.
L6
Why can the same document score differently on two shards?
Answer
BM25 uses local term statistics by default. A rare word on a small shard looks rarer than the same word on a large shard. dfs_query_then_fetch gathers global frequencies at a cost.
L7
Where do kNN and rerankers belong?
Answer
On the hybrid retrieval lesson. This page can name a lexical score as one input. It does not design the fusion.
Failure modes
Scored tenant filter
tenant id sits in must, so every hit from that tenant gets a score bump that is unrelated to the text.
Analyzer mismatch
The query emits tokens the index does not contain, and the bool clause matches nothing.
query_string injection
User input with operators changes the query shape.
should without minimum_should_match
Optional clauses either match nothing useful or, when they are the only clauses, surprise you by becoming required.
One giant shard statistic story
You tune relevance on a single-shard lab index and ship different scores once the index has many shards.
Misconceptions
Filters are just slow queries with the score thrown away.
Filter context skips scoring and can cache bitsets. That is a different execution path.
Adding replicas changes the BM25 formula.
A replica is a copy of a shard. The score is computed on the copy that serves the request. Replica count is an availability and read-capacity knob.
BM25 replaced IDF.
BM25 still has an IDF-like rarity term. What it changes is term-frequency saturation and length normalization.
Interviewer traps
Deriving the full BM25 equation from memory and missing filter context.
Name k1, b, and IDF, then show the bool clause that is not scored.
Designing a cross-encoder reranker for a keyword catalog.
Point at hybrid retrieval. Stay on DSL and BM25.
Catalog search with a brand and a minimum price
Prefer
Score the text, filter the constraints
multi_match on title and description produces the BM25 score. Brand and price live in filter so they do not add a fake relevance bump.
- The ranking answers how well the text matched.
- Repeated brand filters can hit a cached bitset.
- A missing brand returns no hit, without a confusing score.
- Field boosts stay on the text fields you meant to boost.
Alternative
Put brand, price, and title all in must
Every clause contributes a score. A rare brand token can outrank a better title match.
- Tenant or brand frequency leaks into relevance.
- Range clauses are a poor fit for BM25.
- Cache behavior is worse than filter context.
- Tuning becomes a pile of boosts that hide the bug.
How one bool query runs
Filters narrow the set. Scores are computed for the clauses that asked for them.
- 1
Parse the DSL
bool splits clauses into score and no-score buckets. - 2
Apply filters
term, terms, range, and exists build or reuse bitsets. must_not removes hits the same way. - 3
Score must and should
Each text clause asks the inverted index for postings and computes BM25 on that shard. - 4
Combine and sort
The coordinating node merges shard hits. Default sort is score. - 5
Optional global stats
dfs_query_then_fetch spends a round trip so IDF is global. Skip it until scores are actually unstable.
Overview
Query DSL is the JSON language of Elasticsearch and, with later divergence, OpenSearch. The clauses you will say out loud are bool, match, match_phrase, multi_match, term, terms, range, and sometimes function_score.
Senior interviews hinge on filter context versus query context. Filters do not score and cache well. Queries score and drive ranking. The default similarity since Elasticsearch 5 and modern Lucene is BM25, not classic TF-IDF.
Analyzers decide which tokens exist. This page decides which of those tokens are constraints and which are relevance.
Filter context and query context
| Aspect | Filter context | Query context |
|---|---|---|
| Output | Include or exclude | Score, and include |
| Caching | Bitsets are often cached | Scores depend on the query shape |
| Typical clauses | term, terms, range, exists | match, match_phrase, multi_match |
| Use for | Stock, price, tenant id, status | Title and body relevance |
| Failure mode | Fuzzy text stuffed only into a filter | Scoring an exact id with match |
Flow
- 1
bool query
- nextmust, scored and required
- 2
must, scored and required
- nextshould, scored and optional
- 3
should, scored and optional
- nextBM25 adds those scores
- 4
BM25 adds those scores
Lesson map
Query DSL, Relevance Scoring (TF-IDF/BM25) & Filters vs Queries
Query DSL composes bool, match, term, and range clauses. Relevance defaults to BM25. Filters answer yes or no and cache well. Queries compute the score that ranks hits.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB bool["bool query"] must["must, scored and required"] should["should, scored and optional"] bm["BM25 adds those scores"] bool -->|bool query to must, scored and required| must must -->|must, scored and required| should should -->|should, scored and optional| bm
Flow
- 1
bool query
- nextfilter, not scored
- 2
filter, not scored
- nextmust_not, not scored
- 3
must_not, not scored
- nextCacheable bitsets
- 4
Cacheable bitsets
should rules. With a must or a filter, each should is optional extra score. With neither, at least one should must match, unless you set minimum_should_match. Say that sentence. It is a common stumble.
BM25 versus classic TF-IDF
Classic TF-IDF multiplies a term-frequency signal by an inverse-document-frequency signal. Variants take a log of tf. The reward for repeating a word keeps climbing.
BM25 (Okapi) changes two things that interviews care about:
- Term frequency saturates.
k1controls how fast extra occurrences stop helping. In Elasticsearch the defaultk1is 1.2. For an average-length document, that is about the tf where the term has earned half of its maximum tf contribution. - Length is explicit.
bcontrols how much a long document is discounted. The defaultbis 0.75.bof 0 ignores length.bof 1 fully normalizes by length. - Rarity remains. An IDF-like component still boosts terms that are rare in the shard.
| Idea | Classic TF-IDF | BM25 |
|---|---|---|
| Term frequency | Often raw or log, still climbing | Saturating, controlled by k1 |
| Document length | Sometimes light | Explicit, controlled by b |
| Rare terms | IDF boosts them | An IDF-like term remains |
| Default in modern Lucene | Historical | Current default similarity |
Interview line: BM25 reduces the reward for stuffing the same term and adjusts for long documents. It is a better default for products and articles than naive TF. You still evaluate on your queries. k1 and b are per-field settings, not a global magic number you change on a hunch.
Scores are per shard
By default Elasticsearch computes BM25 with that shard's document count and term frequencies. Two shards with different contents give different IDF for the same word. Practical consequences:
- A one-shard lab index will not reproduce production scores.
- Very small shards make IDF noisy.
search_type=dfs_query_then_fetchcollects global frequencies first, then scores. It costs a round trip. Use it when you have measured a real skew, not as a default.
Replicas do not change this formula. They copy a shard so another node can serve the same statistics. Replica count is availability and read capacity. The sharding lesson owns placement.
Patterns you should be able to place
| Pattern | Shape | Strength | Cost |
|---|---|---|---|
match | Analyzed full text | Default relevance | Can be broad |
match_phrase | Ordered proximity | Precision for phrases | Misses reordered words |
multi_match | Several fields, with boosts | Title versus body | Boosts need a reason |
term / terms | Exact, no analysis | Fast filters | Misses analyzed text |
bool | must, should, filter, must_not | The composition tool | Easy to over-should |
function_score | Business signals after the text | Popularity, recency | Easy to drown out BM25 |
query_string and simple_query_string parse operators. Raw user text in query_string is an injection risk: a user can change the query with quotes and operators. Prefer a structured match plus filters.
Hybrid, one sentence
Production search sometimes blends a BM25 score with a vector score and a reranker. That fusion, including reciprocal rank fusion, lives on Hybrid retrieval — BM25, rerank, and filters. This page's job is the lexical score and the filter bitset.
Build the bool body
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Does adding replicas change BM25 scores?
Answer
The formula does not consult replica count. Scoring uses the shard copy that answers. Replicas exist for read capacity and failover. They do add write amplification, which is a sharding topic, not a similarity topic.
When do you use should versus must?
Answer
must is required. should contributes optional relevance when a must or filter already defines the set. If should is the only clause, at least one should must match unless minimum_should_match says otherwise.
Why does tenant id belong in filter?
Answer
It is an exact constraint. It needs no score, it caches well, and a scored tenant term would change ranking for a reason the shopper did not type.
What do k1 and b mean, with defaults?
Answer
k1 saturates term frequency. The Elasticsearch default is 1.2. b controls length normalization. The default is 0.75. Raise k1 when extra repetitions should still matter. Lower b when long and short fields should be treated more alike.
Why did production scores disagree with the lab?
Answer
The lab index was probably one shard. Production BM25 uses per-shard document frequency. Small or uneven shards make IDF noisy. Fix shard balance before you reach for dfs_query_then_fetch.
Is match_phrase always better than match?
Answer
It is stricter. The tokens must appear in order with limited slop. That raises precision for names and lowers recall when shoppers reorder words. Use it when the phrase is the product, not as a global replacement.
What is wrong with query_string for a search box?
Answer
The parser honors operators. A user can change your query with quotes, fields, and wildcards. A structured match or multi_match plus filters keeps the shape you designed.
Where does hybrid RAG scoring go?
Answer
Here you can say the lexical clause is one input. Fusion with vectors and rerankers is the hybrid retrieval lesson. Do not invent a chunking scheme on this whiteboard.
Pitfalls
Take a query that puts tenant id, price, and title all in must. Move the exact clauses to filter, keep a multi_match on title and description, and say what happens to the score of a long description that repeats the query word ten times.