Inverted Index, Analyzers, Tokenization & Mappings
An inverted index maps each term to a posting list. Analyzers turn raw strings into those terms, and mappings decide which analyzer and field type apply. A wrong mapping silently drops recall or explodes the cluster.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is the difference between the inverted index and _source?
Answer
The inverted index answers which documents contain a term. _source, or stored fields, returns the JSON you show to the client.
L2
Why do sorts and aggregations avoid the inverted index?
Answer
They read doc values, a columnar structure per field. Norms carry length factors used at score time.
L3
Name the analyzer stages in order.
Answer
Character filters, then exactly one tokenizer, then token filters such as lowercase, stop, stem, or synonym.
L4
When is index-time analysis allowed to differ from search-time analysis?
Answer
When you mean it. Edge n-grams at index time plus a standard search analyzer is a classic autocomplete shape. An accidental mismatch drops recall.
L5
Why does a term query on a text field miss?
Answer
term skips analysis. The indexed tokens were analyzed, often lowercased and split. Use match on text, or term on a keyword or a .keyword multi-field.
L6
What is a multi-field?
Answer
One JSON field indexed two ways. title as text for relevance, title.keyword as keyword for sorts and aggregations.
L7
What is a mapping explosion?
Answer
Unbounded distinct field names, often from dynamic mapping of user keys, create a huge mapping. Heap grows, cluster state bloats, and indexing stalls. Use strict dynamic or a template.
Failure modes
Term query on analyzed text
The exact string never lands in the index, so the lookup misses.
Analyzer mismatch
Index-time tokens and query-time tokens diverge, and recall falls with no error.
Mapping explosion
A new field per user attribute exhausts heap and cluster state.
Keyword used for prose
The whole sentence is one term, so full-text queries cannot hit the words inside it.
Nested objects queried as flat objects
Correlated fields inside an array match across elements and return the wrong parent.
Misconceptions
The inverted index stores the original JSON.
_source stores the original document. The inverted index stores terms and postings.
Standard analysis is safe for SKUs.
Standard analysis splits and lowercases. SKUs and enums belong on keyword, with a text multi-field only if you also need partial match.
Dynamic mapping is free convenience.
It is convenient for prototypes and dangerous when field names come from tenants.
Interviewer traps
Explaining HNSW when asked how a text field is stored.
Stay on postings, analyzers, and mappings. Point at the RAG lessons for vectors.
Saying term and match are interchangeable.
term is exact and unanalyzed. match runs the search analyzer and scores.
How should a product title and a SKU be indexed?
Prefer
text for the title, keyword for the SKU, multi-field when you need both
The title is analyzed so running matches Running Shoes. The SKU stays one exact term so filters and facets do not split SKU-42A into pieces.
- Recall on prose comes from the analyzer.
- Exact ids, sorts, and low-cardinality facets use keyword.
- title.keyword exists only when you also sort or facet on the raw title.
- The mapping is written down, not inferred from the first document.
Alternative
One dynamic text field for every JSON key
The first document type-checks the cluster. Merchant attributes then become new fields forever.
- SKU-42A may be tokenized into sku and 42a.
- A term filter for the original SKU misses.
- Field count grows with tenants, not with your schema.
- You notice during a heap alert, not during indexing of document one.
From raw string to a posting
Same three stages at index time. Search time should meet those terms on purpose.
- 1
Character filters
Strip HTML or map characters before any split. Over-normalizing an id here is already too late to undo. - 2
Tokenizer
standard, whitespace, keyword, or edge n-gram. This chooses the boundaries. - 3
Token filters
Lowercase, stop words, stemming, synonyms. Aggressive stems raise recall and can glue unrelated words. - 4
Terms hit the inverted index
Each term appends a posting: document id, and frequency or positions if the field asked for them. - 5
Query must emit the same terms
A different search analyzer looks up words that were never stored.
Overview
The inverted index is Lucene's core structure. For each term, a posting list holds document ids plus whatever you configured: term frequency, positions, offsets. Analyzers turn strings into terms at index time and, usually, at query time. Mappings declare the field type and which analyzer applies.
A wrong mapping fails quietly. Recall drops, or the cluster grows a field per user and then falls over. This page is that contract. Scoring math is the next lesson.
Four structures, four jobs
| Structure | Holds | Used for |
|---|---|---|
| Inverted index | Term to postings | Finding candidate documents |
_source or stored fields | Original JSON, or a subset | Returning hits |
| Doc values | Columnar values per field | Sorting, aggregations, scripts |
| Norms | Length-related factors | Scoring |
Interview line: search finds through the inverted index, display uses _source, and sort or aggregate uses doc values.
Flow
- 1
Term shoe
- nextPostings for docs 3, 7, 19
- 2
Postings for docs 3, 7, 19
- nextLoad _source for the hit
- 3
Load _source for the hit
- nextReturn the product JSON
- 4
Return the product JSON
Lesson map
Inverted Index, Analyzers, Tokenization & Mappings
An inverted index maps each term to a posting list. Analyzers turn raw strings into those terms, and mappings decide which analyzer and field type apply. A wrong mapping silently drops recall or explodes the cluster.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB term["Term shoe"] postings["Postings for docs 3, 7, 19"] source["Load _source for the hit"] client["Return the product JSON"] term -->|Term shoe to Postings for docs 3, 7, 19| postings postings -->|Postings for docs 3, 7, 19| source source -->|Load _source for the hit| client
Doc 3's title might be "Running shoe" and its brand "Acme". The term shoe only knows the document id. Positions, if you stored them, are how phrase queries prove the words were neighbors.
Analyzer pipeline
Flow
- 1
1. Raw string
- next2. Character filters
- 2
2. Character filters
- next3. Tokenizer
- 3
3. Tokenizer
- next4. Token filters
- 4
4. Token filters
- next5. Terms in the inverted index
- 5
5. Terms in the inverted index
| Stage | Examples | When it helps | When it hurts |
|---|---|---|---|
| Character filter | HTML strip, character map | Clean markup early | Destroys ids that needed their punctuation |
| Tokenizer | standard, whitespace, n-gram, keyword | Defines token boundaries | A prose tokenizer on a SKU splits the id |
| Token filter | lowercase, stop, stem, synonym | Normalization and recall | Heavy stemming hurts precision |
Index time versus search time. The usual choice is one analyzer family on both sides. An intentional asymmetry, such as edge n-grams at index time and a standard analyzer at search time, is a product decision. OpenSearch documents the lookup order: query analyzer, then search_analyzer, then the index default, then the field analyzer, then standard.
Mapping essentials
| Field type | Use | Watch-out |
|---|---|---|
text | Analyzed full text | Poor for exact ids |
keyword | Exact values, aggregations, sorts | High-cardinality aggregations are expensive |
numeric or date | Ranges and sorts | Prefer a real numeric type over a string |
boolean | Flags | Comfortable in a filter |
object | JSON without correlation | Array elements match across each other |
nested | Arrays where fields stay together | Extra query cost |
dense_vector | A place to store a vector | Index layout for vectors lives on the RAG lessons |
Dynamic mapping infers types from the first value. That is convenient in a prototype. It is dangerous when field names are user input. Unexpected keys create new fields. A mapping explosion, millions of fields, burns heap and stalls the cluster. Disable dynamic mapping or lock it with a strict template and a field limit you have actually set.
What the tokens actually are
| Input | standard plus lowercase | keyword | edge n-gram sketch |
|---|---|---|---|
| Running Shoes | running, shoes | Running Shoes as one term | prefixes such as ru, run |
| SKU-42A | often sku and 42a | SKU-42A | useful only if you meant autocomplete |
Rule: product SKUs and enums go to keyword. Titles and descriptions go to text with a language analyzer. Add a multi-field when one JSON property needs both behaviors.
A multi-field indexes one input twice. title as text serves match. title.keyword as keyword serves a sort or a facet. You pay disk and indexing CPU for the second view. You pay recall if you skip it and then term-query the analyzed side.
Nested, lightly
An object array flattens. A color of red on size small and a color of blue on size large can match a query for red and large even though no single element has both. nested keeps each element as its own hidden document so the correlation holds. Use it when the correlation matters. It costs a heavier query. This is still a mapping choice, not a join engine.
Vectors, one sentence
A dense_vector field can live in the same index as title. How you embed, chunk, and index vectors is RAG and vector databases. This page stops at the mapping type.
A toy analyzer
This is not Lucene. It shows the three beats: clean characters, split, then filter tokens.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The SKU line is the lesson. Punctuation became a split, so the exact id is gone. That string belonged on a keyword path that does not tokenize.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Why can a term query on a text field miss?
Answer
term does not analyze. The index holds the analyzed tokens, often lowercased and split. Query the analyzed field with match, or run term against a keyword field or a .keyword multi-field.
What is a multi-field?
Answer
One JSON field indexed in two ways. A title can be text for relevance and title.keyword for sorting and aggregations.
What is a mapping explosion?
Answer
Unbounded distinct field names, often dynamic mapping on user input, create a huge mapping. Heap and cluster state grow until indexing and cluster updates stall. Set dynamic to strict, or constrain names with a template and a field limit.
Where does the original JSON live?
Answer
In _source, unless you disabled it. The inverted index is not a document store. Returning a hit loads _source or stored fields after the posting list identifies the id.
Why not stem everything?
Answer
Stemming raises recall. It also collapses words a shopper treats as different, which hurts precision. Language analyzers are a choice per field, not a default you apply to SKUs and proper nouns.
When do you store positions?
Answer
When you need phrase or proximity queries. Positions cost disk and indexing time. A field that is only ever a bag of words for a loose match can skip them. Say the cost out loud.
Object or nested?
Answer
Use object when you do not care which array element a field came from. Use nested when color and size must come from the same element. Nested has a query cost. It is still not a relational join.
Does this page own vector search?
Answer
No. A mapping may include dense_vector. Embeddings, chunking, and the vector index live on the RAG lessons. Hybrid scoring lives on the hybrid-retrieval lesson.
Pitfalls
Write the tokens you expect for "The Running Shoes!", "SKU-42A", and "red small / blue large" as an object array. Then mark which field type you would store, and which query would miss if you picked the other type.