Elasticsearch & OpenSearch — Inverted Indexes, Relevance & Ops
Interview hub: inverted indexes, BM25, shards, ILM, and Elasticsearch versus OpenSearch versus Solr. Use a search cluster for full-text, filters, and aggregations, and keep multi-row transactions in an OLTP store.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does an inverted index store that _source does not?
Answer
Terms pointing at posting lists of document ids, plus frequencies, positions, and offsets when you asked for them. _source keeps the original JSON so you can return the hit.
L2
What are the three stages of an analyzer?
Answer
Character filters, then one tokenizer, then token filters. The same family usually runs at index time and at search time unless you chose an asymmetry such as edge n-grams on the way in.
L3
How do filter context and query context differ?
Answer
A filter includes or excludes and can cache a bitset. A query computes a relevance score, BM25 by default, and that score is what sort order uses when you do not override it.
L4
How does BM25 differ from classic TF-IDF?
Answer
BM25 saturates term frequency with k1, so the tenth repeat of a word adds less than the second. Parameter b normalizes for document length. An IDF-like rarity term remains.
L5
What goes wrong with shard count and custom routing?
Answer
Thousands of tiny shards inflate cluster state and heap. A custom routing key such as user id co-locates related docs and can also pin a whale to one shard.
L6
What do bulk, refresh, and ILM rollover each buy?
Answer
Bulk amortizes HTTP. Refresh opens a new searcher so documents become searchable. Rollover bounds the hot index by size or age so retention and shard size stay sane.
L7
When do you pick OpenSearch, Elasticsearch, Solr, or Postgres full text?
Answer
Apache-2 and AWS gravity point at OpenSearch. Elastic-only features point at Elasticsearch. An existing SolrCloud estate stays on Solr. A small corpus with simple ranking can stay on Postgres full text. None of them replace the OLTP source of truth.
Failure modes
Too many tiny shards
Cluster state and heap spend their budget on shard overhead instead of queries.
Mapping explosion
Dynamic fields from user input create unbounded mappings and stall the cluster.
Refresh on every write
refresh=true makes a document searchable immediately and wrecks ingest throughput.
Yellow or red ignored
Missing replicas or missing primaries stay unallocated until a disk watermark or a dead node is fixed.
Untested snapshot repo
A repository that has never been restored is not a backup.
Search as the order database
Near-real-time inverted indexes are a poor primary for multi-row ACID.
Misconceptions
Elasticsearch replaces the relational database.
It is a search and analytics projection. Orders, payments, and constraints stay in an OLTP engine.
More shards are always faster.
Extra shards add cluster-state and heap cost. Size shards toward tens of gigabytes, as a rule of thumb, and measure.
A match query is SQL LIKE.
Match analyzes text and scores it. LIKE is a character pattern with no relevance model.
Replicas speed up indexing.
Every write is copied to the replicas. Replicas buy read capacity and failover, and they add write amplification.
Interviewer traps
Rebuilding embeddings, chunking, and HNSW when the question was lexical search.
Name hybrid retrieval as an adjacent cluster. Stay on analyzers, BM25, shards, and ILM.
Designing S3 bucket IAM when the question was a snapshot repository.
Say the repo is usually object storage, then stay on register, snapshot, and restore.
Design scenario
Same prompt for every reader.
Requirements
5k search QPS, about 200 updates per second, multi-language analyzers, facet aggregations, and 99.9 percent search availability. Stock and price are filters. Title text is BM25. Logs use a separate ILM index. Nightly snapshots go to object storage.
Traffic / scale
80 million products, 5k searches per second, 200 document updates per second.
Latency
Search p99 stays interactive. Index refresh can lag about a second on the catalog path.
Consistency
The catalog service in the OLTP database is the source of truth. Search is a near-real-time projection.
Availability
Losing one data node leaves primaries allocated. Replicas cover the failed copies.
Failure assumptions
- A mapping change can drop recall until you reindex.
- One hot shard can peg two nodes while the rest idle.
- A snapshot that has never been restored may not restore.
Constraints
- Primary shards land in a rough 10 to 50 GB band.
- Heap per node stays in the low tens of GB so the OS can cache files.
- Vector retrieval stays an adjacency, not a second design.
Prompt
Design catalog search for 80 million products.
API
Which HTTP calls index a product and which call runs search with filters?
Data
Which fields are text, keyword, numeric, and where does _source sit?
Architecture
Where do coordinating nodes, primary shards, replicas, ILM, and the snapshot repository sit?
Catalog search for 80 million products
Prefer
OLTP stays the source of truth, search is the projection
The catalog service commits price, stock, and the product row. A pipeline projects a denormalized document into the search index. Filters handle stock and price. BM25 ranks title and description.
- Multi-row rules stay in the transactional database.
- Analyzers and mappings are explicit, including a keyword multi-field for facets.
- Snapshots land in an object-storage repository on a schedule.
- A down search node degrades search, not checkout.
Alternative
Replace Postgres with the search cluster
The cluster will answer text queries. It will not give you cheap multi-document transactions, foreign keys, or read-your-writes on every index call.
- A crash between two document updates leaves a torn catalog.
- Refresh lag means a buyer can miss a product that just committed.
- Dynamic mapping on merchant JSON creates a field explosion.
- Joins across orders and products become application code.
One product becomes searchable
Vertical cards for phones. The sequence diagram below is the same path.
- 1
OLTP commits the product
Price, stock, and the canonical row commit in the transactional database. - 2
Project a search document
Denormalize the fields you will filter, facet, and rank. Keep the mapping explicit. - 3
Bulk index onto a primary
The coordinating node routes by id or by a routing key. The primary writes the translog and the in-memory buffer, then copies to replicas. - 4
Refresh opens a searcher
Near-real-time visibility. Default refresh is about one second. Forcing it on every write burns ingest. - 5
Search scatters and gathers
Filters cut the set. BM25 scores the rest. One slow shard can return a partial result if you allow it.
Overview
Elasticsearch and OpenSearch are distributed search and analytics engines on Apache Lucene. They earn their place when the access pattern is full-text relevance, filters, and aggregations at scale. They are a weak primary store for multi-row transactions.
Interviewers probe four things:
- Can you explain an inverted index and an analyzer without hand-waving?
- Do you put exact constraints in filters and relevance in queries?
- Can you talk about shards, replicas, routing, and health colors?
- Do you know bulk, refresh, ILM or ISM, and snapshots, and when to refuse a search engine?
Ask this out loud: relevance dropped after a mapping change, and two nodes are hot. What do you inspect first?
Search versus OLTP
| Dimension | OLTP (Postgres or MySQL) | Search (Elasticsearch or OpenSearch) |
|---|---|---|
| Primary access | Keys, joins, transactions | Full-text, filters, aggregations |
| Consistency | ACID rows | Near-real-time after refresh |
| Schema | Tables and constraints | Mappings, with a dynamic-mapping hazard |
| Ranking | ORDER BY a column | BM25, a script, or a later rescore |
| Best for | Orders, inventory, payments | Catalog, docs, logs |
| Weak at | Fuzzy relevance | Multi-row transactions and heavy joins |
Rule of thumb: the source of truth stays in OLTP. Search holds a denormalized read model. Cold copies of that read model often live in object storage via snapshots.
What this cluster covers
- Inverted index, analyzers, tokenization, and mappings — postings, the analyzer chain,
textversuskeyword. - Query DSL, BM25, and filters versus queries —
bool, scores, and cacheable bitsets. - Sharding, replicas, routing, and cluster health — shard size, hotspots, green, yellow, red.
- Indexing pipelines, bulk, ILM, and snapshots — refresh versus flush, rollover, repositories.
- OpenSearch versus Elasticsearch versus Solr — license, API, and when to stay on Postgres full text.
Decisions
- ?
Need full text?
- noKeep the OLTP store
- yesLarge facets or fuzzy rank?
- 2
Keep the OLTP store
- ?
Large facets or fuzzy rank?
- noPostgres full text can be enough
- yesSearch cluster
- 4
Postgres full text can be enough
- 5
Search cluster
Lesson map
Elasticsearch & OpenSearch — Inverted Indexes, Relevance & Ops
Interview hub: inverted indexes, BM25, shards, ILM, and Elasticsearch versus OpenSearch versus Solr. Use a search cluster for full-text, filters, and aggregations, and keep multi-row transactions in an OLTP store.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB app["App"] coord["Coordinating node"] shard["Primary shard"] rep["Replica"] app -->|1. Bulk index| coord coord -->|2. Route by id| shard shard -->|3. Replicate the| rep app -->|5. Search with| coord coord -->|6. Scatter to a| shard coord -->|7. Scatter to| rep
Decisions
- 1
Search cluster
- nextLicense and estate?
- ?
License and estate?
- Apache 2 or AWSOpenSearch
- Elastic featuresElasticsearch
- 3
OpenSearch
- 4
Elasticsearch
Decisions
- 1
Search cluster
- nextAlready on SolrCloud?
- ?
Already on SolrCloud?
- yesStay on Solr
- noPick OpenSearch or Elasticsearch
- 3
Stay on Solr
- 4
Pick OpenSearch or Elasticsearch
License, Solr, and scale are three narrow decisions. The ecosystem lesson owns the comparison. This hub only places them on the map.
Index, then search
Sequence
- 1
App → Coordinating node
1. Bulk index documents
- 2
Coordinating node → Primary shard
2. Route by id or routing
- 3
Primary shard → Replica
3. Replicate the write
- 4
Primary shard
4. Refresh lag means not searchable yet
- 5
App → Coordinating node
5. Search with filters
- 6
Coordinating node → Primary shard
6. Scatter to a shard copy
- 7
Coordinating node → Replica
7. Scatter to another copy
- 8
Coordinating node → App
8. Gather, merge scores, sort
- 9
Coordinating node
9. A shard timeout can yield partial hits
The failure in step 4 is refresh lag. The failure in step 9 is a partial result. Both are normal knobs, and both surprise teams that treat the cluster like a synchronous SQL replica.
Preview of the choice
| Aspect | Elasticsearch | OpenSearch | Solr |
|---|---|---|---|
| Lucene core | Yes | Yes, from the fork lineage | Yes |
| License trend | Elastic License and SSPL history. Verify the current terms. | Apache 2.0 | Apache 2.0 |
| API shape | Query DSL | A familiar DSL, plus later divergence | Solr query parsers and JSON Request API |
| Ops gravity | Elastic Cloud or self-managed | OpenSearch Dashboards, including AWS | SolrCloud, historically with ZooKeeper |
| Choose when | You need Elastic features you have verified | You need Apache 2.0 and OpenSearch gravity | The estate is already Solr |
Confirm license and feature needs for the date of the interview. The ecosystem page is the place to practice that answer. Do not memorize plugin SKUs here.
Adjacent lessons
Snapshot bytes usually land in object storage. Bucket consistency, multipart upload, and lifecycle rules live on Object storage — consistency, multipart, and lifecycle. Here the contract is only register a repository, snapshot, restore, and test the restore.
Lexical plus dense retrieval is a different cluster. Start at RAG and vector databases and Hybrid retrieval — BM25, rerank, and filters. A vector field can sit beside BM25. Embeddings, chunking, and graph indexes stay on those pages.
Lucene segments are an inverted-index layout with a translog. OLTP crash recovery is Database storage engines — WAL, B-trees, and LSM trees. A segment is not an InnoDB page.
Query versus filter, in miniature
The query lesson owns scoring. This sketch only shows the shape interviews expect: scored text in must, exact constraints in filter.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Why not replace Postgres with Elasticsearch?
Answer
Elasticsearch is near-real-time search over inverted indexes. Multi-row transactions, foreign keys, and strict relational constraints stay in the OLTP database. The common pattern is a projection: the commit happens once, and search indexes the resulting document.
Filter versus query, in one sentence?
Answer
A filter answers yes or no and can be cached as a bitset. A query computes a relevance score, BM25 by default, and that score drives ranking.
Where do snapshots usually live?
Answer
In a registered snapshot repository, commonly object storage. Bucket IAM and multipart uploads belong on the object-storage lesson. This cluster only owns the search-side contract.
What do you inspect when relevance drops after a mapping change?
Answer
The analyzer and the field type first. A text field that became keyword, or an analyzer that stopped lowercasing, changes the terms in the index. Then check whether the hot nodes own the shards that hold the new mapping, and whether a reindex is still running.
Do replicas make indexing faster?
Answer
No. Each write is copied to every replica. Replicas add read throughput and a failover copy. They add write amplification.
Is a match query the same as SQL LIKE?
Answer
No. match analyzes the text and scores it. LIKE is a character pattern. Exact ids belong on a keyword field with a term query in filter context.
What is the interview-safe shard size story?
Answer
Aim for fewer, larger shards rather than thousands of tiny ones. A common band is about 10 to 50 GB per shard, and under a couple hundred million documents, as a rule of thumb you then measure. Changing primary shard count later means reindex, split, or shrink.
Where do vectors fit?
Answer
Only as an adjacency. Hybrid lexical plus dense retrieval has its own lessons. Say that a search cluster can store a vector field, then stop. Do not rebuild chunking or graph indexes on this whiteboard.
Pitfalls
Draw the OLTP commit, the bulk request, one primary, one replica, refresh, and a search that filters in_stock and scores title. Mark which box is allowed to lag, and which lesson owns analyzers, BM25, shard count, and snapshots.