Ecosystem — OpenSearch vs Elasticsearch vs Solr (and when not to use a search engine)
Elasticsearch, OpenSearch, and Solr all sit on Lucene. They differ in API, license, and operations. The senior answer includes when a search engine is the wrong tool.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What do Elasticsearch, OpenSearch, and Solr share?
Answer
Lucene inverted indexes, analyzers, and BM25-style scoring. If you understand segments and postings, the rest is API and operations.
L2
Why does OpenSearch exist?
Answer
It was forked from Elasticsearch 7.10 after Elastic changed the license. It is Apache 2.0 and has its own releases, plugins, and Dashboards. APIs overlap and then diverge.
L3
Is OpenSearch a permanent drop-in?
Answer
No. Treat compatibility as a version you test. Reindex, security plugins, and query features drift. A client pinned to one dialect will break on the other.
L4
Solr versus Elasticsearch in one line?
Answer
Both are Lucene. SolrCloud historically centers ZooKeeper and collections. Elasticsearch centers its own cluster state and Query DSL. Choose by estate and team skill, not by a slogan.
L5
When is Postgres full text enough?
Answer
A modest corpus, simple ranking, a strong wish to operate one system, and no heavy facet or log-analytics load.
L6
How should the OLTP database and the search index stay in sync?
Answer
The OLTP commit is the source of truth. A change stream or an outbox projects documents into search. Dual-writing from the request is how the two stores diverge.
L7
Where do vectors sit in this choice?
Answer
A search distribution can store a vector field, or a specialist vector store can sit beside it. That design is the RAG and hybrid-retrieval lessons. This page only says a search engine may participate.
Failure modes
Search as the order database
Multi-row invariants move into application code and refresh lag becomes a payments bug.
Assumed drop-in compatibility
A plugin or query that exists on one fork fails on the other during a migration.
Dual-write from the app
The row commits and the index call does not, or the reverse.
License surprise
A feature or a distribution you cannot ship shows up after the design review.
Solr and Elasticsearch side by side with no owner
Two Lucene estates, two sets of shards, and nobody who can restore either.
Misconceptions
OpenSearch is only a rename.
It is a fork. Shared history is not a promise that today's APIs match.
Solr is obsolete because Elasticsearch exists.
SolrCloud is a maintained Lucene distribution. An existing estate is a reason to stay, not a reason to rewrite.
Any text search requires a cluster.
Postgres full text covers modest corpora. Add a search cluster when relevance, facets, or ingest scale justify the second system.
Interviewer traps
A license rant with no workload.
State the constraint you would verify, then choose from corpus size, facets, and who operates the cluster.
Re-teaching HNSW because the word vector appeared.
Say the search engine can participate in hybrid retrieval, then point at that lesson.
Design scenario
Same prompt for every reader.
Requirements
Orders are the source of truth and need transactions. Product search needs facets once the catalog passes a few million documents. Backups of the search index must leave the cluster.
Traffic / scale
Catalog may grow to tens of millions of documents. Order writes stay on the OLTP primary.
Latency
Search can lag the commit by about a second. Checkout cannot lag on the search cluster.
Consistency
The order row is authoritative. The product document in search is a projection.
Availability
Search may degrade while orders continue. A search restore uses a snapshot repository.
Failure assumptions
- The projection can lag or fail without rolling back the order.
- Fork APIs are compatible only for the versions you tested.
- A snapshot repository shares fate with its bucket if you never restore.
Constraints
- Do not make the search cluster the primary for orders.
- Name the license constraint instead of inventing a vendor verdict.
- Leave vector fusion on the hybrid retrieval lesson.
Prompt
Choose a store for product text, order rows, and a nightly backup.
API
Which calls hit OLTP, and which call is the search projection?
Data
What is the source of truth for an order versus a product document?
Architecture
Where do Postgres, the search distribution, and the snapshot repository sit?
A new catalog, no existing search estate
Prefer
OLTP for orders, a Lucene distribution only if search features earn it
Start from the workload. Transactions stay in Postgres or MySQL. If facets and relevance at tens of millions of documents are real, pick OpenSearch or Elasticsearch from the license and the features you have checked, and project documents in.
- One sentence on Lucene covers the engine.
- The license check is a dated fact, not a tribal identity.
- Postgres full text remains the answer under half a million simple documents.
- Snapshots still leave the cluster.
Alternative
Standardize on a search engine for every store
Orders, sessions, and reports move into a cluster that is near-real-time and weak at multi-document transactions.
- Checkout depends on refresh lag.
- You operate shards for a key-value workload.
- A fork migration becomes a rewrite of the system of record.
- Two copies drift if the app dual-writes.
Pick a store from the requirement
Text search is a feature. A search cluster is a system.
- 1
Name the source of truth
Orders and payments that need multi-row transactions stay in OLTP. - 2
Measure the text need
Simple contains-style search on a small catalog can stay in Postgres full text. - 3
Add a search cluster for facets and scale
Relevance, aggregations, and a large ingest are the features that earn the second system. - 4
Pick the distribution
Apache 2.0 and OpenSearch gravity, Elastic features you verified, or the SolrCloud estate you already run. - 5
Project, do not dual-write
The commit produces the document. A change stream or an outbox indexes it. Lag is an SLO.
Overview
Elasticsearch, OpenSearch, and Apache Solr are three ways to operate Lucene. Inverted indexes, analyzers, and BM25 are the previous lessons. This page is the choice: which distribution, and when the answer is to skip a search cluster.
Interview discipline: state tradeoffs, and do not litigate a vendor. Confirm license and feature needs for the date in front of you. A sentence that was true in 2021 about a plugin may be stale.
One engine, three distributions
Flow
- 1
Lucene segments, postings, BM25
- nextElasticsearch APIs and cluster state
- 2
Elasticsearch APIs and cluster state
- nextOpenSearch APIs and Dashboards
- 3
OpenSearch APIs and Dashboards
Lesson map
Ecosystem — OpenSearch vs Elasticsearch vs Solr (and when not to use a search
Elasticsearch, OpenSearch, and Solr all sit on Lucene. They differ in API, license, and operations. The senior answer includes when a search engine is the wrong tool.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB lucene["Lucene segments, postings, BM25"] es["Elasticsearch APIs and cluster state"] os["OpenSearch APIs and Dashboards"] lucene -->|Lucene segments, postings, BM25| es es -->|Elasticsearch APIs and cluster state| os
Decisions
- 1
Same Lucene ideas
- nextExisting SolrCloud?
- ?
Existing SolrCloud?
- yesOperate Solr
- noElasticsearch or OpenSearch
- 3
Operate Solr
- 4
Elasticsearch or OpenSearch
If you can explain postings, analyzers, BM25, and segments, moving between distributions is mostly API and operations translation. The sharding and ingest lessons still apply. The button names change.
| Dimension | Elasticsearch | OpenSearch | Apache Solr |
|---|---|---|---|
| Lineage | Elastic's distribution of Lucene | Fork from the Elasticsearch 7.10 line | Long-running Lucene server |
| License trend | Elastic License and SSPL history. Verify the current terms. | Apache 2.0 | Apache 2.0 |
| Query API | Query DSL | A familiar DSL that has diverged | Query parsers and the JSON Request API |
| UI | Kibana | OpenSearch Dashboards | Admin UI and admin APIs |
| Cluster brain | Elasticsearch cluster state | OpenSearch cluster state | SolrCloud, historically ZooKeeper |
| Cloud gravity | Elastic Cloud | Amazon OpenSearch Service and others | Self-managed and various hosts |
| Best fit | Teams that need Elastic features they have checked | Apache 2.0 and an OpenSearch estate | An existing SolrCloud estate |
OpenSearch is not a rename. Shared clients and a shared DSL get you started. Security plugins, index APIs, and newer query features move on different calendars. Treat a migration as a test matrix: bulk, mappings, a representative query, a snapshot restore, and the security plugin. "It worked on 7.10" is a history fact, not a plan.
Solr in one line. SolrCloud groups cores into collections, elects a leader per shard, and has long used ZooKeeper for cluster state. Elasticsearch keeps that state in its own cluster. Neither fact makes the other wrong. An estate with Solr operators and solrconfig is a reason to stay.
When a search engine is the wrong tool
| Need | Prefer | Why |
|---|---|---|
| Multi-row orders and payments | Postgres or MySQL | Transactions and constraints |
| Session or simple key-value | A key-value store | A search cluster is pure overhead |
| Small corpus, rare text search | Postgres full text (tsvector) | One system, simple ranking |
| Heavy relational reporting | A warehouse | Joins and governance |
| Semantic search as the whole product | A vector path you design on purpose | See the RAG lessons. A search node is only one option |
Postgres full text is enough when the catalog is modest, the rank is "some of these words, newer first", and the team does not want a second cluster. It becomes the wrong tool when facets, typo tolerance, language analyzers, and log-scale ingest are the product.
Decisions
- ?
Real text search?
- noDo not add a search cluster
- yesFacets or large corpus?
- 2
Do not add a search cluster
- ?
Facets or large corpus?
- noPostgres full text
- yesSearch distribution
- 4
Postgres full text
- 5
Search distribution
Decisions
- 1
Search distribution
- nextApache 2.0 required?
- ?
Apache 2.0 required?
- yesOpenSearch
- noElasticsearch if the features hold
- 3
OpenSearch
- 4
Elasticsearch if the features hold
Projection, not a second primary
| Pattern | What you gain | What you owe |
|---|---|---|
| Change stream or outbox into the index | The OLTP commit is the fact | Lag, and an idempotent indexer |
| Dual-write from the request | Fewer moving parts on day one | A crash leaves one store updated |
| Reindex into a new index | A clean mapping change | Temporary disk and CPU |
| Snapshot restore onto another cluster | Disaster recovery and some migrations | Version compatibility |
Dual-write is the trap. The application updates Postgres and then calls the index API. One success and one failure is permanent drift. Prefer a projection from the commit. This page does not re-teach the change-stream machinery.
Reindex is how mappings evolve inside one distribution. Snapshot restore is how you copy committed indices between clusters, after you check versions. Those mechanics are the bulk, ILM, and snapshots lesson. The repository itself usually sits on object storage. Bucket design is Object storage — consistency, multipart, and lifecycle.
Vectors, one sentence
A search distribution can store a vector next to BM25, or a separate vector store can own that job. Embeddings, chunking, and fusion live on RAG and vector databases and Hybrid retrieval — BM25, rerank, and filters. The only decision here is whether the lexical cluster should be in the design at all.
A coarse recommendation
This helper is an interview sketch. It is not a vendor selector and it does not know your license.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Half a million documents without facets is the sketch's cutoff for Postgres full text. Facets, or a large corpus, cross into a search engine. Acid without relevance stays on OLTP even at 10 million rows.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Is OpenSearch a drop-in replacement for Elasticsearch forever?
Answer
They share a lineage and many APIs. The forks diverge. Test bulk, mappings, queries, security, and snapshot restore on the versions you will run. Do not promise drop-in compatibility past the matrix you ran.
Solr versus Elasticsearch in one line?
Answer
Both run Lucene. SolrCloud is a collection plus ZooKeeper-centered operations and its own request shape. Elasticsearch is Query DSL and its own cluster state. Pick the one your estate and your operators already know, unless you are starting clean.
When is Postgres full text enough?
Answer
A modest corpus, simple ranking, and a preference for one operational system, without a heavy facet or analytics load. When those features become the product, add a search cluster and keep Postgres as the source of truth.
Why mention the license at all?
Answer
Because it changes what you are allowed to ship and which hosted service you can buy. State the constraint, say you would read the current license, and move to the workload. A political speech is not an architecture.
What is the dual-write failure?
Answer
The application writes the row and then indexes the document. A crash or a timeout between them leaves one store updated. Project from the commit so a failed index call can retry from a durable fact.
Can you restore an Elasticsearch snapshot into OpenSearch?
Answer
Only when the versions and the repository format say you can. Treat it as a tested migration, not as a right. The snapshot lesson owns the restore steps.
Where do vectors go in this decision?
Answer
Decide first whether you need lexical search. If you also need dense retrieval, the RAG and hybrid lessons choose the index and the fusion. This page stops at "the search cluster may hold a vector field".
What would you operate on day two?
Answer
Mappings, shard size, refresh interval, rollover, snapshot restores, and the projection lag. The distribution choice does not remove those. It only changes the admin UI and the license page.
Pitfalls
For orders, a 200 megabyte help center, an 80 million product catalog with facets, and a session cache, name the store. For the catalog, say which distribution you would evaluate and which fact you would verify before you commit.