Indexing Pipelines, Bulk, ILM & Snapshots
Production search lives or dies on the ingest path: bulk requests, refresh versus flush, ILM or ISM rollover, and snapshots to a repository that is usually object storage.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why is the bulk API the default ingest path?
Answer
One HTTP request carries many actions. Single-document index calls spend the budget on round trips.
L2
What is a partial bulk failure?
Answer
The HTTP call can succeed while individual items fail. You read the item errors, retry the failed ids, and back off when the node returns 429.
L3
What does refresh do?
Answer
It opens a new searcher over the in-memory buffer so documents become visible. The default interval is about one second. refresh=true on every request forces that work immediately.
L4
What do the translog and flush add?
Answer
The translog records operations so a shard can replay them. A flush commits Lucene segments and can truncate the translog. Durability is this pair, not the refresh.
L5
Why roll an index instead of writing one index forever?
Answer
Rollover bounds shard size, lets you delete a whole old index for retention, and keeps the hot write on a small backing index behind an alias.
L6
How do ILM and ISM relate?
Answer
Elasticsearch Index Lifecycle Management and OpenSearch Index State Management tell the same story with different APIs: hot, then cheaper phases, then delete, with rollover on the hot alias.
L7
What proves a snapshot works?
Answer
A restore you have done. Incremental snapshots are the ongoing copy. A repository on the same disk as the cluster is not a disaster plan.
Failure modes
refresh=true in production
Every bulk item opens a searcher. Ingest collapses and segment counts climb.
Ignoring item errors
The bulk response is 200 and half the documents never indexed.
Retry without a stable id
A timeout retry creates a second document.
One index for all time
You cannot drop old data cheaply, and shard size leaves the planning band.
Untested repository
Credentials, permissions, or version skew show up on the day you need the restore.
Heavy script in an ingest pipeline
The cluster does transformation work that belonged upstream.
Misconceptions
Refresh means the write is durable.
Refresh means a searcher can see the document. The translog and flush are the durability path.
A snapshot is a substitute for replicas.
Replicas are online copies for failover. A snapshot is a point-in-time backup in a repository, used for restore and disaster recovery.
ILM is only an Elasticsearch paid feature you can skip.
The idea is rollover plus retention. OpenSearch ships the same idea as Index State Management. Name the product you operate and the phases.
Interviewer traps
Designing S3 bucket policies when asked about snapshots.
Say the repository is usually object storage. Stay on register, snapshot, restore, and a tested drill.
Quoting WAL checkpoint details for a Lucene flush.
Point at the storage-engine lesson for ARIES. Here, name refresh, translog, and flush.
Logs at a steady ingest rate
Prefer
Bulk into a write alias that rolls
Clients index logs-write. ILM or ISM rolls the backing index on size or age. Readers use a broader alias. Refresh stays on an interval. Snapshots copy committed segments to a repository.
- HTTP overhead is amortized.
- Shard size stays inside the band you planned.
- Deleting an old backing index is the retention tool.
- A restore drill has a target that is not the only copy.
Alternative
One index call per event, refresh=true, one index forever
The demo shows the document in search immediately. Production spends its CPU opening searchers, and retention means a delete-by-query across a huge shard.
- Round trips dominate.
- Segment count climbs.
- A retry without an id duplicates the event.
- The only copy of the index sits on the disks that just failed.
From bulk item to off-site copy
Searchable and durable are different timestamps.
- 1
Bulk request lands
Many actions, stable ids, a payload sized in megabytes rather than in wishes. - 2
Primary writes the translog
The operation is durable on the translog before you depend on a Lucene commit. Replicas get the write too. - 3
Refresh publishes a searcher
About one second later the document is visible. Forcing it per request is a test hook. - 4
Flush commits segments
A Lucene commit point exists. The translog can drop operations the commit covers. - 5
Snapshot copies the commit
The repository, usually object storage, stores an incremental copy. Restore is the only proof.
Overview
Search clusters fail in production on the ingest path more often than on a clever query. The pieces are the bulk API, the difference between refresh and flush, light ingest pipelines, ILM or ISM for rollover, and snapshots to a repository you do not keep on the same disk.
Shard count is the previous lesson. This page assumes the primaries exist and asks how bytes arrive, when they become searchable, and how you get them back.
Bulk indexing
| Approach | What you gain | What breaks |
|---|---|---|
| Single-document index API | Simple demos | Throughput |
| Bulk API | Amortized HTTP | Item-level errors, and 429 under load |
| Enormous payloads | Fewer round trips | Memory spikes and timeouts |
| Many parallel bulks | Higher ingest | Hot shards if routing is skewed |
Ops cues. Watch 429 and rejected execution on the bulk thread pool. Back off and retry with jitter. Tune bulk size by bytes as well as document count. A few hundred documents of tiny logs and a few documents of huge payloads are different requests. Give every document a stable id so a timeout retry overwrites instead of duplicating.
The HTTP status of the bulk call is not the status of each item. Parse the items array. Retry the failures. Drop or quarantine a mapping error instead of retrying it forever.
Refresh, flush, translog, snapshot
Sequence
- 1
Client → Primary shard
1. Bulk index
- 2
Primary shard → Translog and Lucene
2. Append the translog
- 3
Primary shard → Translog and Lucene
3. Refresh opens a new searcher
- 4
Primary shard → Translog and Lucene
4. Flush commits segments
- 5
Primary shard → Snapshot repository
5. Snapshot copies committed state
Lesson map
Indexing Pipelines, Bulk, ILM & Snapshots
Production search lives or dies on the ingest path: bulk requests, refresh versus flush, ILM or ISM rollover, and snapshots to a repository that is usually object storage.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB client["Client"] primary["Primary shard"] disk["Translog and Lucene is one of the participants this lesson's sequence actually names."] repo["Snapshot repository is one of the participants this lesson's sequence actually names."] client -->|1. Bulk index| primary primary -->|2. Append the| disk primary -->|3. Refresh opens| disk primary -->|4. Flush commits| disk primary -->|5. Snapshot| repo
| Mechanism | The document becomes | The cost |
|---|---|---|
| Refresh | Searchable on a new searcher | Segment work. Default interval is about one second |
| Translog | Durable across a crash, for replay | Extra disk writes |
| Flush or commit | A Lucene commit point | IO, then the translog can shrink |
| Snapshot | A copy outside the cluster | Repository IO. Later snapshots are incremental |
refresh=true is for tests and the rare read-your-writes call. On a hot ingest path it forces a refresh per request and wrecks throughput. wait_for waits for the next scheduled refresh instead of forcing an extra one. That is still a latency choice. Most pipelines acknowledge the bulk and let the interval do its job.
A refresh does not replace the translog. Searchers can see a document that a crash would replay from the translog if the Lucene commit has not landed yet. If you need the OLTP comparison: a translog is the search engine's redo path. Checkpoint and ARIES details stay on Database storage engines — WAL, B-trees, and LSM trees.
Ingest pipelines, lightly
An ingest pipeline runs processors on a document before it is indexed: grok, rename, set, a small script. Use it for light cleanup that must travel with the document. Heavy enrichment, joins, and large parsing belong upstream, in the worker that builds the bulk body. A script processor on every log line becomes the bottleneck you cannot autoscale independently of the cluster.
ILM and ISM
| Phase | Typical action | Why |
|---|---|---|
| Hot | Rollover on size, age, or document count | The writable alias stays on a bounded index |
| Warm or cold | Shrink, force-merge, move to cheaper nodes | Cost, after the index is read-mostly |
| Delete | Delete the backing index | Retention without a giant delete-by-query |
Aliases. Writers send bulk requests to logs-write. Readers query logs-read, which spans the backing indices. Rollover creates the next backing index, often named with a counter or a date, and moves the write alias. Force-merge down to one segment is for indices that no longer receive writes. Merging a live index fights the next refresh and can make snapshots more expensive.
OpenSearch Index State Management is the sibling of Elasticsearch Index Lifecycle Management. The interview story is the same. The API names differ. Say which one you have operated.
Rollover is also how you keep primary shards inside the size band from the sharding lesson. max_primary_shard_size on an ILM rollover is the direct knob.
Snapshots and object storage
A snapshot repository is registered on the cluster and usually points at object storage (S3, GCS, or Azure Blob) through a repository plugin. Bucket consistency, multipart upload, and lifecycle rules for those objects live on Object storage — consistency, multipart, and lifecycle. The search-side contract is smaller:
| Practice | Why it matters |
|---|---|
| Incremental snapshots | Later snapshots store what changed |
| Restore drills | An untested repository is a hope |
| A repository per environment | A restore cannot see another environment's credentials by accident |
| Watch restore IO | Restoring every shard at once is a thundering herd |
| Version compatibility | A snapshot restores onto a cluster that understands it |
Replicas do not replace snapshots. Replicas die with the failure domain you failed to spread. A snapshot is how you rebuild after the domain is gone, or how you copy an index to another cluster. Take the snapshot from committed data, then restore into an index you can afford to verify.
Bulk body, as text
The bulk format is newline-delimited JSON: an action line, a document line, and a trailing newline.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Two documents produce four lines. The trailing newline is part of the protocol. A client that omits it gets a parse error on the last item.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
What is the difference between refresh and flush?
Answer
Refresh opens a new searcher so documents become visible to search. Flush commits Lucene segments. The translog covers durability between those commits. A snapshot copies committed state to a repository. Three clocks, three meanings.
Why roll with ILM instead of one index forever?
Answer
Rollover bounds shard size, makes retention a delete of an old backing index, and isolates the hot write from historical data. Readers keep using an alias that spans the indices you still want.
Where should a snapshot repository point?
Answer
At durable storage outside the cluster, typically object storage. A directory on the same host that holds the shards shares the failure. Bucket policy details belong on the object-storage lesson.
The bulk HTTP call returned 200 and documents are missing. What happened?
Answer
Bulk reports per item. Some items failed validation, a mapping conflict, or a rejection. Read the items, retry only the ones that can succeed, and back off on 429.
When is refresh=true justified?
Answer
Tests, and a narrow read-your-writes path where the caller must observe the document before the next interval. It is not the production default for logs or catalog updates.
What is wait_for?
Answer
The request waits until the next scheduled refresh makes the write visible. It does not force an extra refresh the way true does. It still couples the client's latency to the interval.
ILM or ISM?
Answer
Same lifecycle idea. Elasticsearch calls it Index Lifecycle Management. OpenSearch calls it Index State Management. Describe hot rollover, a cheaper phase, and delete. Then use the API of the cluster you run.
Replicas or snapshots?
Answer
Replicas are live copies for search and failover inside the cluster. Snapshots are point-in-time backups in a repository. You want both, for different failures.
Pitfalls
For one bulk item, mark when it is durable, when it is searchable, and when it exists outside the cluster. Name the mechanism for each timestamp, and say which one a 429 forces you to retry.