RAG Failure Modes — Hallucination, Stale Indexes, Evals & Grounding
RAG fails loudly, with an empty answer, or quietly, with a confident citation that does not support the claim. This lesson is the catalog: hallucination, stale indexes, tenant leaks, and the evals that catch them before a doc edit rots the demo.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The answer quotes a chunk id and the policy is still wrong
Prefer
Check support, freshness, and tenancy separately
Faithfulness asks whether the claim is in the retrieved text. Lag asks whether that text is the current document. ACL asks whether this caller was allowed to retrieve it. One green dashboard does not cover the other two.
- Citation theater fails a span check, not a vibe check.
- A canary query runs after every index swap.
- User A never receives tenant B, and that test is in CI.
Alternative
RAG is on, so hallucination is solved
The model can ignore the context, cite an id it did not use, or quote a vector you forgot to delete. BLEU against a gold answer will not see any of those.
- Helpful and ungrounded can both score well on relevance.
- Lexical overlap can be gamed.
- A fine-tune does not fix a stale index.
Overview
RAG reduces hallucination when retrieval is strong and the prompt is forced to stay inside it. It does not eliminate hallucination. The model can ignore the window, the window can be the wrong version of the document, and the window can belong to another tenant.
Loud failures are obvious: an empty retrieval, a refusal, a timeout. Quiet failures ship: a fluent answer, bracketed ids, and a claim the chunks do not support. This page is the catalog and the gates.
Packing and the citation contract are RAG Context Assembly. JSON shape, if you need it, is Structured Outputs / Constrained Decoding. Serving the generator is LLM Inference Runtime — Prefill, KV, Batching & Parallelism. Faster decode does not make a stale chunk true.
Failure catalog
| Failure | Symptom | Root cause | Mitigation |
|---|---|---|---|
| Hallucination despite RAG | Fluent wrong answer | Weak retrieval, or the model ignored the context | Force cites, refuse on weak coverage, raise rerank quality |
| Citation theater | Ids are present and unused | The prompt asked politely | Verify spans against the cited chunks |
| Stale index | The answer matches an old PDF | Ingest lag or a failed delete | Lag SLO, content hashes, rebuild |
| Wrong tenant | Another customer’s snippet | Missing ACL filter | Pre-filter, plus an integration test |
| Embedding drift | Recall drops after a “harmless” upgrade | Model swap without a re-embed | Versioned namespaces |
| Chunk boundary miss | A procedure is cut in half | Chunk too small, or no overlap | Overlap, or parent expand |
| Poison or boilerplate | Irrelevant chunks score high | Adversarial text, nav chrome, SEO | Filters, diversity, domain allowlists |
| Eval gaming | High lexical overlap, wrong claims | The metric is not faithfulness | Use a grounded judge, and keep retrieval scores separate |
Control loop
Decisions
- 1
1. Query arrives
- next2. Retrieve with ACL filters
- 2
2. Retrieve with ACL filters
- next3. Coverage strong enough?
- ?
3. Coverage strong enough?
- no4. Refuse or ask to clarify
- yes5. Generate under cite contract
- 4
4. Refuse or ask to clarify
- 5
5. Generate under cite contract
- next6. Claims supported?
- ?
6. Claims supported?
- no7. Tighten context and retry
- yes8. Return answer and citations
- 7
7. Tighten context and retry
- 8
8. Return answer and citations
- next9. Log ids, scores, and model
- 9
9. Log ids, scores, and model
- next10. Alert on index lag
- 10
10. Alert on index lag
Lesson map
RAG Failure Modes — Hallucination, Stale Indexes, Evals & Grounding
RAG fails loudly, with an empty answer, or quietly, with a confident citation that does not support the claim. This lesson is the catalog: hallucination, stale indexes, tenant leaks, and the evals that catch them before a doc edit rots the demo.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Query arrives"] b["2. Retrieve with ACL filters"] c["3. Coverage strong enough?"] d["4. Refuse or ask to clarify"] a -->|1. Query arrives to 2. Retrieve with ACL filters| b b -->|2. Retrieve with ACL filters| c c -->|no| d
The refuse branch is a success if the alternative is a quiet lie. Retry once with a tighter prompt or a wider (still filtered) context. Do not loop the generator until it sounds confident.
Evals that separate
| Eval | Measures | Weakness |
|---|---|---|
| Recall@k or nDCG on a labeled retrieval set | Did the retriever surface the right chunks? | Needs qrels |
| Faithfulness or groundedness | Are the claims inside the context? | A judge model costs money and has its own bias |
| Answer relevance | Is the reply on topic? | An ungrounded reply can still look relevant |
| Citation precision | Do the cited chunks support the claims? | Someone has to define “support” |
| End-to-end task success | Did the user get the job done? | Slow, and sparse |
BLEU or ROUGE against a single gold answer misses open-ended support bots. Always report a retrieval number and a generation number. A prompt change cannot fix Recall@5. A chunker change can wreck both.
Token-overlap faithfulness, below, is a demo. It catches a claim that shares no content words with the context. It does not catch a negation, a swapped subject, or a true sentence that paraphrased past the token set. Use it to understand the shape of the metric. Do not gate production on it alone.
Overlap heuristic (run this)
ProblemScore a supported sentence and an unsupported sentence against one context.
ExpectedThe sentence that repeats the context scores higher than the sentence that contradicts it.
Edge cases
- An empty answer scores 0.
- Words of length 3 or less are ignored, so short claims can look empty.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Freshness guard (run this)
Index lag is now - indexed_at for the documents you claim to serve. This guard is a predicate you can put in front of generation when the product promises a freshness SLO. It does not prove the chunk text matches the source. Pair it with a content hash.
ProblemCompare indexed time to now against a 15 minute SLO.
ExpectedFive minutes of lag passes. One hour of lag fails.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Fixed timestamps keep the demo deterministic. Production passes Date.now() at request time and the index’s high-water mark, not a clock baked into the prompt.
What to log
Log enough to replay an offline judge without storing raw secrets:
- query id, the rewritten query, and the filter predicates
- retrieved ids, scores, and the stage (BM25, dense, rerank)
- token budget used, and whether you refused
- citation ids returned, versus ids that were retrieved
- index generation and embedder model version
Hash or drop free-text fields that are personal. The point of the log is the join between a claim and a chunk version.
CI gates
- Retrieval Recall@k must not drop by more than an agreed epsilon on the golden set after a chunker or index change.
- A faithfulness score on a held-out answer set stays above a threshold.
- ACL tests: user A never retrieves tenant B.
- Canary queries run after every production index swap, before the alias is the default.
Incident: answers flipped after a doc edit
Work the checklist before you edit the prompt:
- Did ingest succeed for the new document?
- Were the old vectors deleted, or are both versions in the index?
- Is an answer cache or a CDN still serving the previous completion?
- Are you querying the wrong namespace or the previous embedder version?
- Did hybrid weights move in the same deploy?
The durable fix is an idempotent upsert keyed by chunk hash, and an atomic alias swap so readers never see a half-published index. Partial deletes are how “both policies” stay in the top-k.
Interview Q&A
Does RAG eliminate hallucination?
Answer
No. It reduces hallucination when retrieval is strong and the prompt refuses to leave the context. The model can still ignore the window, and the window can be stale or forbidden.
How do you detect a stale index?
Answer
Track the gap between document updated_at and indexed_at, count failed ingest jobs, and run canary queries after each publish. A freshness predicate can block generation when lag breaks the SLO.
Offline evals or online evals?
Answer
Offline labeled sets catch regressions in CI. Online signals (thumbs, escalations, citation clicks) catch drift the golden set does not contain. You want both. Neither replaces the other.
What is faithfulness?
Answer
The claims in the answer are supported by the retrieved context. It is orthogonal to “sounds helpful” and orthogonal to “matches a single reference answer.”
What security failure is specific to this path?
Answer
Cross-tenant retrieval when the metadata filter is wrong or applied too late. Treat it as a data leak, not as a bad answer. The test is: user A never receives tenant B’s chunk ids.
When do you fall back to fine-tuning?
Answer
When the failure is behavioral: format, tone, tool use. Keep RAG for facts. A fine-tune will not delete last week’s PDF from the index.
Why is BLEU a weak RAG metric?
Answer
It compares the answer to one reference string. A grounded paraphrase scores poorly. An ungrounded sentence that copies the reference’s wording can score well. Split retrieval quality from whether the claims sit in the context.
What is citation theater?
Answer
The answer contains bracketed ids, and those ids do not support the sentences next to them. The fix is a check that maps claims to spans, not a stronger request to “please cite.”
What do you log so an offline judge can run later?
Answer
Query id, filters, retrieved ids and scores by stage, whether you refused, cited ids versus retrieved ids, and the embedder and index versions. Redact raw personal text.
A doc edit shipped and answers flipped. What do you check first?
Answer
Ingest success, deletion of old vectors, answer caches, namespace and model version, then hybrid weights. Do not start by rewriting the prompt.
Pitfalls
Legal replaced a retention period from 30 days to 90 days. The bot cites both chunks in one answer. Walk ingest, delete, cache, and alias swap, and say which eval would have failed in CI if you had one.