RAG & Vector Databases — Retrieval, Embeddings & Grounding
RAG grounds an LLM on your documents at query time: embed, retrieve, assemble context, and generate with citations. This hub is the decision map for when retrieval beats fine-tuning or long context, and how to judge grounding.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The model must answer from private docs that change
Prefer
Retrieve, then generate
Chunk and embed the corpus, search at query time, and condition the answer on those passages with citation ids. Updating a PDF means re-indexing, not retraining.
- Freshness follows the index lag, not a training job.
- Citations point at chunk ids you can audit.
- The corpus can be larger than the context window.
Alternative
Fine-tune the facts into the weights
Continued training can teach style and stable skills. It is a weak place to store a policy PDF that legal will edit next week.
- Updates wait on a retrain.
- The model cannot point at a page it no longer has as text.
- Private facts sit in weights that are hard to delete per tenant.
What the interviewer is actually asking
If the answer is correct tomorrow after a PDF edit, did we retrain or re-index?
- 1
Name the knowledge
Changing facts, private corpus, and a need to cite. That is retrieval. Stable style or format is a different knob. - 2
Name the query path
Filter, then hybrid or dense retrieve, then rerank, then pack a window. Top-k cosine alone is a prototype. - 3
Name the SLOs
p99 retrieve, a faithfulness bar, and index lag. Fluent text is not a grounding metric. - 4
Name two failures
Stale vectors after an edit, and a missing tenant filter. Both are in the failure-modes lesson.
Overview
Retrieval-augmented generation keeps facts outside the weights. At query time you embed the question, retrieve passages, assemble a context window, and generate an answer that can cite those passages. The corpus can change without a training run.
Senior interviews are not asking you to name a vendor. They are asking four decisions:
- When retrieval beats fine-tuning, and when stuffing a small corpus into a long context is enough.
- How you choose an embedding model, a chunk size, and an approximate nearest-neighbor index.
- When hybrid search (lexical plus dense, then a rerank) beats “top-k cosine and hope.”
- How you measure grounding: faithfulness and citation precision, plus a stale-index alarm.
Ask this out loud: if the model answers correctly tomorrow after we update a PDF, did we retrain, or re-index?
RAG, fine-tuning, and long context
| Approach | Freshness | Citeability | Cost to update | Prefer when |
|---|---|---|---|---|
| Fine-tune or continued pretrain | Slow (retrain) | Weak (weights) | High | Style, format, or a stable skill |
| Long context only | Good if you resend the docs | Possible, but noisy | Per-request tokens | A small corpus that fits the window |
| RAG | Fast (re-index) | Strong (chunk ids) | Index and storage | Large, private, or frequently edited knowledge |
| RAG plus a light fine-tune | Fast facts, plus style | Strong | Medium | You need grounding and a stable task skill |
Rule of thumb: RAG for facts that change. Fine-tune for behavior that should stick. Long context for a small corpus or an agent scratchpad.
JSON shape is a different problem. Structured Outputs / Constrained Decoding covers schemas and constrained decoding. A valid JSON object can still be unfaithful to the retrieved text. Serving the generator (prefill, KV cache, batching) lives in LLM Inference Runtime — Prefill, KV, Batching & Parallelism. This cluster does not re-teach either lesson.
Decision path
Decisions
- 1
1. Private or fresh docs?
- no2. Prompt or tools only
- yes3. Fits in context cheaply?
- 2
2. Prompt or tools only
- ?
3. Fits in context cheaply?
- yes4. Long context only
- no5. Build a RAG pipeline
- 4
4. Long context only
- 5
5. Build a RAG pipeline
- next6. Need exact tokens too?
- ?
6. Need exact tokens too?
- yes7. Hybrid BM25 and vectors
- no8. Dense ANN top-k
- 7
7. Hybrid BM25 and vectors
- next9. Pack windows and cites
- 8
8. Dense ANN top-k
- next9. Pack windows and cites
- 9
9. Pack windows and cites
- next10. Generate and evaluate
- 10
10. Generate and evaluate
- next11. Faithfulness holds?
- ?
11. Faithfulness holds?
- no12. Fix chunks or filters
- yes13. Ship and watch index lag
- 12
12. Fix chunks or filters
- 13
13. Ship and watch index lag
Lesson map
RAG & Vector Databases — Retrieval, Embeddings & Grounding
RAG grounds an LLM on your documents at query time: embed, retrieve, assemble context, and generate with citations. This hub is the decision map for when retrieval beats fine-tuning or long context, and how to judge grounding.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Private or fresh docs?"] b["2. Prompt or tools only"] c["3. Fits in context cheaply?"] d["4. Long context only"] a -->|no| b a -->|yes| c c -->|yes| d
Read it as a filter. If the corpus is small and queries are rare, stop at long context. If error codes and paraphrases both matter, take the hybrid branch. If faithfulness fails, fix retrieval and the prompt before you blame the generator.
What you draw on the board
- 1
Parse and chunk
Keep section metadata and a content hash. One vector per whole PDF is usually the wrong granularity.
- 2
Embed
One model and one normalize convention for the corpus and the query.
- 3
Upsert the index
Store the vector plus doc id, ACL tags, and updated time.
- 4
Filter, then retrieve
Tenant and ACL first. Then hybrid or dense search, then a short rerank.
- 5
Pack and cite
Stay inside a token budget. Force chunk ids on factual sentences.
- 6
Evaluate grounding
Faithfulness and index lag, not only a fluent answer.
What this cluster covers
- Embeddings & Similarity — Dense Vectors, Metrics & Chunking Basics — metrics, model families, overlap.
- Vector Indexes — HNSW, IVF & Product Quantization Tradeoffs — ANN latency, memory, and recall.
- Hybrid Retrieval — BM25 + Vectors, Reranking & Metadata Filters — fusion, rerank, ACL.
- RAG Context Assembly — Windows, Parent-Child Chunks & Citations — budgets, parents, citation contracts.
- RAG Failure Modes — Hallucination, Stale Indexes, Evals & Grounding — quiet failures and CI gates.
Each page is a full lesson. This hub does not re-teach them.
Why “one vector per PDF” fails
Why it shows up. It is a fast prototype. One row per file, one embedding call, a simple mental model.
Why interviews reject it.
- The hit is the whole file, so the top-k is full of unrelated pages.
- Citations have no locality. You cannot point at the paragraph that supported the claim.
- The embedding averages the document. A specific procedure is diluted.
- Any edit re-embeds the entire file.
Better default: chunk with overlap, or retrieve a child and expand to a parent. Store doc_id, page, and section. Re-embed changed chunks, and delete stale ids.
Toy dense retrieve (run this)
No network and no vector database. Three fake 3-D vectors stand in for a real embedder (often 384 to 3072 dimensions). Cosine is computed after L2 normalize, so magnitude does not decide the rank.
ProblemRank chunks by cosine similarity to a query vector.
ExpectedThe HNSW chunk ranks first. The citation chunk is second. The BM25 chunk falls outside k=2 for this query.
Edge cases
- A zero vector must not divide by zero.
- k larger than the corpus returns every row.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Toy context packer (run this)
A character heuristic stands in for a tokenizer (about 4 characters per English token in this demo). The packer stops when the next block would pass the budget. It does not split a chunk in half.
ProblemWalk scored hits in order and keep whole chunks until the budget is spent.
ExpectedTwo citation ids fit. The third hit is left out.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The real packer, parent expansion, and citation checks are in RAG Context Assembly — Windows, Parent-Child Chunks & Citations.
Deliberately not re-taught
| Already taught | Where |
|---|---|
| JSON Schema, constrained decoding, tool-call shape | Structured Outputs / Constrained Decoding |
| Prefill, KV cache, paging, continuous batching | LLM Inference Runtime — Prefill, KV, Batching & Parallelism |
Document ingest can be event-driven. Reliable publish is a messaging problem. This cluster scores and grounds retrieval. It does not re-teach a log or a consumer group.
Interview Q&A
What is RAG?
Answer
Retrieve relevant passages from an external store, then condition the LLM on those passages. Answers can cite the passages, and the facts can change by re-indexing instead of retraining.
When do you pick RAG over fine-tuning?
Answer
Fine-tuning changes weights: style, format, or a stable skill. RAG supplies facts at runtime. Prefer RAG for private or changing knowledge and for citations. Fine-tune when the behavior should stick even if the prompt is short. Many products use a light fine-tune for format and RAG for the corpus.
Why not only long context?
Answer
Cost, latency, and distraction grow with the corpus. Retrieval selects a small relevant window and scales to millions of chunks. Long context still wins when the whole corpus fits cheaply and you rarely ask.
Dense versus sparse retrieval?
Answer
Dense embeddings capture paraphrases and synonyms. Sparse lexical rankers such as BM25 win on exact tokens: ids, error codes, SKUs. Production queries usually need both. That design is the hybrid lesson.
What is an ANN index?
Answer
An approximate nearest-neighbor structure (HNSW, IVF, often with product quantization) trades a little recall for sub-linear search latency once brute force is too slow. You publish a recall-versus-latency curve. You do not trust a blog default.
How do you ground an answer?
Answer
Force citations to chunk ids, refuse or ask a clarifying question when retrieval is weak, and evaluate faithfulness against the retrieved text. Answer quality alone can reward a fluent hallucination.
What is the biggest production failure?
Answer
A stale or incomplete index plus an overconfident generator. Monitor ingest lag, and run retrieval and faithfulness checks when the chunker or index changes. The failure-modes lesson is the catalog.
Where do structured outputs fit?
Answer
After retrieval, if a tool or UI needs JSON. Structured Outputs / Constrained Decoding owns schemas. A schema-valid payload can still cite the wrong chunk. Do not treat validity as grounding.
Where does inference runtime fit?
Answer
The generator still has a prefill and a decode. LLM Inference Runtime — Prefill, KV, Batching & Parallelism owns serving. A faster decode does not fix a bad index.
What belongs on the whiteboard in the first two minutes?
Answer
Ingest: parse, chunk, embed, upsert with metadata. Query: filter, hybrid retrieve, rerank, pack, generate, cite. Then three numbers: p99 retrieve, a faithfulness threshold, and index lag. Then two failures: stale index and ACL miss, each with a mitigation.
Pitfalls
A policy PDF changes section 4. The chat answer still quotes the old rule an hour later. List the checks you run before you touch the prompt: ingest job, delete of the old chunk id, answer cache, embedder version, and hybrid weights.