RAG & Vector Databases
Studies in this cluster, in series order. Each one keeps its own URL.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
RAG & Vector Databases
6 studies- 1.RAG & Vector Databases — Retrieval, Embeddings & GroundingRAG grounds an LLM on your documents at query time: embed, retrieve, assemble context, and generate with citations. This hub is the decision map for when retrieval beats fine-tuning or long context, and how to judge grounding.
- 2.Embeddings & Similarity — Dense Vectors, Metrics & Chunking BasicsEmbeddings map text to dense vectors so similar meaning lands nearby. The usable part is the metric, the model, and the chunk size. A strong embedder with a 4k-token mush window still retrieves mush.
- 3.Vector Indexes — HNSW, IVF & Product Quantization TradeoffsBrute-force top-k dies at millions of vectors. HNSW, IVF, and product quantization trade a measured amount of recall for latency and memory. Pick them the way you pick a B-tree versus a hash index: with a recall curve and a p99.
- 4.Hybrid Retrieval — BM25 + Vectors, Reranking & Metadata FiltersProduction RAG rarely ships vectors alone. Hybrid retrieval fuses BM25 with dense similarity, reranks a shortlist, and applies metadata filters for tenant, ACL, and time before the model ever sees a chunk.
- 5.RAG Context Assembly — Windows, Parent-Child Chunks & CitationsRetrieval returns candidates. Context assembly decides what the model actually sees: a token budget, deduped parents, neighbor windows, and citation ids. A perfect top-k still fails if you dump it raw into the prompt.
- 6.RAG Failure Modes — Hallucination, Stale Indexes, Evals & GroundingRAG fails loudly, with an empty answer, or quietly, with a confident citation that does not support the claim. This lesson is the catalog: hallucination, stale indexes, tenant leaks, and the evals that catch them before a doc edit rots the demo.