Embeddings & Similarity — Dense Vectors, Metrics & Chunking Basics
Embeddings map text to dense vectors so similar meaning lands nearby. The usable part is the metric, the model, and the chunk size. A strong embedder with a 4k-token mush window still retrieves mush.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
A support article mentions the same bug in three phrasings
Prefer
A retrieval embedder plus tight chunks
The same model embeds the corpus and the query. Overlap keeps a sentence that straddles a cut. Metadata records the model version.
- Paraphrases land near each other.
- Citations point at a short span.
- A model swap is a new namespace, not a silent mix.
Alternative
Keyword match, or a chat model’s hidden state
Keywords miss synonyms. A generative model’s last hidden state was not trained to rank passages. Both fail the interview when the question is a paraphrase.
- Exact tokens still matter later, in the hybrid lesson.
- Chat hidden states are the wrong space for the index.
- A 4k-token chunk averages the article into mush.
Overview
An embedding is a dense vector. For retrieval models, geometric nearness approximates the training objective (usually semantic similarity or query-passage match). It is not a guarantee that two nearby vectors “mean the same thing” for every task you wish they did.
Interviews fail on the boring invariants:
- Query and corpus use the same model.
- They use the same normalize convention (cosine, or inner product on unit vectors).
- The window you embed is small enough that a citation is a real span, with overlap so a cut does not hide the answer.
A great model with 4,000-token chunks still retrieves mush. A modest model with overlapping chunks and a measured Recall@5 often wins the production bake-off.
Hub: RAG & Vector Databases. Indexes start on the next page. This page stops at the vector and the chunk.
Metrics and chunk sizes
| Choice | Pros | Cons | Prefer when |
|---|---|---|---|
| Cosine, or normalized dot | Scale-invariant. Common default | Query and corpus must agree on normalize | Most text embedders |
| L2 (Euclidean) | Matches some training losses | Magnitude matters if you skip normalize | Models trained with L2 |
| Small chunks (about 128 to 256 tokens) | Precise citations | Misses cross-sentence context | FAQ and API docs |
| Medium chunks with overlap (about 400 to 800 tokens) | Balance | More vectors to store | General knowledge bases |
| Parent-child, or late chunking | Retrieve small, expand for context | Extra store and a join | Long manuals and legal text |
Invariant: if one side is normalized and the other is not, inner product ranks by magnitude as well as angle. Many vector databases store unit vectors and search with inner product because it matches cosine and is cheaper.
Same model, then rank
Flow
- 1
1. Split text into chunks
- next2. Embed with one model
- 2
2. Embed with one model
- next3. Store the dense vectors
- 3
3. Store the dense vectors
- next4. Embed the query the same way
- 4
4. Embed the query the same way
- next5. Score cosine, dot, or L2
- 5
5. Score cosine, dot, or L2
- next6. Rank the candidates
- 6
6. Rank the candidates
Lesson map
Embeddings & Similarity — Dense Vectors, Metrics & Chunking Basics
Embeddings map text to dense vectors so similar meaning lands nearby. The usable part is the metric, the model, and the chunk size. A strong embedder with a 4k-token mush window still retrieves mush.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Split text into chunks"] b["2. Embed with one model"] c["3. Store the dense vectors"] d["4. Embed the query the same way"] a -->|1. Split text into chunks| b b -->|2. Embed with one model| c c -->|3. Store the dense vectors| d
Mixing two model versions in one list silently destroys recall. There is no error. The neighbors are just wrong. Rebuild, or keep a namespace per model_id.
Model families
| Family | Typical dim | Strength | Weakness |
|---|---|---|---|
| General sentence models (MiniLM, E5, BGE) | 384 to 1024 | Cheap, strong baseline | Domain jargon |
| Large retrieval-tuned models | 1024 to 3072 | Higher recall | Cost and latency |
| Domain fine-tune | Varies | Internal vocabulary | Drift and retrain cost |
| Multimodal | Varies | Image plus text | Extra ops |
Keyword-only search skips synonyms. Using a chat LLM as the embedder for every corpus row is too slow, and the vector was not trained for retrieval. Prefer a bi-encoder trained for similarity. Use the chat model to generate, and optionally to rewrite the query before you embed it.
Query rewrite is preprocessing. A short, ambiguous query (“that timeout bug”) can be expanded with synonyms, then embedded with the same model. That is not fine-tuning, and it is not a second vector space.
Why a chat hidden state is the wrong index
Chat models are trained to generate the next token. Embedding models used for retrieval are trained so a query is near its positive passage and far from negatives. Taking an arbitrary layer from a chat model and calling it an embedding usually loses that geometry.
If you must change models, write model_id and model_version on every row and do not query across versions.
Chunking with overlap (run this)
The demo splits characters so it stays self-contained. Production uses the embedder’s tokenizer. A character window lies about the token budget. Overlap must be smaller than the window or the loop never advances.
ProblemCut a short document into overlapping character windows.
ExpectedEach next window starts size-minus-overlap characters later. The loop rejects overlap that is not smaller than size.
Edge cases
- Overlap equal to size must throw.
- A document shorter than size returns one chunk.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Cosine versus L2 (run this)
Two 2-D vectors. Cosine prefers the neighbor that points the same way. L2 also cares about length. After you normalize, the rank from cosine and the rank from inner product match.
ProblemCompare a query to a near-angle neighbor and an orthogonal neighbor.
ExpectedCosine prefers the near-angle vector. L2 distance is smaller for that same neighbor in this example.
Press Run. Snippets must be self-contained — no network, files, or native modules.
What to store on every row
- Tokenize with the embedder tokenizer. Character splits lie about the budget.
- Keep structure. Do not slice a heading away from the list it introduces when you can avoid it.
- Provenance: source URI, section path,
updated_at, ACL tags. - Model identity:
model_idandmodel_versionon every vector. - Dedup. Hash normalized text and collapse repeated nav boilerplate.
- Measure. Retrieval Recall@5 on a golden query set before you tune the generator prompt.
Updates: content-hash the chunk, re-embed only changed hashes, delete stale ids. A model swap is a new namespace, then a cutover. The failure-modes lesson covers the drift you get if you skip that.
Parent-child layout (retrieve the small child, generate from the parent) is specified in RAG Context Assembly — Windows, Parent-Child Chunks & Citations. Sparse exact-token search is the next concern after indexes, in Hybrid Retrieval.
Interview Q&A
What does an embedding represent?
Answer
A dense vector whose geometry matches the training objective, usually retrieval or semantic textual similarity. Nearness is “similar under that objective,” not a universal meaning.
When are cosine and dot product the same ranking?
Answer
When both vectors are L2-normalized. Otherwise magnitude changes the dot product. Many databases normalize at write time and search with inner product.
Why chunk instead of embedding the whole document?
Answer
The vector averages the window. Smaller chunks improve precision and make a citation a real span. Overlap reduces the chance that the answer sat on the cut.
Can you mix embedding models in one index?
Answer
No. They are different spaces. Rebuild, or keep a namespace per model version, and stamp model_id on every row.
How do you update one edited paragraph?
Answer
Hash the normalized chunk text. Re-embed only changed hashes. Delete ids that no longer exist. Do not leave the old vector in the index beside the new one.
Where do sparse scores fit?
Answer
Dense vectors cover paraphrase. BM25 covers exact tokens. Hybrid fusion is a later lesson. Do not pretend a bi-encoder will reliably rank a brand-new error code it never saw as a token pattern you care about.
Why not embed the corpus with the chat model?
Answer
The chat model is trained to generate. A retrieval bi-encoder is trained so queries sit near positive passages. Using chat hidden states as the index usually underperforms, and it is expensive per document.
What is query rewrite?
Answer
A cheap expansion of a short or ambiguous query before you embed it. You still embed the rewritten text with the same model. It is not a fine-tune and not a second index.
What do you measure before you touch the prompt?
Answer
Retrieval Recall at a small k on a labeled query set. If the right chunk is not in the candidate list, the generator cannot cite it.
Which chunk width do you start with?
Answer
For FAQ and API docs, a short window (about 128 to 256 tokens). For a general knowledge base, a medium window with overlap (about 400 to 800 tokens). For manuals, retrieve a child and expand to the parent. Then measure. Do not ship the first width that looked fine on ten queries.
Pitfalls
You embedded last quarter with model A and this week with model B into the same index. A canary query that used to hit the runbook now returns a glossary page. What do you check first, and how do you cut over without a mixed week?