RAG Context Assembly — Windows, Parent-Child Chunks & Citations
Retrieval returns candidates. Context assembly decides what the model actually sees: a token budget, deduped parents, neighbor windows, and citation ids. A perfect top-k still fails if you dump it raw into the prompt.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The right paragraph was in the top 20, and the answer still missed it
Prefer
Pack a small window with parent context and citation ids
Retrieve children for precision. Expand to the parent section once. Dedup. Stop at a token budget. Order so the best evidence is not lost in the middle. Tell the model to cite ids or say the context is insufficient.
- The generator sees coherent sections, not fragments.
- Cost and distraction stay bounded.
- A later check can map a claim back to an id.
Alternative
Paste the raw top 50 into the prompt
You pay for tokens, you bury the useful chunk, and duplicate children from the same section crowd out a second document.
- Long context is not free attention.
- Duplicate parents waste the budget.
- There is no id to audit.
Overview
The retriever’s job ends at a ranked candidate list. Context assembly decides the bytes in the prompt: which hits survive a budget, which neighbors or parents get pulled in, which duplicates die, which order the model sees, and which ids a citation must use.
Bad assembly looks like a retrieval bug. The right chunk was retrieved and then dropped, duplicated, or buried. Fix the packer before you retrain an embedder.
How the list was built is Hybrid Retrieval. What you do when the model ignores the window is RAG Failure Modes.
Assembly strategies
| Strategy | Pros | Cons | Prefer when |
|---|---|---|---|
| Raw top-k | Simple | Fragments and duplicates | A prototype |
| Token-budget packer | Predictable cost | Can drop a mid-ranked hit that was the real evidence | Production default |
| Parent-child | Precise retrieve, richer context | Extra store and a join | Long documents |
| Neighbor window | Local continuity | Noise beside the hit | Procedures and logs |
| Maximal marginal relevance | Diversity | You must tune the tradeoff | Corpora full of near-duplicate chunks |
Parent, then pack
Flow
- 1
1. Retrieve small child chunks
- next2. Expand to parent sections
- 2
2. Expand to parent sections
- next3. Dedupe shared parents
- 3
3. Dedupe shared parents
- next4. Pack under the token budget
- 4
4. Pack under the token budget
- next5. Attach stable citation ids
- 5
5. Attach stable citation ids
- next6. Generate with a cite contract
- 6
6. Generate with a cite contract
Lesson map
RAG Context Assembly — Windows, Parent-Child Chunks & Citations
Retrieval returns candidates. Context assembly decides what the model actually sees: a token budget, deduped parents, neighbor windows, and citation ids. A perfect top-k still fails if you dump it raw into the prompt.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB q["1. Retrieve small child chunks"] p["2. Expand to parent sections"] d["3. Dedupe shared parents"] b["4. Pack under the token budget"] q -->|1. Retrieve small child chunks| p p -->|2. Expand to parent sections| d d -->|3. Dedupe shared parents| b
Embed and search the child so the hit is specific. Before generation, replace or supplement that child with the parent section, or with a small neighbor window (plus or minus a few chunks by offset). Dedup so three children from one section do not consume the budget three times. Then stop.
Stable ids look like a document id plus a chunk id, or a document id plus page and start offset. The prompt names those ids. The UI can map an id to a URL later. Do not make the model invent URLs.
Citation contract
The instruction is blunt:
- Answer only from the context blocks.
- Cite each factual sentence with the chunk id in brackets.
- If the context is not enough, say so and ask one clarifying question.
- Do not fill gaps from parametric memory.
A post-check maps claim spans back to cited chunks. That can be a heuristic overlap, a human audit, or an NLI-style faithfulness judge. The failure-modes lesson compares those metrics. This page only makes the ids available so a check is possible.
If the product needs a JSON array of citations, that shape is Structured Outputs / Constrained Decoding. Schema validity and faithfulness are different gates. A missing required field is not the same bug as a fluent claim the chunks do not support.
Pack under a budget (run this)
Hits are sorted by score. The first child of a parent wins. Later children of that parent are skipped. A chunk that does not fit is skipped, not split, so a later smaller chunk can still enter. The token count is a 4-character heuristic. Production uses the generator tokenizer, because that is the budget you actually pay.
ProblemPack the highest-scoring hit per parent without crossing the budget.
ExpectedThe sibling child from the same parent is dropped. The second document still fits.
Edge cases
- A single chunk larger than the budget is skipped.
- Equal scores keep input order after the sort is stable.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pull citation ids back out (run this)
The checker only trusts ids that were actually packed. An id the model invented fails the membership test even if the brackets look right.
ProblemCollect ids cited in an answer and keep only ids that were packed.
ExpectedTwo real ids are kept. An invented id is dropped.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Lost in the middle
Models often attend more to the beginning and end of a long context than to the middle. Practical tactics:
- Put the strongest evidence first, or first and last.
- Keep chunks from the same parent contiguous.
- Cap the window. Extra low-score chunks can hurt, even when they are “relevant” in isolation.
Order is an experiment, not a slogan. Some tasks want document order inside a parent. Some want pure relevance order. Measure faithfulness when you flip it. Do not assume the model read line 14.
Windowing patterns
| Pattern | Mechanism | Use |
|---|---|---|
| Fixed top-k | Take k by score | Baseline |
| Score threshold | Drop hits below a cutoff | Noisy corpora |
| Neighbor expand | Take nearby chunks by offset | Procedures |
| Hierarchical | Section, then subsection | Manuals |
| Compression | Summarize low-value hits | Huge passages that still must be represented |
Maximal marginal relevance trades relevance against novelty so five near-copies of one paragraph do not fill the window. The tradeoff coefficient is another knob. Sweep it if the corpus is redundant. Skip it if every chunk is already distinct.
Prompt contract
SYSTEM
Use ONLY the CONTEXT blocks.
Cite each factual sentence with a [chunk_id] that appears in CONTEXT.
If CONTEXT is not enough, reply: Insufficient evidence.
Then ask one clarifying question.
CONTEXT
...packed blocks...Empty retrieval takes the refuse branch. Do not silently answer from parametric memory and hope the missing citation goes unnoticed. That failure mode is catalogued next.
Interview Q&A
Why not dump the top 50 chunks into the prompt?
Answer
Context limits, cost, and distraction. A budget plus dedupe beats a raw dump. Extra low-score chunks can pull the answer off the evidence that mattered.
What does parent-child buy?
Answer
You retrieve a small child for precision, then show the parent section so the generator has a coherent span and a cleaner citation. Three children of one parent should not be pasted three times.
How do you order chunks in the prompt?
Answer
Often relevance order, or document order inside a parent. Models are sensitive to position. Put the strongest evidence at the ends if the window is long, and test the choice. “Lost in the middle” is a measured effect, not a reason to ignore order.
Citation id or URL?
Answer
Stable internal ids for verification. Map those ids to URLs in the UI. Asking the model to invent links produces citations you cannot join back to a chunk.
What do you do when retrieval is empty?
Answer
Refuse, or ask a clarifying question. Do not generate from parametric memory and present it as grounded. The prompt contract has to say that, or the model will fill the gap.
How do structured outputs overlap this page?
Answer
If the UI needs a JSON object with an answer string and a citations array, that shape is Structured Outputs / Constrained Decoding. Validate the schema separately from faithfulness. A well-formed array can still point at a chunk that does not support the sentence.
What is lost in the middle?
Answer
On long contexts, evidence placed in the middle is used less reliably than evidence at the start or the end. Packing order is part of the design. Stuffing more tokens is not a fix.
When do you expand neighbors instead of parents?
Answer
When the source is a procedure or a log and the useful span is “the chunks next to the hit,” not a whole section node. Parents fit manuals with a real hierarchy. Neighbors fit flat sequences.
Why can a score threshold beat a fixed k?
Answer
A fixed k always fills the window, including noise, on a query that only had one good hit. A threshold leaves the window short when the rest of the list is weak. You still need a budget cap for the queries where everything scores high.
What do you check after the model replies?
Answer
Every cited id was in the packed set. Invented ids are dropped. Claims with no citation, or citations that do not overlap the claim, fail the grounding check in the next lesson.
Pitfalls
Your retriever returns three overlapping chunks of the same policy section plus one chunk from a different doc. The budget fits about two parents. Which ids reach the model, in what order, and what does the prompt say if the second document was the one that answered the question?