Agent Memory — Working Context, Long-term Store, Compaction & Sessions
Agent memory is a budgeted systems problem. Working context fits in the window. Episodes and semantic facts live outside it. Compaction must keep goals and commitments and must not leak PII across sessions. Retrieval of documents is a different hub.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is working context?
Answer
The tokens sent on this step: policy, goal, active tool schemas, pins, and the recent tail.
L2
How is that different from retrieval?
Answer
Working context is the live window for this loop. RAG fetches external documents. Episodic and semantic stores are durable agent state, not the corpus index.
L3
What do you compact, and what do you pin?
Answer
Compact chatter and repeated tool dumps. Pin goals, authz decisions, amounts, idempotency keys, human outcomes, and open obligations.
L4
Why do summary errors compound?
Answer
Each summary is the input to the next one. A dropped approval becomes the new truth unless the pin lives outside the prose.
L5
How do you isolate multi-tab users?
Answer
Give each tab a session id. Share semantic preferences only with consent, and always include the tenant id in the key.
L6
What makes a vector memory unsafe?
Answer
Cross-tenant search, no retention limit, and no split between documents and agent state. You build a PII lake.
L7
How do you assemble a window under a token budget?
Answer
Walk a priority order: system, goal, pins, recent turns, then summaries. Skip a piece that does not fit. Never drop the system policy to fit chatter.
Failure modes
Summary amnesia
The prose summary drops the order id or the approval, and the next step invents one.
Window overflow
Instructions or the goal fall off the front of a naive sliding window.
Cross-tenant semantic search
A shared vector store returns another customer's preference or ticket.
Stale semantic fact
A distilled preference survives after the user revoked it.
Procedure drift
A stored playbook no longer matches the tool schema that shipped.
Misconceptions
Vector memory is a free long-term brain.
Without isolation and retention it is a PII lake. Documents belong in the RAG hub.
A longer context window removes compaction.
Cost, latency, and lost-in-the-middle effects remain. Pins still beat a raw dump.
The summary is the source of truth.
Structured pins and the episodic record are. The summary is a view.
Interviewer traps
Redesigning chunking, embeddings, and ANN indexes.
Say document retrieval is the RAG hub. Stay on session state, pins, and compaction.
Claiming the model will remember if you ask it to.
Put ids and amounts in structured fields the compactor is not allowed to rewrite.
Design scenario
Same prompt for every reader.
Requirements
Working context stays inside a token budget. Refunds retain order id, amount, currency, and approval id across compaction. Tenants cannot see each other.
Traffic / scale
About 2000 concurrent sessions, some lasting many turns.
Latency
Assembly of the next window stays well under the acknowledgement budget. Compaction can be asynchronous after the turn.
Consistency
Pins are read-your-writes for the same session. Semantic facts may lag by a distillation job.
Availability
If the semantic store is down, the run still proceeds from the session episode and the pins.
Failure assumptions
- A summary can drop a numeric fact.
- A user can have two sessions open.
- A tenant key can be omitted by a bug.
Constraints
- Every memory key includes a tenant id.
- Compaction cannot rewrite pin fields.
- Retrieved documents are cited, not dumped into the window.
Prompt
Design memory for a support agent that keeps 30 days of episode history.
API
What does the loop load at the start of a turn, and what does it append at the end?
Data
Which fields are pins, which are episodic text, and which are semantic facts?
Architecture
Where do the session store, the compactor, and the document retriever sit?
A refund thread that no longer fits in the window
Prefer
Pins plus a summary
Structured fields survive compaction. The prose is a view of the tail. The next tool reads the pins, not a paragraph that might have dropped the amount.
- Order id, amount, currency, and approval id stay typed.
- Chatter and repeated tool dumps shrink.
- The tenant id is part of every key.
Alternative
Slide the raw transcript, or embed every turn into one index
A sliding window forgets the early commitment. A shared vector index recalls the wrong tenant and the wrong session.
- Early instructions fall off the front.
- Summary-of-a-summary drifts.
- Document retrieval and agent state get mixed.
Overview
The window is a budget, not a database. Working context is what this step sends to the model. Episodic memory is the trajectory you can resume and debug. Semantic memory is a small set of distilled facts. Procedural memory is a versioned playbook or graph, closer to code than to chat.
Document retrieval is a different system. If the fact lives in a PDF, cite it through RAG and Vector Databases. Do not rebuild chunking or ANN indexes on this page.
| Tier | What it holds | Lifetime | Failure if wrong |
|---|---|---|---|
| Working context | Policy, recent turns, tool observations | One request or a few steps | Overflow, or lost instructions |
| Episodic | Trajectory of a session or run | Days to months | Cannot debug or resume |
| Semantic | Distilled facts and preferences | Long, with a TTL | Stale or wrong facts stick |
| Procedural | Playbooks, skills, graphs | Versioned with the tools | Drift versus the code |
Decisions
- 1
1. Incoming event
- next2. Load session key
- 2
2. Load session key
- next3. Fetch facts
- 3
3. Fetch facts
- next4. Fit token budget
- 4
4. Fit token budget
- next5. Run model and tools
- 5
5. Run model and tools
- next6. Append episode
- 6
6. Append episode
- next7. Over budget?
- ?
7. Over budget?
- yes8. Compact and pin
- no9. Persist session
- 8
8. Compact and pin
- next9. Persist session
- 9
9. Persist session
Lesson map
Agent Memory — Working Context, Long-term Store, Compaction & Sessions
Agent memory is a budgeted systems problem. Working context fits in the window. Episodes and semantic facts live outside it. Compaction must keep goals and commitments and must not leak PII across sessions. Retrieval of documents is a different hub.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Incoming event"] b["2. Load session key"] c["3. Fetch facts"] d["4. Fit token budget"] a -->|1. Incoming event| b b -->|2. Load session key| c c -->|3. Fetch facts to 4. Fit token budget| d
Compaction runs because the episode grew, not because the model felt like summarizing. The pin write happens before the prose shrink, so a crash still leaves the amount.
Budget order
Fill the window in this order. Stop when the next piece does not fit. Never steal from an earlier tier to make room for a later one.
- Safety and system policy. Never drop it.
- User goal and constraints.
- Active tool schemas, or a routed subset.
- Pinned facts: ids, amounts, approvals.
- Recent observations, the tail of the trajectory.
- Older history, as summaries only.
- Retrieved documents, cited, not dumped. That last tier is the RAG sibling.
| Strategy | Pros | Cons |
|---|---|---|
| Sliding window | Simple | Forgets early commitments |
| Recursive summary | Fits long sessions | Errors compound |
| Pin, then summarize | Keeps ids and money | Needs explicit pin rules |
| Hierarchical map-reduce | Scales to very long traces | Extra latency and cost |
What compaction may touch
Do compact: chatter, repeated tool dumps, boilerplate.
Do not compact away: goals, authorization decisions, money amounts, idempotency keys, human-gate outcomes, and open obligations such as “the user asked to cancel.”
Press Run. Snippets must be self-contained — no network, files, or native modules.
The summary can lose sentences. It cannot lose order_id or amount, because those were copied into the pin block first.
Sessions and isolation
| Concern | Rule |
|---|---|
| Tenant | Every key includes a tenant id |
| User and session | Separate session id and user id |
| Agent version | Tag memories with prompt and tool versions |
| Privacy | No cross-tenant semantic search by default |
| Retention | A TTL, plus a legal hold when required |
Two browser tabs are two sessions unless the product explicitly shares state. Semantic preferences can be shared only with a consent flag, and still never across tenants.
Assembly under a budget
export type MemoryPiece = {
kind: "system" | "goal" | "pin" | "summary" | "recent";
text: string;
tokens: number;
};
export function assemble(pieces: MemoryPiece[], budget: number): MemoryPiece[] {
const priority = ["system", "goal", "pin", "recent", "summary"] as const;
const chosen: MemoryPiece[] = [];
let used = 0;
for (const kind of priority) {
for (const piece of pieces.filter((item) => item.kind === kind)) {
if (used + piece.tokens > budget) continue;
chosen.push(piece);
used += piece.tokens;
}
}
return chosen;
}Press Run. Snippets must be self-contained — no network, files, or native modules.
With a budget of 50, system, goal, and pin fit (38). Recent needs 30 more, so it is skipped, and the summary is skipped too. To keep the tail, raise the budget or shrink an earlier tier. Do not drop pins to make room for chatter.
Pitfalls
- Letting the model “remember” an amount that is only in prose.
- One vector collection for all tenants because it was easier to demo.
- Dropping the system prompt when the tool dump is large.
- Storing playbooks as free text with no version, then shipping a new tool schema.
- Calling agent memory “RAG” and re-teaching embeddings. Link the hub and move on.
Write five turns: the ask, a tool result with an amount, a human approval, a repeated status dump, and a closing line. Circle what must remain after compaction. If the amount is only inside a sentence, move it to a pin.
Interview Q&A
What is working memory versus RAG?
Answer
Working memory is the live window for this control loop. RAG retrieves external documents so a new generation can cite them. Episodic and semantic stores sit between those: durable agent state, not the corpus index. The document path is RAG and Vector Databases.
How do you prevent summary amnesia on refunds?
Answer
Pin structured fields outside the free-text summary: order id, amount, currency, and approval id. The compactor may rewrite prose. It may not rewrite pins. The next tool call reads the pins.
What about a user with two tabs?
Answer
Isolate by session id. Do not merge transcripts just because the user id matches. Share a semantic preference only when the product has an explicit consent flag, and still scope it by tenant.
What is the fill order inside the window?
Answer
Policy, goal, active schemas, pins, recent tail, summaries, then cited documents if they still fit. If you must cut, cut from the bottom of that list. Cutting the policy to keep a tool dump is how the loop forgets its own rules.
Why is recursive summarization dangerous?
Answer
Each pass conditions on the previous summary. A dropped negation or a dropped amount becomes the source the next pass trusts. Pins and the raw episode remain available for debug even when the window only sees the summary.
Why is vector memory not a free long-term brain?
Answer
Nearest-neighbor search over chat logs ignores tenant boundaries unless you filter first, and it ignores retention. You will surface another customer's ticket and you will keep PII longer than policy allows. Use a narrow semantic store with a key and a TTL. Use the RAG hub for documents.
How do procedures differ from semantic facts?
Answer
A semantic fact is “this user prefers email.” A procedure is the refund graph or the playbook, versioned with the tool schema. When the schema changes, old procedures are stale code, not memories to retrieve blindly.
What do you keep if the semantic store is down?
Answer
The session episode and the pins. The run degrades. It does not block on a preference lookup, and it does not invent a fact to fill the gap.
Go Deeper
- MemGPT paper for hierarchical memory as an operating-system metaphor. The pin rules on this page are the production constraint the metaphor still needs.
- Lost in the Middle for why a longer window does not mean the model uses the early tokens.
- Anthropic — contextual retrieval is about documents. Read it with the RAG hub, not as a session store.
- GDPR Article 5 for storage limitation on chat logs and derived facts.