Caching for Agents — Prompt Cache, KV Reuse, Semantic Cache & Tool Result Cache
The wrong agent cache is a correctness bug. Separate provider prompt cache, runtime KV reuse, semantic response cache, and tool-result memoization. This page is the application policy. PagedAttention and the KV layout stay on the inference hub.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Name the four cache layers.
Answer
Provider prompt cache, runtime KV, semantic response cache, and tool-result cache.
L2
Who owns KV reuse?
Answer
The inference runtime. Application code does not invent PagedAttention. It avoids prefix churn so the engine can reuse state.
L3
When is a semantic hit safe?
Answer
Public FAQ or a generic how-to, with a similarity threshold, a version, and no personal data or hourly prices.
L4
What is the tool-result cache key?
Answer
Tenant, tool name, canonical arguments, auth scope, and schema version, plus a TTL.
L5
How do you order a prompt for cache hits?
Answer
Rarely changing blocks first: policy and tool schemas. Volatile user text last.
L6
What invalidates a cache?
Answer
A schema or prompt version bump, a policy change, an entity write, or TTL expiry.
L7
How is a semantic cache different from RAG?
Answer
RAG retrieves evidence for a new generation. A semantic cache returns a previous answer and skips generation.
Failure modes
Cross-user semantic hit
User A's refund guidance is returned for User B because the embedding ignored tenant and identity.
Stale policy answer
The terms changed and the semantic entry was not versioned.
Cached auth or secret
A token, a cookie, or an account number was stored as a tool result.
Prefix churn
A timestamp in the system prompt defeats prompt cache on every call.
KV treated as an app map
The service tries to share attention state across tenants and isolation breaks.
Misconceptions
Semantic cache is RAG.
RAG grounds a new answer. Semantic cache skips the model and returns an old one.
Prompt cache and KV cache are the same knob.
Prompt cache is a provider billing feature for a stable prefix. KV is runtime attention state.
If the tool is a GET, it is always safe to cache.
Not when the result depends on the caller, the clock, or a secret.
Interviewer traps
Explaining paged KV blocks and prefix trees.
Name the inference hub for PagedAttention. Stay on what the agent runtime is allowed to store.
Caching the final agent answer by default.
Default to none. Opt in for FAQ-like, non-personal, versioned text.
Design scenario
Same prompt for every reader.
Requirements
FAQ may skip the model. Order status may cache for a short TTL per tenant. Refunds, tokens, and personalized answers are never cached. Prompt prefixes stay stable.
Traffic / scale
A large share of turns repeat the same policy preamble and the same order lookup.
Latency
A semantic FAQ hit returns without a model call. A tool hit removes one round trip. Prefill still depends on the provider prefix cache.
Consistency
An order update invalidates that order's read key. A policy publish bumps the semantic version.
Availability
A cache outage falls through to the tool and the model. It does not fail the turn.
Failure assumptions
- A similarity hit can be the wrong tenant.
- A prompt can include a timestamp and kill prefix reuse.
- A tool named like a read can still return a secret.
Constraints
- Banned tools bypass the cache.
- Keys include tenant id and schema version.
- KV layout is not implemented in the app.
Prompt
Add caching to a support agent that answers public FAQ and also reads order status.
API
Which responses can be served from a semantic hit, and which must call the model?
Data
What fields are in a tool-cache key, and what is banned?
Architecture
Where do prompt assembly, the tool cache, and the inference KV sit?
The same refund policy is prepended to every call
Prefer
A stable prefix, and a narrow app cache
Policy and tool schemas stay byte-stable at the front of the prompt. Tool reads cache only when the ban list says so. FAQ answers can skip the model. Everything else misses on purpose.
- Prompt-cache hits are a billing and TTFT win.
- Order status has a short TTL and a tenant key.
- Refunds and tokens never enter the store.
- KV stays inside the inference process.
Alternative
One semantic cache in front of the whole agent
The highest cost win, and the easiest way to serve the wrong customer's balance after a policy change.
- Similarity ignores authorization.
- A timestamp in the prompt wastes the prefix cache.
- The app tries to manage attention state it does not own.
Overview
Caching is why an agent fleet can be affordable. It is also a correctness feature: the default is do not store. You opt in per layer.
| Layer | What is reused | Who owns it | Safe when | Unsafe when |
|---|---|---|---|---|
| Prompt cache | Stable prefix tokens | Provider API | Long static policy and tools | The prefix changes every call |
| KV reuse | Attention state | Inference runtime | The engine isolates sessions | You share state across requests yourself |
| Semantic cache | A near-duplicate answer | Application | FAQ, low staleness | Personalized or time-sensitive |
| Tool-result cache | GET-like outputs | Application | Idempotent reads plus a TTL | Auth, money, PII, or writes |
Decisions
- 1
1. Agent step
- next2. Safe semantic hit?
- ?
2. Safe semantic hit?
- yes3. Return cached answer
- no4. Tool read cached?
- 3
3. Return cached answer
- ?
4. Tool read cached?
- yes5. Inject observation
- no6. Call the tool
- 5
5. Inject observation
- next7. Stable prompt prefix
- 6
6. Call the tool
- next7. Stable prompt prefix
- 7
7. Stable prompt prefix
- next8. Store eligible caches
- 8
8. Store eligible caches
Lesson map
Caching for Agents — Prompt Cache, KV Reuse, Semantic Cache & Tool Result Cache
The wrong agent cache is a correctness bug. Separate provider prompt cache, runtime KV reuse, semantic response cache, and tool-result memoization. This page is the application policy. PagedAttention and the KV layout stay on the inference hub.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Agent step"] b["2. Safe semantic hit?"] c["3. Return cached answer"] d["4. Tool read cached?"] a -->|1. Agent step to 2. Safe semantic hit?| b b -->|yes| c b -->|no| d
A semantic hit returns before tools. A tool hit still goes to the model, with the observation filled in. Step 7 is where prompt cache applies. Step 8 writes only what the ban list allows. Runtime KV is not a box you implement. LLM Inference Runtime owns PagedAttention, prefill, and decode.
Prompt cache versus KV
| Approach | Idea | Interview note |
|---|---|---|
| Provider prompt cache | A marked or automatic stable prefix, cheaper on hit | Design prompts so policy and tools do not churn |
| Local assembly cache | Remember the rendered string | You still pay full inference unless the provider supports reuse |
| Session KV | Keep decode state warm | Belongs to the engine. Do not reinvent it |
Design tip: rarely changing blocks first, volatile user text last. A date stamp or a random request id at the top of the system prompt is an own goal.
Semantic cache
Embed the query, optionally with the tenant, and return a prior answer when similarity crosses a threshold and the entry's version matches.
| Safer | Ban |
|---|---|
| Public FAQ | Account balances |
| Doc explainers that are not personal | Authorization decisions |
| Generic how-tos | Anything with PII in the query or the answer |
| Non-personalized tips | Prices that change hourly |
The failure mode is serving one user's refund guidance to another, or a policy answer from before the terms changed. Version the entry with the prompt version. Isolate by tenant even for FAQ if the wording differs by contract.
This is not retrieval. RAG fetches passages so the model can write a new, cited answer. A semantic cache skips that generation. If you needed evidence, you needed RAG.
Tool-result memoization
Cache key: tenant, tool name, canonical arguments, auth scope, and schema version.
Banned names skip the store even if a caller forgets the flag. Reads that are truly idempotent get a TTL. A write, a refund, or a secret never does.
Press Run. Snippets must be self-contained — no network, files, or native modules.
refund runs the function every time. get_order runs it once inside the TTL. The key includes the tenant (t1) so another tenant's A-1 is a different entry.
Invalidation
| Event | Action |
|---|---|
| Tool schema version bump | New cache-key namespace |
| Policy or prompt change | Bust semantic entries and the prompt-prefix version |
| Entity write, such as an order update | Invalidate the related read keys |
| TTL expiry | Lazy delete on the next read |
A cache miss is a correct outcome. Serving a hit after a write is not.
Decision table
export type CacheDecision = "prompt" | "semantic" | "tool" | "none";
export function decideCache(input: {
personalized: boolean;
toolMutates: boolean;
isFaqLike: boolean;
hasPii: boolean;
}): CacheDecision {
if (input.hasPii || input.toolMutates || input.personalized) return "none";
if (input.isFaqLike) return "semantic";
return "tool";
}Press Run. Snippets must be self-contained — no network, files, or native modules.
Prompt-prefix stability is orthogonal: you still order the prompt even when the app decision is none. The table above is the app store, not the provider prefix.
Cost and latency
| Cache | Latency win | Cost win | Correctness risk |
|---|---|---|---|
| Prompt | Medium, mostly time to first token | High on long prefixes | Low if the prefix is pure |
| KV | High on decode | Infra efficiency | Isolation bugs inside the runtime |
| Semantic | Highest, the model is skipped | Highest | Highest |
| Tool | Medium | Medium | Staleness |
What not to cache
- Auth tokens, session cookies, and API keys.
- Personalized PII and health or financial secrets.
- Non-idempotent tool results.
- Answers that depend on “now” unless the TTL is shorter than the change.
- Cross-tenant embeddings without a hard filter.
Pitfalls
- A cache key that forgets the tenant or the schema version.
- Putting the user id in the system prefix so prompt cache never hits, then “fixing” cost with a semantic cache of personalized answers.
- Invalidating on a timer only, while
update_orderalready changed the row. - Re-explaining paged blocks. Link the inference hub and describe the policy.
For get_order, refund, and “what is your return window?”, say prompt, semantic, tool, or none. Write the cache key for the one that is a tool hit, and the version bump that kills the FAQ entry.
Interview Q&A
Is a semantic cache the same as RAG?
Answer
No. RAG retrieves evidence so the model can write a new answer and cite it. A semantic cache returns a previous answer and skips generation. The retrieval lesson is RAG and Vector Databases. If the user needs a source passage, do not serve a semantic hit.
What is the difference between prompt cache and KV cache?
Answer
Prompt cache is a provider feature that reuses a stable prefix and bills the hit cheaper. KV cache is attention state inside the serving engine, including PagedAttention in LLM Inference Runtime. You design the prefix. You do not implement the pager.
How do you version caches?
Answer
Put prompt version, tool schema version, and tenant id in the key. A policy publish bumps the prompt version and drops semantic entries. A schema publish changes the tool namespace so old argument shapes cannot hit.
What must you never cache?
Answer
Secrets and tokens, personalized PII, non-idempotent tools, and anything whose correctness depends on the current minute unless the TTL is honest. When in doubt, the decision is none.
How do you get prompt-cache hits?
Answer
Keep policy and tool schemas byte-stable and put them first. Put the user text, the date, and the retrieval snippets after that prefix. A random id in the system message is prefix churn.
What is in a tool-result key?
Answer
Tenant, tool name, canonical arguments, auth scope, and schema version. Canonical means sorted keys and normalized ids, so two JSON orders are one entry. The ban list short-circuits before the lookup.
A write just updated the order. What else happens?
Answer
Invalidate read keys that include that order id. Do not wait for the TTL if a user can read-your-writes. Semantic entries that quoted the old status need the same bust, or a version that no longer matches.
Which layer saves the most money, and what does it cost you?
Answer
The semantic layer, because the model does not run. The cost is correctness: a near miss is a wrong answer delivered with high confidence. Use it for FAQ-like text. Use tool cache for a smaller, safer win. Use prompt cache for the prefix you repeat anyway.
Go Deeper
- Anthropic prompt caching and OpenAI prompt caching for the prefix feature.
- PagedAttention paper and vLLM docs for KV. Return here for the application policy.
- MDN HTTP caching for the mental model behind tool GETs: key, TTL, and invalidation.