Production AI Agents — Architecture, Scaling, Caching & Reliability
A production AI agent is a policy loop: model, tools, memory, and guardrails, with durable state, cost and latency SLOs, and cache layers. It is not a chat box glued to one completion. This hub maps architecture, memory, scale-out, caching, and reliability, and only cross-links the RAG, structured-outputs, and inference hubs.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is an agent versus a single LLM call?
Answer
A single call returns one completion. An agent repeats observe, act, and remember until a stop condition or a budget.
L2
When do you pick ReAct, plan-then-act, or a graph?
Answer
ReAct for short exploration. Plan-then-act for a known playbook. A graph when you need explicit edges, audits, and human gates.
L3
How do you budget context and compact a session?
Answer
Pin policy, the goal, and money or id facts. Summarize chatter. Do not drop approvals or open obligations.
L4
How do you scale concurrency without melting provider rate limits?
Answer
Queue work, apply per-tenant quotas, and put a shared limiter and circuit breaker in front of the provider and the tools.
L5
Which caches are safe for agents?
Answer
Stable prompt prefixes, idempotent tool reads with a TTL and a tenant key, and semantic answers only for non-personal FAQ text. Runtime KV stays in the inference engine.
L6
How do you eval trajectories and break stuck loops?
Answer
Grade tool order, arguments, budget stops, and side effects, not only the final string. Cap identical calls and trip a breaker.
L7
How do you design a multi-tenant agent platform with cost SLOs and human gates?
Answer
Isolate memory by tenant, require approval for money tools, kill a run at a dollar cap, and degrade to an answer-only path when tools are unhealthy.
Failure modes
Unbounded tool loop
The same call repeats until the bill or the latency SLO is gone.
Non-idempotent retry
A retried refund or email lands twice.
Cache poisoning
A bad or cross-tenant cached answer is served as if it were fresh.
Cross-tenant memory leak
Semantic search or a shared session key returns another customer's facts.
Silent tool failure
The model treats an error observation as success and continues.
Fan-out cost explosion
Parallel tools or specialist agents multiply tokens and side effects.
Misconceptions
An agent is a chat product with plugins.
A production agent is a budgeted control loop with durable state, policies, and SLOs.
More tools are always better.
Every tool is a product surface with a side-effect class, a schema, and a blast radius.
A semantic cache is free correctness.
It skips generation. It is safe only when the answer is not personal, not time-critical, and versioned.
The KV cache is an application concern.
Application code owns prompt, semantic, and tool-result policy. Runtime KV belongs to the inference engine.
Interviewer traps
Redrawing retrieval, JSON schemas, or PagedAttention when the question was the agent runtime.
Name those hubs in one sentence, then stay on loops, tools, memory, queues, caches, and evals.
Promising autonomy without a stop condition.
State max steps, max tool calls, a dollar cap, and what happens when the budget hits.
Design scenario
Same prompt for every reader.
Requirements
Human approval for refunds above 50 USD. Episodic memory kept for 30 days. Degrade to an answer-only path when tools are unhealthy.
Traffic / scale
About 2000 concurrent sessions.
Latency
Acknowledgement first token within 2 seconds. Resolved-ticket cost around 0.05 USD at p50 and 0.40 USD at p95.
Consistency
A refund is exactly-once for one idempotency key. Session memory is isolated by tenant.
Availability
If billing is down, the agent still acknowledges and queues a human. It does not retry money writes.
Failure assumptions
- A loop can run without a step budget.
- A retry can double a refund.
- A shared cache can leak across tenants.
Constraints
- Money tools require human approval above 50 USD.
- Memory keys include a tenant id.
- Max steps, tokens, and wall-clock are enforced.
Prompt
Design a multi-tenant support agent with CRM and billing tools.
API
What does the client send to start a run, and what acknowledgement returns before tools finish?
Data
Where do the session, the episodic trace, and the idempotency key live?
Architecture
Where do the queue, the worker, the human gate, and the degrade path sit?
A support bot that can refund, or only talk
Prefer
A budgeted policy loop with tools and memory
The runtime owns stop conditions, schemas, tenant isolation, and human gates. The model proposes the next action. Code decides whether that action may run.
- Side effects carry an idempotency key.
- Refunds above a threshold wait for a person.
- Cost and step budgets kill a runaway loop.
- Traces join model calls and tool calls on one run id.
Alternative
One chat completion, plugins optional
Fine for a demo. There is no durable session, no quota, and no story for a retry that sends the refund twice.
- State is whatever fit in the last messages.
- Tool errors look like more chat.
- Cost variance is invisible until the invoice.
- A stuck loop has nothing to trip.
What you say before you draw boxes
The interviewer is asking whether you would ship the loop, not whether you can name a framework.
- 1
Name the SLO
Task success, p95 end-to-end latency, dollars per successful task, and tool error rate. - 2
Name the blast radius
Which tools mutate, send, or move money. Those need allowlists and a human gate. - 3
Name the stop
Max steps, max tool calls, max tokens, and a wall-clock deadline. - 4
Name the degrade path
When tools are unhealthy, answer from a safe path and queue a human. Do not keep calling refund.
Overview
A production AI agent is a control loop. It observes a goal and session state, chooses an action, executes tools, writes memory, and stops. The model is the policy proposal. The runtime is the product: schemas, budgets, isolation, caches, and traces.
Use an agent when the next tool is not known up front, or the task needs several environment steps. Prefer a script or a single structured call when the pipeline is fixed. Agents add variance, cost, and tool risk.
This cluster owns that productization:
- Architecture — ReAct, plan-then-act, graphs, and the tool registry.
- Memory — working context, episodes, compaction, and sessions.
- Scaling — queues, fan-out, quotas, and cost budgets.
- Caching — prompt, semantic, and tool-result policy.
- Reliability — trajectory evals, guardrails, and tracing.
Each page is its own lesson. This hub does not re-teach them.
Toy chatbot versus production agent
| Dimension | Toy chatbot | Production agent |
|---|---|---|
| Loop | One completion | Multi-step tool loop with stop conditions |
| State | Stateless messages | Durable session plus episodic and semantic stores |
| Tools | Optional demos | Registry, schemas, allowlists, sandboxes |
| Scale | One user | Queues, workers, per-tenant quotas |
| Cache | Accidental | Prompt, semantic, and tool caches with invalidation |
| Reliability | Manual review | Offline and online evals, tracing, circuit breakers |
| Cost | Ignored | Budget per task plus a kill switch |
| Blast radius | Low | Email, money, and writes |
Production checklist
- SLO. Task success rate, p95 end-to-end latency, dollars per successful task, tool error rate.
- Blast radius. Which tools mutate. Human approval or dual control for money and deletes.
- Idempotency. Side effects keyed so a retry is safe. A typical key joins the run id and the step id.
- Budgets. Max steps, max tokens, max tool calls, wall-clock deadline.
- Observability. A span per model call and per tool, correlated by run id.
- Evals. Trajectory grading offline, a canary online, and a regression gate in CI.
- Cache policy. What is cached, the TTL, the tenant scope, and a ban on secrets and PII.
- Degrade mode. Fall back to an answer-only path when tools are unhealthy.
The loop
Decisions
- 1
1. Goal and session
- next2. Context budget
- 2
2. Context budget
- next3. Propose action
- 3
3. Propose action
- next4. Action type?
- ?
4. Action type?
- tool5. Validate and run
- answer9. Guardrail check
- 5
5. Validate and run
- next6. Append observation
- 6
6. Append observation
- next7. Budget left?
- ?
7. Budget left?
- yes3. Propose action
- stop8. Persist result
- 8
8. Persist result
- 9
9. Guardrail check
- next8. Persist result
Lesson map
Production AI Agents — Architecture, Scaling, Caching & Reliability
A production AI agent is a policy loop: model, tools, memory, and guardrails, with durable state, cost and latency SLOs, and cache layers. It is not a chat box glued to one completion. This hub maps architecture, memory, scale-out, caching, and reliability, and only cross-links the RAG, structured-outputs, and inference hubs.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Goal and session"] b["2. Context budget"] c["3. Propose action"] d["4. Action type?"] a -->|1. Goal and session| b b -->|2. Context budget| c c -->|3. Propose action| d
Read it as a product, not a prompt. The model proposes. The runtime validates the schema, checks the side-effect class, and either executes, asks a human, or stops. The back edge is the only place a loop is allowed, and the budget node is what cuts it.
Grounding and document retrieval live in RAG and Vector Databases. Argument shape and constrained decoding live in Structured Outputs. Token-level KV and PagedAttention live in LLM Inference Runtime. Classifiers for route and guard live in System One and Jev, including the hybrid control plane. This cluster does not re-teach those pages.
Control patterns
| Pattern | Strength | Weakness | Prefer when |
|---|---|---|---|
| ReAct (reason and act) | Flexible and small | Can wander and is hard to bound | Short exploration |
| Plan-then-act | Clear milestones | A wrong plan stays wrong | Known playbooks and audits |
| Graph or state machine | Explicit edges and human nodes | More engineering | Compliance, refunds, several actors |
| Multi-agent fan-out | Parallel specialists | Coordination tax and cost | Independent subtasks with a cap |
- 1
Receive the goal
Load the session with a tenant key. Reject a missing isolation key.
- 2
Assemble a budgeted context
Pin policy, goal, and ids. Summarize the rest. Leave room for the next tool schema.
- 3
Propose one action
The model returns a tool call or a final answer. The registry is the source of truth for arguments.
- 4
Gate side effects
Reads may proceed. Writes need an idempotency key. Money waits for a person above the threshold.
- 5
Stop or continue
A budget, a terminal answer, or a breaker ends the run. Persist the trace either way.
Runnable sketch — budgeted loop
The policy is a stub. The point is the stop condition: steps and tool calls are counters the runtime owns, not a sentence in the prompt.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Policy knobs
The typed sketch is the interview card. The runnable below is the same refund rule without a class hierarchy.
export type AgentSlo = {
successRateTarget: number;
p95LatencyMs: number;
maxUsdPerSuccess: number;
};
export type AgentPolicy = {
maxSteps: number;
maxToolCalls: number;
hitlToolNames: string[];
cacheableTools: string[];
};
export function assertRefundHitl(
toolName: string,
amountUsd: number,
policy: AgentPolicy,
): void {
if (toolName === "refund" && amountUsd > 50 && !policy.hitlToolNames.includes("refund")) {
throw new Error("refund requires a human gate in policy");
}
}Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
- Shipping a demo loop and calling the step counter a prompt sentence. The runtime must increment it.
- Retrying a refund because the HTTP client retries. The tool must recognize the same idempotency key.
- Stuffing a document corpus into the agent window. That is retrieval. Link the RAG hub and keep this page on agent state.
- Treating provider prompt-cache discounts as if you had designed the KV cache. The inference lesson owns PagedAttention.
- Letting every tenant share one semantic cache. Personal answers and PII do not belong there.
Draw one support run that reads an order, drafts a refund, and stops for a person above 50 USD. Mark the idempotency key, the span names, and the node that refuses a second refund with the same key.
Interview Q&A
When is an agent the wrong abstraction?
Answer
When a deterministic workflow or one structured call solves it. Agents add variance, cost, and tool risk. Prefer a graph or a script for a fixed pipeline, and a single constrained call when you only need a shape. Shape lives in Structured Outputs.
What is blast radius for an agent?
Answer
The worst irreversible side effect a looping or compromised agent can cause: mass email, refunds, deletes. Shrink it with allowlists, human gates, sandboxes, and rate limits. Name the tool class out loud: read, write, money, or external communication.
How do you talk about caching without confusing inference KV?
Answer
Separate the layers. The provider prompt cache reuses a stable prefix. The runtime KV cache is attention state inside the inference engine, covered by LLM Inference Runtime. The app owns a semantic cache and a tool-result cache, with tenant scope and a TTL. The caching lesson is the policy. Do not redesign PagedAttention here.
What belongs on the production checklist?
Answer
SLOs, blast radius, idempotency, step and dollar budgets, spans per model and tool call, trajectory evals, a written cache policy, and a degrade mode. If you cannot say what happens at the dollar cap, you do not have a product yet.
ReAct, plan-then-act, or a graph?
Answer
ReAct for a short exploratory tool sequence. Plan-then-act when a reviewer should see milestones before writes. A graph when refunds, compliance, or several actors need explicit edges and a fail-closed human node. The architecture lesson compares them. Do not spawn extra agents to avoid picking one.
How do you bound a loop?
Answer
Counters in the runtime: max steps, max tool calls, max tokens, wall-clock. A loop detector for the same tool and the same arguments. A circuit breaker on the tool error rate. The prompt may ask the model to stop. The process must stop anyway.
What is a degrade mode?
Answer
When the provider or a tool is unhealthy, stop mutating and return a safe acknowledgement or an answer that does not depend on the failed tool. Queue a human. Shedding work is better than a 429 storm or a partial refund.
How do you keep this cluster from swallowing RAG, schemas, and inference?
Answer
One sentence each. Documents and citations are RAG. JSON shape is Structured Outputs. Prefill, KV, and batching are LLM Inference Runtime. Route and guard classifiers are System One. Then return to the loop, the registry, the queue, and the eval.
Go Deeper
- ReAct paper for the reason-and-act loop, then come back and put a budget on it.
- OpenAI function calling and Anthropic tool use for the wire shape of a tool call.
- Anthropic prompt caching for stable prefixes. Runtime KV stays in the inference hub.
- LangGraph overview and Temporal docs for graph and durable-workflow ideas, not an SDK quiz.
- OpenTelemetry GenAI conventions for span names.
- Hugging Face agents course for a guided tour of the same loop.