Agent Architecture — Planner/Executor, Tool Registry & Control Loops
Production agents need an explicit control architecture: who plans, who executes, how tools are registered and validated, and where humans gate irreversible actions. Compare ReAct, plan-then-act, and a graph, and point at structured outputs for argument shape without re-teaching constrained decoding.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a tool loop actually repeat?
Answer
Think or plan, emit one action, validate it, observe the tool result, and decide whether to stop.
L2
When is ReAct the wrong default?
Answer
When the path must be auditable, the writes are irreversible, or a reviewer must approve a milestone before it runs.
L3
What does the planner return, and what does the executor return?
Answer
The planner returns a plan or the next edge. The executor returns one tool call or a message, then an observation.
L4
What lives in the tool registry besides a name?
Answer
A schema, an auth scope, a side-effect class, a timeout, a retry policy, an idempotency rule, and an allowlist.
L5
Why validate arguments in the runtime if the model emitted JSON?
Answer
Emission is not enforcement. The registry rejects missing fields and refuses money tools that lack approval.
L6
What must a human gate show, and what happens on timeout?
Answer
Show the argument diff and the amount. On timeout, fail closed. Do not execute the refund.
L7
How do you score a planner separately from an executor?
Answer
Grade plan quality against the playbook. Grade the executor on schema-valid calls and exactly-once side effects.
Failure modes
Planner calls the refund tool
The planning model mutates state, so you cannot review the plan before the write.
Schema only in the prompt
The model drifts and the runtime executes whatever JSON arrived.
Human gate fails open
The reviewer times out and the refund still runs.
Mis-routed specialist
A router sends a billing task to a shipping agent and the wrong tools fire.
Graph with no terminal edge
The state machine can sit on a node forever because stop was not a transition.
Misconceptions
The framework is the architecture.
Nodes, edges, registries, and gates are the architecture. The library is an implementation.
Tool calling and structured outputs are the same lesson.
Structured outputs constrain a shape. This page is about who validates and who executes.
Twelve specialist agents are a design.
Fan-out is a cost and a merge problem. Use it when the subtasks are independent.
Interviewer traps
Reciting an SDK method list.
Say planner, executor, registry, and human gate. Mention a library only as an example of graphs or durable resume.
Teaching JSON Schema or constrained decoding in depth.
Point at the Structured Outputs hub, then say the registry still validates before the tool runs.
Design scenario
Same prompt for every reader.
Requirements
Explicit states, an audit edge, and a human approval before any refund above 50 USD. The planner cannot call the money tool.
Traffic / scale
A few hundred refund investigations per hour, each with two or three read tools.
Latency
A draft reaches the reviewer within 5 seconds. Execution waits on the human SLA.
Consistency
A refund runs only after approval, and a second delivery of the same approval does not pay twice.
Availability
If the reviewer does not answer before the SLA, the transition fails closed and the ticket stays open.
Failure assumptions
- The model can emit a refund call during planning.
- The reviewer can time out.
- Tool arguments can miss required fields.
Constraints
- Only the executor invokes mutating tools.
- Money tools require an approval id.
- Unknown tools are rejected by the registry.
Prompt
Design the control path for a banking refund agent.
API
What does the planner emit, and which call is the executor allowed to make?
Data
Where do the plan, the approval, and the tool schema live?
Architecture
Which graph nodes are read, draft, wait-for-human, and execute?
A refund that must be explainable next quarter
Prefer
A graph with a human node
States and edges are data. The executor is the only writer. Approval is a transition, not a sentence in the trace.
- Read tools can run before the draft.
- The money edge does not exist until approval is present.
- Timeout takes the fail-closed edge.
- You can eval the plan separately from the tool call.
Alternative
An open ReAct loop with every tool attached
Fast to demo. The model can wander into refund, email, and a second refund before anyone looks.
- The audit is a chat log.
- Stop conditions are hopes in the prompt.
- A bad route still has the money tool in scope.
Overview
Interviews reward a control architecture, not a vendor tour. You should be able to say who proposes the next step, who is allowed to execute it, where the schema lives, and which edge waits for a person.
Three loops cover most answers:
- ReAct. Interleave a thought, a tool call, and an observation. Good for short exploration. Weak when you need a bound or an audit.
- Plan-then-act. Draft milestones, then execute them. Good for a known playbook. Brittle if the plan is wrong and nobody revises it.
- Graph or state machine. Nodes, edges, and typed state. More engineering. The right default for compliance, refunds, and multi-actor work.
A router plus specialists is a fourth shape: classify, then dispatch. A mis-route cascades, so the router’s tools must be narrower than the union of every specialist.
Decisions
- 1
1. Goal arrives
- next2. Exploratory?
- ?
2. Exploratory?
- yes3. ReAct loop
- no4. Need audit edges?
- 3
3. ReAct loop
- next7. Tool registry
- ?
4. Need audit edges?
- yes5. Graph with HITL
- no6. Plan then execute
- 5
5. Graph with HITL
- next7. Tool registry
- 6
6. Plan then execute
- next7. Tool registry
- 7
7. Tool registry
- next8. Irreversible?
- ?
8. Irreversible?
- yes9. Human gate
- no10. Effect or answer
- 9
9. Human gate
- next10. Effect or answer
- 10
10. Effect or answer
Lesson map
Agent Architecture — Planner/Executor, Tool Registry & Control Loops
Production agents need an explicit control architecture: who plans, who executes, how tools are registered and validated, and where humans gate irreversible actions. Compare ReAct, plan-then-act, and a graph, and point at structured outputs for argument shape without re-teaching constrained decoding.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB ex["Executor"] po["Policy"] hu["Reviewer"] api["Tool API"] ex -->|Propose refund| po po -->|Show the arg| hu hu -->|Approve or deny| po po -->|Run only if| api api -->|Return receipt| ex
The registry sits on every path. The human gate is a property of the tool, not of the loop style.
Planner and executor
Planner. Proposes milestones, a tool sequence, or the next graph edge. A larger, slower model is acceptable. Its output is data: plan JSON or an edge name.
Executor. Runs one step with a tight schema, a cheaper model, and a strict timeout. Its failure mode is bad arguments or a retry, not a bad strategy.
| Concern | Planner | Executor |
|---|---|---|
| Model size | Larger and slower is fine | Smaller and fast |
| Output | Plan or graph edge | One tool call or a message |
| Failure mode | Bad plan | Bad arguments or retries |
| Eval | Plan quality | Step success and side effects |
Rule: if you need to audit a write, the planner must not call mutating tools. The plan is data. The executor is the only writer.
Tool registry
A tool is a product surface, not a function pointer the model discovered.
- Name and description, so the model can select it.
- A JSON schema for arguments. Shape and constrained decoding are the Structured Outputs hub. This runtime still validates before it executes.
- Auth scope: which credential and which tenant.
- Side-effect class: read, write, money, or external communication.
- Timeout, retry policy, and idempotency.
- Allowlist per agent role.
| Side-effect class | Example | Default policy |
|---|---|---|
| read | get order | Automatic, and cacheable if the read is stable |
| write | update ticket | Automatic, with an idempotency key |
| money | refund | Human gate above a threshold |
| external-comm | send email | Rate limit and a template allowlist |
The model emitting JSON is not the same as the runtime accepting it. System One is where a classifier routes or guards. It does not replace the registry.
Human gate
Sequence
- 1
Executor → Policy
Propose refund
- 2
Policy → Reviewer
Show the arg diff
- 3
Reviewer → Policy
Approve or deny
- 4
Policy → Tool API
Run only if approved
- 5
Tool API → Executor
Return receipt id
The gate is not a checkbox in a prompt. Specify what the reviewer sees (the argument diff and the amount), the SLA, and the timeout edge. Timeout fails closed: no receipt, ticket stays open.
Refund investigation, then one write
Reads can be a plan. The money call is a separate transition.
- 1
Plan the reads
Order, payments, and a draft summary. No refund tool is in scope. - 2
Validate each step
The registry checks required fields before the executor runs the tool. - 3
Wait for approval
Above 50 USD the policy desk shows the diff. Silence is a deny. - 4
Execute once
The executor sends the refund with an idempotency key and stores the receipt.
Frameworks, at concept level
| Concept | You should explain | Not the interview goal |
|---|---|---|
| Message and tool loop | How a tool call and an observation alternate | Memorizing SDK names |
| Graph orchestration | Nodes, edges, checkpoints | One library’s method list |
| Durable workflow | Resume after a crash from saved state | Rewriting a workflow tutorial |
| Multi-agent | Fan-out and fan-in contracts | Spawning many agents for a hello-world |
Runnable sketch — registry
type JsonSchema = {
type: "object";
required?: string[];
properties: Record<string, { type: string }>;
};
export type ToolSpec = {
name: string;
sideEffect: "read" | "write" | "money" | "external-comm";
schema: JsonSchema;
execute: (args: Record<string, unknown>) => Promise<unknown>;
};
export class ToolRegistry {
private tools = new Map<string, ToolSpec>();
register(spec: ToolSpec) {
this.tools.set(spec.name, spec);
}
async call(name: string, args: Record<string, unknown>, opts: { hitlApproved?: boolean } = {}) {
const tool = this.tools.get(name);
if (!tool) throw new Error(`unknown tool: ${name}`);
for (const key of tool.schema.required ?? []) {
if (!(key in args)) throw new Error(`missing arg: ${key}`);
}
if (tool.sideEffect === "money" && !opts.hitlApproved) {
throw new Error("human gate required for money tools");
}
return tool.execute(args);
}
}Press Run. Snippets must be self-contained — no network, files, or native modules.
Plan, then execute
Press Run. Snippets must be self-contained — no network, files, or native modules.
The plan has no refund step. That is the architecture, not an accident of the stub.
Pitfalls
- Putting every tool on the planner “so it can be smart.” You lose the review edge.
- Copying a schema into the prompt and never checking it in code.
- Treating a human gate as best-effort. Timeout must be a deny.
- Using a graph for a one-shot FAQ. The extra states are theater.
- Teaching constrained decoding here. Link Structured Outputs and move on.
For each task — “summarize this PDF,” “investigate a refund over 50 USD,” “look up an order status” — say ReAct, plan-then-act, or graph, and name the tool that is not allowed to run until a gate opens.
Interview Q&A
ReAct or a graph for banking refunds?
Answer
A graph with a human node. You need explicit states, audit edges, and a fail-closed transition. Free-form wandering will eventually call refund because the tool was in scope.
Where do schemas live?
Answer
In the tool registry, as the source of truth. The prompt describes the tool. The runtime rejects a missing field. Decoder strictness is Structured Outputs. Both can be true: a valid JSON object can still be the wrong tool.
Why split planner and executor?
Answer
Cost, safety, and evals. The executor can be cheaper and tightly timed. Mutating tools stay on the executor so a plan can be reviewed. You can score the plan without scoring the side effect, and the reverse.
What is a side-effect class?
Answer
A label the policy understands: read, write, money, or external communication. Reads may be automatic and sometimes cacheable. Writes need an idempotency key. Money needs a person above a threshold. Email needs a rate limit and a template allowlist.
What does the human see, and what if they never answer?
Answer
They see the argument diff: tool name, amount, order id, and the draft. If the SLA expires, the edge is deny. The executor does not get a receipt id. Leaving the default as “approve” is a production incident.
Why is a router plus specialists dangerous?
Answer
The router can send billing work to the shipping agent, and that agent may still have a write tool. Constrain each specialist’s allowlist. A mis-route should be a dead end, not a different mutation.
How do you explain a graph without naming a library?
Answer
Nodes are states, edges are allowed transitions, checkpoints are saved state you can resume. A human gate is a node with two outgoing edges, approve and deny. A library can store that graph. The interview is the graph.
Where do route and guard classifiers sit?
Answer
Beside the loop, not inside the tool schema. System One and Jev covers classifiers for intent, toxicity, and confidence. This page still owns who may call the tool after the route is chosen.
Go Deeper
- ReAct paper for the interleaved loop, then put a registry under it.
- OpenAI function calling and Anthropic tool use for the wire format.
- LangGraph overview for nodes and checkpoints as a concept.
- Temporal activity failures for durable resume, not a product pitch.
- OWASP LLM Top 10 for excessive agency and tool abuse.