Agent Reliability — Evals, Guardrails, Tracing & Failure Modes
Agents fail differently from chatbots: stuck loops, poison tool output, partial side effects, and trajectory drift. Reliability means grading whole runs, sandboxing tools, circuit breakers, and span-level traces. A fluent final string is not a pass.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What do you grade besides the final string?
Answer
Tool choice and order, schema-valid and authorized arguments, budget stops, faithfulness to observations, and exactly-once side effects.
L2
Offline versus online evals?
Answer
Offline runs golden trajectories in CI. Online uses shadow, canary, and sampled production traces. Each misses what the other sees.
L3
What is a poison tool observation?
Answer
Untrusted text from a tool or a page that tries to override policy. Never promote it into instructions.
L4
How do you detect a stuck loop?
Answer
The same tool name and the same argument key repeated N times. Stop, even if the model wants another try.
L5
What spans do you emit?
Answer
A root run span, then model, tool, guardrail, and human-gate spans, correlated by run id.
L6
Why does an offline pass fail online?
Answer
Live schema drift, rate limits, cache misses, and a shift in user goals. Sample production trajectories on a schedule.
L7
How do you roll back a tool that already ran?
Answer
Prefer a compensating action and a stored intent. Do not assume every tool has an undo. Atomic workflows belong in a durable engine.
Failure modes
Stuck loop
The same tool is called with the same arguments until the budget, or forever if there is no budget.
Poison tool output
Text inside an observation is obeyed as if it were the system prompt.
Oscillation
The agent does a write and then undoes it, or flips a field back and forth.
Budget exhaustion
The run stops mid-task with no checkpoint and no user-visible degrade.
Partial commit
One of several writes landed. There is no compensating action.
Cache poison
A bad semantic hit is trusted as the answer. Version and personalization bans belong to the caching lesson.
Misconceptions
A high score on the final answer means the agent is safe.
The path can still call the wrong tool, leak data, or double-charge.
Tool output is data the model should follow.
Tool output is untrusted content. Policy stays in the system channel.
A circuit breaker replaces a step budget.
The breaker trips on error rate. The budget trips on a loop that is successfully doing the wrong thing.
Interviewer traps
Quoting BLEU or a single helpfulness score.
Grade the trajectory and the side effects. Helpfulness is one column.
Re-teaching RAG faithfulness metrics in full.
Say faithfulness to observations uses the same idea as grounding, then point at the RAG hub and return to tool order and exactly-once.
Design scenario
Same prompt for every reader.
Requirements
Offline evals sandbox the email tool. Refunds above the threshold still require a person. A stuck-loop detector and a tool circuit breaker are on. Traces cover model, tool, and human spans.
Traffic / scale
CI runs a few hundred golden trajectories per change. Production samples a slice of live runs.
Latency
Guardrail checks sit on the tool path without adding a second model call unless the risk class requires it.
Consistency
Side effects in the eval sandbox are not real sends. Production sends are exactly-once per idempotency key.
Availability
When the email tool's breaker is open, the run degrades and does not retry-storm. A failed eval blocks release.
Failure assumptions
- Tool text can contain instructions.
- The same call can repeat.
- Offline fixtures drift from the live schema.
Constraints
- Email is allowlisted in offline runs.
- Identical tool calls trip a detector.
- Release requires the trajectory suite to pass.
Prompt
Ship an agent that can email customers and issue refunds under 50 USD.
API
What does an offline case assert about a would-send email?
Data
Which fields land on the model span and the tool span?
Architecture
Where do the sandbox, the breaker, and the canary allowlist sit?
The demo answer looks right, and the tool log looks wrong
Prefer
Score the path, then canary
CI replays goals against sandboxed tools and fails the build on a bad tool order or a double write. Production samples traces and canaries the email tool to an allowlist.
- Final text is one check among several.
- Poison strings in fixtures must not change the policy.
- A breaker and a step budget are both on.
- Compensating actions are named before launch.
Alternative
A helpfulness score on the last message
The suite stays green while the agent emails the wrong person or loops on a read. You find out from a customer.
- Side effects are invisible to the metric.
- Live schema drift is invisible until Monday.
- There is no span to join the model call to the tool.
Overview
Chatbots fail by saying the wrong sentence. Agents fail by doing the wrong thing, repeating it, or trusting a tool that lied. Reliability work grades the run, contains the tools, and leaves a trace you can replay.
| Mode | What you run | When | What it misses |
|---|---|---|---|
| Offline | Golden trajectories and simulators | CI, before release | Live tool drift and real rate limits |
| Online | Shadow, canary, sampled feedback | Production | Needs careful labeling and a small blast radius |
| Hybrid | An offline gate plus online monitors | The default you should describe | The cost of keeping both honest |
Trajectory grading
Score the path:
- Did it call the right tools in a valid order?
- Were arguments schema-valid and authorized?
- Did it stop under the budget?
- Was the final answer faithful to observations? Grounding metrics live on the RAG hub. Use the idea. Do not re-teach the retrieval suite.
- Were side effects exactly-once?
Decisions
- 1
1. Goals dataset
- next2. Run in sandbox
- 2
2. Run in sandbox
- next3. Capture spans
- 3
3. Capture spans
- next4. Grade tools and args
- 4
4. Grade tools and args
- next5. Grade the answer
- 5
5. Grade the answer
- next6. Pass thresholds?
- ?
6. Pass thresholds?
- no7. Block the release
- yes8. Canary online
- 7
7. Block the release
- 8
8. Canary online
Lesson map
Agent Reliability — Evals, Guardrails, Tracing & Failure Modes
Agents fail differently from chatbots: stuck loops, poison tool output, partial side effects, and trajectory drift. Reliability means grading whole runs, sandboxing tools, circuit breakers, and span-level traces. A fluent final string is not a pass.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Goals dataset"] b["2. Run in sandbox"] c["3. Capture spans"] d["4. Grade tools and args"] a -->|1. Goals dataset to 2. Run in sandbox| b b -->|2. Run in sandbox| c c -->|3. Capture spans to 4. Grade tools and args| d
A red trajectory blocks release. A green suite still gets a canary, because fixtures are not production.
Guardrails
| Control | Purpose | If it is missing |
|---|---|---|
| Tool allowlist | Limit capability | Excessive agency |
| Argument validators | Schema and business rules | Injection through tool arguments |
| Sandbox for files and network | Contain code tools | The host is in scope |
| Output filters | PII and policy leaks | The user sees the secret |
| Circuit breaker | Trip on error rate | A cascading outage |
| Step budget | Stop loops | The bill melts |
Classifiers that route or block at the edge are the System One and hybrid lessons. This page assumes those gates can exist and focuses on the agent loop around them: allowlists, sandboxes, and the eval that proves the gate was consulted.
Failure modes that are specific to agents
| Failure | Symptom | Mitigation |
|---|---|---|
| Stuck loop | The same tool repeats | Detector plus a max identical count |
| Poison tool | An observation tries to override policy | Treat tool text as untrusted data |
| Oscillation | Do A, then undo A | Monotonic plans for writes |
| Budget exhaustion | Stop in the middle | Checkpoint and a degrade message |
| Partial commit | One of N writes landed | Compensating action, or do not fan out writes |
| Cache poison | A bad semantic hit | Version bust. The caching lesson has the bans |
Poison tools. Text that came back from a browser, a ticket, or a file is untrusted content. The same rule you use for untrusted web pages applies: do not execute instructions found inside the observation. Keep that rule on this page. Do not turn the answer into a general security survey.
Tracing
Minimum span model:
agent.runas the root.agent.modelwith a prompt hash, token counts, and whether the prefix cache hit.agent.toolwith name, latency, error, and the idempotency key.agent.guardrailwith allow or deny.agent.hitlwith wait time and the decision.
Correlate them with one run id. Export with OpenTelemetry GenAI conventions where you can. The point of the span is replay: you should see that the refund was proposed, denied, and never executed.
Loop detector
export function isStuck(
recent: Array<{ name: string; argsKey: string }>,
n = 3,
): boolean {
if (recent.length < n) return false;
const slice = recent.slice(-n);
return slice.every((item) => item.name === slice[0].name && item.argsKey === slice[0].argsKey);
}Press Run. Snippets must be self-contained — no network, files, or native modules.
Three identical calls is a stop, not a hint. A successful read that repeats is exactly the case a circuit breaker will not catch, because the calls are not errors.
Circuit breaker
Press Run. Snippets must be self-contained — no network, files, or native modules.
The breaker watches error rate. The scaling lesson decides what the queue does when the breaker is open: wait, shed, or degrade. Cascading failures in the SRE book are the same shape. An agent adds a twist: the calls that succeed can still be the bug.
Rollback
Not every tool has an undo. Design a compensating action (void against a refund request) and store the intent in episodic memory so a later step can see it. If the business step must be atomic across several systems, use a durable workflow with explicit compensations. That is a concept, not a vendor tutorial, and it is not a reason to fan out writes.
Pitfalls
- Shipping on a vibe check of five chats.
- Letting tool text rewrite the system prompt because it was “just context.”
- One metric for prose quality and no metric for “did we email.”
- Opening the email tool to all recipients on day one of the canary.
- Assuming offline green means the live schema still matches the fixture.
Describe a goal, three tool calls, and a final sentence that looks helpful. Make the second call a repeat, and hide a “ignore your policy” sentence inside the first observation. Say which checks fail, and which span would show the repeat.
Interview Q&A
How do you eval an agent that emails users?
Answer
Sandbox the email tool offline. Score whether it would send, whether the template is allowed, and whether the human gate ran when the policy says so. Online, canary with an allowlisted recipient set. Do not grade only the prose of the email.
What is a poison tool observation?
Answer
Untrusted text from a tool, a web page, or a ticket that tries to override the system policy. Isolate it. Never copy it into the instruction channel. A faithful agent quotes the observation. It does not obey it.
Why does an offline pass fail online?
Answer
The live tool schema drifted, rate limits changed the path, a prompt-cache miss changed latency and retries, or users asked for goals outside the fixture set. Sample production trajectories and refresh the suite. The caching and scaling lessons explain those drifts. This page is why you notice them.
What is on a trajectory scorecard?
Answer
Right tools, valid order, authorized arguments, a budget stop, faithfulness to the observations you actually received, and exactly-once side effects. A pleasant final sentence does not cancel a bad row.
How is a stuck loop different from a circuit breaker?
Answer
The detector trips when the same successful call repeats. The breaker trips when calls fail. You want both. A read that returns 200 forever will not open a breaker and will still burn the budget.
What do you do about a partial commit?
Answer
Do not fan out writes in the first place. If one write landed, run the compensating action you named in the design, and record the intent in the episode so a resume does not guess. “We will fix it by hand” is not a control.
Which spans are non-negotiable?
Answer
Root run, model, tool, guardrail, and human gate, sharing a run id. The model span records tokens and cache-hit. The tool span records the idempotency key and the error. Without those, the eval cannot replay the path it claims to grade.
Where do grounding scores and classifier guards live?
Answer
Faithfulness of an answer to evidence is the RAG cluster, starting at RAG and Vector Databases. Route and guard classifiers are System One and the hybrid architecture. This page checks that the agent consulted its tools and its gates, and that the side effects match the trace.
Go Deeper
- OpenAI evals guide and the Evals repository for harness shape. Your assertions are still trajectory assertions.
- OWASP LLM Top 10 for excessive agency and prompt injection via tool output.
- OpenTelemetry GenAI conventions for span names.
- Addressing cascading failures and CircuitBreaker for the open state. Pair it with the step budget so successful loops still stop.