Tool/Function Calling vs Structured Outputs
Tools execute side effects and fetch live data; structured outputs constrain a final JSON contract; agent loops mix both. Pick by side effects, latency, and security — not by habit.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why a pricing tool wins over schema-only “extract a price”
Prefer
Tool call to a pricing service
The model selects get_price(sku) with typed args. Your runtime hits the catalog. The number is true as of that service, not as of the weights.
- Live data and business logic stay in code you test.
- Audit trail: tool_name, args, latency, result.
- The LLM never has to invent a dollar amount that merely type-checks.
Alternative
Structured final JSON with a price field
Fast, no I/O, closed schema. Perfect for extraction from a document that already contains the price. Lethal if the price is supposed to be live.
- Zero extra round-trips — cheapest hop when the source text is the source of truth.
- Schema-valid 9.99 is still a hallucination if the catalog says 42.50.
- No sandbox needed because nothing executes — and nothing can fetch.
Router: decide, maybe execute, then emit a typed object
Prefer one clear contract per hop. The mermaid flowchart is the same loop.
- 1
Bind tools and/or a final schema
Tool registry: name, arg schema, auth, idempotency. Final schema: closed JSON for the user-visible answer. - 2
Model emits an envelope
Either a tool request (name + args) or a final payload. Constrained decoding can enforce that envelope. - 3
Router inspects type
Validate args against the tool schema before any call. Unknown names fail closed. - 4
If tool: execute in the sandbox
Live I/O, optional parallel calls, log request ids. Feed the observation back. Never run on invalid JSON. - 5
If final: schema-validate and return
No side effects. Business validators still run. This is the cheap path. - 6
Cap the loop
Max hops, timeouts, and a terminal schema. Infinite tool ping-pong is a production incident.
Overview
LLMs can produce structured data in three ways that interviews collapse into “just use function calling”:
- Tool / function calling — the model emits a request (
name,arguments). Your runtime executes something: DB lookup, HTTP, code, payment. The model does not return the business result; the tool does. - Structured outputs (
response_format/json_schema/guided_json) — the model is forced to emit a JSON object that matches a schema. Nothing runs. The string is the product. - Hybrid agent loops — the model repeatedly chooses tool vs final. Each hop should have one contract: either “call this function” or “here is the answer object.”
The hub at /studies/structured-outputs taught token masking. Schema strictness taught closed keys. This page is the product decision: do you need a side effect or a typed blob?
Rule of thumb: if the number must be true in the world right now, it is a tool. If the number must be faithful to the prompt text, it is structured extraction. If the model must plan, fetch, then answer, it is a loop — not a single schema with a pretend price_usd.
Why the distinction exists
A JSON Schema can say price_usd is a number. Constrained decoding will happily emit 9.99. That value is schema-valid and economically false if the catalog lists 42.50.
A tool get_price(sku) returns whatever your pricing engine says. The model’s job is to choose the sku (still a hallucination risk!) and not to arithmetic the list price from memory.
| Need | Prefer | Why |
|---|---|---|
| Typed extract from a document | Structured final JSON | Source of truth is the text; no I/O |
| Classification / routing labels | Structured final JSON | Closed enums, one shot |
| Config or UI props from a description | Structured final JSON | Shape is the product |
| Live price, inventory, authz, time | Tool | World state is outside the weights |
| Write, email, charge, mutate | Tool | Side effects must be in your runtime |
| Multi-step “look up then decide” | Agent loop | One contract per hop; cap retries |
Latency. A tool hop is at least one extra model call after the observation, plus the tool’s own RTT. Pure JSON is one completion. Parallel tool calls amortize wall clock when lookups are independent (price and stock for the same sku).
Security. Tools are handles to the outside world. Arbitrary code execution, raw SQL, and unrestricted HTTP are how agents become incidents. A schema with no tools cannot call your billing API — that is a feature. Expose the smallest tool surface; validate args; sandbox.
Observability. Tool calls give you tool_calls_total, latency histograms, and an audit log of args. Pure JSON leaves you with a completion and a parse. Both need schema_id / tool version in the trace.
Contracts in the API
Provider names differ; the split does not.
- OpenAI function calling — tools array, model returns
tool_calls; you execute; you sendrole: toolresults back. Strict function schemas reuse the same closed-JSON rules as structured outputs. - OpenAI structured outputs —
json_schemaon the response (or on tool args). Final answer is typed JSON without a function dispatch. - Anthropic tool use — tools in the request; model emits a tool_use block; you return tool_result. Same runtime story: you run the function.
Do not confuse “the model produced JSON that looks like a function call” in free text with an actual tool protocol. Free-text call get_price(...) is neither masked nor dispatched.
Architecture
Router. Central dispatcher: parse the model output, distinguish tool vs final, enforce auth, log. Do not scatter if name == "get_price" across handlers.
Tool registry. name → callable plus metadata: rate limits, idempotency keys, required scopes, timeout, whether it is mutating. Unknown names fail closed.
Schema store. Versioned JSON schemas for both tool args and final responses. Tool results should be validated too — a flaky HTTP client can return a 200 with the wrong shape.
Parallel executor. Independent lookups (get_price, get_stock) run concurrently; merge by name. Deterministic merge order in the next prompt so the model does not see a race.
Sandbox. Containers, network allowlists, no raw eval. SQL tools get bound parameters and a read-only user unless writes are the point — and writes need idempotency.
Loop budget. max_hops (often small: 3–8), wall-clock timeout, and a terminal “you must now emit final” instruction. Oscillation (same tool, same args) is a stop condition.
Decisions
- 1
User prompt
- nextLLM with tools and/or final schema
- 2
LLM with tools and/or final schema
- nextEnvelope type
- ?
Envelope type
- toolValidate tool args
- finalValidate closed JSON schema
- 4
Validate tool args
- invalidFail closed - do not execute
- okSandbox execute
- 5
Fail closed - do not execute
- 6
Sandbox execute
- nextObservation
- 7
Observation
- nextLLM with tools and/or final schema
- 8
Validate closed JSON schema
- okTyped response
- badRepair or fail - no side effects
- 9
Typed response
- 10
Repair or fail - no side effects
Lesson map
Tool/Function Calling vs Structured Outputs
Tools execute side effects and fetch live data; structured outputs constrain a final JSON contract; agent loops mix both. Pick by side effects, latency, and security — not by habit.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB user["User prompt"] llm["LLM with tools and/or final schema"] env["Envelope type"] args["Validate tool args"] user -->|User prompt to LLM with tools and/or final schema| llm llm -->|LLM with tools and/or final schema| env env -->|tool| args
Decision matrix
| Situation | Choose | Tradeoff |
|---|---|---|
| Static report from supplied text | Structured outputs | May hallucinate facts that fit the schema |
| Price, inventory, user profile | Tool (sync) | Extra RTT; need timeouts |
| Many independent fetches | Parallel tools | Merge logic; partial failure policy |
| Planner + executor | Structured plan object, then tools | Two hops; keep the plan schema small |
| Legacy model, no tools, no SO | Prompt + repair loop | Cost and non-determinism |
| Untrusted tenant prompts | Schema-only or tightly sandboxed tools | Tools expand attack surface |
Hybrid envelope (what the playground uses): the model returns either { "type": "tool", "name", "args" } or { "type": "final", "data" }. One discriminated union, one router. Slightly larger payload; much clearer than “sometimes we parse functions out of prose.”
Prefer one clear contract per hop. A final schema that also contains a nested tool_hint blob is how you get half-executed side effects.
Deep dive · The envelope is a CFG, the tool is not
You can constrain { type: "tool", name, args } vs { type: "final", data } with the same closed-schema rules as the rest of this cluster. That only guarantees the request is well-typed. get_price("SKU-123") can still be a sku the user must not see. Authz, rate limits, and the catalog live in the runtime. Think of the grammar as the socket shape; the sandbox is the process that is allowed to plug in. Schema-only hops need no socket.
Typed parse vs tool args
Tool args: the SDK maps JSON to a function signature. Less boilerplate. You still validate (ranges, tenant id vs auth context). Never trust the model’s account_id if you already know the session user.
Typed parse of a final object: you own the deserializer and business validators. Use it when the product is the object (extraction, classification, form fill).
Typed parse of tool args plus execute: the common agent path. Invalid args → do not call. Invalid result → do not treat as truth; maybe retry the tool, not the model.
Parallel calls
When the model asks for several independent tools, run them concurrently and pack observations in a stable order. Partial failure: decide fail-closed (any miss aborts) vs partial (null the missing field). Payments and deletes are not “fire all in parallel and hope.” Mutating tools are sequential unless you have a real saga.
Provider APIs that emit multiple tool_calls in one completion are the parallel hint. You still own timeouts and cancellation.
Security, briefly
- Least privilege. A retrieval tool should not write. A write tool should not take raw SQL.
- Arg allowlists. Enums and ids you can check against the tenant.
- No tool on invalid JSON. Same rule as repair loops: syntax first, then execute.
- Human in the loop for expensive or irreversible actions.
- Schemas as a sandbox. If the feature does not need I/O, do not register tools “for flexibility.”
In-memory router (run this)
No SDK, no network. The model output is already an envelope: tool or final. The mock catalog knows SKU-123 → 42.5. The schema-only path accepts any number that looks like a price — including a confident 9.99 hallucination.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Run it. The tool path always prints 42.5 from the catalog. The schema-only 9.99 passes type checks and fails the catalog. That is the whole lesson: validity is not truth; tools are how truth gets into the loop when the world is outside the prompt.
Interview Q&A
When do you prefer response_format / structured outputs over a tool call?
Answer
When the data is in the prompt (extraction, classification, static config) and you need a typed object with no side effects. You skip tool RTT, shrink the attack surface, and let constrained decoding guarantee shape. You still validate business rules.
When does a tool win?
Answer
Live data, business engines, or mutations: pricing, inventory, authz, payments, search. The model should select arguments; code should compute the value. A schema-valid price is still a guess.
How do you safely expose a code-execution tool?
Answer
Isolate the process (container), cap CPU/memory/time, allowlist imports and network, validate the exact function name and args before run, audit inputs and outputs. Do not pass a free-form script from the model to eval. Prefer purpose-built tools over a generic shell.
How would you implement parallel tool calls?
Answer
If the completion contains multiple independent tool requests, dispatch them concurrently, wait with a timeout, merge observations in a stable order, then call the model again. Define partial-failure policy. Do not parallelize irreversible mutations without a saga.
What is a hybrid envelope and why use one?
Answer
A discriminated union: type tool with name and args, or type final with data. The router has one parse path. Constrained decoding can enforce the union. Clearer than scraping function-call prose out of a markdown blob.
Typed parse vs tool args — which?
Answer
Tool args when the hop is an action and the runtime should dispatch. Typed parse when the hop is the product (the JSON is stored or rendered). Many systems use both: structured plan object, then tool args, then structured summary. One contract per hop.
Does structured output make tools unnecessary for extraction from PDFs?
Answer
Usually yes for the extract hop: the PDF text is the source of truth, so a closed invoice schema is the right contract. Use a tool if you must join to a live SKU database after extract. Do not make the extractor invent catalog prices.
What do you log?
Answer
schema_id or tool version, envelope type, tool name, arg hashes (careful with PII), latency, success, hop index, finish reason. Pure JSON needs the same request id story or you cannot debug a bad object.
How do you stop infinite agent loops?
Answer
max_hops, wall-clock timeout, detect repeated tool+args, force a final schema on the last hop, fail closed to a human. Unbounded loops are a cost and safety bug, not “the model is thinking.”
Security difference in one sentence?
Answer
Tools are capabilities; schemas are not. Register the minimum tools, validate args, never execute invalid JSON, and use structured finals when I/O is unnecessary.
Pitfalls
Prompt: “What should we charge customer C for SKU-123 today?” Write two boxes. Box A: get_price + get_customer_tier tools, then a final quote object. Box B: one schema { sku, price_usd, tier }. Say which can hallucinate the dollar amount. Circle side effects, extra RTTs, and the sandbox. That is the interview.
Go Deeper
Docs:
Cluster: