Structured Outputs / Constrained Decoding
Strict JSON Schema vs JSON mode; CFG token masking; required fields + additionalProperties:false; Pydantic/Zod; refusals/truncation still break validity.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why strict schema beats “please return JSON”
Prefer
Structured Outputs (json_schema + strict)
The sampler can only emit tokens that stay inside your schema. Parsers stop dying on missing keys and extra fields.
- Required keys and enums are enforced at decode time.
- Refusals can be a first-class branch instead of malformed JSON.
- Retries for format errors become rare; you spend budget on content quality.
Alternative
JSON mode or regex repair
Valid JSON is not a typed contract. Post-hoc salvage is non-deterministic and expensive.
- JSON mode allows any object shape — missing fields, extra keys, wrong enums.
- Regex “JSON repair” fails on truncation and invented wrappers.
- Business validators still needed either way; format retries should not be the first line of defense.
Decode path — shape first, meaning second
Use this vertical flow on a phone. Mermaid below is the same pipeline.
- 1
Author a closed schema
Root object, every property required, additionalProperties: false, nullables for absence. - 2
Compile to a CFG / automaton
Provider or local backend (XGrammar / Outlines / vLLM) builds the token mask. - 3
Mask illegal next tokens
Sampler never leaves the language of your schema. - 4
Branch on refusal / truncation
Do not stuff a partial string into Pydantic/Zod. - 5
Run business validators
CFG cannot check invoice totals, SKU existence, or policy. Fail closed on semantic errors.
Overview
Structured Outputs (also called constrained decoding) force a model to emit text that conforms to a developer-supplied JSON Schema (or grammar / regex). Instead of “please return JSON,” the inference stack masks illegal next tokens so the sampler can only choose tokens that keep the string valid against the schema.
Why it matters in production:
- Deterministic shape for downstream code. Parsers, typed SDKs, and agent tool routers stop dying on missing keys, wrong enums, or trailing commas.
- Fewer retries. JSON mode only guarantees valid JSON; strict schema mode guarantees schema adherence. Retries for format errors become rare.
- Safer agent loops. Tool arguments and extraction payloads become typed contracts; failures shift to content quality, not syntax.
- Programmatic refusals. With OpenAI Structured Outputs, safety refusals are surfaced as a distinct refusal path instead of malformed JSON that crashes your deserializer.
Mental model: JSON mode constrains the language to JSON. Strict Structured Outputs constrains the dialect to your schema. Schema validity is by construction — semantic correctness (truthfulness, business rules, numeric ranges beyond what the grammar encodes) is not guaranteed. You still validate.
OpenAI explicitly positions Structured Outputs as the evolution of JSON mode: both produce parseable JSON; only Structured Outputs guarantees schema adherence. Prefer strict schema whenever the model / provider supports it.
JSON mode vs Structured Outputs
| JSON mode | Structured Outputs (strict) | |
|---|---|---|
| Valid JSON | Yes | Yes |
| Matches your schema | No | Yes (supported schema subset) |
| Required fields / enums | Not enforced | Enforced via token masking |
| Typical enablement | type: "json_object" | type: "json_schema" + strict: true |
| When to use | Legacy models; any JSON blob | Production extraction, tool args, typed APIs |
Provider notes (high level)
- OpenAI:
response_format/ Responses APItext.formatwithjson_schema+strict: true; SDKs map Pydantic / Zod → schema. - Gemini: schema-constrained JSON / response schema APIs (provider-specific schema subset).
- Anthropic: tool-use / structured helpers; treat as a tool schema contract rather than OpenAI-identical CFG masking APIs — validate provider docs for current guarantees.
- Local (Outlines / vLLM / XGrammar): guided decoding backends (
json,regex,choice,grammar) apply the same idea at the sampler.
How constrained decoding works (CFG / token masking)
At each decode step the model scores the full vocabulary. Constrained decoding inserts a mask between logits and sampling:
- Compile the JSON Schema (or EBNF / regex) into a context-free grammar (CFG) / automaton.
- Track the current parse state given tokens emitted so far.
- Compute the set of legal next tokens for that state.
- Zero out (mask) illegal token logits → sample only from the legal set.
- Advance the automaton with the chosen token; repeat.
OpenAI’s launch write-up describes converting JSON Schema → CFG and using a cached structure so mask updates stay cheap per step. Local stacks (XGrammar, Guidance / llguidance, Outlines) do analogous work inside vLLM / SGLang-style engines.
Decisions
- 1
JSON Schema or EBNF
- nextCompile to CFG / automaton
- 2
Compile to CFG / automaton
- nextParse state from tokens so far
- 3
Parse state from tokens so far
- nextZero illegal token logits
- 4
Model logits over full vocab
- nextZero illegal token logits
- 5
Zero illegal token logits
- nextSample from legal set
- 6
Sample from legal set
- nextAdvance automaton
- 7
Advance automaton
- nextComplete and valid?
- ?
Complete and valid?
- noModel logits over full vocab
- yesSchema-valid JSON string
- 9
Schema-valid JSON string
Lesson map
Structured Outputs / Constrained Decoding
Strict JSON Schema or EBNF vs JSON mode; CFG token masking; required fields + additionalProperties:false; Pydantic/Zod; refusals/truncation still break validity.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB schema["JSON Schema or EBNF"] cfg["Compile to CFG / automaton"] state["Parse state from tokens so far"] logits["Model logits over full vocab"] schema -->|JSON Schema or EBNF| cfg cfg -->|Compile to CFG / automaton| state
Important consequence: schema validity is by construction, not “usually.” The model can still invent values that type-check. Ground with RAG / tools, prefer nullable unknowns, and run business validators.
Schema rules that trip people
Strict mode typically requires (OpenAI-compatible rules; check your provider):
- Root must be an object (not a bare array /
anyOfroot in many APIs). - Every property must be listed in
required. “Optional” fields are usually modeled as nullable (type: ["string","null"]or a union) while still required. Absence isnull, not a missing key. additionalProperties: falseon every object. Extra keys are forbidden; forgetting this is a common rejection reason.- Keep schemas flat. Deep nesting, huge property counts, and recursive shapes increase compile cost, latency, and unsupported-feature risk. Prefer shallow objects plus arrays of simple items.
- Supported type subset only.
string/number/boolean/integer/object/array/enum/anyOfare common; manypattern,format,minLength,$refgymnastics are limited or unsupported depending on backend. - Key order follows schema order in several implementations — design schemas intentionally if consumers care about order.
- Parallel tool calls + strict function schemas may need parallel calling disabled depending on provider constraints.
Anti-pattern: stuffing business logic into an ultra-deep schema. Better: flat extract → code validates invariants.
Deep dive · Why every field is required plus additionalProperties false
The grammar needs a closed world of keys and a definite production for each property. Optional-as-absent keys complicate the automaton (the parser cannot know whether a key will appear later). Providers typically require listing every property as required and modeling absence with nulls / unions, while forbidding undeclared keys. That is why additionalProperties: false is not a style nit — it is how the CFG stays finite and closed.
System design: extraction / agent pipeline
Components
- Ingest — raw text, PDF OCR, tickets, emails, tool traces.
- Schema registry — versioned Pydantic / Zod models → JSON Schema artifacts (
schema_id, semver). - Prompt + schema binder — system instructions + few-shots that complement (not replace) the grammar.
- Constrained decoder — provider Structured Outputs or vLLM guided backend (XGrammar / Guidance / Outlines).
- Typed parse — SDK
.parse()/ Zod parse → domain objects. - Business validators — ranges, cross-field rules, referential integrity the CFG cannot express.
- Router / agent — either (a) structured answer objects or (b) structured tool calls; never free-form glue.
- Observability — log
schema_id, finish reason, refusal, token usage, validation failures.
Data flow
Decisions
- 1
Raw input
- nextOptional chunk or retrieve
- 2
Optional chunk or retrieve
- nextLLM call with strict schema S at version N
- 3
LLM call with strict schema S at version N
- nextToken-masked generation
- 4
Token-masked generation
- nextOutcome
- ?
Outcome
- typed objectBusiness validation
- refusalBranch UI plus audit - do not force-parse
- truncated incompleteRaise max_tokens or split schema
- 6
Business validation
- nextPersist and act
- failReprompt, degrade schema, or human review
- 7
Branch UI plus audit - do not force-parse
- 8
Raise max_tokens or split schema
- nextLLM call with strict schema S at version N
- 9
Persist and act
- 10
Reprompt, degrade schema, or human review
- nextLLM call with strict schema S at version N
Failure modes (design for these)
| Mode | Symptom | Mitigation |
|---|---|---|
| Refusal | Safety refusal field / empty structured body | Branch UI + audit; do not force-parse content |
| Truncation | finish_reason=length, incomplete object mid-generation | Raise max_tokens; stream + early abort; split schema into smaller objects |
| Invented required fields | Model fills required keys with plausible junk when unknown | Nullable + explicit unknown enums; instruct “null if absent”; validate against source evidence |
| Schema drift | Clients on v1, server on v2 | Schema registry + compatibility tests; pin schema_id in requests |
| Unsupported schema | API 400 / backend fallback / silent looseness | Lint schemas in CI against provider allowlist; keep flat |
| Semantic hallucination | Valid JSON, wrong facts | Ground with RAG / tools; separate extractive vs generative fields |
Tradeoffs vs tool calling
| Contract | Best when | Side effects |
|---|---|---|
Structured response (json_schema) | You consume a typed answer (extraction, classification, UI props) | None — single shot |
| Strict tool / function calling | The model must choose an action with typed args | Live in your tool runtime |
| Hybrid | Planner emits a structured plan object; executor maps steps to tools | Prefer one clear contract per hop |
Cost / latency: grammar compile + mask overhead is usually small vs retries. Complex / dynamic schemas favor backends with fast TTFT (for example Guidance) or cached schemas (XGrammar).
Teaching code: Pydantic + OpenAI SDK
SDK helpers compile the Pydantic model to a strict JSON Schema (additionalProperties: false, required fields). Always branch on refusal before using parsed. Re-validate business rules (totals, SKU existence) after parse. This fence talks to the OpenAI API — not runnable in the sandbox.
from typing import Literal, Optional
from pydantic import BaseModel, Field
from openai import OpenAI
class LineItem(BaseModel):
sku: str
qty: int = Field(ge=1)
unit_price_usd: float
class InvoiceExtract(BaseModel):
vendor: str
invoice_id: Optional[str] # still required in schema; null if unknown
currency: Literal["USD", "EUR", "GBP"]
line_items: list[LineItem]
confidence: Literal["high", "medium", "low"]
client = OpenAI()
completion = client.beta.chat.completions.parse(
model="gpt-4o-2024-08-06",
messages=[
{
"role": "system",
"content": (
"Extract invoice fields from the user text. "
"Use null for unknown optional strings. "
"Do not invent SKUs."
),
},
{
"role": "user",
"content": "Acme Corp invoice AC-991: 2x WIDGET-1 at $9.50, 1x GADGET-9 at $20.",
},
],
response_format=InvoiceExtract,
)
msg = completion.choices[0].message
if msg.refusal:
raise RuntimeError(f"refused: {msg.refusal}")
invoice: InvoiceExtract = msg.parsed
print(invoice.model_dump_json(indent=2))Notes:
Optional[str]becomes a required key whose value may benull— that is the strict-mode encoding of “unknown.”Field(ge=1)is a business constraint. Some backends will not encodegeinto the grammar; always re-check in code.- Local serving: pass the same schema to vLLM
StructuredOutputsParams/guided_json. Cache repeated schemas for throughput; pick backend for TTFT vs long-generation tradeoffs.
Teaching code: Zod + OpenAI SDK
Same invoice contract. zodResponseFormat names the schema (invoice_extract) so the API can cache the compiled grammar.
import OpenAI from "openai";
import { z } from "zod";
import { zodResponseFormat } from "openai/helpers/zod";
const LineItem = z.object({
sku: z.string(),
qty: z.number().int().positive(),
unit_price_usd: z.number(),
});
const InvoiceExtract = z.object({
vendor: z.string(),
invoice_id: z.string().nullable(),
currency: z.enum(["USD", "EUR", "GBP"]),
line_items: z.array(LineItem),
confidence: z.enum(["high", "medium", "low"]),
});
const client = new OpenAI();
const completion = await client.beta.chat.completions.parse({
model: "gpt-4o-2024-08-06",
messages: [
{
role: "system",
content:
"Extract invoice fields. Use null when a string field is unknown. Do not invent SKUs.",
},
{
role: "user",
content:
"Acme Corp invoice AC-991: 2x WIDGET-1 at $9.50, 1x GADGET-9 at $20.",
},
],
response_format: zodResponseFormat(InvoiceExtract, "invoice_extract"),
});
if (completion.choices[0].message.refusal) {
throw new Error(completion.choices[0].message.refusal);
}
const invoice = completion.choices[0].message.parsed!;
console.log(JSON.stringify(invoice, null, 2));In-memory schema + token mask (run this)
No network, no SDK. First half is the closed-world check production schemas actually need (required + additionalProperties: false + nullable + enum). Second half is a tiny CFG: given a JSON prefix, only a few tokens stay legal — everything else is masked. That is constrained decoding, shrunk to a whiteboard.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Run it and notice: JSON-mode-shaped objects with an extra debug key fail the closed schema; invoice_id: null passes; "maybe" is never a legal next token once the automaton is in the enum state.
Interview Q&A
JSON mode vs Structured Outputs — what is the real difference?
Answer
JSON mode guarantees syntactically valid JSON. Structured Outputs (strict) additionally guarantees adherence to a supplied JSON Schema via constrained decoding. Use strict schemas for production contracts; reserve plain JSON mode for older models or intentionally schemaless blobs.
How does constrained decoding actually work?
Answer
The schema is compiled to a CFG / automaton. After each token, illegal next tokens are masked to probability 0 before sampling. Validity is enforced at generation time, not by post-hoc regex salvage.
Why require all fields plus additionalProperties: false?
Answer
The grammar needs a closed world of keys and a definite production for each property. Optional-as-absent keys complicate the automaton; providers typically require listing every property as required and modeling absence with nulls / unions, while forbidding undeclared keys.
Does constrained decoding stop hallucinations?
Answer
No. It stops format errors. The model can still invent values that type-check. Mitigate with grounding, nullable unknowns, evidence fields, and business validators. Valid JSON with a plausible SKU is still a bug if the SKU was not in the source.
When would you prefer tool calling over a structured final answer?
Answer
When the model must select and parameterize side-effecting actions (search, DB write, payment). Use structured final answers for pure extraction / classification. Many systems combine both: structured plan → tools → structured summary. Prefer one clear contract per hop.
How do you handle refusals and truncation?
Answer
Treat refusal as a first-class branch (do not deserialize). For truncation, monitor finish reasons, increase token budgets, shrink schemas, or emit multi-step partial objects. Never assume json.loads of a truncated string is recoverable.
How would you run this locally?
Answer
Serve with vLLM structured-output backends (XGrammar / Guidance; Outlines historically). Pass guided_json / StructuredOutputsParams with your schema. Cache repeated schemas for throughput; pick backend for TTFT vs long-generation tradeoffs. Lint against that backend’s supported JSON Schema subset — mixing OpenAI-strict rules with a local backend’s subset is a silent footgun.
What is the schema-design smell in interviews?
Answer
Deeply nested, open-ended objects with dozens of optional fields and regex-heavy constraints. Prefer flat schemas, enums, nullables, a versioned registry, and push complex invariants to code.
How do you version a schema without breaking clients?
Answer
Pin a schema_id / version on every request and persist it with the output. Additive optional-as-nullable fields can be compatible; renaming, tightening enums, or flipping additionalProperties is not. Run compatibility tests in CI. Dual-write old and new for a window, then drop. Treat the schema like an API contract, because it is one.
Streaming a structured object — what is the failure mode?
Answer
Tokens arrive before the automaton has closed the object. UI that JSON.parses a prefix will throw; repair regexes will lie. Stream into the typed parser only when the grammar says the value is complete, or emit field-at-a-time with a schema that allows incremental objects. Truncation mid-array is the common production bug — budget max_tokens for the largest legal payload. See streaming structured output.
When is an enum better than a free-text string?
Answer
Whenever downstream is a switch (status, intent, category). Enums shrink the legal token set, make constrained decoding cheaper, and stop “almost the right label” drift. Free text is for quotes and explanations. If you need both, two fields: label: enum plus rationale: string.
If the schema is already strict, why keep a repair loop?
Answer
Constraints fix syntax. Totals that do not add up, impossible dates, and missing SKUs are semantic. Cap retries, never repair a refusal, never call a mutating tool on invalid JSON. Validation and repair loops.
Can you JSON.parse every streamed chunk?
Answer
Usually no — wait for a complete object or use an incremental parser, and only commit side effects after the final schema check. Token masks still apply per token.
Pitfalls
Explain CFG token masking in 60 seconds. Contrast JSON mode vs strict schema on one table. Then sketch the extraction pipeline with three outbound arrows from the decoder: typed object, refusal, truncated incomplete. Name the repair policy on each failure path. That sequence answers most interview probes on this topic.
Go Deeper
Docs and posts:
- OpenAI — Structured model outputs
- OpenAI — Introducing Structured Outputs in the API
- OpenAI Cookbook — Introduction to Structured Outputs
- vLLM — Structured Outputs
- XGrammar (MLC)
- Outlines
- Pydantic
- Zod
YouTube: