JSON Schema Strictness for Structured Outputs
Strict structured outputs need a closed schema: root object, every property required, additionalProperties false, provider-supported subset, schema_version, null for absence.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why a closed schema wins over “JSON Schema, but optional everywhere”
Prefer
Closed strict schema
Root object, every key required, extra keys forbidden, unknown values as null. The automaton always knows what must appear next.
- Missing keys and invented helper fields fail at decode time, not in your parser.
- Absence is explicit (null + confidence), so the model cannot skip a field and hope you default it.
- schema_version travels with the payload so downstream code can migrate.
Alternative
Loose JSON Schema or JSON mode
Valid JSON with optional keys and open objects looks friendly and reintroduces the format bugs you thought you killed.
- Optional properties tempt omission; parsers then guess defaults.
- additionalProperties defaults true — models emit debug, notes, tax_hint.
- You pay in retries, regex repair, and silent extra-key drops.
Author → compile → parse — the strict path
Same pipeline as the mermaid sequence below. Fail closed on extra keys, missing keys, and refusals.
- 1
Author a closed schema
Root type object. List every property in required. additionalProperties false on every object. Nullables for absence. Enum for closed strings. - 2
Lint against the provider subset
Drop unsupported pattern, format, deep oneOf, recursive $ref. Keep shapes flat. Pin schema_version. - 3
Compile to a CFG / token mask
SDK or vLLM backend builds the automaton. Cache repeated schemas — compile is the expensive step. - 4
Sample only legal tokens
The model cannot emit an extra key or skip a required one. Shape is by construction. - 5
Typed parse, then business rules
Pydantic extra=forbid / Zod .strict() should match the grammar. Totals, SKUs, and policy still run in code. - 6
Branch on refusal / truncation
Do not stuff a partial string into the typed parser. Raise max_tokens or split the schema.
Overview
JSON Schema is a large language. Structured Outputs only compile a subset, and that subset is stricter than the Schema you would write for a human-facing API.
The job of this page is not “how to write JSON Schema.” It is why production extractors die unless the schema is closed:
- Root is an object. Many APIs reject a bare array or
anyOfat the root. Wrap lists:{ "items": [ ... ] }. - Every property is required. “Optional” in the English sense is a required key whose value may be
null(or a discriminated union), not a key the model may omit. additionalProperties: falseon every object, including nested ones. Without it, the grammar is an open world and the model invents helper fields.- Provider-supported keywords only.
string/number/integer/boolean/object/array/enum/anyOfare the usual core.pattern,format,$refrecursion, and cleveroneOfnests are the first things to 400. schema_versionlives in the object. Downstream code migrates on a field, not on “whatever the prompt used to say.”- Absence is
null, not a missing key. Instruct “null if unknown.” Invented fillers for required strings are a semantic bug the CFG cannot see.
Mental model from the hub: JSON mode constrains the language to JSON. A strict schema constrains the dialect to your keys and types. This page is the dialect.
If you only remember one sentence for an interview: the automaton needs a closed world of keys and a definite production for each one.
Why optional keys break the mask
Constrained decoding tracks parse state from tokens so far. After {, the grammar must know the set of legal keys. After "vendor":, it must know whether a string or null comes next. After that value, it must know whether a comma-plus-next-key or } is legal.
If a key is optional-as-absent, the parser cannot decide “will invoice_id appear later?” without looking into the future. Implementations therefore require every property and encode “we do not know” as a present key with a null (or union) value.
That is why additionalProperties: false is not a style nit. Extra keys are an unbounded set of identifiers. A CFG that allows any identifier after a comma is no longer your schema — it is JSON mode with extra steps.
Deep dive · Closed world, First sets, and why extras explode the automaton
Think of each object as a record type, not a dictionary. The compiler builds FIRST sets for “which key tokens can start the next property.” A finite list of required keys plus a fixed order (several providers emit keys in schema order) keeps those sets small and cacheable. An open object means “any string key,” which is huge, and optional keys mean “this production might be epsilon,” which makes FIRST/FOLLOW messy. Providers pick the boring closed record because mask updates have to stay cheap per token. If you need a bag of unknown keys, you do not want strict Structured Outputs for that hop — extract a closed envelope and parse the bag in code, or use a tool.
The rules, with failure modes
| Rule | If you skip it | What to do instead |
|---|---|---|
Root type: "object" | API 400, or a backend that silently loosens | Wrap arrays: { "items": [...] } |
Every property in required | Model omits keys; your parser defaults or crashes | Required + null; or a tagged union found / missing |
additionalProperties: false on every object | Extra debug, notes, tax_hint; some parsers drop them, others throw | Forbid extras in the schema and in Pydantic/Zod |
| Supported subset only | Compile error, or worse: fallback to unconstrained JSON | Lint schemas in CI against the provider allowlist |
| Flat-ish shapes | Compile latency, truncation, unsupported nesting | Split tasks; arrays of simple items; tools for the long tail |
schema_version field | Silent drift between prompt v1 and parser v2 | Literal enum "v1" inside the object; pin in logs |
max_tokens sized for the payload | Truncation mid-array; invalid prefix | Budget tokens; paginate large lists |
Nested objects inherit the rules. A strict root with a loose line_items[] object is a hole. Every object in the tree gets required + additionalProperties: false.
Enums are closed strings. If the source can say something you did not list, either add "unknown" or use a free string plus a separate normalized enum you fill in code. Do not pretend a short enum is a classifier for unbounded language.
JSON mode vs strict vs “normal” JSON Schema
People mix three layers:
- JSON mode (
type: "json_object") — syntactically valid JSON. Any keys. Any nesting. Missing fields allowed. - JSON Schema as documentation — optional properties,
additionalPropertiesdefault true,format: email,$refgraphs. Great for OpenAPI humans. Hostile to a token mask. - Strict Structured Outputs (
json_schema+strict: true, or vLLMguided_jsonwith a closed schema) — the subset this page describes.
Pydantic model_config = ConfigDict(extra="forbid") and Zod .strict() / .passthrough() choices must match layer 3. If the decoder forbids extras and your parser silently strips them, you will not notice schema drift. If the decoder allows extras and your parser forbids them, you will fail closed after paying for the tokens — better than silent drop, worse than constraining at decode time.
Sequence
- 1
App → SDK
closed json_schema plus strict
- 2
SDK
Compile schema to CFG / token mask
- 3
SDK → Decoder
generate with mask
- 4
Decoder → SDK
JSON string in the dialect
- 5
SDK → App
branch UI - do not force parse
- 6
SDK → Parser
Pydantic or Zod forbid extras
- 7
Parser → App
InvoiceExtract
Lesson map
JSON Schema Strictness for Structured Outputs
Strict structured outputs need a closed schema: root object, every property required, additionalProperties false, provider-supported subset, schema_version, null for absence.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB app["App"] sdk["SDK"] decoder["Decoder"] parser["Parser"] app -->|closed| sdk sdk -->|generate with| decoder decoder -->|JSON string in| sdk sdk -->|branch UI - do| app sdk -->|Pydantic or Zod| parser parser -->|InvoiceExtract| app
Modeling absence without optional keys
English “optional” is three different designs. Pick one per field:
| Intent | Schema | Why |
|---|---|---|
| Value or unknown | type: ["string", "null"], key required | Closed production; null is a token the mask allows |
| Present vs missing as a status | { "status": "found" | "missing", "value": T | null } | Makes “we looked” explicit; good for audit |
| Field does not exist in this version | New schema_version, additive change | Do not toggle optionality to version a contract |
Anti-pattern: required vendor: string when OCR often has no vendor. The model will invent "Acme" because the grammar demands a string. That is a semantic hallucination with a valid shape. Prefer null + confidence: "low".
Provider subsets and compile cost
OpenAI-style strict mode, Gemini response schemas, and local XGrammar / Outlines / vLLM guided_json all compile some JSON Schema, not all of it. Differences that bite:
- Root shape — object-only is common.
- Key order — some implementations emit properties in schema order. Do not depend on sorted keys unless you designed for that.
- Numeric constraints —
minimum/maximum/gemay be checked in the SDK after parse, not in the grammar. Re-validate in code. format,pattern,minLength— often unsupported or only partially enforced. Treat them as documentation unless your backend says otherwise.$ref/ recursion — recursive types and huge$defsgraphs are compile-time landmines.- Parallel tool calls + strict function schemas — some stacks want parallel calling off.
Ops concern: first request with a new schema pays compile + cache fill. Repeated schemas should hit a cache (OpenAI and XGrammar both optimize this). A 200-property nested monster will hurt TTFT even if it eventually compiles.
If the schema is too large: split the task (vendor header first, line items second), use a tool for the long tail, or paginate arrays. Do not “just raise max_tokens” as the only lever — truncation still leaves an invalid prefix.
Schema versioning
Put a literal in the object:
{ "schema_version": "v1", "vendor": null, "total_cents": 100 }Log schema_id / semver with the request. Compatibility tests: old parser vs new payload, new parser vs old payload. Additive fields need a version bump and a required+nullable story, not a surprise optional key.
Downstream should switch on schema_version, not on “the prompt in git at the time.”
Teaching schema (Pydantic + JSON Schema)
Illustrative — this is the contract the playground validates. Optional[str] is a required key that may be null. extra="forbid" matches additionalProperties: false. Field(ge=0) is a business constraint: re-check it after parse.
from typing import Literal, Optional
from pydantic import BaseModel, ConfigDict, Field
class LineItem(BaseModel):
model_config = ConfigDict(extra="forbid")
description: str
quantity: float
unit_cents: int
currency: Literal["USD", "EUR", "GBP"]
class InvoiceExtract(BaseModel):
model_config = ConfigDict(extra="forbid")
schema_version: Literal["v1"] = "v1"
vendor: Optional[str]
total_cents: Optional[int]
line_items: list[LineItem]
confidence: float = Field(ge=0, le=1)Zod mirror: .strict() is the extra-key wall; .nullable() is absence; .literal("v1") is the pin.
import { z } from "zod";
const LineItem = z.object({
description: z.string(),
quantity: z.number(),
unit_cents: z.number().int(),
currency: z.enum(["USD", "EUR", "GBP"]),
}).strict();
const InvoiceExtract = z.object({
schema_version: z.literal("v1"),
vendor: z.string().nullable(),
total_cents: z.number().int().nullable(),
line_items: z.array(LineItem),
confidence: z.number().min(0).max(1),
}).strict();SDK helpers that compile these models to json_schema + strict: true still need a refusal branch before you read .parsed. Syntax is solved; meaning is not. Cross-field rules (line totals vs total_cents) belong in validation loops, not in a deeper schema.
In-memory strict validator (run this)
No network, no SDK. A closed invoice schema: extra keys fail, missing keys fail, vendor: null passes, a hallucinated tax_hint fails. Nested line_items objects are closed too.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Run it. JSON-mode-shaped objects with debug or tax_hint fail the closed schema. Omitting vendor fails. vendor: null passes. A nested invented sku_guess fails even when the root looks clean.
That last case is the interview follow-up: strictness is recursive.
Interview Q&A
Why additionalProperties: false?
Answer
It closes the world of keys so the CFG cannot emit undeclared fields. Without it, models invent helper keys (debug, notes, tax_hint). Some parsers ignore them; others throw; all of them mean your typed contract was a suggestion. Forbid extras in the schema and in Pydantic extra=forbid / Zod .strict().
Why require every property instead of optional keys?
Answer
Optional-as-absent keys make the automaton guess whether a key will appear later. Providers therefore list every property as required and model absence with null or a union. Omission is a common production failure; null is explicit and still schema-valid.
Is JSON mode enough for a typed extractor?
Answer
No. JSON mode guarantees syntactically valid JSON, not your keys, types, or enums. Valid JSON can miss vendor, add extras, or use the wrong currency string. Use json_schema + strict (or local guided_json with a closed schema) for production contracts.
Can the root be an array?
Answer
Often no in strict mode. Wrap the list in an object: a required items array. Same story for a bare anyOf root — many APIs want a single object type at the top.
The source has no vendor. What should the model emit?
Answer
A required vendor key with value null, and usually a low confidence. Do not hallucinate a plausible company name just to satisfy type string. Invented fillers are valid JSON and wrong facts.
What if the schema is too large?
Answer
Split the extraction (header vs lines), call a tool for the long tail, or paginate arrays. Watch compile latency, cache misses, and max_tokens. Deep recursive $ref and huge enums are the first unsupported-feature failures.
How do Pydantic and Zod relate to the decoder?
Answer
They compile to JSON Schema and parse the string afterwards. extra=forbid / .strict() should match additionalProperties false. Field ge=1 and Zod .positive() are often business checks — re-run them even if the grammar allowed the number. Always branch on refusal before reading parsed.
Why put schema_version inside the object?
Answer
So the payload is self-describing when it hits a queue, a log, or a second service. Pinning only in the request metadata is easy to drop. Silent drift is clients on v1 and a parser that assumed v2 required keys.
Does a strict schema stop hallucinations?
Answer
No. It stops format errors. The model can still invent a SKU that type-checks. Ground with tools or RAG, prefer null for unknowns, and run business validators. Shape is by construction; truth is not.
What is the schema-design smell?
Answer
Deeply nested open objects, dozens of optional fields, regex-heavy format constraints, and no version pin. Prefer flat records, enums, nullables, a registry, and push cross-field invariants to code. See /studies/structured-outputs for the decode-time mask; this page is the schema that mask compiles.
Pitfalls
Draw { schema_version, vendor, total_cents, line_items, confidence }. Mark every key required. Write additionalProperties: false on the root and on a line item. Change vendor to string | null and say out loud why you did not omit the key. Then name two unsupported keywords you would strip before sending the schema to a provider. That is this lesson in one board.
Go Deeper
Docs and specs:
- OpenAI — Structured model outputs
- Understanding JSON Schema
- Pydantic — extra=forbid / model config
- Zod — .strict()
- vLLM — Structured Outputs
YouTube:
Cluster: