Validation & Repair Loops for LLM Outputs
Prefer constrained decoding for syntax. Use a bounded validate→feedback→regenerate loop for business rules or legacy models. Never repair refusals; never execute tools on invalid JSON.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why mask syntax and loop on meaning
Prefer
Constrained decode + bounded semantic repair
The sampler guarantees a parseable object. A 2–3 attempt loop only runs when business validators fail, or when the model has no schema mask at all.
- Format retries drop near zero when a CFG is available.
- Pydantic/Zod errors become a small, structured feedback string.
- Hop budget and oscillation checks keep cost bounded.
Alternative
Regex JSON repair or unbounded regenerate
Looks scrappy and works on a demo. In production it oscillates, mutates meaning, and still cannot add two line items.
- No provider support required — last resort for legacy models.
- Regex fixers fail on truncation and invented wrappers.
- Infinite loops burn tokens; repairing a refusal is a safety bug.
Validate → feedback → regenerate (capped)
Syntax errors and semantic errors share the retry budget. Refusals do not enter this loop.
- 1
Prompt + schema (or grammar)
Prefer strict structured outputs / GBNF so the first sample is already JSON in your dialect. - 2
Generate
One completion. Record attempt index, token usage, finish reason. - 3
Parse + validate
json.loads / typed parse first. Then business rules: totals, ranges, referential integrity. - 4
OK → commit
Only now: return to the user, write the DB, or run tools. Side effects wait for a passing object. - 5
Refusal → return
Do not ask the model to rephrase the disallowed content into schema-valid JSON. - 6
Error and attempts left → feedback
Append a short error list. Ask for corrected JSON only. Same error twice: stop. - 7
No attempts left → fail closed
Human, smaller schema, a tool, or a typed error. Do not silently default fields.
Overview
Constrained decoding makes shape cheap. It does not make invoices add up. Production systems therefore layer two checks:
- Syntactic / schema — can we parse this into the typed object? Prefer a token mask so this almost always passes. If the model has no mask (legacy, some local builds), the repair loop must also catch missing keys and extra fields.
- Semantic / business — do line items sum to
total_cents? Does the SKU exist? Is the date real? No grammar in this cluster encodes that. You validate in code.
The loop is: generate → validate → if fail and budget remains, show the errors, generate again. Bound it. Unbounded repair is a cost incident and a way to oscillate forever on the same total_cents mismatch.
Libraries like Instructor wrap this pattern (Pydantic errors back into the next message). You still own max_attempts, refusal short-circuit, and “do not call checkout on garbage JSON.”
Syntax vs semantics vs refusal
| Class | Example | First response |
|---|---|---|
| Format | Trailing comma, missing vendor, extra debug | Prefer not to see this: enable strict SO / CFG. Else repair with schema errors. |
| Semantic | Lines sum to 80, total_cents is 100 | Repair with the rule in the feedback, or call a tool that computes the total. |
| Refusal | Safety block, empty structured body | Return the refusal. Do not “fix” it into a compliant invoice. |
| Truncation | finish_reason=length | Raise max_tokens, split the schema, do not regex-close braces. |
Treat truncation as incomplete, not as a validation error you can comment-fix. A half-closed array is not “almost valid” — see streaming.
Why regex JSON repair loses
Post-hoc salvage (please close the braces, brace counters, “extract the first {…}”) fails on:
- Truncation in the middle of a string (
"vendor": "Ac) - Invented wrappers (
```jsonfences,Sure! Here is...) - Extra
}the model added after you “fixed” a missing one - Semantic nonsense that was already valid JSON
A bounded regenerate with validator text at least asks the model for a new sample under the same schema. A CFG prevents the bad token. Repairing bytes you already sampled can change meaning (you might drop a line item to make JSON parse). Prefer regenerate over mutate.
Instructor-style loops regenerate. They are not brace rewriters.
Deep dive · Two error classes, one budget — why you still split them in logs
A format error after you enabled strict Structured Outputs is a pipeline bug (unsupported schema, not actually strict, truncation). A semantic error is a model or product bug (the CFG cannot add). If you lump them as validation_failed, you will raise max_attempts to fix a compile problem. Log error_class=format|semantic|refusal|truncation separately. Repair success rate is only meaningful on semantic fails. Format fails should go to zero when the mask is on; if they do not, fix the schema lint, not the prompt.
Architecture
Decisions
- 1
Prompt plus schema
- nextGenerate
- 2
Generate
- nextParse and validate
- ?
Parse and validate
- okCommit / return
- refusalReturn refusal
- format errorAttempts left?
- semantic errorAttempts left?
- 4
Commit / return
- 5
Return refusal
- ?
Attempts left?
- yes, new errorAppend short errors; regenerate
- same error twiceFail closed
- noFail closed
- 7
Append short errors; regenerate
- nextGenerate
- 8
Fail closed
Lesson map
Validation & Repair Loops for LLM Outputs
Prefer constrained decoding for syntax. Use a bounded validate→feedback→regenerate loop for business rules or legacy models. Never repair refusals; never execute tools on invalid JSON.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB prompt["Prompt plus schema"] gen["Generate"] kind["Parse and validate"] commit["Commit / return"] prompt -->|Prompt plus schema| gen gen -->|Generate to Parse and validate| kind kind -->|ok| commit
Hybrid. Strict SO (or GBNF) for shape; repair only for business rules. Two layers to monitor: format_fail_rate should be tiny; semantic_fail_rate is the interesting curve.
Feedback shape. Send the validator’s field paths and messages, not a stack trace, not the entire previous JSON unless it is small. “total_cents 100 != sum(line_items) 80. Return corrected JSON only.”
Idempotent side effects. The loop may run three times. Tools and DB writes run once, after success. A charge on attempt 1 plus a “fixed” object on attempt 2 is a double-bill.
Metrics. format_fail_rate, semantic_fail_rate, repair_success_rate, tokens_per_success, attempts_histogram, oscillation count. If repair never succeeds, you have a prompt/schema/policy problem, not a “needs max_attempts=8” problem.
How many retries?
2–3 is the usual cap for a user-facing hop. Measure pass@1 / pass@3 on a frozen eval set. Each extra attempt is latency and money. After the cap: smaller schema, a tool that computes the invariant, or a human.
Oscillation. If attempt 2 repeats the same total_cents error as attempt 1, a third identical prompt will not magically learn arithmetic. Stop, or switch strategy (tool computes the total; ask for lines only).
Temperature. Repair is not a reason to crank temperature. You want a corrected object, not a more creative one.
Pydantic / Zod / Instructor
Pydantic ValidationError and Zod safeParse errors are the right feedback if you truncate them. .error.format() can be large; send the first N issues.
# Illustrative — not the sandbox. Instructor and SDK parse helpers
# wrap this; you still pass max_attempts and skip refusals.
def repair_loop(generate, base_prompt, max_attempts=3):
prompt = base_prompt
last = None
for attempt in range(1, max_attempts + 1):
raw = generate(prompt)
if is_refusal(raw):
return Refusal(raw) # do not continue
parsed, err = try_parse_and_business(raw)
if parsed:
return parsed
last = err
prompt = (
base_prompt
+ "\nPrevious JSON failed validation.\n"
+ err
+ "\nReturn corrected JSON only."
)
raise RuntimeError("validation failed after %s: %s" % (max_attempts, last))Zod’s .strict() aligns with closed schemas. A repair loop that uses a loose parser will “succeed” by dropping extra keys — you will never teach the model to stop emitting them.
In-memory repair loop (run this)
Fake generate() is a tape: (1) not JSON, (2) JSON that fails the totals rule, (3) correct. repair_loop caps at 3 and prints every attempt. No network.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Run it. Attempt 1 dies in JSON.parse. Attempt 2 parses and fails the semantic totals rule. Attempt 3 passes. If you drop the cap to 2, the function fails closed with the totals error — that is the production behavior you want, not a fourth silent retry.
Interview Q&A
Repair loop vs constrained decoding — which first?
Answer
Constrain syntax at decode time whenever the provider allows it. Use a bounded repair loop for business rules, or for models that cannot mask tokens. Do not choose regex salvage as the default if a CFG is on the table.
How many retries?
Answer
Typically 2–3. Measure pass@k and tokens per success on an eval set. If pass@3 is still poor, change the schema, prompt, or compute the invariant with a tool — do not set max_attempts to 12.
Why not regex-fix JSON?
Answer
Truncation, fences, and extra braces defeat brace counters. Mutating a sampled string can drop fields and change meaning. Prefer parse + regenerate under the same schema. Regex extractors are a last resort for logs, not a product path.
When do side effects run?
Answer
After a fully valid object — schema and business rules. Never on invalid JSON, never mid-repair, never on a refusal. Agent tools follow the same rule as DB writes.
What is oscillation and how do you stop it?
Answer
The same validation error on consecutive attempts. Stop; switch strategy (tool computes the total, split schema, human). A third copy of the error message is not new information.
Do you repair refusals?
Answer
No. A refusal is a first-class outcome. Asking the model to emit schema-valid JSON that bypasses the refusal is a safety bug. Return the refusal to the product and audit it.
What do you put in the repair prompt?
Answer
Short field paths and messages from Pydantic/Zod plus the rule that failed. Ask for corrected JSON only. Do not paste a novel of stack frames. Do not change the original task mid-loop.
Which metrics?
Answer
format_fail_rate, semantic_fail_rate, repair_success_rate, attempts histogram, tokens_per_success, oscillation count, time to valid object. Format fails after enabling strict SO should collapse; if they do not, your schema is unsupported or you are not actually in strict mode.
Can repair fix hallucinated facts?
Answer
Only if the validator can see the lie (totals, missing SKU in a table you check). A plausible vendor name that is not in the source will pass schema and pass a totals check. Ground with evidence fields or tools; do not expect retries to invent honesty.
How does Instructor fit?
Answer
It automates validate-and-retry with Pydantic errors in the next message. You still set the attempt cap, handle refusals yourself, keep side effects out of the loop, and prefer a provider schema mask so Instructor is not your JSON fixer.
Pitfalls
On a board, draw generate → validate with three arrows: commit, refusal (no retry), feedback if attempts remain. Write max_attempts = 3 on the feedback arrow. Add a sticky: “same error ×2 → stop.” Then say which of the playground’s first two failures is format vs semantic. That is the interview.
Go Deeper
Docs:
Cluster: