Streaming Structured Output
Streaming tokens of JSON improves UX but partial JSON is invalid. Incremental parsers paint completed fields; commit side effects only after final schema validation. Constrained decoding still applies per token.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why incremental field UX wins over parse-every-chunk
Prefer
Buffer tokens; emit values when a field closes
An incremental JSON parser (or a careful scanner) tells you vendor is done while line_items is still open. The UI paints; the database waits.
- Users see progress on large extracts without lying about completeness.
- JSON.parse is reserved for the finished buffer — or the parser’s end event.
- CFG masks still apply per token, so the prefix stays in the language.
Alternative
JSON.parse each SSE event, or write on first key
Throws on every prefix, or worse: persists a partial invoice and charges a half-array.
- Simplest code for tiny payloads if you buffer-all then parse once.
- Naive per-chunk parse is all exceptions until the last byte.
- Side effects mid-stream are the production incident this lesson exists to prevent.
Stream tokens, paint fields, commit once
Same sequence as the mermaid diagram. Side effects live in the last box only.
- 1
Start completion with stream plus strict schema
json_schema / guided_json still compiled. stream=true does not turn the mask off. - 2
Each token updates the prefix
Append to a buffer. Incremental parser advances. Partial JSON remains un-parseable as a document. - 3
Paint completed scalar fields
When a top-level string or number closes, the UI may show it. Arrays stay skeleton until each element closes. - 4
Do not run tools or writes
A completed vendor string is not an invoice. Mutating APIs wait. - 5
On done: full schema + business validate
finish_reason stop (or equivalent). Closed-schema parse, then totals/SKUs. Then persist. - 6
On length / disconnect: incomplete
Non-empty parse stack, JSON.parse still throws. Raise max_tokens, split schema, or repair as a new request — do not brace-close.
Overview
Streaming tokens of a JSON object is good UX. It is not a sequence of valid JSON documents. "{\n \"ven" is a prefix of a sentence in your JSON CFG. JSON.parse is a parser for complete documents. Those two facts fight if you wire SSE straight into JSON.parse.
Three layers:
- Token stream — what the model emits (and what the CFG mask allows) one piece at a time.
- Value stream — events from an incremental parser: “key
vendorstarted,” “string value closed:Acme,” “array element 0 closed.” - Commit — one typed object after the stream ends successfully, run through the strict schema and business validators.
Product UIs want layer 2 (vendor appears before line items finish). Billing, tools, and database rows want layer 3. Mixing them is how you store a truncated line_items array and call it an invoice.
Streaming does not disable constrained decoding. The mask still runs every token. The final string, if generation finishes in-grammar, still conforms. Truncation (finish_reason=length) is the usual way you end outside a complete object — treat it as incomplete, not as “almost JSON.”
Token stream vs value stream
A tokenizer might emit ", Ac, me, ". The value Acme exists only after the closing quote (and escapes). Painting Ac in a vendor field is a flicker and a lie if the next tokens are me Corp.
Rules of thumb:
- Strings: show when the closing unescaped quote lands, or when the incremental parser emits
valuefor that key. - Numbers: show when the parser knows the number ended (next token is
,or}), not after the first digit. - Enums: do not highlight
USas currencyUSDuntil the enum token is complete — a constrained decoder will only allow legal enum tokens, but the UI should still wait for the value event if you stream sub-tokens. - Arrays: show
item[i]when that element object closes; keep a skeleton fori+1. - Nested objects: same as arrays — completed children only.
Deep dive · Why JSON.parse cannot be your incremental parser
JSON.parse requires a complete value. A prefix is not a JSON value except in the accidental cases where an object closed early (it did not — you are still generating). Incremental parsers (event-based JSON, SAX-style, or a hand-rolled scanner for a known shallow schema) emit events on transitions: key end, string end, number end, object end. They keep a stack, which is the same stack the CFG is keeping. You can implement a tiny scanner for a known flat schema (the playground’s vendor-quote detector) without a full SAX library. You cannot implement “parse whenever” with JSON.parse unless you buffer until the stream ends.
Wire formats
| Approach | Use when | Notes |
|---|---|---|
| Buffer-all, parse once | Small objects, simplest correctness | Still stream to the client as a spinner if you want; parse on done |
| Incremental field UX | Large objects, form-fill, “vendor first” | Parser events → UI; commit later |
| NDJSON / JSONL | Many independent records | Each line is a complete JSON value; JSON.parse per line is valid |
SSE (text/event-stream) | Browser-friendly token push | Events are chunks, not objects. Apply backpressure; do not block the event loop on parse |
| Web Streams / fetch body | Same idea in Fetch | Read UTF-8 incrementally; watch surrogate pairs at chunk boundaries |
NDJSON is the underrated alternative: if you can shape the product as a list of records, stream one complete object per line. Then JSON.parse per line is correct, and you still validate each record. Do not force a 500-line invoice into one object if the UI wants rows.
Protobuf / binary rarely helps LLM hops — models emit text. Constrain text, then encode binary at your service boundary if you must.
Sequence
- 1
UI → Edge
POST extract stream
- 2
Edge → Model
stream plus strict schema
- 3
Model → Parser
next token
- 4
Parser → UI
field vendor completed
- 5
Parser
Arrays stay buffered until each element closes
- 6
Model → Edge
done
- 7
Edge → Parser
final schema plus business validate
- 8
Parser → UI
InvoiceExtract or incomplete error
Lesson map
Streaming Structured Output
Streaming tokens of JSON improves UX but partial JSON is invalid. Incremental parsers paint completed fields; commit side effects only after final schema validation. Constrained decoding still applies per token.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB ui["UI"] edge["Edge"] model["Model"] parser["Parser"] ui -->|POST extract| edge edge -->|stream plus| model model -->|next token| parser parser -->|field vendor| ui model -->|done| edge edge -->|final schema| parser
Truncation and disconnects
finish_reason=length (or a dropped SSE connection) leaves:
- A string that
JSON.parserejects - A CFG stack that is not empty
- A UI that may already show a vendor
Product policy:
- Mark the extract incomplete; do not persist as success.
- Raise
max_tokens, split the schema (header vs lines), or paginate arrays. - Optional: a new request that continues with a smaller schema (
remaining_line_items). Do not splice bytes onto the truncated prefix with a regex closer. - If you streamed to the UI, show a retry/error state on the skeleton fields — do not leave
Acmelooking like a committed invoice.
Provider differences: not every stack streams schema-constrained tokens identically (some buffer internally until a JSON value boundary). Read the current docs; still treat your commit point as “full object validated.”
Side effects
Never:
INSERTthe invoice on first completed field- Call a payment tool because
total_centsappeared (the array might still grow) - Fire emails from a partial
customer.emailstring (ada@is not an address yet)
Do:
- Optimistic UI on completed fields
- Validate the final buffer with the closed schema + business rules
- Then one commit (DB, queue, tools)
Idempotency: the user may refresh mid-stream. The commit key is the finished object (or a request id), not “we painted vendor.”
Backpressure and UTF-8
SSE and fetch streams can outrun a slow tab. If you apply a parser on the main thread that does O(n) over the whole buffer on every chunk, a large extract will hitch. Prefer incremental parsers that eat new bytes only, and pause the reader when the UI is behind (backpressure).
Chunk boundaries can split a multibyte UTF-8 character or a surrogate pair. Decode with a streaming TextDecoder (stream: true), not JSON.parse on random byte slices.
In-memory chunks (run this)
Build a closed object, slice it into ugly chunks, and try JSON.parse after each append. It fails until the last chunk. A small scanner reports when the top-level vendor string has closed — that is the paint signal, not a commit.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Run it. Early chunks throw. vendor may flip from (open) to Acme before JSON.parse succeeds — that is the UI paint. The last line is the commit sermon: parse OK is necessary, not sufficient; still run the closed schema and totals check.
In the snippet, \\\\ in the scanner is how we write a real backslash inside this playground string. The demo payload has no escapes; the branch is there so a split \\ across chunks does not pretend the string closed.
Interview Q&A
Can you JSON.parse every SSE event?
Answer
Usually no. Each event is a chunk of a single JSON value. Parse when the document is complete, or use an incremental parser. JSON.parse per event only works for NDJSON (one object per line) or for a buffer-all-then-parse design.
Does streaming turn off constrained decoding?
Answer
No. Masks apply per token. A finished in-grammar stream still matches the schema. Truncation and disconnects are how you get an incomplete prefix anyway.
When do you paint the UI?
Answer
On completed values from an incremental parser (closed string, closed number, closed array element). Not on raw tokens. Enums and emails wait until the value event so you do not flash US as USD or ada@ as an address.
When do side effects run?
Answer
After the stream completes successfully and the full object passes schema plus business validation. Never on a partial vendor field. Tools and INSERTs are commit-time, same as a non-streaming extractor.
Truncation mid-array?
Answer
Detect via finish_reason length, a non-empty CFG stack, or JSON.parse failure at end-of-stream. Increase max_tokens, paginate the schema, or start a follow-up extract. Do not brace-close the buffer and persist it.
Why is NDJSON easier?
Answer
Each line is a complete JSON value, so JSON.parse per line is valid. Use it when the product is a list of records. A single huge object is the case that needs incremental parsers.
SSE vs fetch streams?
Answer
Both push bytes. SSE is event-framed for browsers; fetch body is a raw byte stream. Neither frames JSON values for you unless you chose NDJSON. Handle backpressure and UTF-8 chunk edges in both.
What if the provider only emits tokens at JSON value boundaries?
Answer
Then your UI events may already align with completed fields. You still must not commit until the final validated object. Provider buffering is an optimization, not a substitute for a commit point.
How do refusals appear on a stream?
Answer
Often as a distinct event or an empty structured body plus a refusal field at the end. Do not parse a partial object and then ignore the refusal. Short-circuit like the repair-loop lesson: refusals are not incomplete JSON.
Buffer-all vs incremental — which default?
Answer
Buffer-all for small payloads (simplest correctness). Incremental paint for large extracts and form-fill UX. NDJSON when you control the shape toward many records. All three still validate once per complete value before side effects.
Pitfalls
Write {"vendor":"Acme","total_cents":1} on a board. Split it after {"ven, dor":"Ac, me","total_, and the rest. At each split, mark JSON.parse throws/ok and whether vendor is paint-able. Circle the only split where you would INSERT. That is this lesson.
Go Deeper
Docs:
Cluster: