LLM Token Streaming vs Parallel Decisions — SSE, TTFT & When Not to Stream
Staff interviews split “stream tokens for UX” from “await a structured decision.” Mixing them causes bad architectures: streaming a classifier, or blocking the UI for a paragraph that should have streamed. This lesson covers TTFT, SSE/chunked tokens, cancellation, backpressure, and when a System One parallel call is enough.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
User asks “explain fees” vs “route this ticket”
Prefer
Gate first; stream only if a human is reading prose
Await System One (or equivalent) for toxicity / intent. If the job is a branch, return a typed action. If the job is copy, stream tokens with AbortSignal.
- Streaming is UX and control, not a discount on billable tokens.
- Dropping tokens corrupts prose — pull-based reads, bounded buffers, cancel.
- Server-to-server pipelines can await the full result unless they need early abort.
Alternative
Stream the classifier, or block the paragraph
Token-drip on a closed decision is architecture theater. Blocking a long explanation is the UX failure streaming exists to hide.
- Tiny outputs: TTFT ≈ total time — skip the stream machinery.
- Incremental JSON is still an LLM string problem — live cluster, not Jev.
- Do not recap MQTT / WebSocket guts here.
Overview
Staff interviews split “stream tokens for UX” from “await a structured decision.” Mixing them causes bad architectures: streaming a classifier, or blocking the UI for a paragraph that should have streamed.
This page covers TTFT, SSE / chunked tokens, cancellation, backpressure, and when a System One parallel call is enough. Messaging already covers WebSocket vs SSE vs long-poll — point there. Light-link Streaming Structured Output for incremental JSON — do not re-teach schema / CFG.
You should be able to:
- Pick stream vs await from the consumer (human reading vs code branching).
- Cancel in-flight generation without inventing a partial System One answer.
- Refuse “streaming saves tokens” as a cost argument.
TTFT and why streaming exists
- TTFT — time to first token. Autoregressive models emit tokens sequentially; streaming lets the UI paint early.
- Total latency still grows with output length; streaming hides wait, it does not remove compute.
- System One vendor claim: ~70–500 ms end-to-end for structured decisions — often no progressive UX needed because the useful unit is the typed answer bundle, not tokens. Treat that range as a claim; measure yours.
Transport (light)
- SSE / chunked HTTP — common for LLM token APIs (server → client event stream). Details of WS vs SSE vs long-poll: WebSocket vs SSE vs long-polling vs MQTT. Here: prefer SSE for one-way token push from your BFF.
- Client: fetch + ReadableStream, EventSource, or SDK async iterators; pass AbortSignal for cancel-on-navigate.
- Backpressure: if UI / processor is slow, apply pull-based reads; dropping tokens corrupts prose (buffer with limits + cancel).
When NOT to stream
- Closed routing / moderation / scoring — System One or non-stream structured call; gate then act.
- Tiny outputs where TTFT ≈ total time.
- Server-to-server pipelines with no human watching — await full result unless you need early abort on partial signals.
- Incremental structured JSON — different problem (Streaming Structured Output); still not System One.
Architecture (choose path)
Single-column consumer fork. Three exits: stream, decide, or incremental JSON (live link).
Decisions
- ?
1 Human reading prose?
- yes chat/copy2 LLM token SSE
- no branch in code2 System One POST
- partial JSON object2 Streaming structured JSON
- 2
2 LLM token SSE
- next3 AbortSignal cancel
- 3
2 System One POST
- next3 Route / escalate
- 4
2 Streaming structured JSON
- 5
3 AbortSignal cancel
- 6
3 Route / escalate
Lesson map
LLM Token Streaming vs Parallel Decisions — SSE, TTFT & When Not to Stream
Staff interviews split “stream tokens for UX” from “await a structured decision.” Mixing them causes bad architectures: streaming a classifier, or blocking the UI for a paragraph that should have streamed. This lesson covers TTFT, SSE/chunked tokens, cancellation, backpressure, and when a System One parallel call is enough.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["1 Human reading prose?"] stream["2 LLM token SSE"] dec["2 System One POST"] sso["2 Streaming structured JSON"] u -->|yes chat/copy| stream u -->|no branch in| dec u -->|partial JSON| sso
Sandbox: SSE-ish async chunks + abort (Python)
Guardrail / route first. Stream only if prose is requested.
Press Run. Snippets must be self-contained — no network, files, or native modules.
fetch stream + AbortSignal (TypeScript)
Mock System One first. Real LLM path would fetch(llmUrl, { signal, headers: { Accept: "text/event-stream" } }).
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comparative
| Path | Pros | Cons |
|---|---|---|
| Token streaming | Perceived speed, cancel mid-flight, UX for long answers | Engineering cost; partial failure states; not for hard automation branches |
| Parallel System One | Typed multi-answer, confidence, vendor latency claims | No prose; jagged edges; second round-trip if dependent |
| Non-stream full LLM completion | Simpler parsers | Worse UX for long text; still stringly for decisions unless Structured Outputs (live cluster) |
Pitfalls
Three jobs: (1) moderate a comment, (2) explain a 800-token fee policy, (3) emit a growing JSON object from an LLM. For each: stream, System One POST, or Streaming Structured Output? Where does AbortSignal sit?
Interview Q&A
Does streaming reduce billable tokens?
Answer
Usually no — same tokens; you pay for generation either way (vendor pricing differs). Streaming is UX / control.
TTFT vs e2e System One?
Answer
Different products: first token vs finished typed answers. Compare apples-to-apples on your task shape. TypeSafe’s ~70–500 ms is a vendor claim.
Where does SSE fit?
Answer
One-way server push of tokens. For bidirectional agent protocols see messaging (WebSocket) — not re-taught here.
Cancel a System One call?
Answer
Abort the HTTP request; there is no partial answer stream to trim. Contrast: LLM streams can stop mid-sentence and show what you have.
Streaming structured JSON vs Jev?
Answer
Former incrementally builds schema-constrained text from an LLM. Jev never generates strings — parallel decisions. Cross-link Streaming Structured Output only.
When is streaming the wrong default?
Answer
Closed routing, moderation, scoring; tiny completions; no human watching. Gate with System One (or a non-stream structured call), then act.
Backpressure if the UI is slow?
Answer
Pull-based reads. Bound the buffer. Cancel rather than drop tokens — dropped tokens corrupt prose. Composition details: generators.
Client APIs in one breath?
Answer
fetch + ReadableStream, EventSource, or SDK async iterators. Always pass AbortSignal for cancel-on-navigate.
Do you stream then classify?
Answer
Prefer ingress checks before the stream (hybrid). Egress may buffer or scan rolling windows — that is a TTFT tradeoff, not a reason to skip the gate.
Non-stream full LLM completion?
Answer
Simpler parsers, worse UX for long text, still stringly for decisions unless you use the live Structured Outputs cluster. For branches, prefer System One-shaped questions.