System One Models & Jev — Fast Structured Decisions for Software
Chat models optimize for preferred strings. Software needs typed, probabilistic decisions it can branch on without parsing prose. System One Models (TypeSafe, announced 2026-09-15) take unstructured state in and emit typed probabilistic decisions. Jev is the first public System One model — this hub maps the cluster.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Interviewer: we have chat LLMs — why isn't automation done?
Prefer
Typed parallel decisions when code must branch
Unstructured state in, Choice / Score / Noul out, with probabilities and confidence. One round-trip, then your code routes.
- Software needs a closed decision, not a preferred paragraph.
- Parallel questions share one state — speculative fan-out is cheap vs N sequential LLM calls.
- Confidence is an automation dial. Depth: the routing sibling.
Alternative
Chat LLM, then parse the prose
Autoregressive strings (chat, code, refusals, hallucinations). You still parse, validate, and retry. Latency often seconds.
- Structured Outputs helps schema-constrained text — different cluster; do not recap it here.
- Tool/Function Calling is an agent loop, not calibrated multi-label.
- If humans need progressive copy, stream an LLM on the hybrid path.
Interview order of operations
Job shape first. Typed questions second. Prose only if a human is reading.
- 1
Name the job
Does code need a closed branch, or does a human need copy / explanation? - 2
Build state
Ticket, policy, log — filtered fields. Nested paths in instructions, not giant dumps. - 3
Name questions
Choice / Score / Noul with stable IDs. Independent judgments go in one POST. - 4
Gate on confidence
High-conf + low stakes can auto. Low-conf or high stakes escalate. Policy lives in code. - 5
Generate only if needed
Stream an LLM for prose after the decision. Keep arithmetic and dates off Jev.
Overview
Interviewers and staff engineers keep asking: "We have chat LLMs — why isn't automation done?" Chat models optimize for preferred strings. Software needs typed, probabilistic decisions it can branch on without parsing prose.
System One Models (TypeSafe, announced 2026-09-15 by Diogo Almeida) are a vendor-defined class: unstructured state in → typed probabilistic decisions out. Jev is the first public System One model. This hub contrasts System One vs System 2-ish LLMs, names RLCD, sketches the API, and maps the cluster.
You should be able to:
- Draw the job-shape fork: generate copy vs classify / route / guard.
- Point at primitives, confidence, generators, streaming, and hybrid as separate full lessons.
- Refuse homepage latency multiples as proof — reproduce on your traffic.
What this cluster covers
- Choice, Score & Noul — state, parallel questions, typed answers.
- Confidence-gated routing — composite scoring and workflow decomposition.
- Generators, yield & yield* — composing streams and decision workflows.
- LLM token streaming vs parallel decisions — SSE, TTFT, and when not to stream.
- Hybrid architecture — System One for route / guardrail, LLMs for prose.
System One vs System 2-ish LLMs
- Inspiration: Kahneman's Thinking, Fast and Slow — fast System 1 vs deliberate System 2. Vendor naming: "System One Models" emphasize fast, focused judgments for code, not chat.
- Typical frontier LLMs: RLHF / RLVR; autoregressive token sampling; outputs are strings (chat, code, refusals, hallucinations). Software must parse + validate. Latency often seconds.
- System One + Jev (vendor framing): Reinforcement Learning for Calibrated Decisions (RLCD); parallel sampler; outputs constrained to Choice / Score / Noul shapes with probabilities and confidence. Gives up string generation.
- Naming: Jev after William Stanley Jevons (efficiency → more demand for intelligence). Schema matching is claimed guaranteed (vendor; mathematically impossible to emit off-schema types). Still jagged on arithmetic, dates, counting — see jaggedness notes.
Vendor-reported performance (claims)
As of jev-1.13.0 / jev-latest (glossary + TypeSafe blog), TypeSafe reports:
- End-to-end ~70–500 ms (vendor claim).
- ~40–200× faster on System One-shaped tasks vs frontier LLMs (vendor claim).
- Input ~$0.042 / MTok; output free (vendor claim).
- Context ~64k total with state + longest-question ceilings (vendor claim).
Workflow eval ratios on the marketing site are vendor-reported; independent benchmarks may lag. Reproduce on your traffic and question mix. West Coast laptop numbers are not your p99.
API shape (one round-trip, parallel questions)
POST https://api.typesafe.ai/v1/systemone — Authorization Bearer; body: state + model + questions map. All questions evaluate in parallel against the same state. Mocks below — no real keys in Study Docs.
Architecture (decision vs generation)
Single-column job-shape fork. Generate copy on the hybrid path; classify / route / guard here.
Decisions
- ?
1 Need prose or a branch?
- generate copy / explain2 Generative LLM
- classify / route / guard2 Build state
- 2
2 Generative LLM
- 3
2 Build state
- next3 Named Choice Score Noul
- 4
3 Named Choice Score Noul
- next4 Parallel System One POST
- 5
4 Parallel System One POST
- next5 Confidence / noul gates
- ?
5 Confidence / noul gates
- high conf6 Code routes / scores
- low conf6 Human or reasoning LLM
- 7
6 Code routes / scores
- 8
6 Human or reasoning LLM
Lesson map
System One Models & Jev — Fast Structured Decisions for Software
Chat models optimize for preferred strings. Software needs typed, probabilistic decisions it can branch on without parsing prose. System One Models (TypeSafe, announced 2026-09-15) take unstructured state in and emit typed probabilistic decisions. Jev is the first public System One model — this hub maps the cluster.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB j["1 Need prose or a branch?"] llm["2 Generative LLM"] s["2 Build state"] q["3 Named Choice Score Noul"] j -->|generate copy /| llm j -->|classify / route| s s -->|2 Build state to 3 Named Choice Score Noul| q
- Primitives: Choice, Score & Noul
- Gates: confidence-gated routing
- Prose after a decision: hybrid architecture
Sandbox: mock System One client (Python)
Stub only. Real client POSTs to api.typesafe.ai/v1/systemone. No API key here.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same mock shapes (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comparative pros / cons
| Path | Pros | Cons |
|---|---|---|
| System One / Jev | Typed answers; parallel multi-question; confidence for gating; vendor claims of low latency / cost | No string generation; jaggedness (math / dates); closed API; vendor benchmarks |
| Chat LLM + Structured Outputs | Flexible schemas | Autoregressive cost; validation loops. Light: Streaming Structured Output |
| Tool / Function Calling | Agent loops | Tool hallucination; not calibrated multi-label. Light: Tool vs Structured Outputs |
Pitfalls
Support ticket: "Charged twice — refund." Policy says duplicates are refundable. Sketch state, two parallel questions (refund Noul + department Choice), and what your code does at high vs low confidence. When would you stream an explanation instead of returning a typed action?
Interview Q&A
Is Jev just a smaller LLM?
Answer
Vendor position: different training objective (RLCD), parallel sampler, no string generation — optimized for System One tasks, not chat. Skeptical follow-up: demand task-matched evals on your workflows. Do not accept “small LLM” as the whole story, and do not accept homepage speedups as yours.
What is RLCD vs RLHF?
Answer
RLHF optimizes human preference over text. RLVR uses verifiable rewards. RLCD (vendor term) targets calibrated decisions with honest uncertainty on System One tasks. Interview answer: name the objective, then name the output contract (Choice / Score / Noul).
Why do parallel questions matter?
Answer
One state, many independent judgments in one round-trip. Speculative fan-out is nearly free vs N sequential LLM calls. If B truly depends on A’s value, make a second request — depth: primitives and fan-out.
Can it hallucinate?
Answer
Vendor: cannot emit off-schema types (schema matching claimed guaranteed). It can still be wrong within the schema; jagged on arithmetic / dates; adversarial state can steer answers. Keep math and calendars in code.
When do you still want a chat LLM?
Answer
When a human is reading progressive prose, drafting, or exploring. Typed routing and guardrails belong on System One (or equivalent). Wire them in the hybrid lesson — do not generate the branch condition as a paragraph.
How is this different from Structured Outputs?
Answer
Structured Outputs constrains an LLM’s token stream to a schema. System One does not generate strings — it returns typed decisions from a parallel sampler. Light-link Structured Outputs; do not recap schema strictness here.
What does Jevons have to do with it?
Answer
Vendor naming: efficiency can increase demand for intelligence (Jevons paradox metaphor). Say it in one sentence, then pivot to API, primitives, and confidence. Depth: hybrid Q&A.
What belongs in the first sentence of a senior answer?
Answer
“If code must branch, I want a typed decision with a probability and a confidence gate; if a human must read, I stream an LLM after that gate.” A model name without the job shape is junior.
Where do latency and price numbers come from?
Answer
TypeSafe blog / glossary as of jev-1.13.0 — vendor claims. Independent benchmarks may lag. Compare apples-to-apples on your task shape, region, and question mix.
What is the rest of this cluster for?
Answer
Primitives, confidence / composite / fan-out, generators, streaming vs parallel, hybrid. This page is the map, not the encyclopedia.
Go Deeper
- TypeSafe blog: Introducing System One Models & Jev
- Docs: System One
- Docs index
- Glossary: Jev
- Jaggedness jev-1.13
- Kahneman, Thinking, Fast and Slow (System 1 / System 2 framing)
- Next: Choice, Score & Noul — state, parallel questions & typed answers