Hybrid Architecture — System One for Route/Guardrail, LLMs for Prose
Production systems rarely pick one model class. The winning shape: System One (or equivalent classifiers) for route, score, and guardrail; generative LLMs for explanations, drafts, and open-ended tools; deterministic code as the source of truth. This capstone wires primitives, confidence, generators, and streaming into an interview-ready architecture.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Explain-fees bot with toxicity and PII risk
Prefer
System One on the control plane, LLM on the data plane
Parallel Nouls/Scores at ingress. Intent Choice picks a handler. Stream prose only if needed, then egress-check. Confidence is the kill-switch.
- Code owns thresholds, weights, retries, audit logs, arithmetic, and dates.
- Humans take low confidence, high stakes, and policy exceptions.
- Two eval harnesses: decision nodes vs prose quality.
Alternative
LLM-only agent, or rules-only NLP
LLM-only is flexible and stringly: tool hallucinations, weak latency SLOs. Rules-only is deterministic and brittle on the edges System One targets.
- One Structured Outputs call can work — compare on your workflow evals, not homepage multiples.
- LLM verbalized confidence ≠ System One confidence.
- Tool/Function Calling is the live agent-loop cluster — light-link only.
Overview
Production systems rarely pick one model class. The winning shape:
- System One (or equivalent classifiers) for route, score, and guardrail.
- Generative LLMs for explanations, drafts, and open-ended tools.
- Deterministic code as the source of truth.
This capstone wires primitives, confidence, generators, and streaming into an interview-ready architecture. Mark vendor speed / cost claims; note jaggedness.
You should be able to:
- Draw ingress → route → (stream | JSON action) → egress.
- Answer “where does Jev sit in an agent?” in one sentence.
- Separate decision evals from prose evals.
Division of labor
- System One / Jev — intent Choice, toxicity / PII Nouls, severity Scores, citation support checks, RAG passage relevance, confidence gates.
- Generative LLM — user-facing prose, code gen, multi-hop reasoning, open extraction proposals (then Choice among candidates).
- Code — thresholds, weights, schema checks if you use Structured Outputs (live cluster), retries, audit logs, arithmetic / dates (keep off Jev).
- Humans — low confidence, high stakes, policy exceptions.
Reference pipeline
- Ingress guardrails (System One parallel Nouls / Scores) → block / review / pass.
- Intent + entity routing (Choice) → handler selection.
- Optional speculative questions (fan-out) consumed by code.
- If prose needed → stream LLM with AbortSignal; else return a structured action.
- Egress guardrails on model output → deliver or redact.
- Escalate when confidence is low or noul is ambiguous.
Architecture (hybrid request path)
Four participants, short labels. Lifelines stay inside the card.
Sequence
- 1
User / Client
Step 1 - Guard + route
- 2
User / Client → BFF / API
request
- 3
BFF / API → System One
state + parallel questions
- 4
System One → BFF / API
noul/choice/score + confidence
- 5
BFF / API → User / Client
block / human queue
- 6
BFF / API → User / Client
typed action JSON
- 7
BFF / API → Generative LLM
stream completion
- 8
Generative LLM → BFF / API
tokens
- 9
BFF / API → System One
egress checks
- 10
System One → BFF / API
ok / redact
- 11
BFF / API → User / Client
SSE tokens or final text
Lesson map
Hybrid Architecture — System One for Route/Guardrail, LLMs for Prose
Production systems rarely pick one model class. The winning shape: System One (or equivalent classifiers) for route, score, and guardrail; generative LLMs for explanations, drafts, and open-ended tools; deterministic code as the source of truth. This capstone wires primitives, confidence, generators, and streaming into an interview-ready architecture.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["User / Client"] bff["BFF / API"] s1["System One"] llm["Generative LLM"] u -->|request| bff bff -->|state + parallel| s1 s1 -->|noul/choice/scor| bff bff -->|block / human| u bff -->|typed action| u bff -->|stream| llm
Sandbox: hybrid stub (Python)
Policy object first. Stream only on the explain path. Egress is a mocked noul.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same policy object (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comparative architecture choices
| Shape | Pros | Cons |
|---|---|---|
| Hybrid System One + LLM | Automation where typed; UX where generative; confidence as kill-switch | Two vendors / systems; dual eval harnesses |
| LLM-only agents | Flexible | Stringly control flow; tool hallucinations; harder latency SLOs |
| Rules-only | Deterministic | Brittle NLP edge cases System One targets |
Failure modes & skepticism
- Vendor latency / cost — reproduce on your region and question mix; West Coast laptop numbers ≠ your p99.
- Jaggedness — arithmetic, dates, counting, adversarial state: keep in code or generative + verify.
- Over-automation — high confidence on the wrong schema option still ships; monitor with labeled shadow traffic.
- Dual calibration — don’t assume LLM verbalized confidence equals System One confidence.
Pitfalls
Banking assistant: “Why fees?” plus a possible wire. Draw ingress Nouls, intent Choice, stake-scaled confidence, stream vs structured action, egress PII, and the human queue. Which numbers are vendor claims you would refuse to quote as your SLO?
Interview Q&A
Where does System One sit in an agent?
Answer
Before tools (intent / skill pick), around tools (arg Choice), and as guardrails — not as the prose generator.
Why not one Structured Outputs call for everything?
Answer
Works for many apps. System One’s pitch is parallel calibrated decisions, speed, and no string sampling. Compare on your workflow evals; don’t accept homepage multiples blindly. Light: Structured Outputs.
How do you test hybrids?
Answer
Freeze workflows in code; score decision nodes vs human / LLM ensemble (vendor approach); separately grade prose with raters / rubrics. Two harnesses, two dashboards.
Streaming + guardrails order?
Answer
Prefer ingress checks before stream; egress may buffer or scan rolling windows — product tradeoff vs TTFT. Depth: streaming.
Jevons / Kahneman names in an interview?
Answer
Brief: efficiency → demand metaphor; System One ≈ fast judgments for software — then pivot to API / primitives / confidence. Do not spend the loop on the book.
What do you keep in code no matter how good Jev gets?
Answer
Thresholds, weights, retries, audit logs, arithmetic and dates, and any policy you must A/B. Jaggedness is not a rounding error.
LLM-only vs hybrid in one sentence?
Answer
LLM-only is a flexible stringly control plane; hybrid puts typed decisions and confidence on the control plane and keeps the LLM on the data plane.
How do extraction and Choice combine?
Answer
LLM (or regex) proposes candidates; Choice selects. Do not chain Choices to force generation (primitives).
Function-calling cookbook vs the live tool-calling page?
Answer
TypeSafe’s cookbook uses questions for typed args. The live Tool vs Structured Outputs page is the LLM-tool cluster — light-link, do not recap tool hallucination loops here.
What is the cluster recap?
Answer
Hub: job shape. Primitives: Choice / Score / Noul. Confidence: gates, weights, fan-out. Generators: yield vs round-trip. Streaming: TTFT / SSE / when not to stream. This page wires them.
Go Deeper
- How to build with System One
- LLM guardrails cookbook
- Intent routing
- Function calling cookbook
- TypeSafe blog
- Docs index
- Jaggedness
- Glossary
- Kahneman, Thinking, Fast and Slow