Confidence-Gated Routing, Composite Scoring & Workflow Decomposition
Typed answers alone do not make automation safe. Confidence (Choice/Score) and noul magnitude (yes/no) are the second axis: what vs whether to act. This lesson teaches stake-scaled thresholds, composite scores owned in code, and speculative fan-out — without re-teaching Structured Outputs validation loops.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Payment intent looks like approve_transfer — do you auto-execute?
Prefer
Answer chooses path; confidence chooses automation
check_balance can auto at modest confidence. approve_transfer needs a high bar or a confirm step. Uncertain dimensions fail the composite.
- Thresholds are a product policy, versioned in code.
- Atomic Scores / Nouls plus weights beat a mega-prompt.
- Ask severity even if it might not be a bug — ignore it when the noul is low.
Alternative
One narrative LLM call, or always-human
A single mega-prompt is opaque and hard to reweight. Always-human is safe and has no leverage when high-conf volume is large.
- Wrong criteria + high confidence still automates the wrong policy.
- Do not port Noul thresholds onto Choice blindly (structural invariance).
- Calibration is a vendor aggregate claim — not a guarantee on one sample.
Overview
Typed answers alone do not make automation safe. Confidence (Choice / Score) and noul magnitude (yes / no) are the second axis: what vs whether to act.
Interviewers probe thresholds by risk, composite scores owned in code, and speculative fan-out. This page stays on those patterns. Do not re-teach Structured Outputs validation loops.
You should be able to:
- Draw a high / mid / low gate on a Choice and on a Noul.
- Write a weighted composite that escalates if any atomic confidence is poor.
- Explain why extra parallel questions are usually cheaper than extra round-trips (vendor claim — measure).
Confidence vs probability
- Probabilities: full distribution over Choice options or Score levels.
- Confidence: vendor-derived 0–1 summary of how peaked that distribution is (flatter → lower). Approximate demo formula for 3 options:
(3×peak − 1)/2— TypeSafe also returns raw probabilities so you can define your own statistic. - Noul: only
noul∈ [0, 1]; no separate confidence — gate on distance from 0.5 or dual thresholds. - Calibration claim (vendor): higher confidence correlates with higher accuracy in aggregate — not a guarantee on one sample.
Three confidence bands (starting pattern)
- High — act automatically
- Medium — confirm, flag, or gather more state
- Low — human, clarification, or fallback model
Thresholds scale with risk: wrong balance screen ≠ wrong wire transfer.
Core patterns
- Confidence-gated routing — answer chooses path; confidence chooses automation level.
- Composite scoring — atomic Scores / Nouls + weights in code; change product priorities without prompt rewrites.
- Speculative fan-out — ask questions you might need; ignore unused answers (vendor cookbooks: batching many questions ≫ sequential calls on cost / latency).
- Intent routing — Choice over intents → deterministic handler / specialist LLM / human.
- LLM guardrails — screen ingress / egress with Nouls / Scores (hazard, jailbreak, PII) before / after generative calls — hybrid.
Architecture (confidence gate)
Single-column path from parallel ask to act / confirm / human.
Decisions
- 1
1 State
- next2 Choice + Scores + Nouls
- 2
2 Choice + Scores + Nouls
- next3 Confidence / noul
- ?
3 Confidence / noul
- high + low stakes4 Auto path
- mid or high stakes4 Confirm / soft gate
- low4 Human or LLM fallback
- 4
4 Auto path
- 5
4 Confirm / soft gate
- 6
4 Human or LLM fallback
Lesson map
Confidence-Gated Routing, Composite Scoring & Workflow Decomposition
Typed answers alone do not make automation safe. Confidence (Choice/Score) and noul magnitude (yes/no) are the second axis: what vs whether to act. This lesson teaches stake-scaled thresholds, composite scores owned in code, and speculative fan-out — without re-teaching Structured Outputs validation loops.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s["1 State"] q["2 Choice + Scores + Nouls"] conf["3 Confidence / noul"] auto["4 Auto path"] s -->|1 State to 2 Choice + Scores + Nouls| q q -->|2 Choice + Scores + Nouls| conf conf -->|high + low| auto
Sandbox: risk-scaled gates + composite (Python)
Weights owned by product — not the model.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Speculative fan-out consume (TypeScript)
Asked severity even if it is not a bug — ignore when the noul is low.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Workflow decomposition checklist
- List every atomic judgment a knowledgeable person would make in a second given state.
- Map each to Choice / Score / Noul; name stable IDs for logs.
- Ask independent ones together; only chain when state or option sets truly depend on prior answers.
- Encode product policy (weights, stake thresholds) in code — version them.
- Shadow-mode: log would-act vs human labels before enabling auto.
- Revisit thresholds when the model alias changes (e.g. jev-1.13.x).
Comparative pros / cons
| Pattern | Pros | Cons |
|---|---|---|
| Confidence gating | Explicit uncertainty channel; stake-scaled policy | Thresholds need offline tuning; overtrust if calibration fails on your domain |
| Single mega-prompt LLM | One-call narrative | Opaque; hard to reweight; weak automation contract |
| Always-human | Safe | No leverage; System One shines when high-conf volume is large |
Pitfalls
Intents: check_balance, approve_transfer, dispute_charge. Sketch high/mid/low bands per intent, a composite of severity + evidence, and which extra questions you would fan out speculatively. What do you log in shadow mode?
Interview Q&A
Confidence vs max probability?
Answer
Related but not identical; confidence summarizes distribution shape. Prefer returned confidence for gates; keep probabilities for custom metrics / expected value. Demo formula for 3 options: (3×peak − 1)/2 — still prefer the vendor field plus your own statistic.
Why composite in code?
Answer
Domain weights, auditability, A/B without retraining; the model supplies atomic judgments. Product can reweight severity vs frustration without rewriting instructions.
Is fan-out wasteful?
Answer
Vendor: extra questions add mostly question tokens; the parallel sampler keeps latency flatter vs N round-trips. Still measure on your mix. Do not yield N sequential System One calls unless dependency forces it — generators.
Noul gating without confidence?
Answer
Use bands, e.g. act if noul > 0.85, human if 0.4–0.6, reject if < 0.15 — tune on a labeled set. Near 0.5 is uncertain, not “medium skill” (primitives).
Structural invariance trap?
Answer
Separate Noul and Choice on the “same” question need not match; don’t port thresholds across types blindly (jaggedness).
What is intent routing?
Answer
A Choice over intents, then a deterministic handler, a specialist LLM, or a human. The Choice is not the prose generator. Depth: hybrid.
Where do LLM guardrails sit?
Answer
Ingress / egress Nouls and Scores (hazard, jailbreak, PII) around generative calls. Cookbook-level here; architecture on the hybrid page.
Why shadow-mode?
Answer
High confidence on the wrong option still ships. Log would-act vs human labels, then enable auto per intent when precision holds. Revisit when jev-1.13.x aliases move.
Low-stakes vs high-stakes example?
Answer
Wrong balance screen is recoverable UX. Wrong wire transfer is money. Same Choice family, different confidence floors and confirm steps — that is the interview answer.
Always-human vs System One?
Answer
Always-human is the correct default until shadow traffic says otherwise. System One earns its keep when high-conf volume is large and the criteria are right.