Learn Jev, the first "System One" model — an AI that returns typed, structured decisions instead of generated text.
For developers who already know LLM APIs (OpenAI, Anthropic, etc.) — Python examples throughout.
If you've ever written json.loads() inside a try/except and prayed, this tutorial is for you.
We're going to learn TypeSafe — a new kind of AI that doesn't generate text. It returns typed answers, probability distributions, and confidence scores. No parsing. No prompt engineering. No hallucinated categories.
The one-sentence version: Jev is TypeSafe's flagship model and the first System One model — you send it content (state) plus typed questions, and it returns typed answers, probability distributions, and confidence. No text generation, no parsing.
How to Read This Tutorial
This is written for engineers who have already wired up a chat-completions call and parsed a JSON response. You know what a system prompt is. You've fought with "reply in valid JSON" instructions. That's exactly the pain TypeSafe removes.
You don't need any machine learning background. Every concept is introduced from scratch.
| If you want to... | Read... | Time |
|---|---|---|
| Understand why this model exists | Part 0–1 | 15 min |
| Make your first API call | Part 2 | 20 min |
| Master the three question types | Part 3–4 | 45 min |
| Build systems you can trust | Part 5–7 | 45 min |
| Ship something real | Part 8 (mini-project) | 60 min |
| Test yourself | Exercises + Solutions | 60 min |
Table of Contents
- Part 0 — What is TypeSafe? (System One vs. the LLMs you know)
- Part 1 — The mental model: State, Questions, Answers
- Part 2 — Getting started: Playground, API, and your first Python call
- Part 3 — The three primitives in depth: Choice, Score, Noul
- Part 4 — The art of asking: atomic questions and batching
- Part 5 — Confidence: turning uncertainty into control flow
- Part 6 — Architecture patterns from the docs
- Part 7 — The Python SDK properly: sync, async, errors, retries
- Part 8 — Mini-project: the Resume Screener
- Part 9 — Porting an LLM workflow to TypeSafe
- Part 10 — Best practices, limits, and pitfalls
- Exercises (by part) · Solutions appendix · Cheat sheet · Further reading
Part 0 — What is TypeSafe?
The Problem Every LLM Developer Knows
Large language models are built to produce text for humans to read. That's a feature when your user is a human. It becomes a bug when the consumer of the output is your code.
Think about the last time you used an LLM to make a decision inside an application — classifying a support ticket, detecting toxicity, routing an intent. You did something like this:
- Wrote a prompt saying "Classify this message. Reply ONLY with valid JSON:
{"category": "billing|technical|sales"}" - Prayed the model didn't add a friendly sentence before the JSON.
- Stripped markdown fences.
- Called
json.loads()in atry/except. - Handled the case where the model invented a category you never listed.
- Retried on parse failure, maybe with a sterner system prompt.
Steps 2–6 exist because you're coercing a text-generation system into outputting structured decisions, then parsing the results back into something your code can depend on.
That is the exact mismatch TypeSafe was built to eliminate.
System One vs. System Two
The name comes from Daniel Kahneman's Thinking, Fast and Slow.
- System 1 thinking is fast and intuitive (a gut read).
- System 2 is slow and deliberate (step-by-step reasoning).
Current frontier LLMs — including the o-series and reasoning models — are System Two machines: powerful, but slow and expensive, because they generate long chains of text.
Jev is a System One machine. It's trained to make the fast, focused judgments a knowledgeable person could make in a few seconds given the right context:
- "Does this message convey urgency?" → yes, 0.98
- "Which team should handle this ticket?" → technical, 0.85
- "How frustrated is this customer, 0–2?" → 1.4
Like an LLM, Jev understands natural-language input. Unlike an LLM, it returns typed decisions and probabilities rather than generated text. It does not write replies, produce code, or explain its reasoning.
You define the space of possible answers; the model picks within it.
How the Training Differs (in One Minute)
The docs' AI primer frames three post-training paths for pretrained language models:
| Approach | What it trains for | Result |
|---|---|---|
| RLHF — reinforcement learning from human feedback | Responses people prefer | Chatbots (GPT-class assistants). Can reward sycophancy and confident-sounding hallucinations. |
| RLVR — reinforcement learning with verifiable rewards | Tasks with checkable answers | Reasoning models: strong at math, but slower and more expensive. |
| RLCD — reinforcement learning for calibrated decisions | Decisions + calibrated probabilities | TypeSafe's Jev. |
Calibration is the key word. In a calibrated model, probabilities are optimized against real outcomes: across many predictions, outcomes assigned 0.2 occur about 20% of the time, and outcomes assigned 0.8 occur about 80% of the time.
That makes uncertainty usable by software — you can threshold it, route on it, and escalate on it.
Calibration is measured across groups of predictions; it is not a guarantee that any single answer is correct.
RLHF also causes mode dropping: the model narrows toward the outputs humans preferred during training. Fine for chat; bad when you need honest probability distributions.
Fun fact: RLHF was co-invented by Diogo Almeida, cofounder of TypeSafe. Their bet is called Machine Native Intelligence — AI with software-like properties: structure, reliability, observability, testability, speed, consistency, and low cost.
Side-by-Side: What You Know vs. What You Get
| Aspect | LLM chat API (OpenAI/Anthropic) | TypeSafe System One API |
|---|---|---|
| Output | Free-form text | Typed values constrained to your options |
| Structure | Prompt-engineered JSON mode / tool calls | Native — every answer is typed |
| Valid values | Model can hallucinate anything | Guaranteed within the options you supplied |
| Uncertainty | "Confidence: high" in prose | Numeric probabilities + confidence (0–1) |
| Latency | Seconds (generation-bound) | ~150 ms class — built for real-time |
| Cost | Every token of the answer bills | Answers are tiny — judged content is the cost |
| Adding more outputs | Bigger, riskier prompts | More questions run in parallel, barely changing latency |
| Failure mode | Parse errors, invented values | Low confidence (which you can act on) |
| Best at | Writing, reasoning, summarizing | Classifying, scoring, detecting, routing, verifying |
The one-paragraph takeaway: Keep the LLM for generating language; use TypeSafe when code needs a judgment. They compose beautifully — Part 6 shows LLMs and Jev working in the same pipeline.
Part 1 — The Mental Model: State, Questions, Answers
Every TypeSafe request has the same shape. Three words to internalize:
one request → one response: state + questions → typed answers + probabilities + confidence
state + questions ──▶ TypeSafe AI model ──▶ typed answers
(evaluates each + probabilities
question against + confidence
the state in parallel)
│
▼
your code
branch · sort · route
1.1 State — The Thing Being Judged
State is the content you want evaluated: a support message, a resume, a passage of text, the current state of your application. It can be:
| Format | Useful for | Example |
|---|---|---|
| String | A message, article, or passage | "My card was charged twice." |
| Object | Named fields, related records, app state | {"message": "My card was charged twice.", "order_id": "A-104"} |
| Array | A sequence of messages or records | ["Hi", "My customer number is TS1337.", "My card was charged twice."] |
The docs recommend an object for most requests: each part gets a descriptive name, and relationships stay clear.
Two input rules to remember:
- Text only (for now). Jev evaluates strings, JSON objects, and arrays of text. Images, audio, and video are not supported (yet).
- English is the primary training language. Other languages, including CJK scripts, are accepted but currently have lower accuracy.
1.2 Questions — The Judgments You Want
A question has four parts (only three are always required):
-
ID — the key you pick, e.g.
refund_requested. It labels the answer in the response. (The ID is for your code only — it is not sent to the model.) -
type — one of
choice,score, ornoul(the three primitives; Part 3). - instructions — the actual question or statement to judge. This is where your evaluation logic lives. Write it as a complete, self-contained question, even if the ID seems self-explanatory.
- criteria — the possible answers: options for a Choice, ordered levels for a Score, or optional true/false definitions for a Noul.
A common beginner bug: Writing
instructions: "refund"and relying on the IDrefund_requestedto carry the meaning. The model never sees the ID. Write the full question ininstructions.
1.3 Answers — Typed Values, Not Prose
Each question returns one typed answer under the ID you chose. Two properties make these answers composable in ways LLM output is not:
- Constrained: every answer is a probability distribution over your options or levels. The model cannot return a value outside them. There is nothing to parse and nothing to hallucinate outside the space you defined.
- Independent: every question is evaluated in isolation against the state. One question's answer never becomes hidden context for another, so adding or removing questions never changes the others' results. This is the "no context-rot" guarantee.
1.4 The Panel-of-Experts Picture
The docs offer a mental model worth memorizing: you are briefing a panel of experts.
- The state is the case file you put on the table.
- Each question is one expert, asked one narrow question, who answers independently of the others.
- Your code is the chair of the panel: it collects the typed answers and decides what happens.
This picture immediately suggests the design rule that separates good TypeSafe systems from bad ones:
Make each question a gut-check judgment, and do the complex reasoning by composing answers in code.
That rule gets a full treatment in Part 4.
Part 2 — Getting Started: Playground, API, and Your First Python Call
You can get hands-on in three ways. We'll do all three, fastest first.
2.1 Try It: The Playground (Zero Code)
The quickest way to feel what "typed answers" means is the Playground at console.typesafe.ai/playground:
- Open the Playground and log in.
- Paste any text as the state, for example:
Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.
- Add a Noul question (a yes/no probability):
{
"urgency": {
"type": "noul",
"instructions": "Does this message express urgency?"
}
}
- Add more questions. Mix Noul, Choice, and Score in one call and watch all the typed answers appear at once.
Notice how the answers arrive: a number for the Noul, an option plus probabilities for the Choice, a position on your scale for the Score. No prose to parse anywhere.
2.2 Call It: The Raw HTTP API
Get an API key from the dashboard, then POST to a single endpoint:
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer <API_KEY>
Content-Type: application/json
A complete cURL request — one state, three mixed questions:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}
EOF
And the actual response:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
},
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": { "0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language" },
"probabilities": { "0": 0.0, "1": 1.0, "2": 0.0 }
},
"is_urgent": { "type": "noul", "noul": 1.0 }
},
"usage": { "input_tokens": 392, "output_tokens": 65 }
}
Map this against what you know from chat APIs: there is a model field, a messages-analogue (state), a tools-analogue (questions), and a usage block counting tokens. The answers object is keyed by the question IDs you chose.
Notice what is missing: no choices[0].message.content, no markdown, no "As an AI..." preamble.
2.3 Code It: The Python SDK
Requires Python ≥ 3.10.
pip install typesafe-sdk # or: uv add typesafe-sdk
export TYPESAFE_API_KEY="sk-..." # the client reads this automatically
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient()
ticket = ("Hi, I've been trying to connect my Stripe account for 3 days "
"and the integration keeps failing. I'm losing sales. Please help ASAP.")
response = client.system_one(
state=ticket,
questions={
"department": Choice(
instructions="Which team should handle this",
criteria={
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
),
"frustration": Score(
instructions="How frustrated the customer appears",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language",
],
),
"is_urgent": Noul(
instructions="The message conveys urgency or time-sensitivity",
),
},
)
print(response.answers["department"].choice) # "technical"
print(response.answers["frustration"].score) # 1.0
print(response.answers["is_urgent"].noul) # 1.0
Three things to notice:
-
No prompt. The questions are the prompt.
instructionsis your evaluation logic;criteriais your output space. -
No parsing.
.choice,.score,.noulare typed fields on typed objects. - One call, three question types. They all evaluated in parallel against the same state.
2.4 Request/Response Anatomy, Annotated
For your mental model, every request you will ever write has this shape:
client.system_one(
state=<string | dict | list>, # the thing being judged
model="jev-latest", # optional, this is the default
questions={
"<your_id>": Choice(instructions=..., criteria={...}), # → choice, probabilities, confidence
"<your_id>": Score(instructions=..., criteria=[...]), # → score, legend, probabilities, confidence
"<your_id>": Noul(instructions=...), # → noul (0–1)
},
)
# → response.answers["<your_id>"] (typed answer objects)
# → response.usage (token counts)
Tip for the impatient: The docs also offer an "agent skill" so coding agents (Claude Code, Cursor, etc.) can write TypeSafe integrations for you:
claude plugin marketplace add typesafe-ai/skillsthenclaude plugin install typesafe@typesafe-ai, ornpx skills add typesafe-ai/skills --skill typesafe-ai. Handy once you know the basics yourself.
Part 3 — The Three Primitives in Depth: Choice, Score, Noul
TypeSafe calls its question types primitives — small, typed, composable building blocks, like software primitives. Each pairs a question shape with an answer shape:
| Type | What it answers | Returns |
|---|---|---|
| Choice | Which of these options? |
choice, probabilities, confidence
|
| Score | Which level on a spectrum? |
score, legend, probabilities, confidence
|
| Noul | Is this true? |
noul (0–1 probability of yes) |
Pick the type whose answer your code can act on directly: a Choice maps onto a
matchstatement, a Score maps onto a threshold, a Noul maps onto anif.
3.1 Choice — One of a Known Set of Options
Use a Choice when the answer is one of a fixed set of options with no order between them: which team handles a ticket, which category a product belongs to, which language a code snippet is in.
"language": Choice(
instructions="What programming language is this code written in?",
criteria={
"python": None,
"javascript": None,
"typescript": None,
"go": None,
"rust": None,
"other": "Anything else, including mixed or unclear snippets",
},
)
Request fields: type: "choice", instructions, and criteria — a map of {option: description}. Descriptions are optional (None = just the label), but a good description sharpens the boundary between options.
Response fields:
{
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
}
-
choice— the selected option (always one of your keys). -
probabilities— the full distribution across every option, summing to 1. -
confidence— how peaked that distribution is, 0–1 (formula in Part 5).
Two practical rules from the docs:
-
Add an escape hatch. Include an
other/none of the aboveoption when your list might not cover every input, so the model is never forced to lie. - Low confidence often means your option list is ambiguous — two options overlap, or the state lacks the information needed to separate them.
3.2 Score — A Position on a Spectrum You Define
Use a Score when the answer is a position on an ordered spectrum and you can describe what each step means: bug severity, customer frustration, skill level, urgency.
"bug_severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
)
Request fields: criteria is an ordered list of level descriptions, low end to high end. Minimum 2, maximum 10 levels. A level's number is its index in the array — the model sees the descriptions, not the numbers.
Response fields (real example from the docs, for the bug report "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari."):
{
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": { "0": 0.0, "1": 0.57, "2": 0.43 }
}
Read it like this:
-
score: 1.43is a position on the level line (0 to 2 here). It is the probability-weighted level:0×0.0 + 1×0.57 + 2×0.43 = 1.43. It can land between two levels — the model is saying "between 'workaround exists' and 'no workaround', leaning toward workaround". This interpolation is a feature, not rounding error. -
legendmaps level numbers back to your descriptions. -
confidence: 0.35is low because the probability is split across levels 1 and 2 — exactly matching the report's ambiguity.
Designing good levels: make them one ordered dimension, mutually exclusive, and describe observable behavior ("strong language" rather than "quite upset"). If you cannot order the options, that is a Choice, not a Score.
3.3 Noul — The Probability That a Statement Is True
Use a Noul for a clean yes/no question where the probability itself is the useful signal: does this message contain PII, is the customer requesting a refund, does this resume mention distributed systems.
"is_human_escalation": Noul(
instructions="Is the customer asking for a human agent?",
),
"is_repeat_contact": Noul(
instructions="Has the customer contacted support about this before?",
criteria=NoulCriteria( # optional, defines yes and no
true="Mentions a prior attempt, ticket, or that they have asked before",
false="No sign of any previous contact",
),
),
Response: a single number.
{
"model": "jev-1.13.0",
"answers": {
"is_human_escalation": { "type": "noul", "noul": 0.99 },
"is_repeat_contact": { "type": "noul", "noul": 0.93 }
}
}
Reading a Noul: the number is both the answer and the certainty in one value.
- Near 1 = strong yes.
- Near 0 = strong no.
- Near 0.5 = the model gives yes and no equal probability — it is telling you it cannot tell.
Real recorded values for is_human_escalation across different messages (from the docs):
| State | noul |
|---|---|
| Thanks, that fixed it! | 0.02 |
| How do I reset my password? | 0.07 |
| I need this sorted today, whatever it takes. | 0.26 |
| Are you a bot? | 0.40 |
| Is there any way to speak to someone about my invoice? | 0.84 |
| I have asked three times now. Can I please just talk to a real person? | 0.99 |
"I need this sorted today" is urgent but never asks for a person → 0.26. "Are you a bot?" hints at wanting a human without asking → 0.40, an almost even split. Those middle values are exactly where your code should be making threshold decisions (Part 5).
There is no separate confidence field on a Noul — with only two outcomes, the single probability describes everything. (If you want a confidence-style number anyway, use |2p − 1|; see Part 5.)
3.4 The Classic Mix-Up: Noul vs. Score
This is the #1 conceptual trap, so the docs call it out explicitly. "Is this candidate strong in Python?" is a bad Noul question if what you want is a skill level.
- A Noul of 0.5 means "yes and no are equally likely" — not "medium skill". If "strong" is undefined, the probability is hard to interpret at all.
- Want a level? Use a Score with defined levels: no experience → some familiarity → daily use → deep expertise.
- Want a yes/no decision? Define the condition precisely: "Does the resume state that the candidate has used Python at work?" — then a Noul is perfect.
Rule of thumb: Noul judges a claim; Score measures a position.
3.5 Choosing Between Them — The Cheat Table
| Your answer shape | Use | Code shape it maps to |
|---|---|---|
| One of N unordered buckets | Choice |
match / dispatch table / route |
| A point on an ordered scale | Score |
if score >= threshold / sort key |
| True or false, with a useful probability | Noul | if noul > 0.9: ... elif noul < 0.1: ... |
Part 4 — The Art of Asking: Atomic Questions and Batching
Getting the primitives right is half the skill. The other half is how you phrase questions and how many you ask per call.
4.1 Ask for One Snap Judgment per Question
System One models are built for the kind of judgment a knowledgeable person makes in a second given the right context:
✅ Good: "Does this message convey urgency?" · "Which team should handle this?" · "Does the resume mention Kubernetes?"
❌ Bad: "Analyze this message and determine the best course of action."
The bad one needs slow reasoning — that is a System Two job, or better: a decomposition problem. If your desired judgment weighs several independent factors, ask about each factor separately and combine the answers in code.
The docs' example: instead of "rate this startup pitch", ask separately about market size, technical feasibility, and differentiation — then weight them in your own formula. When priorities shift, you change a coefficient in code instead of rewriting a prompt. Try doing that with a monolithic LLM prompt.
4.2 Point at Parts of the State with Paths
Real states are structured objects. When a question targets one part of the state, name it with a dot-and-index path in backticks inside instructions:
state = {
"ticket": {
"subject": "Duplicate charge",
"messages": [
{"from": "customer", "text": "I was charged twice for order A-104. Please refund the duplicate."},
{"from": "support", "text": "We are checking the charges."},
],
},
"order": {
"id": "A-104",
"charges": [
{"amount_usd": 49, "status": "captured"},
{"amount_usd": 49, "status": "captured"},
],
},
"refund_policy": "Duplicate charges are eligible for a refund.",
}
questions = {
"refund_requested": Noul(
instructions="Does `ticket.messages[0].text` request a refund?",
),
"policy_supports_refund": Noul(
instructions=(
"Does `refund_policy` support the refund requested "
"in `ticket.messages[0].text`, given `order.charges`?"
),
),
}
Explicit paths tell the model exactly which parts of the state inform each judgment — this is your familiar "grounding" problem from RAG, solved by convention instead of prompt tricks.
4.3 Ask Many Questions Together — They Are Nearly Free
This is where TypeSafe's economics get fun. Every question in a request is evaluated in parallel and in isolation against the same state:
- Adding questions barely changes response time.
- Extra questions cost only their own (cheap) question tokens.
- Answers stay independent — no context rot, no cross-contamination.
So the docs recommend speculative fan-out: ask every question your code might need, including ones that only matter for some inputs, and let the code ignore the answers it does not need.
The docs' parallel-questions cookbook quantifies it: batching 13 questions into one call is 11.5× cheaper and 9.6× faster than 13 separate calls, with no change in the answers. If you have written the "loop over items, one API call each" version of an LLM pipeline, this flips the economics.
# One request: classify + urgency + frustration + PII + refund + sentiment...
# ...all in the same latency ballpark as one question.
response = client.system_one(
state={"ticket_message": "My flight was cancelled. Can I get a refund?",
"refund_policy": "Cancelled flights are eligible for a full refund."},
questions={
"refund_requested": Noul(instructions="Does `ticket_message` request a refund?"),
"request_type": Choice(
instructions="What is the main request in `ticket_message`?",
criteria={
"refund": "The customer wants money returned.",
"rebooking": "The customer wants a replacement flight.",
"information": "The customer is asking for information only.",
},
),
"frustration": Score(
instructions="How frustrated does the customer appear in `ticket_message`?",
criteria=[
"Calm and neutral.",
"Concerned but civil.",
"Very angry or using strong language.",
],
),
},
)
4.4 When One Question Genuinely Depends on Another
Questions in the same request are independent — one answer never becomes context for another. So when do you make a second request? Only when your code cannot build it without the first answer:
- The answer is needed to fetch more data for the state.
- The answer decides what the state is made of.
- The answer picks the next question's options.
The docs' rule: two requests are the exception, not the rule. If the second request's questions could have been asked against the original state, ask them in the first request and ignore the extras.
The cookbooks show the legitimate cases: skill suggestion ranks 182 skills in one call, then fetches the full text of the top 3 and re-judges against better evidence; structure recovery merges lines into blocks that only exist after the first pass; hierarchical classification uses each Choice answer to pick the next request's options.
Part 5 — Confidence: Turning Uncertainty into Control Flow
If Part 3 gave you answers, Part 5 gives you trust. Confidence is the feature that makes TypeSafe usable in systems that run unattended.
5.1 Where Confidence Comes From
Every Choice and Score answer already includes probabilities — a distribution across your options or levels. The shape of that distribution tells you how certain the model is:
- Concentrated on one outcome → confident.
- Spread out → uncertain.
The confidence property collapses that shape into one number, 0 to 1: 1.0 when all probability sits on one outcome, 0 when it is spread evenly. TypeSafe computes it for you, so you can threshold on it from your very first call without doing any math. (Noul answers carry no separate confidence — the probability is the uncertainty.)
5.2 The Formulas (For When You Want Them)
You can survive perfectly well without these, but they are exact and short — and every answer includes the full probabilities, so you can compute your own statistic instead.
Noul — two outcomes, so use distance from 0.5 if you ever need a confidence-style number:
confidence = |2p − 1| # 0 at p=0.5, 1 at p=0 or p=1
Choice — how far the top probability sits above an even split, for n options:
confidence = (p_max − 1/n) / (1 − 1/n)
With three options: (3 × p_max − 1) / 2. So (0.6, 0.3, 0.1) and (0.6, 0.2, 0.2) both give 0.4 — only the top probability counts.
def choice_confidence(probabilities: list[float]) -> float:
n = len(probabilities)
return (max(probabilities) - 1 / n) / (1 - 1 / n)
Score — because levels are ordered, distance between levels matters. Probability on a neighboring level hurts less than probability far away:
def score_confidence(probabilities: list[float]) -> float:
n = len(probabilities)
m = probabilities.index(max(probabilities)) # most likely level
spread = sum(p * abs(i - m) for i, p in enumerate(probabilities))
even_spread = sum(abs(i - (n - 1) / 2) for i in range(n)) / n
return max(0.0, 1 - spread / even_spread)
Worked example from the docs: probabilities (0, 0.57, 0.43) on three levels → most likely level 1, spread 0.43, uniform average 2/3 → confidence 1 − 0.43/(2/3) ≈ 0.35.
And the ordering effect: (0, 0.5, 0.5) (torn between adjacent levels) gives 0.25, while (0.5, 0, 0.5) (torn between opposite ends) gives 0 — the plain Choice formula cannot tell those apart.
Two simpler statistics the docs recommend trying alongside confidence:
-
Top probability
p_max— reads directly as "how likely is the selected option" (but its meaning depends on option count; set thresholds per question). -
Top-to-second ratio
p_max / p_second— how clearly the winner beats the runner-up; many real decisions come down to the top two.
5.3 "I Don't Know" Is a Feature
The docs' framing is worth quoting:
If an intelligent system — human or machine — cannot express honest uncertainty, the system cannot be trusted.
Confidence is a built-in mechanism for the model to say "I'm not sure about this one", which lets your code behave differently at different levels of certainty.
5.4 The Three-Zone Pattern
A useful starting pattern divides confidence into three ranges, each producing different behavior:
| Zone | Meaning | System behavior |
|---|---|---|
| High | Clear read | Act automatically — no human involvement |
| Medium | Reasonable answer, not certain | Proceed with caution: ask the user to confirm, flag for review, or gather more info |
| Low | Genuinely unsure | Do not act — route to a human, request clarification, or fall back |
5.5 Thresholds Scale with Risk
One number is never right for every action. Gate destructive operations harder than read-only ones — this banking example is straight from the docs:
response = client.system_one(
state=user_message,
questions={
"action": Choice(
instructions="What is the user trying to do?",
criteria={
"check_balance": "View account balance",
"approve_transfer": "Approve the pending withdrawal request",
"support": "Get help with an issue",
},
),
},
)
action = response.answers["action"]
if action.confidence < 0.5:
route_to_human(user_message) # model is genuinely unsure — don't guess
elif action.choice == "check_balance":
show_balance(account_id) # low stakes: wrong screen is recoverable
elif action.choice == "approve_transfer":
if action.confidence > 0.9:
confirm_then_execute(account_id) # high stakes + high confidence
else:
ask_user_to_confirm(account_id) # high stakes + moderate confidence
The 0.5 floor catches anything the model reports as genuinely uncertain. Above that, the action itself sets the bar.
The docs' tuning advice: start conservative, test with your own data, and adjust as you observe results — the right thresholds depend on your domain and the model's measured performance on your use case.
Calibration, again: because the model is trained for calibrated probabilities, "0.8 confidence" is meaningful in aggregate — not a guarantee for any single answer. Treat thresholds as risk policy, and log everything (confidence values make AI failures wonderfully traceable).
Part 6 — Architecture Patterns from the Docs
TypeSafe is designed to sit inside a larger system, powering decisions while your code owns control flow. The key skill shift: think in terms of discrete, atomic decisions that compose into complex system behavior.
| Pattern | What it does | Benefits |
|---|---|---|
| Speculative Fan-Out | Send many questions in one call, including speculative ones; code decides what's relevant | Cost, speed |
| Confidence-Gated Routing | Use confidence as a second decision axis | Reliability, safety |
| Composite Scoring | Combine several scored dimensions into one number | Cost, reliability, speed |
| Intent Routing | Classify intent, dispatch to a handler | Cost, speed |
6.1 Speculative Fan-Out
You met this in Part 4.3 — ask everything, ignore what you don't need. Because extra questions are nearly free in latency and cheap in tokens, the optimal number of questions per call is "all of them that share this state".
r = client.system_one(state=ticket, questions=ALL_QUESTIONS) # ~everything
if r.answers["is_bug_report"].noul > 0.8:
handle_bug(severity=r.answers["severity"]) # only read severity for bugs
else:
handle_general(r) # severity answer simply unused
6.2 Confidence-Gated Routing
Confidence becomes a second decision axis alongside the answer itself. The banking snippet in Part 5.5 is this pattern verbatim: answer picks the path, confidence picks the assurance level (auto / confirm / human). This is how you build AI systems that fail safe instead of failing weird.
answer = r.answers["intent"]
if answer.confidence < 0.5: ...route_to_human()
elif answer.choice == "refund": ... # per-action thresholds below
6.3 Composite Scoring
Split one complex judgment into several Scores, normalize, weight, and combine in code. The docs work through ticket priority as three separate scores — bug severity, customer frustration, report quality — combined with weights:
# severity: 0..2 (3 levels), frustration: 0..2, report_quality: 0..2
priority = (
0.5 * r.answers["severity"].score / 2 # normalize to 0..1
+ 0.3 * r.answers["frustration"].score / 2
+ 0.2 * r.answers["report_quality"].score / 2
) # → 0..1 priority; tune the weights without touching any prompt
Why not ask for "priority" directly? Because the decomposition keeps each evaluation reliable, gives you auditability (you can see which dimension dragged a ticket down), and lets product decisions live in code where they belong.
6.4 Intent Routing
Classify the user's intent with a Choice, then dispatch on it — the LLM-era "semantic router", minus the parsing:
r = client.system_one(state=user_message, questions={
"intent": Choice(instructions="What does the user want?",
criteria={"book": ..., "cancel": ..., "complain": ..., "other": ...}),
})
HANDLERS[r.answers["intent"].choice](user_message) # KeyError only if you forgot "other"
6.5 The Cookbooks Worth Stealing From
The docs' cookbook section is a set of worked recipes. The ones that most expand what you think is possible:
- Parallel questions — the 11.5× cheaper / 9.6× faster batching study from Part 4.
- Classifying RAG passages — score each retrieved chunk for relevance before it reaches the generator; cheaper than embedding round-trips for precision filtering.
- LLM guardrails — semantic checks on every LLM input, output, and tool call (jailbreaks, prompt injection, sensitive data, policy violations) at a fraction of the cost of the LLM call itself. This is the "Universal Verification" use case: verify other AIs' work.
- Reranking — score query-to-candidate relevance and reorder search results.
- Hierarchical classification — a Choice at the top level picks which options the next request offers; deep taxonomies without one gigantic prompt.
- Skill suggestion — rank 182 skills in one call, fetch full text for the top 3, re-judge with better evidence (a real two-request dependency).
- Structure recovery (autoformat) — ask per-line whether a line break split a sentence, merge into blocks from the answers, then classify the blocks (which didn't exist until pass one).
- Model routing — classify difficulty/risk/domain of each incoming prompt and choose which LLM serves it. A Jev-powered router in front of your expensive models.
- Consistency cookbooks — measure/ensure answer stability across runs.
The docs' use-case map adds categories by industry — recruiting, support triage, insurance claims, financial crime, legal/compliance, moderation, e-commerce, gaming — and a task taxonomy you can map directly to primitives:
Classification → Choice · Detection → Noul · Scoring → Score · Routing → Choice + dispatch · Ranking/Retrieval → Scores · Verification → Nouls
Part 7 — The Python SDK Properly: Sync, Async, Errors, Retries
7.1 Sync vs. Async Clients
The SDK ships both a synchronous and an asynchronous client with identical method surfaces:
# Sync — scripts, simple services
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None}),
"urgency": Score(instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"]),
},
)
print(response.nouls["billing"].noul) # typed collections by kind...
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)
# Async — FastAPI, aiohttp pipelines, high-fanout services
import asyncio
from typesafe_sdk import AsyncTypeSafeClient, Choice, Noul, Score
async def main() -> None:
async with AsyncTypeSafeClient() as client:
response = await client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None}),
"urgency": Score(instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"]),
},
)
print(response.nouls["billing"].noul)
asyncio.run(main())
Note the two access styles: response.answers["id"] gives you any answer generically, while response.nouls / response.choices / response.scores give you typed collections by kind — the SDK equivalent of a tagged union you didn't have to build.
7.2 Typed Responses
Each answer object carries typed fields (not dicts of strings):
| Answer kind | Fields |
|---|---|
ChoiceAnswer |
choice: str, confidence: float, probabilities: dict[str, float]
|
ScoreAnswer |
score: float, confidence: float, probabilities, legend (keys are ints in the SDK) |
NoulAnswer |
noul: float |
7.3 Errors and Retries
The SDK exposes a typed exception hierarchy you already know the shape of from modern API clients: TypeSafeError (base), with AuthenticationError, BadRequestError, NotFoundError, PermissionDeniedError, RateLimitError, UnprocessableEntityError, InternalServerError, APIConnectionError, APITimeoutError, and APIUserAbortError.
from typesafe_sdk import TypeSafeClient
from typesafe_sdk import AuthenticationError, RateLimitError, UnprocessableEntityError
try:
with TypeSafeClient() as client:
response = client.system_one(state=state, questions=questions)
except AuthenticationError:
... # bad/missing TYPESAFE_API_KEY
except RateLimitError:
... # back off — the SDK can also retry for you (see retries in the usage guide)
except UnprocessableEntityError as e:
... # e.g. malformed question (Score with 1 level, >10 levels, etc.)
Retry behavior is configurable via a retry policy on the client (transient RateLimitError / InternalServerError / connection errors are the typical retry set). The http2 extra (pip install "typesafe-sdk[http2]") enables HTTP/2.
7.4 Model Selection and Usage
-
modeldefaults tojev-latest— an alias that tracks the current release; responses pin the concrete version (e.g.jev-1.13.0). -
response.usagereportsinput_tokens/output_tokens. Output tokens are tiny (tens, not thousands), because answers are typed values, not prose. The state dominates your bill, which rewards sending clean, relevant state.
7.5 Structured Instructions (the 20% Case)
instructions and criteria entries can be a string (start here), but also an object or array. Reach for structure when a question needs data alongside it — e.g. a record to compare against — or when part of the question is assembled by your code. The rule: strings first; escalate to objects only when a level/question needs a description plus examples or companion data.
Part 8 — Mini-Project: The Resume Screener
Time to build something real. Goal: a service that screens a resume against a job's requirements and returns a defensible shortlist decision.
8.1 Design the Decision Before the Code
Write down the decision your hiring team actually makes, then decompose it into atomic, snap-judgment questions:
| Business question | Decomposed into | Why |
|---|---|---|
| "How strong is this candidate overall?" | 3 × Score (skills, experience, education) + weighted formula | Composite Scoring — weights live in code |
| "Do they meet the hard requirement?" | 1 × Noul per knock-out criterion | Maps to an if
|
| "What's their seniority for this role?" | 1 × Choice | Maps to a dispatch |
| "Should a human look at this?" | Confidence gates | Confidence-Gated Routing |
This gives us all three primitives, two of the four patterns, and one honest two-request option to talk about.
8.2 Define the State
Structured state, per Part 1 — the resume content plus the job requirements, because the judgment always compares the two:
state = {
"job": {
"title": "Senior Backend Engineer",
"must_have": ["5+ years building backend services in Python or Go",
"Production experience with PostgreSQL",
"Experience designing public APIs"],
"nice_to_have": ["Kubernetes", "Event-driven systems", "Fintech domain"],
"knockouts": ["Requires active work authorization in the US"],
},
"resume": (
"Maya Chen — Backend Engineer, 7 years experience.\n"
"At Northwind (2021–now): designed and ran public REST + gRPC APIs in Python "
"(FastAPI) handling 40k req/s; owns PostgreSQL schema and query tuning; "
"migrated services to Kubernetes on EKS.\n"
"At Brightline (2018–2021): Go microservices for payments; built event-driven "
"pipelines on Kafka. MSc Computer Science, University of Washington. "
"US citizen. Work authorization: US."
),
}
8.3 Define the Questions
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
SKILL_LEVELS = [
"No evidence of this skill",
"Mentioned once, no depth shown",
"Used in real projects with concrete detail",
"Clearly deep expertise; led or scaled it",
]
EXPERIENCE_LEVELS = [
"Junior: under 2 years of relevant work",
"Mid-level: 2–5 years of relevant work",
"Senior: 5–8 years of relevant work",
"Staff+: 8+ years, with scope or leadership signals",
]
questions = {
# -- Composite scoring dimensions (Scores) --------------------------
"skill_depth": Score(
instructions=("Considering `job.must_have`, how deep is the candidate's "
"evidence of the required skills in `resume`?"),
criteria=SKILL_LEVELS,
),
"experience_fit": Score(
instructions="How well does the candidate's experience match `job.title`?",
criteria=EXPERIENCE_LEVELS,
),
"nice_to_have_fit": Score(
instructions=("How many of `job.nice_to_have` does `resume` show real "
"evidence for?"),
criteria=[
"None of them",
"One of them",
"Two of them",
"All of them",
],
),
# -- Knock-out criteria (Nouls) -------------------------------------
"work_authorized": Noul(
instructions=("Does `resume` satisfy the requirement in `job.knockouts`?"),
criteria={
"true": "Resume clearly states US work authorization or citizenship",
"false": "No statement, or a statement that contradicts it",
},
),
# -- Seniority bucket (Choice) ---------------------------------------
"seniority": Choice(
instructions=("Which seniority best matches the evidence in `resume` "
"for `job.title`?"),
criteria={
"junior": "Under 2 years relevant experience",
"mid": "2–5 years relevant experience",
"senior": "5–8 years relevant experience",
"staff": "8+ years with leadership/scope signals",
},
),
# -- Speculative extras (fan-out: nearly free) ------------------------
"mentions_fintech": Noul(instructions="Does `resume` mention fintech or payments experience?"),
"resume_sufficient": Noul(
instructions=("Does `resume` contain enough information to evaluate "
"all of `job.must_have` fairly?"),
),
}
Design notes worth internalizing:
- Every
instructionsis a complete question — IDs never reach the model. - Scores measure one ordered dimension each; the levels describe observable evidence.
- The knockout is a Noul with criteria defining what counts as yes/no.
-
resume_sufficientis the meta-question: if the model says "no", you know the low skill score is a data problem, not a candidate problem. Speculative questions earn their keep. - One call — all of it in parallel, ~same latency as a single question.
8.4 Combine in Code — The Screener
from dataclasses import dataclass
@dataclass
class Screening:
composite: float
decision: str # "shortlist" | "review" | "reject"
reasons: list[str]
def screen_resume(client: TypeSafeClient, state: dict) -> Screening:
r = client.system_one(state=state, questions=questions)
skills = r.answers["skill_depth"] # ScoreAnswer
exper = r.answers["experience_fit"] # ScoreAnswer
nice = r.answers["nice_to_have_fit"] # ScoreAnswer
auth = r.answers["work_authorized"] # NoulAnswer
senior = r.answers["seniority"] # ChoiceAnswer
enough = r.answers["resume_sufficient"] # NoulAnswer
reasons: list[str] = []
# 1) Composite score — weights are a product decision, not a prompt.
composite = (
0.45 * skills.score / 3 # 4 levels → normalize 0..3 to 0..1
+ 0.40 * exper.score / 3
+ 0.15 * nice.score / 3
)
# 2) Knockouts are hard gates — but only act when the model is sure.
if auth.noul < 0.2:
reasons.append("Fails work-authorization knockout")
return Screening(composite, "reject", reasons)
# 3) Data sufficiency: don't judge candidates on missing data.
if enough.noul < 0.5:
reasons.append("Resume lacks information for a fair evaluation")
return Screening(composite, "review", reasons)
# 4) Confidence-gated routing on the composite.
avg_conf = (skills.confidence + exper.confidence + nice.confidence) / 3
if composite >= 0.65 and avg_conf >= 0.6:
decision = "shortlist"
reasons.append(f"Composite {composite:.2f} with solid confidence {avg_conf:.2f}")
elif composite >= 0.65:
decision = "review"
reasons.append(f"Composite {composite:.2f} but low confidence {avg_conf:.2f}")
elif composite >= 0.40:
decision = "review"
reasons.append(f"Borderline composite {composite:.2f}")
else:
decision = "reject"
reasons.append(f"Composite {composite:.2f} below bar")
# 5) Seniority routing — a Choice maps straight onto dispatch.
if senior.choice == "staff" and decision == "shortlist":
reasons.append("Route to staff-track interview loop")
return Screening(round(composite, 3), decision, reasons)
8.5 Wire It Up and Read the Output
if __name__ == "__main__":
with TypeSafeClient() as client:
result = screen_resume(client, state)
print(f"Decision: {result.decision.upper()} (composite {result.composite})")
for reason in result.reasons:
print(f" • {reason}")
# Example output for the state above:
# Decision: SHORTLIST (composite 0.78)
# • Composite 0.78 with solid confidence 0.81
# • Route to staff-track interview loop
8.6 Where You Would Take It Next
-
Batch mode: screen 500 resumes by looping
screen_resume— or better, notice each resume is a different state, so one request per resume is correct (fan-out is within a state's questions, not across states). -
Second-request escalation (the honest kind): for review decisions, fetch the candidate's portfolio/GitHub and re-run just the
skill_depthScore against the richer state. - Fairness hygiene: keep criteria job-related and explicit, log every probability distribution and confidence value per decision, and keep a human in the loop for every review.
- Tuning loop: hiring says "Maya should have gone straight to staff loop"? Raise the staff weight or the composite bar — in code, not in a prompt.
Part 9 — Porting an LLM Workflow to TypeSafe
Let's convert something you have definitely written: an OpenAI-style classification call with JSON mode.
9.1 Before: The Chat-Completions Version
# The way we've all done it (sketch, OpenAI-style SDK)
SYSTEM = """You are a support ticket classifier. Classify the user's message.
Reply ONLY with valid JSON: {"department": "billing|technical|sales",
"frustration": 0-2, "is_urgent": true|false}"""
resp = llm.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "system", "content": SYSTEM},
{"role": "user", "content": ticket}],
response_format={"type": "json_object"},
temperature=0,
)
data = json.loads(resp.choices[0].message.content) # ← the praying step
dept = data["department"] # ← could be anything
9.2 After: The TypeSafe Version
r = client.system_one(
state=ticket,
questions={
"department": Choice(
instructions="Which team should handle this ticket?",
criteria={"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"},
),
"frustration": Score(
instructions="How frustrated does the customer appear?",
criteria=["Calm", "Frustrated but civil", "Very angry, strong language"],
),
"is_urgent": Noul(instructions="Is this ticket time-sensitive?"),
},
)
dept = r.answers["department"].choice # guaranteed one of 3 strings
frust = r.answers["frustration"].score # 0..2, interpolates
urgent = r.answers["is_urgent"].noul # 0..1 probability
9.3 What Actually Changed
| Concern | LLM version | TypeSafe version |
|---|---|---|
| Instruction tuning | System prompt + response_format + temperature 0
|
The question is the contract |
| Output validity | Parsed, hoped-for JSON | Typed answers constrained to your options |
| Hallucinated categories | Possible ("support" wasn't an option…) | Impossible — distribution is over your keys |
| Uncertainty | Absent (or prose) |
probabilities + confidence per answer |
| Latency | Generation-bound | Real-time class (~150 ms per the docs) |
| Adding a question | Fatter prompt, more drift risk | One more entry; answers stay independent |
| Cost profile | Output tokens bill like any text | Output tokens are tiny; state dominates |
What you lose: generation. TypeSafe will not write the apology email, summarize the thread, or explain its reasoning. That's not a limitation to work around — it's the boundary of the tool.
The healthy architecture is hybrid:
- Jev at the edges making decisions: route, detect, score, verify, gate.
- The LLM in the middle when language must be produced: draft, summarize, converse.
- Confidence deciding when the LLM (or a human) gets involved at all — e.g., low-confidence intent → expensive reasoning model; high-confidence intent → cheap path.
This is exactly the "Harness Engineering" and "Model routing" use cases from the docs: a fast, cheap calibrated model in front of the expensive generative one.
Part 10 — Best Practices, Limits, and Pitfalls
10.1 Do / Don't
| ✅ Do | ❌ Don't |
|---|---|
Write complete, self-contained questions in instructions
|
Rely on the question ID to carry meaning (IDs never reach the model) |
| One snap judgment per question | Ask for multi-factor analysis in one question |
| Batch all questions sharing a state into one call | Loop one-question-per-call (11.5× cost, 9.6× latency, per the docs) |
Add an other option to open-ended Choices |
Force every input into a closed list |
| Define Score levels as observable behaviors on one ordered dimension | Use vague adjectives ("quite upset") or mix two dimensions into one scale |
| Use Noul for a precisely defined claim; Score for a position | Use Noul("is this candidate strong in Python?") expecting a skill level |
| Gate actions by confidence, scaled to the stakes | Treat confidence 0.7 as "70% guaranteed correct" for one answer |
| Threshold Noul directly (it has no confidence field) | Look for answer.confidence on a Noul and crash |
Reference state parts with backtick paths (`ticket.messages[0].text`) |
Dump an unstructured blob and hope |
| Start conservative, evaluate on your data, tune thresholds | Ship production thresholds from day one without measurement |
10.2 Hard Limits to Remember
- Text only. State must be a string, JSON object, or array of text. No images/audio/video (yet).
- Language. English is the primary training language; other languages (including CJK) work but currently have lower accuracy.
- Score size. 2–10 levels per Score.
- No explanations. The model will not justify answers; if you need reasoning, that's a System Two tool's job — or ask additional Noul questions about specific evidence.
- Calibration is aggregate. Group-level trustworthiness, not per-answer guarantees.
10.3 When Not to Use TypeSafe
Be honest with yourself on these:
- The task needs generation (write text/code).
- The task needs multi-step deliberative reasoning over many interacting constraints (a System Two task).
- The answer space is genuinely unbounded (open information extraction with unknown schema).
For those, keep the LLM. The sweet spot is the huge middle band of product engineering: any place your code currently asks an LLM a question and parses the reply.
10.4 Going Further with Agents
If you build with coding agents, install the TypeSafe agent skill so the agent writes well-shaped requests from the start (claude plugin marketplace add typesafe-ai/skills + claude plugin install typesafe@typesafe-ai, or npx skills add typesafe-ai/skills --skill typesafe-ai).
The docs note agents fall into the one-question-per-call habit more than humans do — the skill steers them to fan out.
Exercises
Work these with a real API key where a call is required. Solutions in the appendix.
Part 0 — Concepts
- [2 pts] In one sentence each: what do RLHF, RLVR, and RLCD optimize for?
- [3 pts] Your PM asks: "Why can't we just ask GPT to reply with strict JSON?" Give the two-sentence answer that names the structural difference and the failure mode JSON mode can't fix.
Part 1 — Mental Model
- [2 pts] Name the three parts of every TypeSafe request and the rule about how many states each request evaluates.
- [2 pts] Why is the question ID not a substitute for
instructions?
Part 2 — Getting Started
- [4 pts] In the Playground, paste a real email you've sent and ask two questions of different types. Write down both answers with their probabilities/confidence.
- [4 pts] Convert the cURL request in §2.2 so the state is an object
{"message": ..., "account_age_days": 3}andinstructionsreferences the message via a backtick path.
Part 3 — Primitives
- [3 pts] For each, pick Choice / Score / Noul and justify in one line: (a) "Which of our 6 product lines does this review complain about?" (b) "How drunk is this support chat text, on a scale we define?" (c) "Does this refund request cite a duplicate charge?"
- [4 pts] A Score has 4 levels and returns
probabilities = {0: 0.1, 1: 0.2, 2: 0.4, 3: 0.3}. Computescoreby hand. Which two levels is the model mostly "between"? - [4 pts] Write the Noul question (with
NoulCriteria) that decides "should we page the on-call engineer?" — such that a value of 0.5 is meaningfully interpretable.
Part 4 — Asking
- [3 pts] Refactor this into atomic questions: "Read this application log excerpt and decide whether we have a serious incident and what the on-call should do."
- [3 pts] You need "If the complaint is about billing and the customer is angry, offer a discount." A colleague proposes asking that whole conditional as one question. What's the better shape?
Part 5 — Confidence
- [4 pts] Compute Choice confidence for 4 options with
p_max = 0.55. Then for 2 options withp_max = 0.55. Explain in one sentence why the same top probability yields different confidence. - [4 pts] A Score has 3 levels and probabilities
(0, 0.5, 0.5). Compute the confidence using the Score formula. Then compute confidence for(0.5, 0, 0.5). Explain in one sentence why the Score formula distinguishes "torn between adjacent levels" from "torn between opposite ends".
Part 6 — Patterns
- [4 pts] Name the pattern: (a) 14 questions in one call, code reads 5 of them; (b) confidence > 0.9 → execute, else → confirm; (c) three Scores weighted 0.5/0.3/0.2 in Python.
- [4 pts] For the resume screener's two-step escalation, explain why it legitimately needs a second request when most "second calls" don't.
Part 7 — SDK
- [4 pts] Rewrite §7.1's sync example using
AsyncTypeSafeClientand additionally print each answer's confidence for Choice and Score.
Part 8 — Project
- [10 pts] Extend the resume screener: add an
education_fitScore, add a knockout Noul for "states required years of experience", and change weights to 0.35/0.35/0.15/0.15. Show the new composite formula. - [10 pts] Build it end-to-end for a real job posting and 3 real (anonymized) resumes. Report decisions + composite scores + the confidence story. Where did the model want a human?
Solutions Appendix
1. RLHF optimizes for responses humans prefer (chatbots). RLVR optimizes tasks with verifiable answers (reasoning models — strong, slow, expensive). RLCD optimizes for decisions with calibrated probabilities (Jev).
2. Structural difference: the LLM is a text generator constrained by prompting, while Jev is trained to return decisions and probabilities within an answer space you define. Failure mode JSON mode can't fix: nothing stops the model from emitting a value outside your option set (or prose around the JSON) — and when it does, you have no calibrated uncertainty to fall back on, only a parse error.
3. state (the content to judge — string, object, or array), questions (typed judgments with ID/type/instructions/criteria), and model (which System One model, default jev-latest). One request evaluates one state against one or more questions.
4. The ID is for your code's bookkeeping and is never sent to the model — the model only sees instructions (and criteria). A terse ID like refund with instructions: "refund" gives the model nothing to judge.
5. No single correct answer — the point is observing that every answer arrives as a number/option plus a distribution, not prose. If either answer felt vague (confidence below ~0.5), try rewording the question more precisely.
6.
state = {"message": "Hi, I've been trying to connect my Stripe account for 3 days...",
"account_age_days": 3}
# questions:
"is_urgent": Noul(instructions="Does `message` express urgency?")
7. (a) Choice — six unordered buckets; the answer routes to a product team. (b) Score — it's an ordered scale you define with level descriptions. (c) Noul — a yes/no claim about the message, where the probability is the useful signal.
8. score = 0×0.1 + 1×0.2 + 2×0.4 + 3×0.3 = 0 + 0.2 + 0.8 + 0.9 = 1.9 — mostly between levels 2 and 3 (0.4 + 0.3 = 0.7 of the mass), i.e. "above level 2, leaning below the top".
9.
"page_oncall": Noul(
instructions=("Does `log_excerpt` show an outage-class failure that is "
"currently affecting production traffic?"),
criteria={
"true": "Active errors/5xx affecting real user traffic right now, or a service down",
"false": "Warnings only, scheduled maintenance, or errors in a non-production component",
},
)
# code: page if noul > 0.8; ignore if noul < 0.2; review the 0.2–0.8 band
10. Decompose: is_customer_impacting (Noul) · error_rate_level (Score: none→spike) · component (Choice: api/db/auth/other) · is_recurring (Noul) · has_recent_deploy (Noul). The "what should on-call do" part is your code: if is_customer_impacting and error_rate_level >= 1.5: page().
11. Ask two independent questions — is_billing_complaint (Noul or Choice) and is_angry (Score) — and apply the conjunction in code: if complaint == "billing" and frustration.score > 1.2: offer_discount(). The conditional logic is deterministic and auditable; the model only makes the two judgments it's good at.
12. Choice confidence with n options and top probability p_max is (p_max − 1/n)/(1 − 1/n). For n=4: (0.55 − 0.25)/0.75 ≈ 0.40. For n=2: (0.55 − 0.5)/0.5 = 0.10. Same 0.55 is more impressive among four options because it sits further above the even split.
13. For (0, 0.5, 0.5): most likely level m = 1; spread = 0.5; MAD_unif = 2/3; confidence = 1 − 0.5/(2/3) = 0.25. For (0.5, 0, 0.5): m = 0; distance = 1.0; confidence = 1 − 1.0/(2/3) = 1 − 1.5, floored at 0. Adjacent-level disagreement is mild uncertainty; opposite-ends disagreement is maximal. The Choice formula ignores ordering and would give both 0.25.
14. (a) Speculative Fan-Out. (b) Confidence-Gated Routing. (c) Composite Scoring.
15. Because the second state doesn't exist yet until code acts on the first answer: the portfolio text must be fetched based on which candidates were flagged. That's the legitimate dependency test — the answer determines what the state is made of.
16.
import asyncio
from typesafe_sdk import AsyncTypeSafeClient, Choice, Noul, Score
async def main() -> None:
async with AsyncTypeSafeClient() as client:
r = await client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None}),
"urgency": Score(instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"]),
},
)
print(r.nouls["billing"].noul)
print(r.choices["tone"].choice, "conf:", r.choices["tone"].confidence)
print(r.scores["urgency"].score, "conf:", r.scores["urgency"].confidence)
asyncio.run(main())
17.
"education_fit": Score(
instructions=("Does `resume` show the education the role needs (see `job.must_have`)?"),
criteria=["No relevant degree or equivalent evidence",
"Related field, below requirement",
"Meets the requirement",
"Exceeds it (advanced degree or equivalent depth)"]),
"years_knockout": Noul(
instructions=("Does `resume` clearly evidence 5+ years of backend "
"experience, per `job.must_have[0]`?"),
criteria={"true": "Dates/scope clearly add to 5+ years of relevant work",
"false": "Dates don't add up, or evidence is missing"}),
New composite:
composite = (0.35 * skills.score + 0.35 * exper.score
+ 0.15 * nice.score + 0.15 * edu.score) / 3
Treat years_knockout like the authorization gate: noul < 0.2 → reject, else continue.
18. No single correct answer — grade on: one request per resume; all questions batched; knockout gates before composite; confidence-aware decisions logged; and an honest paragraph about low-confidence answers. If everything came back at confidence 1.0, your questions are probably too easy — sharpen the criteria.
Cheat Sheet
# SETUP
pip install typesafe-sdk # Python >= 3.10; or uv add typesafe-sdk
export TYPESAFE_API_KEY=... # console.typesafe.ai/keys
from typesafe_sdk import TypeSafeClient, AsyncTypeSafeClient, Choice, Score, Noul, NoulCriteria
# ONE CALL — state + questions → typed answers
with TypeSafeClient() as client:
r = client.system_one(
state=..., # str | dict | list (text only, English-best)
model="jev-latest", # optional; default
questions={
"id1": Choice(instructions="Which ...?", criteria={"opt": "desc", ...}),
"id2": Score(instructions="How ... on scale?", criteria=["low", ..., "high"]), # 2–10 levels
"id3": Noul(instructions="Is ... true?"), # → noul only
"id4": Noul(instructions="Is ... true?",
criteria=NoulCriteria(true="...", false="...")),
},
)
# READING ANSWERS (response.answers["id"] works for all kinds)
r.answers["id1"].choice; r.answers["id1"].probabilities; r.answers["id1"].confidence
r.answers["id2"].score; r.answers["id2"].legend; r.answers["id2"].confidence
r.answers["id3"].noul # no .confidence on Noul
r.nouls["id3"].noul; r.choices["id1"].choice; r.scores["id2"].score # typed collections
r.usage.input_tokens; r.usage.output_tokens
# ENDPOINT (raw)
POST https://api.typesafe.ai/v1/systemone Authorization: Bearer $TYPESAFE_API_KEY
# CONFIDENCE GATES
answer.confidence < 0.5 → don't act (human / clarify / fallback)
per-action thresholds → risky actions gate higher (0.9 for destructive)
noul > 0.8 act · noul < 0.2 ignore · in between → review band
Primitives: Choice = unordered buckets → match · Score = ordered spectrum → threshold · Noul = yes-probability → if
Patterns: Fan-out (batch everything) · Confidence-gated routing (2nd decision axis) · Composite scoring (weights in code) · Intent routing (Choice + dispatch)
Further Reading
- Introduction: https://docs.typesafe.ai/introduction
- Quick start: https://docs.typesafe.ai/introduction/quickstart
- Primitives: https://docs.typesafe.ai/primitives
- Confidence: https://docs.typesafe.ai/confidence
- Patterns: https://docs.typesafe.ai/patterns
- Cookbooks: https://docs.typesafe.ai/cookbooks
- Python SDK: https://docs.typesafe.ai/sdk/python
- Use-case map: https://docs.typesafe.ai/concepts/use-case-map
- Playground: https://console.typesafe.ai/playground
If you found this useful, the best next step is to open the Playground and ask your first Noul question. It takes about 30 seconds to feel the difference.
Top comments (0)