Most "AI" calls in production aren't conversations. They're decisions: which queue should this ticket go to, is this action safe, how urgent is this lead. We still send them to generative LLMs and then parse JSON that sometimes arrives with a chatty preamble or an invented field.
Jev, released by TypeSafe AI on September 15, 2026, takes a different route: it can't write text at all. It reads a state and a schema, and returns typed answers with calibrated probabilities in 70 to 500 ms. Here is how it works, where it saves money, and where it breaks.
Why decisions don't need generation
Even with constrained decoding, an LLM builds a structured answer one token at a time. A 60-token JSON routing decision means 60 sequential passes through the model. And the logprobs it exposes describe how likely the next token is, not how likely the decision is to be correct.
Jev is a prefill-only model. The whole request goes through the transformer once, and specialized heads score the options in parallel. It is trained on synthetic data with a method TypeSafe calls RLCD, which penalizes miscalibrated probabilities instead of rewarding answers that humans like. As a result, asking ten questions about the same state costs barely more than asking one.
Three primitives
• Choice: pick one label from a set. Returns the winner, the full probability map and a confidence value.
• Score: place something on an ordered rubric, such as severity or frustration. Returns a score, probabilities and confidence.
• Noul: check whether a statement is true. Returns one probability between 0 and 1, with no separate confidence, because the probability is the certainty.
Pricing is $0.042 per million input tokens, and output tokens are free.
Is it really 40-400x cheaper?
TypeSafe markets that range, so let's test it. Take 1 million classifications a month, at about 400 input and 50 output tokens each. Jev costs roughly $16.80. A mini-class LLM costs about $90, and a frontier model about $1,950.
That's 116x cheaper than a frontier model but only 5.4x cheaper than a small one. Most teams already route high-volume, low-context tasks to small models, so the realistic saving is closer to the second number. Note also that TypeSafe's speed claims come from internal benchmarks it admits may be biased.
Where it fits
• Routing before an agent runs. Send docs questions to search, numeric ones to SQL and creative ones to an LLM in under 150 ms. This keeps expensive models out of the loop and is a natural building block for AI agent orchestration.
• Safety gates. Ask Noul whether a proposed command could drop data or leak credentials, and block it above a threshold like 0.15.
• Output guardrails. Check generated text for PII, secrets and tone without a second LLM as judge.
• Ticket triage. Department, frustration level and executive urgency come from one call.
• Lead scoring. Score company size, buyer authority and urgency separately, then combine them in code with coefficients. Business priorities change without prompt rewrites, which makes this a good fit for AI automation workflows.
• RAG filtering and reranking. Screen 20 retrieved chunks in parallel. On TypeSafe's own legal retrieval benchmark, reranking lifted top-1 precision from 5% to 18% and top-10 from 38% to 62%.
Where it breaks
TypeSafe documents 9 failure modes. The ones that matter most:
• Literal reading. Unstated implications are invisible, so spell out edge cases in the criteria.
• No arithmetic, counting or date logic. Do the math in code. To count, ask a yes/no question per item and tally in software.
• Multi-hop relations. Flatten nested permissions or relationships in code before sending them.
• Diluted context. A huge state full of irrelevant prose degrades decisions, so send only the relevant chunks.
• Prompt injection. Malicious text inside the state can override your instructions, so sanitize inputs.
• Option order. Probabilities can shift when you reorder choices. Sort options deterministically.
• No generation. If you need an explanation, use an LLM.
The pattern that works: deterministic prep in code, the decision model in the middle, and confidence thresholds on top. For example, execute at 0.85 or higher, fall back to a generative model between 0.50 and 0.85, and send anything lower to human review.
Alternatives that appeared within weeks
• Cloudflare Clef and Clef-flash (Oct 1): open weights under Apache 2.0, the same API shape as Jev, plus support for up to four images. Clef-flash reports a median of about 39 ms on Workers AI. Its benchmark wins over Jev are self-reported.
• Convai Laya (around Sept 22): open source, about 421M parameters, runs locally in roughly 1 GB of RAM without a GPU, with 9-33 ms latency.
• OpenAI Decisions API (announced Sept 29): no public docs or pricing as of early October. One independent test reported 403 errors on standard keys and about 1.46 s median when simulated with structured outputs, versus the 150 ms OpenAI claims.
Critics point out that classifiers like DeBERTa and SetFit already existed, and for stable, high-volume tasks with fixed labels a fine-tuned small encoder is still cheaper and more accurate. What's new is a zero-shot, typed, calibrated API with no training pipeline.
Takeaway
For classification, routing, scoring and policy checks at volume, a decision model can cut latency and spend. Keep an LLM for anything that needs text, multi-step reasoning or explanations. The real engineering work is the wiring around it: confidence gates, fallbacks and deterministic preprocessing. That's where solid AI integration into your existing stack pays off.
Top comments (0)