Most automation pipelines do not need a paragraph. They need a verdict.
Is this support ticket about billing or about login. Is this comment a policy violation. Does this retrieved chunk actually contain the fact required to answer the question. Each of those is a discrete judgment with a small, known set of outcomes, and most teams currently route them to a general purpose language model that was built to write prose.
Jev, from TypeSafe AI, takes the opposite position. It is a System One structured decision model built for machine consumption, and it removes text generation from the architecture entirely. Nothing it returns is a sentence.
A decision model, not a smaller language model
Where the name comes from
The naming follows the System 1 and System 2 split in cognitive psychology, the fast intuitive judgment and the slow deliberate one. Jev occupies the fast side. It inherits the lightweight semantic classification work that currently sits on fragile rule sets or on an expensive model, and it does that work with a different training objective. There is no autoregressive text generation anywhere in the loop. The model is trained with parallel sampling and reinforcement learning for calibrated decisions, RLCD.
Why generation had to go
Enterprise software architecture rarely consumes continuous generated strings. What the automation chain actually needs is a discrete judgment that is deterministic, type safe, and carries a calibrated probability. A pipeline that branches on an intent label has no use for a fluent explanation of that intent, and every token spent generating one is latency and cost added to a decision that could have been a single forward pass.
The three primitives and the type contract
Choice, Score, Noul
Three primitives cover the decision shapes that appear most often in production code.
Choice performs single selection from a set, with support for up to 255 candidates. This replaces the router pattern where a model is asked to output one label from a long list.
Score performs ordered grading and degree scoring across 2 to 10 levels. This replaces the pattern where a model is asked for a number between one and five and then asked to explain itself.
Noul performs binary yes or no probability judgment and returns a value between 0.0 and 1.0. This replaces the pattern where a model is asked to answer yes or no and the surrounding code has to guess whether the answer is firm.
The return shape is fixed before the call
The calling mechanism is a single forward pass. The caller supplies the program state and a typed question, and receives a typed data structure with probability and confidence scores attached. Because the returned structure is strictly defined at request time, JSON syntax parse errors and type hallucination are ruled out at the level of the mathematical structure rather than caught afterwards by a schema validator.
The operating numbers matter as much as the shape. Average end to end latency sits in the 70 to 500 millisecond range. Input is billed at roughly 0.042 USD per million tokens, and decision output is not counted toward token charges. For a classification step that fires on every incoming request, that combination changes what is affordable.
Where generation-first automation breaks down
Format drift and type hallucination
Autoregressive models emit one token at a time, and output format drifts as a consequence of how they are built. Adding a JSON Schema check catches some of the damage but not all of it. Syntax parse errors and type hallucination still reach production, and the retry logic written to contain them becomes part of the latency budget.
Latency outside the real-time budget
A general purpose model typically lands somewhere between 3 and 30 seconds end to end. Real-time request paths and interactive tooling cannot absorb that. The workaround is usually to move the judgment off the critical path, which is another way of saying the automation is no longer automatic.
Rules and large models are both the wrong tool
At one end, brittle rule sets that break on phrasing nobody anticipated. At the other, a frontier model invoked for a three-way classification. For lightweight semantic classification, neither option pays for itself, and teams end up maintaining both.
Agent tool routing and RAG context bloat
Two specific failure modes show up repeatedly in agent stacks. The first is tool selection. Handing dozens of tool definitions to an expensive model in a single request produces a large context bill and frequent wrong picks. The second is retrieval. The top fragments returned by a vector search can be semantically generic and irrelevant to the question, and feeding them into a generation call introduces hallucination risk while raising cost.
Production patterns that hold up
Ticket triage and severity
A support pipeline can use Choice to assign the business group, Score to quantify how frustrated the customer is, and Noul to decide whether the customer's core business is down. All three judgments complete inside 200 milliseconds, which is fast enough to sit directly in the dispatch path.
Moderation funnel
Community content review uses Noul to probe violation probability, then routes on that score as a tiered funnel. Cheap probability first, expensive human or model review only for the band that needs it.
Lead scoring
Sales qualification maps well onto the primitives. Score from 1 to 5 on purchase intent, Choice to match an industry solution, Noul to verify whether a budget is explicitly stated.
RAG fragment filtering
Before retrieved chunks reach a generation call, Noul judges whether each chunk contains the factual information necessary to answer the question. Low scoring chunks are dropped in local application code, so the expensive model only ever sees context that survived a filter.
Calling conventions worth adopting
Parallel speculative questions
A single request asking one question and a single request asking ten questions take roughly the same amount of time, and the output is not billed. The practical consequence is that every judgment the downstream chain will need should be packaged into one call up front rather than issued as a series of dependent round trips.
Confidence gating by risk
Thresholds should follow business risk rather than a single global cutoff. Low risk read queries can pass automatically at a base confidence level. High risk write operations, the ones that change data or move assets, should demand a very high confidence before proceeding without review. This is where RLCD calibration earns its place. A general model reporting its own certainty fluctuates and skews overconfident, while a calibrated probability distribution supports actual statistical thresholds.
Tool selection in front of the large model
Jev can act as the routing gate for tool use. It picks the single label from the tool inventory first, and the large model receives only that one tool definition for parameter extraction. Context cost drops and the wrong-tool failure mode mostly disappears.
Four limits to design around
Literal reading only
The model performs probability inference strictly on the words given. It does not invent premises that were never written into the prompt. When classification output looks wrong, the first thing to check is whether the instruction supplied mutually exclusive definitions for each category.
No arithmetic and no counting
Symbolic arithmetic and entity counting are outside its capability. To count occurrences, the business code has to iterate and call Noul per item, then sum the results locally.
No temporal reasoning
Date strings are treated as ordinary characters. Time zone conversion and chronological ordering must be computed in code before the values are passed into the State payload.
Long input dilutes attention
Padding State with unrelated text drags down judgment accuracy. Text cleaning before the call is part of using the model correctly, not an optional optimization.
Jev against a general purpose LLM
Comparison
| Dimension | General purpose LLM | Jev |
|---|---|---|
| Output form | Free text, validated afterwards with JSON Schema | Native strongly typed discrete data |
| Response time | 3 to 30 seconds | 70 to 500 milliseconds |
| Billing | Input and output both charged by token, long context is costly | 0.042 USD per million input tokens, output not charged |
| Sampling method | Serial token by token generation | Parallel forward pass |
| Confidence | Self reported, fluctuates, leans overconfident | Probability distribution calibrated by RLCD, statistically usable |
| Suitable tasks | Text generation, code writing, multi-step reasoning, open dialogue | State routing, tool routing, intent classification, conditional filtering |
A three-layer division of labor
The arrangement that holds up is local ordinary code handling exact computation and data movement, Jev handling millisecond natural language probability decisions, and a general model handling deep content work and complex reasoning. Each layer is used for what it is structurally good at, and none of them is asked to cover for another.
Runtime requirements and setup
SDKs and runtimes
Official Python and Node.js SDKs are available, and both synchronous and asynchronous calls are supported. The runtime floor is Python 3.10 or above, and Node.js 20 or above.
Key handling
The API key belongs in an operating system environment variable rather than in source. The official SDK reads TYPESAFE_API_KEY by default, so no credential string needs to appear in the codebase.
If the runtime is the annoying part
For anyone who does not want to hand manage runtime versions, ServBay's Software Packages panel installs the required Python and Node.js instances in one click, which removes environment preparation from the setup path.
ServBay is a one stop AI development management tool that runs on macOS and Windows. Teams that already have a large volume of yes or no questions, score this, and which category does this belong to scattered through their stack can pull a batch of those judgments out of the large model and hand them to Jev, then watch what happens to latency and token spend.
Top comments (0)