TL;DR: Jev (TypeSafe AI) and Laya (Convai Innovations) are decision models: instead of generating text, they answer bounded questions, a choice, a score, or a yes/no probability, that an application can act on directly. Jev is a managed API with natural-language-defined questions; Laya is open-weight under Apache 2.0 and can run on your own hardware. Both suit agent routing, ticket triage, and other classification jobs that don't need a full LLM response.
Two recent releases have put a familiar machine-learning task back in focus: making a bounded decision from messy information. Jev and Laya offer different ways to do it, and their differences are as useful as the idea they share.
TypeSafe AI introduced Jev on September 15 as the first of its "System One Models." Within 24 hours of its arrival on Vercel's AI Gateway, nearly 13% of paid teams had used it, making it the gateway's fastest-adopted model launch. That figure measures early interest rather than lasting demand, but it is still unusual attention for a model that cannot write a response to a user.
Within days, Convai Innovations released Laya, an open-weight decision model that accepts similar questions and can run on a developer's own hardware. The two models arrive at a useful moment for AI applications. Agents and other automated workflows repeatedly need to classify an input, assess a result or make decisions based on available data. Developers have often assigned those jobs to a large language model (LLM) because an LLM can understand context and follow instructions. The application may then discard nearly everything the LLM generated except a label or number.
Jev and Laya are purpose-built to make those decisions.
What a decision model does
Suppose a customer writes, "I've been charged twice, and I can't get anyone to respond." A support system might need to decide which team receives the ticket, how urgent it is and whether the customer appears likely to request a refund.
An LLM can answer those questions, perhaps returning a JSON-structured answer which an application can read. A decision model takes a different route: the application supplies the customer's message as state and defines the possible answers in advance. The model returns choices, scores and probabilities that the application can use directly. It does not generate a paragraph that must then be interpreted or reduced to a label.
Interactive demo: the original post on the Runware blog applies this scenario to all three of Jev's question types at once, comparing an LLM's single generated guess against Jev's returned distributions.
Both Jev and Laya support three question types:
- A
choiceselects from defined options, such as billing, technical support or sales. - A
scoreplaces the input on an ordered scale, such as levels of urgency. - A
noulestimates the probability that a yes-or-no statement is true.
Several questions can be submitted together against the same state. TypeSafe's API documents these primitives, while Laya exposes a closely related interface through its Python package and a Jev-compatible, self-hosted HTTP server.
Defining those questions is part of the work. In Every's account of a personal inbox filter, editor Jack Cheng asked Jev whether an email came from a person, whether someone he knew was waiting for him, and whether ignoring it would cost him something. The first version surfaced newsletters he did not consider urgent, so he revised the criteria against messages he had already judged. The filter improved when its questions more closely matched what "needs my attention" meant to him.
That interface makes the models particularly relevant to agent systems. An agent may need to select a tool, decide whether a result satisfies the user's request, route a task to a more capable model or flag an action for review. LangChain has demonstrated Jev for model routing and tool-risk checks, and added it as a judge in LangSmith Evals. These are bounded questions embedded in larger workflows, rather than requests for new prose.
"System One" is TypeSafe's name for this approach, borrowed from the idea of fast, intuitive judgment. It is a useful description of the intended role, though decision model is the broader term. Image and audio are modalities in the usual sense; a decision is a kind of output and a job within an application.
The idea has a longer history than the label
Models have classified text without generating it for decades. BERT helped make bidirectional encoders a standard tool for language understanding and classification. Research on zero-shot text classification showed that a model could evaluate labels described in language, allowing some classifications without training a new model for each set of labels. Reward models, moderation classifiers and model routers have since made other kinds of judgments from text.
LLMs became attractive for this work because they are flexible. A developer can describe a new task in a prompt instead of collecting examples and training a dedicated classifier. Strict structured outputs also address the problem of malformed JSON or unexpected fields. But the model still generates its answer token by token, even when the only useful result is one of three known options.
Jev's contribution is to package flexible, natural-language-defined questions as a dedicated decision service. TypeSafe says it uses a parallel sampler and a training method it calls Reinforcement Learning for Calibrated Decisions, or RLCD. It has not published enough architectural detail to independently assess all of its claims about the underlying model, but what developers can examine today is the interface, the observed behavior and its performance on their own tasks.
Laya makes a different part of the picture inspectable. Convai publishes its weights and code under the permissive Apache 2.0 license. Its English checkpoint uses ModernBERT-large with a decision head, while a separate checkpoint uses mmBERT for multilingual inputs. Developers can run the models locally, examine the implementation and fine-tune a checkpoint for their own decisions. Laya also offers a specialized checkpoint trained for four synthetic workflow types.
An application can frame its questions in much the same way for either model, while the choice of model changes who operates it, how it can be adapted and which limitations the developer must manage.
What the early results tell us
TypeSafe reports response times of roughly 70 to 500 milliseconds for Jev. Its most dramatic speed comparisons come from selected workflows. TypeSafe acknowledges that those examples use short, dense inputs that favor Jev and represent the high end of the gains it expects developers to see.
Convai reports local inference times of 39.5 milliseconds for one question on Laya's English checkpoint and 32.8 milliseconds on its multilingual checkpoint, measured on a Tesla T4 GPU. Ten questions in a batch took 158.6 and 72.3 milliseconds respectively. Those numbers describe warm local model inference on specified hardware. Comparing them directly with a hosted Jev request, which includes service and network time, would overstate what they tell us about the models themselves.
Accuracy is more revealing than either headline speed. On a synthetic benchmark covering customer service, invoices, security incidents and agent traces, Laya's specialized checkpoint reports 76.6% accuracy across 2,000 decisions. Its general English checkpoint scored 36.2% on the same test, below the benchmark's 46.1% majority-class baseline. The specialized checkpoint was fine-tuned on that benchmark's training split, so its result shows the value of adaptation to those workflows. It does not establish that the general Laya model can handle an unfamiliar decision equally well.
The models have different practical limits. Jev's API accepts up to 255 options for a choice question. Laya's options share a fixed input budget, and its documentation advises care above roughly 20 options at default settings. Its English checkpoint also defaults to a 512-token context, with 1,024 for the multilingual and specialized checkpoints. Those limits may be unimportant for a short support ticket and a few routes, but they could determine the design for a long document or dozens of closely related categories.
Laya's multilingual model is another useful option, but language coverage needs testing rather than assuming every supported language performs equally. Convai's own evaluation found that the English checkpoint could remain highly confident while failing on some non-English inputs. Its router attempts to select the appropriate checkpoint before inference. The model cards also warn that shipped probabilities can be overconfident and recommend calibration on held-out data before using them to control consequential actions.
A valid answer can still be wrong
Both projects describe a benefit of constrained output: the model cannot add an invented fourth option to a three-option question or fabricate a citation in a free-text explanation. TypeSafe markets this as Jev being unable to hallucinate, and Runware staff engineer Félix Sanz has examined what that claim actually covers: restricted to distributing probability across a fixed set of answers, Jev cannot invent an option or a fact outside that space. What it removes is the hallucination that comes with free-form generation, not errors in judgment. A confidently wrong choice from the options it was given is still wrong, and that class of error remains.
If a ticket belongs with billing and the model confidently routes it to sales, the response is perfectly valid at the API level and wrong at the business level. A probability is useful only when its relationship to actual outcomes has been checked on the kind of input the application receives. Laya's published calibration caveats make this especially explicit, while TypeSafe advises setting thresholds according to the consequences of each decision.
Sanz also found a practical example of how much the question itself matters. On a synthetic phishing-email dataset, Jev reached 62.6% accuracy when asked directly whether a recipient should click a link, compared with 81.3% for Claude Haiku 4.5. When the task was split into five specific signals and ordinary code combined them, Jev reached 95% on held-out data. Simple domain-based rules already reached 91.8% on that dataset. The signals were designed with knowledge of the synthetic data, so this is an engineering example rather than proof of production phishing performance. It shows why testing the whole decision process matters more than timing an isolated model call.
Code should still enforce exact requirements such as permissions, price ceilings and account balances. A decision model is useful where the input requires interpretation but the possible outcomes are known. A generative model remains useful when the application needs to create an explanation, draft a response or reason through an open-ended problem.
Where Jev and Laya fit
Jev offers a managed way to try decision models without operating a checkpoint. Its broad option support and natural-language-defined questions make it a reasonable candidate for routing, scoring and evaluation tasks that change frequently. Laya offers control over deployment and training. It is especially interesting when data must stay within an organization's environment or when a team has examples it can use to specialize a model for a repeated task. Its open weights also make the implementation available for inspection and further work.
Neither choice removes the need for an application-level evaluation. A useful test set should contain real inputs, expected outcomes, ambiguous cases and examples where the right answer is to escalate. Measure decision quality, calibration and total request latency. Compare the result with a small LLM using structured output and with ordinary rules. For some workloads, one of those simpler options will win.
Runware has added both Jev and Laya to give developers that choice.
Implementing Jev and Laya with Runware's API
Both models are available through Runware's /v1/systemone endpoint, using your Runware API key. Reference each model by its AIR:
| Model | AIR | Description |
|---|---|---|
| Jev (TypeSafe AI) | typesafe:jev@latest |
Managed decision model. Returns choices, scores and yes/no probabilities for questions defined in natural language. |
| Laya (Convai Innovations) | runware:laya@1 |
Open-weight decision model under Apache 2.0, answering the same three question types. |
A request pairs a state, the material being judged, with one or more questions. Each question uses one of the three primitives described above: a choice, a score or a noul. Because the application defines the possible answers in advance, the response returns probabilities and confidence values that code can act on directly, with no generated text to parse.
Pricing
Prices are in USD per 1M tokens.
| Model | Input | Output |
|---|---|---|
| Jev (TypeSafe AI) | $0.042 | Free |
| Laya (Convai Innovations) | $0.020 | Free |
Neither model charges for output tokens. Laya is free until October 12, 2026, after which the input price above applies.
Decision model concepts glossary
Decision models give developers a new kind of model output to work with: a bounded judgment that software can act on. That brings a vocabulary of its own. A developer supplies the state, defines the question and its possible answers, then decides how much uncertainty the application can tolerate. The concepts below follow that path from input to decision.
| Concept | Description |
|---|---|
| Decision model | A model that evaluates supplied information and returns a bounded judgment, such as a category, a position on a scale, or the probability of a yes answer. The application defines what answers are allowed. |
| System One model | TypeSafe's name for its class of fast decision models. It accepts natural-language context but returns typed decisions and probabilities instead of generated prose. Jev is its first model. |
| State | The material being judged. It can be a message, a passage, or a structured collection of related facts. In a request with several questions, each question sees the same state. |
| Question | One specific judgment the model is asked to make about the state. Questions in the same request are evaluated independently, so one answer does not automatically inform another. |
| Question ID | A key chosen by the application to identify an answer in the response. The ID helps code retrieve the answer; the model reads the question's instructions, not the ID. |
| Instructions | The words that state what should be judged. Clear, narrow instructions make a decision easier to interpret and test. |
| Criteria | The definition of the answer space: named options for Choice, ordered descriptions for Score, or optional definitions of yes and no for Noul. |
| Primitive | A typed question and answer form. TypeSafe defines three: Choice for categories, Score for ordered levels, and Noul for a yes-or-no probability. |
| Choice | Selects one item from a fixed set of options with no inherent order, such as billing or technical support. The answer also includes the probability assigned to each option and a confidence value. |
| Score | Places a case on an ordered scale whose levels the developer describes. The returned number can fall between levels because it is calculated from the probabilities assigned to them. |
| Level | A described point on a Score scale. Levels run from low to high and are numbered from zero. |
| Legend | The mapping returned with a Score answer from each level number back to its description. It makes a numeric score readable in the context of that question. |
| Noul | Answers a yes-or-no question with a number from 0 to 1: the model's probability of yes. Around 0.5 means the two outcomes are similarly likely. It does not measure the degree of a quality and has no separate confidence field. |
| Probabilities | The model's distribution across all Choice options or Score levels. Those values sum to 1. For Noul, the returned value itself is the probability of yes. |
| Confidence | A 0-to-1 summary of how concentrated a Choice or Score probability distribution is. A clear peak gives higher confidence; a spread-out distribution gives lower confidence. It is not a guarantee that the answer is correct. |
| Calibration | How closely reported probabilities match actual outcomes across many examples. If events given a probability near 0.8 occur about 80% of the time, those predictions are well calibrated as a group. |
| Threshold | A cutoff the application sets for acting on a probability or confidence value. The appropriate cutoff depends on the cost of a wrong decision and should be checked against real examples. |
| Escalation | Sending an uncertain or consequential case to a person, another model, or a further check instead of acting automatically. It is a workflow decision made around the model's answer. |
| Composite scoring | Breaking a broad judgment into separate Score questions, then combining the results with weights chosen in application code. This makes the contribution of each factor visible. |
Originally published on the Runware blog.
Top comments (0)