AI agents are becoming increasingly complex. They often need to decide whether a message is urgent, select a model to process a task, or assess the risk of an operation all of which are essentially classification and scoring tasks yet they are frequently handled by invoking a costly large language model (LLM).
Jev by TypeSafe AI is designed to solve exactly this problem. It's not a traditional LLM, it's a System One decision model: you send a state and a set of typed questions, and it returns structured answers with calibrated probabilities, without generating any text.
But a question follows: how does this actually work inside a real agent? In this post, we'll walk through a concrete open-source example JevSample to show how Jev fits into a LangGraph.js agent, handling routing and risk assessment.
The problem with making the LLM decide everything
A typical AI agent might look like this:
The LLM is responsible for almost everything.
It generates content, decides what tool to call, evaluates the result, decides whether it should continue, and sometimes even makes safety decisions.
That can work, but it also creates a problem.
Some of these decisions have a relatively small and well-defined answer space.
For example:
Should this request be allowed? YES / NO
or:
What should happen next? EXECUTE / REVIEW / VERIFY / RETRY
or:
How risky is this operation? LOW / MEDIUM / HIGH
These are structured decisions.
We don't necessarily need another generative response from an LLM for every one of them.
Jev is designed around this kind of problem. Instead of asking it to generate a paragraph, we provide a state and ask a typed question. Jev supports three main question types: Choice, Score, and Noul.
A Choice selects from predefined options, a Score evaluates an ordered scale, and a Noul produces a probability for a yes/no statement.
That makes the output much easier for application code to consume.
Introducing JevSample
JevSample is a small example of an AI agent built with LangGraph.
The interesting part isn't the size of the project. It's the architecture.
The sample separates three different responsibilities:

The agent graph contains several nodes:

This gives the agent a much more explicit control flow.
Building the graph
The central graph construction in JevSample looks like this:
export function buildGraph(deps) {
return new StateGraph(GraphState)
.addNode('router', routerNode(deps))
.addNode('generator', generatorNode(deps))
.addNode('prefilter', prefilterNode(deps))
.addNode('safety', safetyNode(deps))
.addNode('dispatch', dispatchNode(deps))
.addNode('human_review', humanReviewNode(deps))
.addNode('executor', executorNode(deps))
.addNode('verifier', verifierNode(deps))
// ...
.compile({ checkpointer: deps.checkpointer
});
}
The important thing here is that the graph doesn't represent one simple linear pipeline. It represents a state machine.
The agent can move forward, stop, ask for human input, execute an action, verify the result, or return to the generator and try again.
That's an important characteristic of real-world agents.
Where Jev fits
There are two particularly interesting places where Jev can be used in this sample.
The first is the router.
The router can make a structured decision about which path the agent should take.
Conceptually:

Instead of asking an LLM to produce something like:
"I think we should probably execute the request..."
and then parsing that text, the application can ask Jev a Choice question whose possible results are already defined.
For example:
Which workflow should handle this request?
- generate
- verify
- execute
- review
The application can then directly use the returned choice to select the graph transition.
This is a much cleaner contract.
Jev for safety decisions
The second interesting use is the safety node.
The flow is:

The generator produces the candidate result.
The prefilter performs deterministic checks in application code.
Then Jev evaluates the higher-level safety decision.
This separation is important.
The prefilter doesn't need to be an AI model at all.
For example, deterministic checks might handle:
if (!result) {
return "blocked";
}
if (containsInvalidStructure(result)) {
return "blocked";
}
Jev can then handle a judgment that is harder to express with a simple rule.
A Noul question is a natural fit for this kind of decision:
Is this generated result safe to execute?
The result is a value between 0 and 1 representing the probability that the statement is true.
The application can then define its own policy.
For example:
if (safety.noul >= 0.9) {
return "execute";
}
if (safety.noul >= 0.6) {
return "review";
}
return "blocked";
The important point is that Jev doesn't execute anything.
It provides the judgment.
The application remains responsible for deciding what that judgment means operationally.
Why the prefilter still matters
It might be tempting to put everything into Jev.
I don't think that's a good idea.
There are many things that ordinary code does better.
For example:
if (!input) {
return "invalid";
}
There is no reason to use an AI model for this.
Likewise:
if (result.length > MAX_LENGTH) {
return "blocked";
}
is a deterministic rule.
The architecture therefore becomes:

Each component has a different job.
This is one of the main ideas demonstrated by JevSample.
The dispatch stage
After the safety decision, the dispatch node determines what happens next.
The possible paths include:

This is where the value of structured decisions becomes particularly obvious.
The dispatcher doesn't need to interpret a paragraph produced by a model.
It can work with explicit states.
For example:
switch (decision) {
case "execute":
return "executor";
case "review":
return "human_review";
case "verify":
return "verifier";
default:
return "dispatch";
}
The AI model is responsible for judgment.
The application is responsible for control flow.
Human review is still part of the architecture
Another important part of the sample is human_review.
An AI agent shouldn't necessarily make every decision automatically.
Some situations should become:

This is especially useful when a decision has consequences that are too important to automate completely.
The architecture therefore doesn't treat Jev as an authority.
Instead:

That distinction is important when designing production agents.
Verification creates the feedback loop
After execution, the sample doesn't simply assume that everything worked.
The verifier checks the result.
Conceptually:

If the task is complete, the graph terminates.
If it isn't, the agent can return to the generator and try another approach.
This creates the familiar agent loop:

The important difference is that the decisions in this loop don't all have to be made by the same model.
Jev is not an LLM replacement
This is probably the most important point of the example.
Jev and an LLM have different jobs.
An LLM is useful when we need:
- text generation
- reasoning
- summarization
- code generation
- explanations
- open-ended problem solving
Jev is useful when we have a decision that can be expressed as:
Which option?
How severe?
Is this true?
That gives us a useful architecture:

Instead of asking one model to do everything, we can choose the right tool for each part of the workflow.
Choice, Score and Noul
Jev's three question types map nicely to different parts of an agent.
Choice
Use Choice when the answer is one of several predefined options.
For example:
Which workflow should handle this request?
generate
execute
review
verify
The application can directly branch on the returned value.
Score
Use Score when the answer belongs on an ordered scale.
For example:
How risky is this operation?
low
medium
high
critical
This is useful when the application needs a threshold rather than a simple category.
For example:
if (risk.score >= HIGH_RISK_THRESHOLD) {
return "review";
}
Noul
Noul is useful for a yes/no question where the probability itself is useful.
For example:
Is this generated operation safe to execute?
The result can then be used as part of an application policy.
The important distinction is that a Noul value isn't simply a replacement for a boolean.
A value around 0.5 means that the model is uncertain. Your application needs to decide how uncertainty should be handled.
For an agent, that can naturally lead to:
Why this architecture is interesting
The interesting part of JevSample isn't simply that it calls Jev.
It's that Jev becomes one component in a larger system.
The architecture separates:
Generation
Handled by the LLM.
"What should we produce?"
Decisions
Handled by Jev.
"Which path should we take?"
"Is this safe?"
"Is this complete?"
Deterministic rules
Handled by application code.
"Does this value have the correct format?"
"Is it within the allowed range?"
Actions
Handled by tools and application code.
"Execute the operation."
Oversight
Handled by humans when necessary.
"Approve or reject this action."
That separation makes the agent easier to reason about.
A small mental model
When designing an AI agent, I find it useful to ask five questions:
1. What needs to be generated?
2. What needs to be decided?
3. What can be checked deterministically?
4. What action should be executed?
5. Where does a human need to be involved?
The answers can then map to different components:
This is the architectural idea that JevSample is intended to demonstrate.
Summary
AI agents are becoming more complex.
As soon as an agent has multiple tools, retries, verification, safety checks, and human approval, it becomes more than an LLM wrapped in a loop.
It becomes a workflow.
And workflows require decisions.
Jev provides an interesting way to make some of those decisions explicit and structured, rather than asking a generative model to produce text and then parsing that text back into a decision.
This doesn’t mean that every decision should be moved to Jev.
The best architecture is usually a combination of deterministic code, generative models, decision models, tools, and human oversight.
The most interesting aspect of Jev is that it gives an AI agent another specialized component instead of trying to make one model responsible for everything.



Top comments (0)