DEV Community

Cover image for Building an AI Agent Decision Layer with Jev and LangGraph
Peter Saktor
Peter Saktor

Posted on

Building an AI Agent Decision Layer with Jev and LangGraph

AI agents are becoming increasingly complex. They often need to decide whether a message is urgent, select a model to process a task, or assess the risk of an operation all of which are essentially classification and scoring tasks yet they are frequently handled by invoking a costly large language model (LLM).

Jev by TypeSafe AI is designed to solve exactly this problem. It's not a traditional LLM, it's a System One decision model: you send a state and a set of typed questions, and it returns structured answers with calibrated probabilities, without generating any text.

But a question follows: how does this actually work inside a real agent? In this post, we'll walk through a concrete open-source example JevSample to show how Jev fits into a LangGraph.js agent, handling routing and risk assessment.

The problem with making the LLM decide everything

A typical AI agent might look like this:

Agent flow

The LLM is responsible for almost everything.

It generates content, decides what tool to call, evaluates the result, decides whether it should continue, and sometimes even makes safety decisions.

That can work, but it also creates a problem.

Some of these decisions have a relatively small and well-defined answer space.

For example:

Should this request be allowed? YES / NO
Enter fullscreen mode Exit fullscreen mode

or:

What should happen next? EXECUTE / REVIEW / VERIFY / RETRY
Enter fullscreen mode Exit fullscreen mode

or:

How risky is this operation? LOW / MEDIUM / HIGH
Enter fullscreen mode Exit fullscreen mode

These are structured decisions.

We don't necessarily need another generative response from an LLM for every one of them.

Jev is designed around this kind of problem. Instead of asking it to generate a paragraph, we provide a state and ask a typed question. Jev supports three main question types: Choice, Score, and Noul.

A Choice selects from predefined options, a Score evaluates an ordered scale, and a Noul produces a probability for a yes/no statement.

That makes the output much easier for application code to consume.

Introducing JevSample

JevSample is a small example of an AI agent built with LangGraph.

The interesting part isn't the size of the project. It's the architecture.

The sample separates three different responsibilities:
different responsibilities
The agent graph contains several nodes:
agent nodes
This gives the agent a much more explicit control flow.

Building the graph

The central graph construction in JevSample looks like this:

export function buildGraph(deps) {
 return new StateGraph(GraphState)
 .addNode('router', routerNode(deps))
 .addNode('generator', generatorNode(deps))
 .addNode('prefilter', prefilterNode(deps))
 .addNode('safety', safetyNode(deps))
 .addNode('dispatch', dispatchNode(deps))
 .addNode('human_review', humanReviewNode(deps))
 .addNode('executor', executorNode(deps))
 .addNode('verifier', verifierNode(deps))
 // ...
 .compile({ checkpointer: deps.checkpointer
 });
}
Enter fullscreen mode Exit fullscreen mode

The important thing here is that the graph doesn't represent one simple linear pipeline. It represents a state machine.
The agent can move forward, stop, ask for human input, execute an action, verify the result, or return to the generator and try again.
That's an important characteristic of real-world agents.

Where Jev fits

There are two particularly interesting places where Jev can be used in this sample.

The first is the router.
The router can make a structured decision about which path the agent should take.
Conceptually:

router
Instead of asking an LLM to produce something like:

"I think we should probably execute the request..."
Enter fullscreen mode Exit fullscreen mode

and then parsing that text, the application can ask Jev a Choice question whose possible results are already defined.

For example:

Which workflow should handle this request?

- generate
- verify
- execute
- review
Enter fullscreen mode Exit fullscreen mode

The application can then directly use the returned choice to select the graph transition.

This is a much cleaner contract.

Jev for safety decisions

The second interesting use is the safety node.

The flow is:

decisions flow
The generator produces the candidate result.

The prefilter performs deterministic checks in application code.

Then Jev evaluates the higher-level safety decision.

This separation is important.

The prefilter doesn't need to be an AI model at all.

For example, deterministic checks might handle:

if (!result) {
  return "blocked";
}

if (containsInvalidStructure(result)) {
  return "blocked";
}
Enter fullscreen mode Exit fullscreen mode

Jev can then handle a judgment that is harder to express with a simple rule.

A Noul question is a natural fit for this kind of decision:

Is this generated result safe to execute?
Enter fullscreen mode Exit fullscreen mode

The result is a value between 0 and 1 representing the probability that the statement is true.

The application can then define its own policy.

For example:

if (safety.noul >= 0.9) {
  return "execute";
}

if (safety.noul >= 0.6) {
  return "review";
}

return "blocked";
Enter fullscreen mode Exit fullscreen mode

The important point is that Jev doesn't execute anything.

It provides the judgment.

The application remains responsible for deciding what that judgment means operationally.

Why the prefilter still matters

It might be tempting to put everything into Jev.

I don't think that's a good idea.

There are many things that ordinary code does better.

For example:

if (!input) {
  return "invalid";
}
Enter fullscreen mode Exit fullscreen mode

There is no reason to use an AI model for this.

Likewise:

if (result.length > MAX_LENGTH) {
  return "blocked";
}
Enter fullscreen mode Exit fullscreen mode

is a deterministic rule.

The architecture therefore becomes:

architecture
Each component has a different job.

This is one of the main ideas demonstrated by JevSample.

The dispatch stage

After the safety decision, the dispatch node determines what happens next.

The possible paths include:

paths
This is where the value of structured decisions becomes particularly obvious.

The dispatcher doesn't need to interpret a paragraph produced by a model.

It can work with explicit states.

For example:

switch (decision) {
  case "execute":
    return "executor";

  case "review":
    return "human_review";

  case "verify":
    return "verifier";

  default:
    return "dispatch";
}
Enter fullscreen mode Exit fullscreen mode

The AI model is responsible for judgment.

The application is responsible for control flow.

Human review is still part of the architecture

Another important part of the sample is human_review.

An AI agent shouldn't necessarily make every decision automatically.

Some situations should become:

situations
This is especially useful when a decision has consequences that are too important to automate completely.

The architecture therefore doesn't treat Jev as an authority.

Instead:

architecture
That distinction is important when designing production agents.

Verification creates the feedback loop

After execution, the sample doesn't simply assume that everything worked.

The verifier checks the result.

Conceptually:

verifier
If the task is complete, the graph terminates.

If it isn't, the agent can return to the generator and try another approach.

This creates the familiar agent loop:

agent loop
The important difference is that the decisions in this loop don't all have to be made by the same model.

Jev is not an LLM replacement

This is probably the most important point of the example.

Jev and an LLM have different jobs.

An LLM is useful when we need:

  • text generation
  • reasoning
  • summarization
  • code generation
  • explanations
  • open-ended problem solving

Jev is useful when we have a decision that can be expressed as:

Which option?

How severe?

Is this true?
Enter fullscreen mode Exit fullscreen mode

That gives us a useful architecture:

architecture
Instead of asking one model to do everything, we can choose the right tool for each part of the workflow.

Choice, Score and Noul

Jev's three question types map nicely to different parts of an agent.

Choice

Use Choice when the answer is one of several predefined options.

For example:

Which workflow should handle this request?

generate
execute
review
verify
Enter fullscreen mode Exit fullscreen mode

The application can directly branch on the returned value.

Score

Use Score when the answer belongs on an ordered scale.

For example:

How risky is this operation?

low
medium
high
critical
Enter fullscreen mode Exit fullscreen mode

This is useful when the application needs a threshold rather than a simple category.

For example:

if (risk.score >= HIGH_RISK_THRESHOLD) {
  return "review";
}
Enter fullscreen mode Exit fullscreen mode

Noul

Noul is useful for a yes/no question where the probability itself is useful.

For example:

Is this generated operation safe to execute?
Enter fullscreen mode Exit fullscreen mode

The result can then be used as part of an application policy.

The important distinction is that a Noul value isn't simply a replacement for a boolean.

A value around 0.5 means that the model is uncertain. Your application needs to decide how uncertainty should be handled.

For an agent, that can naturally lead to:

score decisions

Why this architecture is interesting

The interesting part of JevSample isn't simply that it calls Jev.

It's that Jev becomes one component in a larger system.

The architecture separates:

Generation

Handled by the LLM.

"What should we produce?"
Enter fullscreen mode Exit fullscreen mode

Decisions

Handled by Jev.

"Which path should we take?"
"Is this safe?"
"Is this complete?"
Enter fullscreen mode Exit fullscreen mode

Deterministic rules

Handled by application code.

"Does this value have the correct format?"
"Is it within the allowed range?"
Enter fullscreen mode Exit fullscreen mode

Actions

Handled by tools and application code.

"Execute the operation."
Enter fullscreen mode Exit fullscreen mode

Oversight

Handled by humans when necessary.

"Approve or reject this action."
Enter fullscreen mode Exit fullscreen mode

That separation makes the agent easier to reason about.

A small mental model

When designing an AI agent, I find it useful to ask five questions:

1. What needs to be generated?
2. What needs to be decided?
3. What can be checked deterministically?
4. What action should be executed?
5. Where does a human need to be involved?
Enter fullscreen mode Exit fullscreen mode

The answers can then map to different components:

components

This is the architectural idea that JevSample is intended to demonstrate.

Summary

AI agents are becoming more complex.

As soon as an agent has multiple tools, retries, verification, safety checks, and human approval, it becomes more than an LLM wrapped in a loop.

It becomes a workflow.

And workflows require decisions.

Jev provides an interesting way to make some of those decisions explicit and structured, rather than asking a generative model to produce text and then parsing that text back into a decision.

This doesn’t mean that every decision should be moved to Jev.

The best architecture is usually a combination of deterministic code, generative models, decision models, tools, and human oversight.

The most interesting aspect of Jev is that it gives an AI agent another specialized component instead of trying to make one model responsible for everything.

Links

TypeSafe AI
JevSample source code

Top comments (0)