DEV Community

Cover image for Stop Sending Every Decision to an LLM: Code vs. Jev vs. Claude
Seenivasa Ramadurai
Seenivasa Ramadurai

Posted on

Stop Sending Every Decision to an LLM: Code vs. Jev vs. Claude

Why bounded semantic decisions may become an important layer in agentic AI architecture.

๐—๐—ฒ๐˜ƒ, ๐—”๐—ฝ๐—ฝ๐—น๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฆ๐˜๐—ฎ๐˜๐—ฒ, and Why I Think of It as ๐—›๐—ผ๐—บ๐—ฒ ๐——๐—ฒ๐—น๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜† vs. a ๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ ๐—•๐˜‚๐—ณ๐—ณ๐—ฒ๐˜

For the last few years, our default AI architecture has been surprisingly simple:

Have a problem? Send it to an LLM.

Need to classify something? LLM.

Need to choose a tool? LLM.

Need to decide whether an agent should retry? LLM.

Need to determine whether a document contains enough evidence? LLM.

Need to choose one option from three possibilities? Again, LLM.

General-purpose models such as Claude are extremely capable, but while experimenting with TypeSafe AI's ๐—๐—ฒ๐˜ƒ, I started asking a different architectural question:

Why hire a chef when all I need is someone to pick the right item from an already-defined menu?

That is where my ๐—›๐—ผ๐—บ๐—ฒ ๐——๐—ฒ๐—น๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜† vs. ๐—•๐˜‚๐—ณ๐—ณ๐—ฒ๐˜ analogy came from.

But before getting to the analogy, it is important to understand what Jev actually is.

What Is ๐—๐—ฒ๐˜ƒ?

Jev is TypeSafe AI's first public ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ ๐—ข๐—ป๐—ฒ ๐— ๐—ผ๐—ฑ๐—ฒ๐—น. Unlike a general-purpose LLM whose primary interface is generated text, Jev is designed for decisions inside software.

The basic idea is remarkably simple:

๐—”๐—ฝ๐—ฝ๐—น๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฆ๐˜๐—ฎ๐˜๐—ฒ โ†’ ๐—ง๐˜†๐—ฝ๐—ฒ๐—ฑ ๐—ค๐˜‚๐—ฒ๐˜€๐˜๐—ถ๐—ผ๐—ป๐˜€ โ†’ ๐—ฃ๐—ฟ๐—ผ๐—ฏ๐—ฎ๐—ฏ๐—ถ๐—น๐—ถ๐˜€๐˜๐—ถ๐—ฐ ๐——๐—ฒ๐—ฐ๐—ถ๐˜€๐—ถ๐—ผ๐—ป๐˜€

TypeSafe describes the model as taking program state and returning typed decisions that software can consume directly. The possible outputs are defined ahead of time instead of allowing the model to generate an arbitrary string.

That difference is important.

With Claude, I might ask:

Analyze this candidate's experience and explain whether this person would be appropriate for an enterprise AI architecture role.

The output could be several paragraphs.

With Jev, I can instead define a bounded question:

Which role best matches this state?

A. Traditional Backend Developer
B. Generative and Agentic AI Architect
C. Database Administrator
D. Manual Test Engineer
Enter fullscreen mode Exit fullscreen mode

Now I am not asking the model to generate language.

I am asking it to make a decision within a vocabulary that my application already controls.

That is why, as my own mental model, I think of Jev as:

โ€œ๐—๐˜‚๐˜€๐˜ ๐—˜๐—ป๐—ฟ๐—ถ๐—ฐ๐—ต๐—ฒ๐—ฑ ๐—ฉ๐—ผ๐—ฐ๐—ฎ๐—ฏ๐˜‚๐—น๐—ฎ๐—ฟ๐˜†โ€ โ€” not the official expansion of Jev, but a useful way for me to remember what the architecture is doing.

The vocabulary belongs to my application.

Jev adds semantic judgment around it.


๐—๐—ฒ๐˜ƒ Starts With the ๐—ฆ๐˜๐—ฎ๐˜๐—ฒ of the Application

This is the part I find most interesting.

Instead of beginning with a chat prompt such as:

โ€œYou are an expert AI architect. Analyze the following informationโ€ฆโ€

I start by defining the ๐—ฆ๐˜๐—ฎ๐˜๐—ฒ that exists right now.

For example:

{
  "name": "Sreeni Ramadurai",
  "myboss": "Krishna K",
  "skills": "GenAI, AgenticAI, AWS, Azure"
}
Enter fullscreen mode Exit fullscreen mode

This is not really a prompt in the traditional chatbot sense.

It is a snapshot of what the application currently knows.

The state might contain a customer message, an account status, a transaction, retrieved documents, an agent trace, tool results, user permissions, workflow history, or any other information required to make the next decision.

TypeSafe's own description emphasizes this distinction: Jev inputs are oriented around structured program state, while conventional LLM interfaces tend to be organized around sequential conversational messages.

That gives me a useful mental model:

State is not what I want the model to say. State is what the system currently knows.

Then I separately define what I want the system to decide.

๐—ฆ๐˜๐—ฎ๐˜๐—ฒ First, ๐—ค๐˜‚๐—ฒ๐˜€๐˜๐—ถ๐—ผ๐—ป๐˜€ Second

Using the small state above, I experimented with questions such as:

Is there sufficient evidence of architecture readiness?

Which enterprise AI role best matches the available evidence?

How strong is the evidence?

Does this state support a production deployment claim?

Does this state establish that Krishna K is Sreeni's manager?

Which conclusion would require information that is not present?
Enter fullscreen mode Exit fullscreen mode

Notice the architecture.

I am not putting everything into one large prompt and asking:

โ€œThink about all of this and tell me what you think.โ€

Instead, the application provides one state and several independent decision questions.

That is very close to normal software engineering.

We define the data.

We define the contract.

We define the permitted outputs.

Then intelligence evaluates the state against those contracts.

Three Kinds of Decisions

In the current Jev interface, these bounded judgments are expressed through typed question forms such as ๐—ก๐—ผ๐˜‚๐—น, ๐—–๐—ต๐—ผ๐—ถ๐—ฐ๐—ฒ, and ๐—ฆ๐—ฐ๐—ผ๐—ฟ๐—ฒ. TypeSafe's public materials describe Jev broadly as returning typed decisions with probabilities and confidence values rather than open-ended generated text.

For example, a binary judgment might ask whether the state contains enough evidence for a particular claim.

A ๐—–๐—ต๐—ผ๐—ถ๐—ฐ๐—ฒ might ask:

Traditional Backend Developer
Generative and Agentic AI Architect
Database Administrator
Manual Test Engineer
Enter fullscreen mode Exit fullscreen mode

A ๐—ฆ๐—ฐ๐—ผ๐—ฟ๐—ฒ might ask how strongly the available evidence supports enterprise AI capability on a predefined scale.

The important point is not the names of these primitives.

The important point is that the shape of the answer exists before inference begins.

The application owns the decision boundary.

The model judges what belongs inside it.

Jev Playground with my own State


Now the Restaurant Analogy Becomes Clear

Imagine opening a food-delivery application.

The restaurant has only three dishes available:

Pizza
Burger
Biryani

Then I provide some state:

Vegetarian
Likes spicy food
Wants something filling
Enter fullscreen mode Exit fullscreen mode

The menu has already been established.

Jev does not need to invent another meal.

It does not need to write a recipe.

It does not need to explain the history of biryani.

Its job is simply to evaluate the state against the known choices.

The result might look conceptually like:

Biryani    82%
Pizza      14%
Burger      4%
Enter fullscreen mode Exit fullscreen mode

This is why I call Jev ๐—ถ๐—ป๐˜๐—ฒ๐—น๐—น๐—ถ๐—ด๐—ฒ๐—ป๐˜ ๐—ต๐—ผ๐—บ๐—ฒ ๐—ฑ๐—ฒ๐—น๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜†.

The intelligence is real, because choosing correctly may require semantic understanding.

But the destination is bounded.

๐—›๐—ผ๐—บ๐—ฒ ๐——๐—ฒ๐—น๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜† = ๐—ธ๐—ป๐—ผ๐˜„๐—ป ๐—บ๐—ฒ๐—ป๐˜‚ + ๐—ฐ๐—ผ๐—ป๐˜๐—ฒ๐˜…๐˜ + ๐—ท๐˜‚๐—ฑ๐—ด๐—บ๐—ฒ๐—ป๐˜.


๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ Is the ๐—•๐˜‚๐—ณ๐—ณ๐—ฒ๐˜

Now imagine walking into a buffet.

I can ask:

What should I eat?

Build me a vegetarian dinner.

Compare these dishes.

Explain why one is healthier.

Suggest something I haven't considered.

Create a completely different meal.

Plan my meals for the entire week.

Now I am no longer choosing from one narrowly defined application vocabulary.

I want exploration.

I want synthesis.

I want explanation.

I may not even know the answer space before I start asking questions.

That is where a general-purpose model such as Claude becomes valuable.

So in my architecture analogy:

๐—๐—ฒ๐˜ƒ is ๐—›๐—ผ๐—บ๐—ฒ ๐——๐—ฒ๐—น๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜†. ๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ is the ๐—•๐˜‚๐—ณ๐—ณ๐—ฒ๐˜.

Home delivery does not mean unintelligent.

And buffet does not mean better.

They solve different problems.


But Before Jev, There Is Still ๐—–๐—ผ๐—ฑ๐—ฒ

There is an even simpler layer.

Suppose my application already has this business rule:

IF vegetarian
    remove all meat dishes
Enter fullscreen mode Exit fullscreen mode

I do not need Jev.

And I certainly do not need Claude.

The business has already determined exactly what should happen.

Use code.

That gives me three architectural layers:

CODE
Explicit rule is already known
        โ†“
JEV
Options are known,
but choosing requires semantic judgment
        โ†“
CLAUDE
Answer space is open,
requiring reasoning, synthesis or generation
Enter fullscreen mode Exit fullscreen mode

Or, using my restaurant analogy:

CODE
"The restaurant rule already decides."

        โ†“

JEV โ€” HOME DELIVERY
"The menu is fixed.
Choose the best item for this context."

        โ†“

CLAUDE โ€” BUFFET
"Explore, compare, combine,
explain and create."
Enter fullscreen mode Exit fullscreen mode

That distinction is much more useful to me than asking which model is โ€œsmarter.โ€


Where Jev Gets Really Interesting: ๐—”๐—ด๐—ฒ๐—ป๐˜๐—ถ๐—ฐ ๐—”๐—œ

Now move away from restaurants and consider an enterprise agent.

During one workflow, an agent may need to decide:

Which tool should I call?

Is this retrieved document relevant?

Did the tool actually satisfy the task?

Should I retry?

Should I continue?

Should I escalate to a human?

Is this transaction low, medium or high risk?

Does the evidence support this claim?
Enter fullscreen mode Exit fullscreen mode

Many of these are semantic decisions.

They cannot always be expressed as simple if/else statements.

But they also do not necessarily need paragraphs of generated reasoning.

Their answer spaces often look like this:

Tool
Salesforce | ServiceNow | SharePoint | None

Action
Continue | Retry | Escalate

Risk
Low | Medium | High

Evidence
Supported | Partial | Unsupported
Enter fullscreen mode Exit fullscreen mode

This is where I see Jev fitting naturally inside an ๐—”๐—ด๐—ฒ๐—ป๐˜ ๐—›๐—ฎ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€.

                     AGENT HARNESS
                          โ”‚
           โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
           โ”‚              โ”‚              โ”‚
         CODE            JEV          CLAUDE
           โ”‚              โ”‚              โ”‚
         Rules         Semantic       Reasoning
                       Judgment       Generation
           โ”‚              โ”‚              โ”‚
           โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                          โ”‚
                        TOOLS
                          โ”‚
                        ACTION
Enter fullscreen mode Exit fullscreen mode

Code controls deterministic behavior.

Jev makes bounded semantic judgments.

Claude handles the parts where broader reasoning or generation is actually necessary.

That architecture interests me much more than simply routing everything through one giant model.


Is Jev's โ€œ๐—”๐—ฝ๐—ฝ๐—น๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฆ๐˜๐—ฎ๐˜๐—ฒโ€ Related to ๐—›๐—”๐—ง๐—˜๐—ข๐—”๐—ฆ?

When I started thinking about Jev in terms of application state, another architecture immediately came to mind:

๐—›๐—”๐—ง๐—˜๐—ข๐—”๐—ฆ โ€” Hypermedia as the Engine of Application State.

There is definitely a conceptual connection, but they are not the same thing.

In REST, HATEOAS means the server representation provides hypermedia controls that tell the client what resources or transitions are available next. Rather than hard-coding every possible navigation path, the client can discover permitted next actions from the representation it receives. This is part of the REST architectural style described by Roy Fielding.

For example, an order might be represented as:

{
  "orderId": 1001,
  "status": "pending",
  "links": [
    {
      "rel": "cancel",
      "href": "/orders/1001/cancel"
    },
    {
      "rel": "pay",
      "href": "/orders/1001/payment"
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

The current state determines which transitions are available.

That sounds somewhat similar to what we are doing with Jevโ€”but there is an important difference.

๐—›๐—”๐—ง๐—˜๐—ข๐—”๐—ฆ exposes the allowed next actions.

๐—๐—ฒ๐˜ƒ can semantically judge which allowed action best fits the current context.

That leads to an architecture I find fascinating.

Imagine the application state says:

{
  "orderStatus": "pending",
  "customerMessage": "I accidentally placed this order twice.",
  "paymentCaptured": false,
  "availableActions": [
    "cancel",
    "continue",
    "escalate"
  ]
}
Enter fullscreen mode Exit fullscreen mode

HATEOAS could tell the application:

These are the actions currently available.
Enter fullscreen mode Exit fullscreen mode

Jev could then answer:

Given the complete state,
which available action best fits the situation?

cancel      94%
escalate     5%
continue     1%
Enter fullscreen mode Exit fullscreen mode

Then deterministic code executes the selected transition subject to whatever confidence threshold, permissions, policy checks, and guardrails the application requires.

So I would summarize the relationship this way:

๐—›๐—”๐—ง๐—˜๐—ข๐—”๐—ฆ tells the client what it may do next.
๐—๐—ฒ๐˜ƒ can help the application judge which permitted action makes sense next.

That is not part of the formal definition of HATEOAS, nor is Jev an implementation of HATEOAS.

It is simply a useful architectural connection.

HATEOAS gives us state-driven discoverability.

Jev gives us state-driven semantic judgment.

Put together, they suggest an interesting pattern for intelligent APIs and agents.


๐—ฆ๐˜๐—ฎ๐˜๐—ฒ โ†’ ๐—ฃ๐—ผ๐˜€๐˜€๐—ถ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐—ถ๐—ฒ๐˜€ โ†’ ๐—๐˜‚๐—ฑ๐—ด๐—บ๐—ฒ๐—ป๐˜ โ†’ ๐—”๐—ฐ๐˜๐—ถ๐—ผ๐—ป

This may actually be the most useful way for me to think about Jev.

Not:

Prompt โ†’ LLM โ†’ Text
Enter fullscreen mode Exit fullscreen mode

But:

CURRENT STATE
      โ†“
AVAILABLE DECISIONS
      โ†“
SEMANTIC JUDGMENT
      โ†“
PROBABILITY / CONFIDENCE
      โ†“
BUSINESS RULES + GUARDRAILS
      โ†“
ACTION
      โ†“
NEW STATE
Enter fullscreen mode Exit fullscreen mode

And then the cycle repeats.

Stateโ‚€
   โ†“
Decision
   โ†“
Action
   โ†“
Stateโ‚
   โ†“
Decision
   โ†“
Action
   โ†“
Stateโ‚‚
Enter fullscreen mode Exit fullscreen mode

Now Jev starts looking less like a chatbot and more like an intelligence primitive inside a stateful software system.

That is the part I find most compelling.


Why Not Just Give Claude Structured Output?

This is the obvious question.

Claude can absolutely be given:

Choose exactly one:

cancel
continue
escalate
Enter fullscreen mode Exit fullscreen mode

It can return JSON.

It can use tool calling.

It can conform to a schema.

So the question is not:

Can Claude make bounded decisions?

Of course it can.

The architecture question is:

If my application needs thousands of small semantic judgments whose output spaces are already known, should every one of them go through a general-purpose text-generation architecture?

Jev was explicitly designed around this narrower machine-facing task. TypeSafe describes it as optimizing for structured, calibrated decisions rather than free-form strings.

This becomes especially interesting inside agentic systems where a single user request may cause many internal decisions.

One model call may be insignificant.

Thousands of decisions across many agents, tools, users, and workflow steps are a different architectural problem.


๐—•๐—ผ๐˜‚๐—ป๐—ฑ๐—ฒ๐—ฑ Does Not Mean ๐—–๐—ผ๐—ฟ๐—ฟ๐—ฒ๐—ฐ๐˜

There is one distinction I think is essential.

Suppose Jev produces:

Biryani 82%
Enter fullscreen mode Exit fullscreen mode

That does not make Biryani objectively correct.

Likewise:

{
  "risk": "high"
}
Enter fullscreen mode Exit fullscreen mode

being perfectly valid structured output does not prove that the risk assessment itself is correct.

๐—ง๐˜†๐—ฝ๐—ฒ ๐˜€๐—ฎ๐—ณ๐—ฒ๐˜๐˜† โ‰  ๐—ฆ๐—ฒ๐—บ๐—ฎ๐—ป๐˜๐—ถ๐—ฐ ๐—ฐ๐—ผ๐—ฟ๐—ฟ๐—ฒ๐—ฐ๐˜๐—ป๐—ฒ๐˜€๐˜€.

So we still need evaluations, thresholds, observability, guardrails, and human review where the consequence demands it.

The difference is that the decision surface is constrained and machine-consumable.

That can make the surrounding system easier to reason about.


How I Would Use Jev in an Enterprise Application

My pattern would be simple:

1. Build the application state
        โ†“
2. Define the decisions the application actually needs
        โ†“
3. Express those decisions as bounded typed questions
        โ†“
4. Let Jev evaluate the state
        โ†“
5. Read probabilities/confidence
        โ†“
6. Apply deterministic thresholds and policies in code
        โ†“
7. Execute a tool, ask Claude for deeper reasoning,
   or escalate to a human
        โ†“
8. Update the application state
        โ†“
9. Repeat
Enter fullscreen mode Exit fullscreen mode

The model should not own the entire workflow.

๐—ง๐—ต๐—ฒ ๐—ต๐—ฎ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€ ๐—ผ๐˜„๐—ป๐˜€ ๐˜๐—ต๐—ฒ ๐˜„๐—ผ๐—ฟ๐—ธ๐—ณ๐—น๐—ผ๐˜„.

The model contributes intelligence at specific decision points.

That separation matters.


My Mental Model After Using Jev

I now think about AI architecture using three increasingly flexible layers.

๐—–๐—ข๐——๐—˜ = ๐—ฅ๐—จ๐—Ÿ๐—˜๐—ฆ

The correct behavior is already explicitly known.

If payment has already been captured, do not cancel automatically.

๐—๐—˜๐—ฉ = ๐—๐—จ๐——๐—š๐— ๐—˜๐—ก๐—ง / ๐—›๐—ข๐— ๐—˜ ๐——๐—˜๐—Ÿ๐—œ๐—ฉ๐—˜๐—ฅ๐—ฌ

The options are known, but the correct choice depends on interpreting context.

Cancel, Continue, or Escalate?

๐—–๐—Ÿ๐—”๐—จ๐——๐—˜ = ๐—ฅ๐—˜๐—”๐—ฆ๐—ข๐—ก๐—œ๐—ก๐—š + ๐—š๐—˜๐—ก๐—˜๐—ฅ๐—”๐—ง๐—œ๐—ข๐—ก / ๐—•๐—จ๐—™๐—™๐—˜๐—ง

The problem requires exploration, explanation, planning, synthesis, or creation.

Analyze the entire situation, explain the trade-offs, and propose a resolution strategy.

This is not a competition between models.

It is separation of concerns.


The Bigger Question: Are We Overusing Generation?

Software architecture has always been about choosing the appropriate abstraction.

We do not use a database where a queue belongs.

We do not use Kubernetes to execute one shell script.

And perhaps we should not use open-ended generation where all we really need is a bounded semantic decision.

That is the larger lesson I took from experimenting with Jev.

The future of AI applications may not be:

Application โ†’ One giant LLM โ†’ Everything
Enter fullscreen mode Exit fullscreen mode

It may look more like:

Application State
        โ”‚
        โ–ผ
      Harness
        โ”‚
   โ”Œโ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   โ”‚    โ”‚            โ”‚
 Code  Jev         Claude
   โ”‚    โ”‚            โ”‚
Rules Judgment    Reasoning
                  Generation
   โ””โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
        โ”‚
       Tools
        โ”‚
      Actions
        โ”‚
     New State
Enter fullscreen mode Exit fullscreen mode

And once I saw it this way, my restaurant analogy finally clicked.

๐—–๐—ผ๐—ฑ๐—ฒ already knows the restaurant rules.

๐—๐—ฒ๐˜ƒ is intelligent ๐—›๐—ผ๐—บ๐—ฒ ๐——๐—ฒ๐—น๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜† from a defined menu.

๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ gives me the ๐—•๐˜‚๐—ณ๐—ณ๐—ฒ๐˜ when I genuinely need exploration and creation.

So before making the next expensive generative call, perhaps the architecture question should not be:

Which model is smartest?

It should be:

Does this decision need a ๐—ฟ๐˜‚๐—น๐—ฒ, intelligent ๐—ต๐—ผ๐—บ๐—ฒ ๐—ฑ๐—ฒ๐—น๐—ถ๐˜ƒ๐—ฒ๐—ฟ๐˜†, or the entire ๐—ฏ๐˜‚๐—ณ๐—ณ๐—ฒ๐˜?

Sometimes we need the chef.

But sometimes all we needed was someone intelligent enough to pick the right item from the menu.

Sreeni Ramadorai
AI Architect

Top comments (0)