Beyond LLMs: Jev vs OpenAI Decisions API — The Emerging Decision Layer for Agentic AI
For the last few years, most of the AI conversation has been around one question:
How intelligent can LLMs become?
- Bigger models.
- More reasoning.
- Longer context.
- Better agents.
- More tools.
But while working on Agentic AI systems, I have increasingly started thinking about a different question:
Do we really need an LLM to reason and generate tokens for every decision inside an AI system?
Think about some very common enterprise decisions.
- Should I approve or escalate?
- Which agent should handle this request?
- Is this transaction suspicious?
- Does this agent execution require human review?
- Which workflow should run next?
- Should I call a tool or ask for more information?
These are not really generation problems.
These are decision problems.
And this is why two recent developments caught my attention:
TypeSafe AI's Jev and OpenAI's Decisions API.
I think both are pointing towards something much bigger than another model or API.
They could be early signals of a new architectural layer in Agentic AI:
👉 The Decision Layer.
🧠 From Generating Tokens to Making Decisions
Most LLM architectures today follow a familiar pattern.
Context → Prompt → LLM → Tokens → Parse → Validate → Action
This works extremely well when the outcome itself requires language, code, reasoning or planning.
But consider something simpler.
An enterprise support agent completes a customer interaction.
Now the platform needs to determine:
- Was the interaction successful?
- Was policy followed?
- Is there a compliance concern?
- Does this require human review?
- What should happen next?
We could send the entire trace to a large reasoning model and ask it to generate an assessment.
But maybe that is not always the right abstraction.
What the software actually needs could simply be:
PASS | REVIEW | ESCALATE
along with confidence.
That changes the architecture considerably.
⚡ TypeSafe Jev: Intelligence Designed for Software
TypeSafe describes Jev as its first System One Model.
What I find interesting is not just another model benchmark.
It is the programming model.
Jev is not primarily designed to generate strings for humans.
It is designed to produce typed decisions that software can consume directly.
Instead of:
Context → Prompt → Tokens → Parse → Validate → Action
we move closer to:
State → Typed Decision → Probability → Action
TypeSafe describes this nicely as:
Unstructured state in → typed probabilistic decisions out.
Jev currently supports decision primitives such as:
Yes / No
Does this transaction require manual review?
Choice
Which workflow should execute?
REFUND | REPLACEMENT | SUPPORT | FRAUD_REVIEW
Score
How risky is this transaction?
Now combine this with confidence.
High confidence → Execute automatically
Medium confidence → Invoke deeper reasoning
Low confidence / high risk → Human review
For enterprise Agentic AI, this becomes very interesting.
🔄 OpenAI Enters the Decision Layer
OpenAI's Decisions API makes this space even more interesting.
The basic idea is to focus model intelligence on user-defined questions with finite predefined answers.
From an application architecture perspective, that means something like:
Give the model context.
Define the possible decisions.
Receive a bounded answer that software can act on.
This naturally fits problems such as:
- classification
- request routing
- workflow selection
- agent selection
- next-action decisions
But there is an important distinction.
Jev is positioned by TypeSafe as a model architecture specifically designed around machine-consumable decisions.
OpenAI is bringing constrained decision-making into its broader model and agent ecosystem.
So I don't really see this as:
Jev vs GPT.
I see it as:
Two approaches to building the decision layer of AI-native software.
📊 Jev vs OpenAI Decisions API
Based on the information available today, this is how I currently see them.
| Dimension | TypeSafe Jev | OpenAI Decisions API |
|---|---|---|
| Primary abstraction | Decision model | Decision API |
| Philosophy | Intelligence specifically designed for software decisions | Model intelligence focused on constrained decisions |
| Input | Program/application state | Context including text/images |
| Output | Typed probabilistic decisions | Finite predefined answers |
| Open-ended generation | Not the primary objective | Not the objective of the Decisions API |
| Probability / confidence | Core part of the architecture | Decision-oriented response; details still emerging |
| Decision types | Yes/No, Choice, Score | User-defined questions with finite answers |
| Architecture | TypeSafe System One Model | OpenAI model-based decision capability |
| Natural fit | High-volume machine decisions and automation | Classification, routing and agent/workflow decisions |
| Ecosystem | Specialized decision infrastructure | Broader OpenAI model/agent ecosystem |
| Maturity | Early access | Newly announced |
There are still many questions that only real production usage will answer.
- Accuracy.
- Calibration.
- Latency under load.
- Cost at scale.
- Observability.
- Governance.
- Versioning.
- Failure behaviour.
And one that I think matters a lot:
Can I reliably use the model's confidence to decide whether software should take an action?
🏗️ The Bigger Shift: A Hierarchy of Intelligence
This is where I think things become much more interesting architecturally.
Many Agentic AI architectures today effectively look like:
Request → LLM → Agent → Tools
We are putting LLMs everywhere.
I am not convinced that will remain the optimal architecture.
I think enterprise AI will increasingly move towards a hierarchy of intelligence.
Deterministic Rules
↓
Decision Models
↓
Reasoning Models
↓
Agents + Tools
↓
Humans
Each layer has a job.
Deterministic Rules
If something can reliably be expressed as code, use code.
There is no reason to ask an AI model whether:
invoice_amount > approval_limit
Decision Models
When judgment is required but the output space is constrained, use a decision model.
For example:
APPROVE | REVIEW | REJECT
Reasoning Models
When the problem involves ambiguity, planning, synthesis or deeper reasoning, bring in a reasoning model.
Agents
When we need multiple steps, tools and actions across systems, use an agent.
Humans
When confidence is insufficient or business risk crosses a threshold, escalate.
To me, this is a much more practical enterprise architecture.
And potentially a much cheaper one.
🤖 Think About Agent Routing
Imagine an enterprise platform with several specialized agents:
- Customer Support Agent
- Billing Agent
- Fraud Agent
- Subscription Agent
- Order Agent
- Human Escalation
A customer says:
"I was charged twice, one order never arrived, and I want to cancel the subscription associated with it."
One option is to send everything to a powerful reasoning model and let it orchestrate the entire workflow.
But another architecture is possible.
Customer Context
↓
Decision Layer
↓
Billing: 0.94
Order: 0.87
Subscription: 0.91
Fraud: 0.18
↓
Policy / Orchestration Layer
↓
Billing Agent + Order Agent + Subscription Agent
Now the expensive reasoning and agentic execution happens only where it is actually needed.
This is where I see the decision layer becoming part of the Agentic AI control plane.
🔍 Agent Evaluation Could Be Another Big Use Case
This is something I am particularly interested in.
When an agent completes a workflow, we need to answer questions such as:
- Did the agent achieve the expected outcome?
- Did it select the right tools?
- Did it violate a policy?
- Was there suspicious behaviour?
- Should the trace be reviewed?
- Should an alarm be triggered?
- Should the workflow automatically retry?
Again, these are decisions.
Today we use LLM-as-a-Judge extensively for this.
And it is extremely useful when semantic reasoning is required.
But do we really need a large generative model for every evaluation?
Maybe not.
Imagine:
Agent Trace
↓
Decision Model
↓
Outcome achieved: 0.96
Policy compliant: 0.98
Tool misuse: 0.07
Human review required: 0.12
↓
AgentOps Policy Engine\
↓
PASS | ALERT | RETRY | HUMAN REVIEW
This could make continuous evaluation cheaper, faster and much easier to operationalize.
And instead of generating another paragraph explaining what happened, we get something that software can directly act upon.
🎯 Confidence May Become More Important Than the Answer
This is probably the part I find most interesting.
Enterprise AI should rarely operate like this:
AI says X → therefore execute X.
A better architecture is:
AI believes X with confidence Y → enterprise policy decides what happens next.
Imagine a payment workflow.
Approve = 0.97
→ Process automatically.
Another transaction:
Approve = 0.71
→ Invoke deeper reasoning.
Another:
Approve = 0.51
→ Human review.
Now we have separated two responsibilities.
AI makes the judgment.
Software decides what level of confidence is acceptable for action.
That separation matters a lot when we start putting Agentic AI into real enterprise workflows.
Does This Make LLMs Less Important?
I actually think the opposite.
It means we can use powerful models where their intelligence really matters.
A reasoning model shouldn't necessarily spend compute answering thousands of repetitive questions like:
Is this request billing-related?
Should this workflow go to Agent A or Agent B?
Does this transaction cross a semantic risk threshold?
Instead:
Use deterministic software where possible.
Use decision intelligence where judgment is needed.
Use reasoning models where reasoning is needed.
Use agents where action is needed.
Use humans where accountability is needed.
That feels much more sustainable than:
Send everything to the biggest LLM available.
🚀 Jev vs Decisions API May Actually Be the Wrong Question
Naturally, engineers will compare Jev and OpenAI Decisions API.
- Accuracy.
- Latency.
- Calibration.
- Cost.
- Developer experience.
- Production reliability.
Those comparisons will be useful.
But I think there is a bigger signal here.
Decision Intelligence may be emerging as its own architectural layer.
For the last few years, our Agentic AI architecture diagrams have increasingly included:
Models | RAG | Agents | Tools | Memory | MCP | Guardrails | Observability
Maybe we need to add another box:
Decision Layer
And once we do that, the architecture starts looking very different.
I don't think the future of enterprise AI is:
One extremely intelligent model doing everything.
I think it could be:
Different forms of intelligence being invoked at the right point in the workflow.
- Rules for certainty.
- Decision models for judgment.
- Reasoning models for complexity.
- Agents for action.
- Humans for accountability.
And a control plane deciding when to use each.
That, for me, is the interesting part of what Jev and OpenAI Decisions API are signalling.
Not every problem needs another generated answer.
Sometimes software simply needs a decision.
Top comments (0)