Hey developers π
Remember when we first started using LLMs in development?
A few years ago, getting an AI model to understand a piece of text and generate a useful response already felt like magic. Then the models kept getting better. They started writing code, debugging applications, summarizing documents, reasoning through problems, calling tools, working with images, and eventually acting more like agents.
AI slowly went from something we experimented with to something we could actually build products around.
And now we're at an interesting point. We've made these models incredibly capable at generating things. But here's the question I've started thinking about:
Do we really need all that capability for every AI task inside our applications?
Because sometimes, our application doesn't need to return essay every time. Sometimes we need an answer in just yes or no.
We've Made LLMs Really, Really Good
When large language models first became popular, the idea was pretty simple:
Give the model some text β get some text back
Over the years, that changed dramatically.
Models got better at understanding context, following instructions, writing code, reasoning through problems, using tools, working with images and eventually acting more like agents.
We went from:
"Write me an email."
to:
"Analyze this document, search the web, write some code, call an API and figure out why my production deployment is broken."
It is pretty impressive but here's where things get interesting. Most applications don't need an AI model to solve a PhD-level problem every time they call one. Sometimes, we just need a small answer. For example:
Is this message spam?
Does this transaction look suspicious?
Should this user receive this offer?
Which category does this product belong to?
Does this customer look like they're going to churn?
These aren't necessarily conversations. They're decisions. And that brings us to a question.
Do We Really Need a Huge LLM For This?
Imagine your application needs to answer:
Q: Is this transaction fraudulent?
A: Yes
A traditional LLM can obviously do this.
You send it a prompt:
Analyze the transaction and determine whether it is fraudulent.
Return only "yes" or "no".
And hopefully you get:
yes
But the model is still fundamentally a generative model. It is generating a response. Your application then has to deal with that response. Maybe it returns:
Yes.
Maybe:
Yes, this transaction appears to be fraudulent.
Maybe:
Based on the transaction details, I believe there is a high
probability that this transaction is fraudulent...
You can obviously solve a lot of this with structured outputs, schemas and validation but the underlying question remains:
Why are we using a system designed to generate language when our application only needs a predefined decision?
What if the model was designed around the decision instead? That's where System One Models come into the picture.
So, What Is JEV?
Think about how we normally use AI in an application.
We send some input, ask the model something, and get a generated response back. But what if we already know the question our application needs to answer? That's the problem TypeSafe AI is trying to approach differently with System One Models.
JEV is TypeSafe AI's first public System One Model, a company founded by a former OpenAI researcher, Diogo Almeida, who is known for RLHF work. The problem he is trying to solve is pretty simple:
If I already know the question I want my AI to answer, why do I need a traditional generative LLM to generate a whole response?
That's the core idea behind JEV.
Traditional LLMs are built around generation. You give them an input and they generate a natural-language response. JEV takes a different approach. Instead of generating an open-ended response, it is designed to extract a predefined result from the input.
Think of the difference like this:
Traditional LLM
Input
β
Generate a response
β
Natural language
β
Your application interprets it
JEV / System One Model
Input + predefined question
β
JEV
β
Predefined decision
β
Your application uses it
So instead of asking an LLM:
"Analyze this transaction and explain whether you think it is fraudulent."
You can frame the problem as:
"Is this transaction fraudulent?"
And that's the kind of predefined decision JEV is designed to handle.
This doesn't mean JEV is trying to replace traditional LLMs. It's solving a different problem: using AI for small, predefined decisions where generating a long natural-language response isn't really what the application needs. That distinction is what makes System One models interesting.
What Kind of Questions Are We Talking About?
There are quite a few. And they're actually pretty common in software.
1. Yes / No Decisions
Probably the simplest example.
Is this transaction fraudulent?
Result:
Yes
Or:
Should this user receive this promotion?
Result:
No
Your application can directly use that decision.
2. Classification
Sometimes you don't need a boolean, you need one result from a predefined set. For example:
What category does this product belong to?
["Electronics", "Clothing", "Food", "Furniture"]
The model can return the relevant category. This is useful for things like:
- ticket classification
- product categorization
- support routing
- content moderation
- lead classification
3. Scoring
You might want a score rather than a simple yes/no. For example:
How likely is this customer to churn?
Instead of generating an essay about the customer, the application can work with a predefined result such as a probability or score. That makes it much easier to plug the model into an existing decision-making system.
4. Routing
Another interesting use case is deciding where something should go. For example:
Which team should handle this support request?
- Billing
- Technical Support
- Sales
The model makes the decision. Your application then handles the rest:
decision β route request β workflow
No chatbot required.
A live example
Jev answers typed questions about the state you send it:
- Choice picks one option from a set you define and returns a probability for each option.
- Noul answers a yes or no question and returns the probability of yes.
- Score places the input on an ordered scale you define and returns a probability-weighted position.
Because the answer is a typed value with an associated probability, your implementation can branch on it directly. There is no free-form text to parse. TypeSafe's own benchmarks put Jev at roughly 40β200x faster than frontier LLMs on comparable tasks, at about $0.042 per million input tokens with output free β the gap that makes running millions of small decisions through it actually make sense.
You can try Jev on https://openrouter.ai
Let's take an example of a support ticket agent. It asks Jev three independent questions in one request, whether the ticket is a bug (Noul), which team owns it (Choice), and how urgent it is (Score).
curl https://openrouter.ai/api/alpha/decisions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe/jev-1.13",
"state": {
"customer_tier": "enterprise",
"ticket": "My checkout page shows a blank screen after I click Pay. I have tried two browsers."
},
"questions": {
"is_bug": {
"type": "noul",
"instructions": "Is the customer reporting a software defect?",
"criteria": {
"true": "The customer describes broken or unexpected product behavior.",
"false": "The customer is asking a question or requesting a feature."
}
},
"team": {
"type": "choice",
"instructions": "Which team should own this ticket?",
"criteria": {
"payments": "Checkout, billing, or payment processing issues.",
"frontend": "Rendering, layout, or browser compatibility issues.",
"account": "Login, permissions, or profile issues."
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": [
"Can wait for the next release",
"Should be fixed this week",
"Blocking revenue right now"
]
}
}
}'
The response contains one typed answer per question, plus usage. This is an actual response to the request above:
{
"id": "gen-dec-1790015143-AIaTutprXsJ5EwohRSjb",
"model": "typesafe/jev-1.13-20260917",
"provider": "TypeSafe",
"answers": {
"is_bug": { "type": "noul", "noul": 0.96 },
"team": {
"type": "choice",
"choice": "payments",
"confidence": 0.67,
"probabilities": { "payments": 0.78, "frontend": 0.22, "account": 0 }
},
"urgency": {
"type": "score",
"score": 1.99,
"confidence": 0.99,
"probabilities": { "0": 0, "1": 0, "2": 1 },
"legend": {
"0": "Can wait for the next release",
"1": "Should be fixed this week",
"2": "Blocking revenue right now"
}
}
},
"usage": { "input_tokens": 476, "output_tokens": 70, "cost": 0.000019992 }
}
The model field in the response names the dated snapshot that served your request. Sending typesafe/jev-1.13 resolves to the current 1.13 release,
Why Not Just Use a Traditional LLM?
You absolutely can. And for many problems, you should.
If you want to generate an article, write code, summarize a document, have a conversation or reason through a complicated problem, a traditional LLM makes a lot of sense. But consider an application that needs to make thousands or millions of small decisions.
For example:
Is this order suspicious?
Should this notification be sent?
Does this support ticket need escalation?
Which workflow should process this request?
You're not really looking for a conversation. You're looking for a decision.
And why it matters going forward
I don't think the future of AI is simply going to be:
Bigger model = better application.
We've already seen how capable large models can become. The next interesting step could be figuring out where we actually need that intelligence. Maybe an application won't use one giant model for everything. Instead, it could look something like:
βββ LLM β Generation
β
Application βββββββββΌββ System One β Decisions
β
βββ Vision Model β Images
β
βββ Embeddings β Retrieval
Use the right model for the right job.
Need to generate something? : Use an LLM.
Need complex reasoning? : Use a reasoning model.
Need to classify or make a predefined decision? : Use a specialized decision model.
Need semantic search? : Use embeddings.
This makes AI feel less like one giant magical box and more like another part of the software stack.
AI Could Become More Specialized
This is probably the part I'm most curious about. For the last few years, we've been obsessed with making models more capable. But software doesn't always need maximum capability. Sometimes it needs speed, sometimes predictability, and sometimes simply cost efficiency. And sometimes, it just needs a "true" instead of 500 tokens explaining why the answer is true.
That's where specialized AI models become interesting. Instead of trying to make one model capable of doing everything, we can start using different models for different kinds of problems. JEV is one example of this direction.
Does JEV Replace LLMs?
A simple answer: No.
At least, that's not how I'd look at it. A model designed for predefined decisions isn't going to replace a model whose job is to generate and reason through open-ended problems. If I'm building an AI coding assistant, I still want a powerful generative model. If I'm asking AI to analyze a 200-page document and explain it to me, I still want a powerful LLM.
But if my backend simply needs to answer questions like "Should I approve this?", "Which workflow should handle this?", or "Does this belong to category A, B or C?", then maybe I don't need the same model for that job.
The Bigger Picture
We've spent the last few years asking:
How do we make AI models smarter?
Maybe the next question is:
How do we make AI fit better into software?
That's why I find JEV interesting. Not because it's another model trying to be the biggest or smartest model out there, but because it represents a different way of thinking about AI. Instead of always asking AI to generate something, we can start asking it to make a decision.
If AI-driven applications become more modular, specialized, and purpose-built over the next few years, I think this distinction is going to matter a lot. Maybe 2027 won't just be about bigger models. Maybe it'll be about using the right model for the right job.
And honestly, that sounds a lot more interesting to me.
I hope you found this helpful! If you have any feedback or suggestions for improvement, please feel free to reach out. I'm always looking to learn and improve, and your input is invaluable. If you have read it till now, thank you so much for reading!, please leave your comments if any βοΈ
Don't forget to bookmark this blog for the future π
Connect with the author:
Top comments (1)