I began with a simple thesis:
For decisions, accuracy matters more than speed.
JEV is fast because it constrains a task to a bounded decision contract. I assumed that frontier LLMs would be more accurate - given their world knowledge - and that their extra latency would be an obvious trade worth making.
So I built a reproducible benchmark to test that assumption against JEV on the same decision contracts:
- Intent routing (Choice)
- Moderation / guardrail classification (Noul)
- Search relevance (Score)
The result was more interesting than "JEV wins" or "frontier models win."
My initial case: accuracy should beat speed
On intent routing, the best tested frontier models did outperform JEV:
That supports my original concern. If a decision is wrong, being fast does not make it useful.
But the advantage did not generalize cleanly.
On HateCheck moderation, JEV reached 98.2%; the tested models ranged from 95.1% to 100.0%. On TREC search relevance, JEV's 47.4% nearest-tier accuracy tied the best tested result.
So JEV was right about something I had underestimated: a bounded decision system can be very fast without automatically giving up meaningful accuracy.
The actual contract matters
For the relevance task, the input is intentionally small and explicit: a query, a candidate passage, and a rubric.
{
"state": {
"query": "what is the capital of france",
"passage": "Paris is the capital and most populous city of France."
},
"levels": [
"Not Relevant",
"Related",
"Highly Relevant",
"Perfect"
]
}
The model has to return both a continuous score and its belief across the four levels:
{
"score": 2.99,
"probabilities": {
"0": 0.0,
"1": 0.0,
"2": 0.01,
"3": 0.99
}
}
That response is coherent: the probability-weighted score is also approximately 2.99.
The unexpected failure mode: plausible, valid JSON that disagrees with itself
This is where my original framing broke down.
A model can return valid JSON, a plausible final score, and a plausible-looking probability distribution - yet have those two fields contradict one another:
{
"score": 2.99,
"probabilities": {
"0": 0.10,
"1": 0.60,
"2": 0.25,
"3": 0.05
}
}
The score says "almost Perfect." The distribution says "mostly Related." Both fields cannot be true at once.
That is not merely a formatting issue. It means a downstream system has to decide which answer to trust: the final decision, or the model's stated uncertainty.
The benchmark changed the question
On the relevance experiment, score/distribution consistency was:
These are results for the tested configurations, not universal claims about the models. Output mode and effort settings matter.
But they exposed a problem I was not measuring before.
We were both wrong, in different ways
A system needs at least three things:
- Quality: Does it select the right answer?
- Consistency: Do its score, probabilities, and structured fields agree?
- Operational performance: Can it deliver that answer within acceptable latency and token use?
I set out to show that bounded decisions were too limiting. Instead, I found that the hard problem is not only choosing the right answer - it is producing an answer you can consistently trust.
Check the results or run your own benchmarks
The benchmark, decision contracts, published comparison summaries, chart source, and export tooling are available here:
https://github.com/Tenkei/jev-decision-bench
A reproducible, self-hosted benchmark for comparing JEV and LLM decision-making on your own datasets, models, and…github.com
The published results are intended to be inspectable and reproducible - not taken as a universal leaderboard. You can use the same workflow with your own model configurations or decision datasets.


Top comments (0)