DEV Community

Keivan Esbati
Keivan Esbati

Posted on Originally published at Medium

Should You Let JEV Make Every Decision? Accuracy Matters More Than Speed.

I began with a simple thesis:

For decisions, accuracy matters more than speed.

JEV is fast because it constrains a task to a bounded decision contract. I assumed that frontier LLMs would be more accurate - given their world knowledge - and that their extra latency would be an obvious trade worth making.

So I built a reproducible benchmark to test that assumption against JEV on the same decision contracts:

  • Intent routing (Choice)
  • Moderation / guardrail classification (Noul)
  • Search relevance (Score)

The result was more interesting than "JEV wins" or "frontier models win."

My initial case: accuracy should beat speed

On intent routing, the best tested frontier models did outperform JEV:

Results from bench tests

That supports my original concern. If a decision is wrong, being fast does not make it useful.

But the advantage did not generalize cleanly.

On HateCheck moderation, JEV reached 98.2%; the tested models ranged from 95.1% to 100.0%. On TREC search relevance, JEV's 47.4% nearest-tier accuracy tied the best tested result.

So JEV was right about something I had underestimated: a bounded decision system can be very fast without automatically giving up meaningful accuracy.

The actual contract matters

For the relevance task, the input is intentionally small and explicit: a query, a candidate passage, and a rubric.

{
  "state": {
    "query": "what is the capital of france",
    "passage": "Paris is the capital and most populous city of France."
  },
  "levels": [
    "Not Relevant",
    "Related",
    "Highly Relevant",
    "Perfect"
  ]
}
Enter fullscreen mode Exit fullscreen mode

The model has to return both a continuous score and its belief across the four levels:

{
  "score": 2.99,
  "probabilities": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.01,
    "3": 0.99
  }
}
Enter fullscreen mode Exit fullscreen mode

That response is coherent: the probability-weighted score is also approximately 2.99.

The unexpected failure mode: plausible, valid JSON that disagrees with itself

This is where my original framing broke down.

A model can return valid JSON, a plausible final score, and a plausible-looking probability distribution - yet have those two fields contradict one another:

{
  "score": 2.99,
  "probabilities": {
    "0": 0.10,
    "1": 0.60,
    "2": 0.25,
    "3": 0.05
  }
}
Enter fullscreen mode Exit fullscreen mode

The score says "almost Perfect." The distribution says "mostly Related." Both fields cannot be true at once.

That is not merely a formatting issue. It means a downstream system has to decide which answer to trust: the final decision, or the model's stated uncertainty.

The benchmark changed the question

On the relevance experiment, score/distribution consistency was:

Results from bench test

These are results for the tested configurations, not universal claims about the models. Output mode and effort settings matter.
But they exposed a problem I was not measuring before.

We were both wrong, in different ways

A system needs at least three things:

  • Quality: Does it select the right answer?
  • Consistency: Do its score, probabilities, and structured fields agree?
  • Operational performance: Can it deliver that answer within acceptable latency and token use?

I set out to show that bounded decisions were too limiting. Instead, I found that the hard problem is not only choosing the right answer - it is producing an answer you can consistently trust.

Check the results or run your own benchmarks

The benchmark, decision contracts, published comparison summaries, chart source, and export tooling are available here:
https://github.com/Tenkei/jev-decision-bench

A reproducible, self-hosted benchmark for comparing JEV and LLM decision-making on your own datasets, models, and…github.com
The published results are intended to be inspectable and reproducible - not taken as a universal leaderboard. You can use the same workflow with your own model configurations or decision datasets.

Top comments (0)