<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Keivan Esbati</title>
    <description>The latest articles on DEV Community by Keivan Esbati (@tenkei).</description>
    <link>https://dev.to/tenkei</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4144960%2Fcbb00472-6baa-4393-859f-b89094ef4788.jpeg</url>
      <title>DEV Community: Keivan Esbati</title>
      <link>https://dev.to/tenkei</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tenkei"/>
    <language>en</language>
    <item>
      <title>Should You Let JEV Make Every Decision? Accuracy Matters More Than Speed.</title>
      <dc:creator>Keivan Esbati</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/tenkei/should-you-let-jev-make-every-decision-accuracy-matters-more-than-speed-3mcf</link>
      <guid>https://dev.to/tenkei/should-you-let-jev-make-every-decision-accuracy-matters-more-than-speed-3mcf</guid>
      <description>&lt;p&gt;I began with a simple thesis:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For decisions, accuracy matters more than&amp;nbsp;speed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;JEV is fast because it constrains a task to a bounded decision contract. I assumed that frontier LLMs would be more accurate - given their world knowledge - and that their extra latency would be an obvious trade worth making.&lt;/p&gt;

&lt;p&gt;So I built a reproducible benchmark to test that assumption against JEV on the same decision contracts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Intent routing (Choice)&lt;/li&gt;
&lt;li&gt;Moderation / guardrail classification (Noul)&lt;/li&gt;
&lt;li&gt;Search relevance (Score)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The result was more interesting than "JEV wins" or "frontier models win."&lt;/p&gt;

&lt;h2&gt;
  
  
  My initial case: accuracy should beat&amp;nbsp;speed
&lt;/h2&gt;

&lt;p&gt;On intent routing, the best tested frontier models did outperform JEV:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkz05cq61j8zq2jkjdey0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkz05cq61j8zq2jkjdey0.png" alt="Results from bench&amp;nbsp;tests" width="799" height="184"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That supports my original concern. If a decision is wrong, being fast does not make it useful.&lt;/p&gt;

&lt;p&gt;But the advantage did not generalize cleanly.&lt;/p&gt;

&lt;p&gt;On &lt;strong&gt;HateCheck&lt;/strong&gt; moderation, JEV reached &lt;strong&gt;98.2%&lt;/strong&gt;; the tested models ranged from &lt;strong&gt;95.1% to 100.0%&lt;/strong&gt;. On &lt;strong&gt;TREC&lt;/strong&gt; search relevance, JEV's &lt;strong&gt;47.4%&lt;/strong&gt; nearest-tier accuracy tied the best tested result.&lt;/p&gt;

&lt;p&gt;So JEV was right about something I had underestimated: a bounded decision system can be very fast without automatically giving up meaningful accuracy.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual contract&amp;nbsp;matters
&lt;/h2&gt;

&lt;p&gt;For the relevance task, the input is intentionally small and explicit: a query, a candidate passage, and a rubric.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{
  "state": {
    "query": "what is the capital of france",
    "passage": "Paris is the capital and most populous city of France."
  },
  "levels": [
    "Not Relevant",
    "Related",
    "Highly Relevant",
    "Perfect"
  ]
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model has to return both a continuous score and its belief across the four levels:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{
  "score": 2.99,
  "probabilities": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.01,
    "3": 0.99
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That response is coherent: the probability-weighted score is also approximately 2.99.&lt;/p&gt;

&lt;h2&gt;
  
  
  The unexpected failure mode: plausible, valid JSON that disagrees with&amp;nbsp;itself
&lt;/h2&gt;

&lt;p&gt;This is where my original framing broke down.&lt;/p&gt;

&lt;p&gt;A model can return valid JSON, a plausible final score, and a plausible-looking probability distribution - yet have those two fields contradict one another:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{
  "score": 2.99,
  "probabilities": {
    "0": 0.10,
    "1": 0.60,
    "2": 0.25,
    "3": 0.05
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The score says "almost Perfect." The distribution says "mostly Related." Both fields cannot be true at once.&lt;/p&gt;

&lt;p&gt;That is not merely a formatting issue. It means a downstream system has to decide which answer to trust: the final decision, or the model's stated uncertainty.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark changed the&amp;nbsp;question
&lt;/h2&gt;

&lt;p&gt;On the relevance experiment, score/distribution consistency was:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhreo48xqmados6gdsqh1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhreo48xqmados6gdsqh1.png" alt="Results from bench test" width="799" height="184"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These are results for the tested configurations, not universal claims about the models. Output mode and effort settings matter.&lt;br&gt;
But they exposed a problem I was not measuring before.&lt;/p&gt;

&lt;h2&gt;
  
  
  We were both wrong, in different ways
&lt;/h2&gt;

&lt;p&gt;A system needs at least three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Quality: Does it select the right answer?&lt;/li&gt;
&lt;li&gt;Consistency: Do its score, probabilities, and structured fields agree?&lt;/li&gt;
&lt;li&gt;Operational performance: Can it deliver that answer within acceptable latency and token use?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I set out to show that bounded decisions were too limiting. Instead, I found that the hard problem is not only choosing the right answer - it is producing an answer you can consistently trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the results or run your own benchmarks
&lt;/h2&gt;

&lt;p&gt;The benchmark, decision contracts, published comparison summaries, chart source, and export tooling are available here:&lt;br&gt;
&lt;a href="https://github.com/Tenkei/jev-decision-bench" rel="noopener noreferrer"&gt;https://github.com/Tenkei/jev-decision-bench&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A reproducible, self-hosted benchmark for comparing JEV and LLM decision-making on your own datasets, models, and…github.com&lt;br&gt;
The published results are intended to be inspectable and reproducible - not taken as a universal leaderboard. You can use the same workflow with your own model configurations or decision datasets.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>jev</category>
      <category>agents</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
