<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: TuringCorp</title>
    <description>The latest articles on DEV Community by TuringCorp (@turingcorp).</description>
    <link>https://dev.to/turingcorp</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4131616%2Fbdbe1583-6168-429d-994c-221eac3deb28.png</url>
      <title>DEV Community: TuringCorp</title>
      <link>https://dev.to/turingcorp</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/turingcorp"/>
    <language>en</language>
    <item>
      <title>Who Wrote the Ballot: The Judgment That Happens Before the Judge Is Called</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Fri, 09 Oct 2026 01:10:11 +0000</pubDate>
      <link>https://dev.to/turingcorp/who-wrote-the-ballot-the-judgment-that-happens-before-the-judge-is-called-3j6g</link>
      <guid>https://dev.to/turingcorp/who-wrote-the-ballot-the-judgment-that-happens-before-the-judge-is-called-3j6g</guid>
      <description>&lt;h1&gt;
  
  
  Who Wrote the Ballot: The Judgment That Happens Before the Judge Is Called
&lt;/h1&gt;

&lt;p&gt;Two sites are on the shortlist for a second location. The rent is modeled, the foot traffic has been counted for a month, the fit-out has three quotes, and the spreadsheet has been open for six weeks. Both sites work. The argument in the room is genuinely about which one, and it is a real argument, decided on real numbers.&lt;/p&gt;

&lt;p&gt;What nobody says out loud is that the option &lt;em&gt;do not open a second location this year, and fix the throughput problem at the first one&lt;/em&gt; was never written on the list.&lt;/p&gt;

&lt;p&gt;That is not a failure of analysis. The analysis was careful. It is a failure at a step that happens before the analysis, that is almost never recorded as a step, and that has already determined more of the outcome than every row of the spreadsheet beneath it. The step is deciding what gets to be on the ballot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two acts, and only one of them can be bought
&lt;/h2&gt;

&lt;p&gt;Split any decision into two acts.&lt;/p&gt;

&lt;p&gt;The first is nomination: which options are admissible. The second is ranking: of the options admitted, which is better. They are performed by different processes, at different times, with different evidence. Only the second one is what every judgment product on the market, ours included, actually does.&lt;/p&gt;

&lt;p&gt;A judge receives the option set as an input, the way a calculator receives numbers. It cannot widen the set, because the set is not in its input space. It cannot report that the set is wrong, because nothing in the interface has a slot for that finding. What it can do is answer the question it was handed, precisely and quickly. That is not a limitation anybody is concealing. It is the shape of the category.&lt;/p&gt;

&lt;p&gt;Earlier in this series I argued that it is a virtue of typed judgment that the set of possible answers is defined by the caller. This piece is the invoice for that virtue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every confidence number is conditional on the list
&lt;/h2&gt;

&lt;p&gt;Read a probability the way a statistician reads it and the missing part becomes obvious. A confidence of 0.9 is not a property of the world. It is a claim that among cases resembling this one, &lt;strong&gt;and among the options you supplied&lt;/strong&gt;, the answer named is right about nine times in ten. The conditioning event is doing half the work, and it is invisible in the output. The number prints to two decimals either way.&lt;/p&gt;

&lt;p&gt;So a figure can be honestly measured, published, and still describe a comparison that should never have been run. The instrument is not lying. The question was malformed before the instrument was switched on, and no amount of care inside the comparison recovers an option that was never compared.&lt;/p&gt;

&lt;p&gt;Our published measurements are a clean illustration of how tightly the conditioning binds. They are self-run on named benchmarks, with the failures disclosed rather than quietly retried. On JudgeBench: 620 judgments, with the 6 first-verdict failures disclosed. Raw accuracy came out at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we publish it as a tie, with no accuracy claim resting on it. The bands are the part worth keeping: calls we reported at 90% confidence or above came back right 99.6% of the time, and the 80–90% band 94.0%. On ContextualJudgeBench, self-run under the official pairwise protocol in which a pair counts as correct only if both presentation orders are judged correctly, which is why the random floor is 25% rather than 50%, consistent accuracy is 67.1% against the benchmark's official reference of 65.4%, with 12 orders excluded after repeated platform failures and that exclusion stated beside the result. The deliberately constructed near-tie splits sit at 46–60%.&lt;/p&gt;

&lt;p&gt;Now look through that paragraph for the row that would tell you whether the right answer was among the two candidates. It is not there, and it cannot be. Every figure measures the ranking act, on questions whose admissible answers were fixed in advance by somebody who does not appear in the measurement at all. That is a genuine measurement, and it covers roughly half of what a decision is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Better ranking makes a bad ballot harder to challenge
&lt;/h2&gt;

&lt;p&gt;There is a comfortable story in which framing is simply the next unmeasured frontier, and better instruments will eventually reach it. The uncomfortable version is worse than that.&lt;/p&gt;

&lt;p&gt;Improving the ranking act does not merely fail to catch a bad ballot. It launders it. A pick reported at 0.91, supported by a published band table, an exclusion count and an owner's name, is a much harder thing to challenge in a review than a shrug. The measured apparatus confers authority on the pair it was pointed at, and the pair inherited that authority by being the only pair on the page. The better the ranking, the more legitimate the framing looks.&lt;/p&gt;

&lt;p&gt;That is a risk of our own product, not only of somebody else's, and it is the reason this piece exists. A careful comparison is the best disguise an unexamined list has ever had. A series that has spent two weeks asking judges to publish their curves should say plainly what a curve measures: the comparison, not the choice of what to compare.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the options actually go missing
&lt;/h2&gt;

&lt;p&gt;Omissions are not random, and naming the patterns is most of the defense.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The status quo is not printed.&lt;/strong&gt; Doing nothing, keeping the current system, waiting a quarter, running a two-month pilot before committing — the most common real outcomes are rarely written on a ballot, because they look like the absence of a decision rather than a decision. Once they are off the page, the deliberation proceeds as though one of the listed items must win. And it usually does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The option that costs somebody in the room is the one missing.&lt;/strong&gt; A missing item is disproportionately the one that would require a person present to conclude that an earlier call of theirs was wrong, or that a capability they own has quietly become the problem. Omission is a political act performed in analytical vocabulary, which is exactly why it survives review: no scoring rubric can flag a row that was never created.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The third way gets read as a compromise and dismissed.&lt;/strong&gt; Real options often sit in a different frame rather than between the two on the list: not vendor A against vendor B, but neither, plus one internal quarter spent learning the domain. A two-item ballot converts a framing question into a preference question, and a false dilemma is durable because it looks like focus.&lt;/p&gt;

&lt;h2&gt;
  
  
  An artifact for the part nobody measures
&lt;/h2&gt;

&lt;p&gt;A ballot is a document, so it can be written, dated and signed like other documents.&lt;/p&gt;

&lt;p&gt;Write the list before the criteria. Criteria written first quietly determine which options are capable of scoring well.&lt;/p&gt;

&lt;p&gt;Force the null onto every ballot: do nothing, keep the status quo, decide later. Print it. If it is genuinely not viable, the reason belongs on the page beside it.&lt;/p&gt;

&lt;p&gt;Give the omission an owner. Before the comparison starts, ask who pays if something is missing from the list, and give that person a documented way to add a row. A line on the page, not a veto.&lt;/p&gt;

&lt;p&gt;Commission the list from somebody who disagrees with the likely conclusion, and ask what they would add. Two people independently producing the same list is the cheapest framing check available.&lt;/p&gt;

&lt;p&gt;Then ask what the options have in common. If both candidates assume the same thing — that the capability is needed at all, that the market is where it was last year, that this is a hiring problem rather than a scoping problem — the shared assumption is the decision, and the comparison is downstream of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the judge's responsibility ends
&lt;/h2&gt;

&lt;p&gt;A pairwise judge takes one question and two candidate answers. It can tell you which of the two is stronger and how much confidence it places in that. It cannot tell you the pair was incomplete, and it should not pretend otherwise. The ballot is not an optional preamble to the work. It is the input, and it is the half of the decision the instrument is silent about.&lt;/p&gt;

&lt;p&gt;So the honest division is this. Counting carefully is a solved problem, and it keeps getting cheaper; that is the whole achievement of the fast end of this field. What stays expensive, and stays unmeasured, is who was allowed onto the ballot. Write the list down first, put the null option on it, let the person who would pay for an omission sign the page, and only then run the comparison.&lt;/p&gt;

&lt;p&gt;If you are the one who has to defend a list as well as a pick — one question, two candidate answers you cannot separate by inspection — the judge that spends minutes on the ranking is here: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-13" rel="noopener noreferrer"&gt;https://api.turingcorp.net/platform/go/decider?src=jev-13&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>No Reference Class: The Decisions Your Past Has Never Made</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Thu, 08 Oct 2026 01:10:06 +0000</pubDate>
      <link>https://dev.to/turingcorp/no-reference-class-the-decisions-your-past-has-never-made-ao8</link>
      <guid>https://dev.to/turingcorp/no-reference-class-the-decisions-your-past-has-never-made-ao8</guid>
      <description>&lt;h1&gt;
  
  
  No Reference Class: The Decisions Your Past Has Never Made
&lt;/h1&gt;

&lt;p&gt;An underwriter who has quoted a thousand cargo policies will quote the thousand-and-first before lunch. Ask the same underwriter to quote the first hull built to a new design, with no voyages behind the design, no claims record, and no sister ship anywhere in the book, and something changes. They go quiet. They ask for time. Nobody in that room would say the second question is harder to &lt;em&gt;compute&lt;/em&gt;. It is harder because there is nothing behind it to lean on.&lt;/p&gt;

&lt;p&gt;That pause is the subject of this piece, because it is the one thing a judge that answers in milliseconds cannot be improved into doing. Not because it is slow, and not because it is costly. Because the material it is made of runs out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two different reasons a judge says it does not know
&lt;/h2&gt;

&lt;p&gt;Take any model that returns a choice with a confidence value attached, ours included, and separate the ways it can be unsure.&lt;/p&gt;

&lt;p&gt;The first is statistical. The case is ordinary, the precedents point in different directions, both candidates survive a careful reading, and the reported number lands near the middle. You have seen this shape before and so has the model. The reference class exists and it is crowded.&lt;/p&gt;

&lt;p&gt;The second is structural. The case has no precedent at all. Not a contested one — none. Your company has never made this call, the industry has no settled answer, and the closest thing in the training data is a different kind of decision that happens to share vocabulary.&lt;/p&gt;

&lt;p&gt;On a dashboard these look identical. A low number is a low number. They are not the same finding, and only one of them can be fixed by collecting more data.&lt;/p&gt;

&lt;h2&gt;
  
  
  A confidence value is a claim about a population
&lt;/h2&gt;

&lt;p&gt;Read the number slowly. When a system reports 0.9, it is claiming that among cases that look like this one, roughly nine in ten turn out the way it says. The phrase doing all the work is &lt;em&gt;cases that look like this one&lt;/em&gt;. A probability is not a property of a question. It is a lookup into a set of comparable questions that have already been answered.&lt;/p&gt;

&lt;p&gt;When that set is large, the number is meaningful and testable: count the outcomes and compare. That is what a reliability table is for, and it is why a vendor that will not publish one has not measured anything. When the set is empty, the number is still printed, still formatted to two decimals, and still a number. It has simply stopped referring to anything. What it inherits is the shape of the nearest population the model has seen — a different question, with a different base rate, and consequences that land somewhere else.&lt;/p&gt;

&lt;p&gt;Be precise about how this differs from the more familiar complaint about distributions. That complaint says the population you are running on is not the population the number was measured on: the mix is different, the base rate has moved, and the curve is stale. It is a serious problem and it is fixable in principle, by re-measuring on your own cases. The case here is not that. It is not a mismatched population, which can be measured again. It is an absent one, which cannot be measured at all.&lt;/p&gt;

&lt;p&gt;So the boundary is not "the model is not smart enough yet." A larger model, trained on ten times as much of your history, still has no answer for a case that has never occurred, because the answer is not in the history. It has not been made yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the reference class exists, we can be held to it
&lt;/h2&gt;

&lt;p&gt;Everything above would be cheap rhetoric if the people making the argument refused to be measured on the part that is measurable.&lt;/p&gt;

&lt;p&gt;Ours is self-run under named benchmarks, with failures disclosed rather than quietly retried. On JudgeBench: 620 judgments, with the 6 first-verdict failures disclosed. Raw accuracy came out at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we report it as a tie; the claim here is not a sharper judge. The bands are the part worth keeping: calls we reported at 90% confidence or above came back right 99.6% of the time, and the 80–90% band 94.0%. On ContextualJudgeBench, self-run under the official pairwise protocol where a pair counts as correct only if both presentation orders are judged correctly — which is why the random floor is 25% rather than 50% — consistent accuracy is 67.1% against the benchmark's official reference of 65.4%, with 12 orders excluded after repeated platform failures and that exclusion stated next to the result. The deliberately constructed near-tie splits sit at 46–60%.&lt;/p&gt;

&lt;p&gt;Put that last figure where this argument needs it. A near-tie is a case whose comparable set is internally contradictory: the population exists, and it points both ways at once. At 46–60% the instrument is reporting that the reference class has dissolved into noise. That is not a defect to be hidden. It is the only honest signal a frequency-based judge can send at the moment frequency stops carrying information — and it is why the useful output there is not a better number but a person.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fast end got its benchmark within a week
&lt;/h2&gt;

&lt;p&gt;There is a telling asymmetry in how quickly the two kinds of judgment acquired measuring instruments.&lt;/p&gt;

&lt;p&gt;Within days of Jev appearing, the ecosystem was building comparisons for the precedent-bound part. "Show HN: JevBench, a reproducible benchmark for typed decision models" reached 154 points on Hacker News on September 22, and "Jeeves. Reasoning improves Jev-like decision models" reached 242 points on September 29. Those are per-post scores read off the public search index, and they drift by a few points over time; the pattern does not. Where the question is "which of these labels fits," a corpus is free, the labels are already attached, and a benchmark can be assembled in an afternoon by someone with no access to the original model.&lt;/p&gt;

&lt;p&gt;For unprecedented decisions there is no equivalent, and the absence is not laziness. You cannot assemble a corpus of decisions that have not been made. You cannot hold a benchmark constant when the ground truth is the thing being selected. The fast end has a leaderboard because it has a yardstick. The slow end has neither, which is why the slow end is still argued about in prose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Unfamiliar is not the same as unprecedented
&lt;/h2&gt;

&lt;p&gt;The mistake in the other direction is more common, and it burns more minutes than the first one.&lt;/p&gt;

&lt;p&gt;Most decisions that feel novel are merely unfamiliar to the person holding them. The pricing question that has never come up inside your company has been answered a thousand times in your industry. The build-or-buy call that feels unprecedented is in somebody's public post-mortem. Novelty that exists in your head but not in the world is a millisecond decision wearing a costume, and treating it as a minute decision is how a team spends a quarter rediscovering a base rate that was already published.&lt;/p&gt;

&lt;p&gt;So the discipline has to run in both directions, and it comes down to three questions.&lt;/p&gt;

&lt;p&gt;Does a comparable case exist anywhere — your own history, a competitor's, a dataset, a public record? If it does, this is a lookup, and a lookup should be done cheaply and at volume.&lt;/p&gt;

&lt;p&gt;Would more data change the answer? If the answer is a fact about the world, then data helps and you should go get it. If the answer depends on what you want to become — which of two futures you are willing to own — then no additional data settles it, because the data describes the world you have, not the one you are choosing.&lt;/p&gt;

&lt;p&gt;Will somebody cite this decision later? If the next person facing this exact call will point at what you did as their precedent, then you are not looking up an answer. You are writing the first entry in a record that will be read by people who were not in the room.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can hand to a decision that has no precedent
&lt;/h2&gt;

&lt;p&gt;The second kind of decision needs a different artifact from the first kind.&lt;/p&gt;

&lt;p&gt;It does not need a probability, because there is no population to be right about. It needs to be written as a first case rather than a verdict: the reasoning spelled out in sentences that a stranger could carry to the next instance of the same question and land in the same place. It needs both candidates laid out, not only the one you already prefer. It needs the argument written down, because the written argument is the part that can be disagreed with, corrected, and cited by the next person who arrives here. And it needs a name, because the first entry in a record is also the moment liability attaches.&lt;/p&gt;

&lt;p&gt;That is the honest division of labor. The cheap end is a lookup, and lookups should be fast, numerous, and boring. The expensive end is a precedent being set, and a precedent is not retrieved — it is authored, once, by someone who can be asked about it later.&lt;/p&gt;

&lt;p&gt;We build for the second case, so it is fair to state plainly what we can and cannot put on the table. Where a reference class exists, we publish bands with counts and with exclusions, which is how you can tell whether a number we report is worth acting on. Where one does not exist, we do not have a calibrated probability to offer, and we will not dress one up. What we can produce is a pick between two candidate answers, a confidence stated together with the population it was measured against, and a written argument long enough to disagree with — which, for a decision whose entire difficulty is that no precedent covers it, is the artifact you actually need.&lt;/p&gt;

&lt;p&gt;If the call in front of you is unprecedented rather than merely hard — one question, two defensible answers, and a record that will be cited — that is the case this exists for: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-12" rel="noopener noreferrer"&gt;Decider&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>One Page: What to Hand to a Millisecond Judge, and What to Keep for the Minutes</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Wed, 07 Oct 2026 01:10:12 +0000</pubDate>
      <link>https://dev.to/turingcorp/one-page-what-to-hand-to-a-millisecond-judge-and-what-to-keep-for-the-minutes-m31</link>
      <guid>https://dev.to/turingcorp/one-page-what-to-hand-to-a-millisecond-judge-and-what-to-keep-for-the-minutes-m31</guid>
      <description>&lt;h1&gt;
  
  
  One Page: What to Hand to a Millisecond Judge, and What to Keep for the Minutes
&lt;/h1&gt;

&lt;p&gt;Nearly every decision tool sold today comes with one instruction attached: route the easy cases to a machine and keep the hard ones for a person. Nobody disagrees with it. Almost nobody can say where the line falls in their own work, and that gap is where the expensive mistakes live.&lt;/p&gt;

&lt;p&gt;So here is the whole thing on one page, as a list you can argue with and a table you can forward.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one check that settles most cases
&lt;/h2&gt;

&lt;p&gt;Before the mood, the stakes, or the meeting, ask one question about the decision itself:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If this turns out wrong, what do I have to spend to undo it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the honest answer is a re-run, a rollback, or a customer who gets an apology and the correct answer, this is a millisecond decision. Hand it over, let code act on the result, and do not put a person near it.&lt;/p&gt;

&lt;p&gt;If the honest answer is a signature, a relationship, a quarter, or a story you will have to tell later, this is a minute decision, no matter how fast an answer can be produced. Speed is available for both. It only helps one of them.&lt;/p&gt;

&lt;p&gt;Everything below is detail on that question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two lists
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Signals that the decision belongs in the millisecond regime&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It is reversible at roughly what it cost to make.&lt;/li&gt;
&lt;li&gt;The option set is closed, and you defined it.&lt;/li&gt;
&lt;li&gt;The same call happens hundreds or thousands of times, so a wrong branch is a statistic rather than an incident.&lt;/li&gt;
&lt;li&gt;A cheap downstream check would catch a wrong call before it reaches anyone.&lt;/li&gt;
&lt;li&gt;The cost of being wrong lands on a system: a retry, a log line, a bad row.&lt;/li&gt;
&lt;li&gt;The question can be written as a predicate over fields you already have.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Signals that it belongs in the minutes&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Undoing it costs more than doing it did.&lt;/li&gt;
&lt;li&gt;Both candidates survive an adversarial reading by someone who disagrees with you.&lt;/li&gt;
&lt;li&gt;A named person has to live with the outcome, and it is not you.&lt;/li&gt;
&lt;li&gt;The decider is a value that has not been written down yet, because the choice is between things measured in different units.&lt;/li&gt;
&lt;li&gt;It happens once or twice, and you will still be explaining it in five years.&lt;/li&gt;
&lt;li&gt;You cannot state the criterion that would settle it, only the feeling that one branch is heavier.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A decision needs one box from the first list to be a candidate for the millisecond regime, and it needs none from the second. When lists conflict, the second list wins. That is not caution; it is arithmetic. The cost of one wrong minute decision is usually larger than the saving from a thousand fast ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  The table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The decision&lt;/th&gt;
&lt;th&gt;Regime&lt;/th&gt;
&lt;th&gt;What you get&lt;/th&gt;
&lt;th&gt;Wrong-branch cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Which of four data sources to read&lt;/td&gt;
&lt;td&gt;Millisecond&lt;/td&gt;
&lt;td&gt;A value to switch on&lt;/td&gt;
&lt;td&gt;A re-run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whether to send this contract to the partner or back to the analyst&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;A pick, a written argument, a name&lt;/td&gt;
&lt;td&gt;The deal, or the relationship&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Which of six labels fits this ticket&lt;/td&gt;
&lt;td&gt;Millisecond&lt;/td&gt;
&lt;td&gt;A label plus a confidence&lt;/td&gt;
&lt;td&gt;A misrouted queue item&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whether to accept the terms on a partnership&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;A record you can defend&lt;/td&gt;
&lt;td&gt;A year of your capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whether the cache entry is still usable&lt;/td&gt;
&lt;td&gt;Millisecond&lt;/td&gt;
&lt;td&gt;A yes, remembered&lt;/td&gt;
&lt;td&gt;One stale render&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Which of two clients to keep when capacity is gone&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;A criterion you have written down&lt;/td&gt;
&lt;td&gt;The one you kept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whether to flip the feature flag&lt;/td&gt;
&lt;td&gt;Millisecond&lt;/td&gt;
&lt;td&gt;A branch taken, logged&lt;/td&gt;
&lt;td&gt;One page of noise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whether to hire for the gap or narrow the roadmap&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;A decision with an owner&lt;/td&gt;
&lt;td&gt;Twelve months&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whether to answer this support ticket or escalate it&lt;/td&gt;
&lt;td&gt;Millisecond&lt;/td&gt;
&lt;td&gt;A route&lt;/td&gt;
&lt;td&gt;Some customer patience&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whether to take the retainer that would make you one client's vendor&lt;/td&gt;
&lt;td&gt;Minutes&lt;/td&gt;
&lt;td&gt;A signed basis, in writing&lt;/td&gt;
&lt;td&gt;Optionality&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The same labels, the same ticketing systems and often the same models show up in both columns. What differs is the third and fourth columns, and those are the only ones a machine has never seen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trap: minute decision, millisecond authority
&lt;/h2&gt;

&lt;p&gt;This is the failure I would flag if you only remember one thing from this page.&lt;/p&gt;

&lt;p&gt;Teams often do the slow part correctly and then hand the result to the fastest actor in the building. The criteria get written down, the review happens, the reasoning is solid — and then the call is delegated to whoever is on call, because the deadline arrived. Now the paperwork says the decision was deliberated and the actual pick was made by somebody with three minutes, no context, and no ownership. You have paid the cost of a minute decision and received the quality of a millisecond one.&lt;/p&gt;

&lt;p&gt;The same trap has a second door. Escalate the &lt;em&gt;question&lt;/em&gt; to a person but not the &lt;em&gt;authority&lt;/em&gt; to answer it, and you get a queue of items that were produced because the machine could not decide, resolved by people who also cannot decide. That queue grows, ages, and eventually gets a decision made for it by the calendar. An unanswered question is not pending. It is being decided, by time, badly.&lt;/p&gt;

&lt;p&gt;The fix is small and structural. For every minute decision, name the picker before the deliberation starts, and give them either the authority or an explicit hand-off to whoever holds it. A minute decision without an owner is not slow. It is absent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the millisecond regime owes you, and what the minutes owe you
&lt;/h2&gt;

&lt;p&gt;The two regimes produce different artifacts, and confusing them is how good routing turns into an unmanageable queue.&lt;/p&gt;

&lt;p&gt;A millisecond call owes you a value, a confidence, and a log line with the confidence, the options and the threshold written into it. That log is the only way you will ever be able to tune the cut-off, and it is also your only evidence if the boundary was wrong in a way you cannot see. Nothing here needs to be readable by a person, and nothing here should be escalated because it feels important. If it feels important, it was filed in the wrong list.&lt;/p&gt;

&lt;p&gt;A minute call owes you a record. Not a better opinion — a record: the two options, the pick, how confident the pick is, the reasoning in prose, and the name of whoever owns it. The point of the record is not documentation for its own sake. It is that a decision made once, affecting named people, will be read again by somebody who was not in the room. That reading is when most slow decisions are actually revealed to have been a coin flip.&lt;/p&gt;

&lt;p&gt;A short way to hold all of this together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Millisecond decisions are judged by their rate.&lt;/strong&gt; One wrong call is noise. A wrong rate is a bug.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minute decisions are judged by their record.&lt;/strong&gt; One wrong call is a fact of life. No record is a failure of the process.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Neither regime produces the thing the other needs.&lt;/strong&gt; A verdict with a number attached will not survive to a review six weeks later, and a thoughtful memo will not route a ticket at four in the morning.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Treat a low confidence from the fast layer as a finding rather than a fault. A near-tie means the candidates do not contain the answer, and the criterion lives with the person. That is the signal to move the decision to the other list, and it arrives before you have spent a single minute on the deliberation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a minute-grade judge fits
&lt;/h2&gt;

&lt;p&gt;We build the second list, so we are the biased party on this page, and the table above is the honest version of the argument anyway.&lt;/p&gt;

&lt;p&gt;Our numbers, self-run with the failures disclosed rather than quietly retried. On JudgeBench, 620 judgments with 6 first-verdict failures disclosed, we measured 92.5% against 92.2% for a plain direct baseline. That is a tie, and we report it as a tie — this page is not a claim of a sharper judge. The bands are where the measurement earns its keep: 90% confidence or above came back right 99.6% of the time, and the 80–90% band 94.0%. On a second self-run corpus, under a scoring rule strict enough that a pair counts as correct only if both presentation orders are judged correctly, consistent accuracy is 67.1% against a 65.4% reference, with 12 excluded orders stated next to the result. The deliberately constructed near-tie splits sit at 46–60%, which is the most useful row on the page for the list above: it is the machine telling you to move the decision to the slow column, including when the machine is us.&lt;/p&gt;

&lt;p&gt;If you are holding something from the second list — one question, two candidate answers, both of which survive an adversarial reading — the minute-grade judge is here: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-11" rel="noopener noreferrer"&gt;https://api.turingcorp.net/platform/go/decider?src=jev-11&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line version
&lt;/h2&gt;

&lt;p&gt;Sort by what a wrong branch costs and who pays it, not by what the model reports. Hand the reversible, enumerable, high-volume, system-paid calls to the milliseconds. Keep the rest, write down the criterion, and put a name on the pick — because that is what the minutes are for, and they are the cheapest thing you will buy for a decision like that.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Show Me the Diagram</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Tue, 06 Oct 2026 01:10:18 +0000</pubDate>
      <link>https://dev.to/turingcorp/show-me-the-diagram-dm0</link>
      <guid>https://dev.to/turingcorp/show-me-the-diagram-dm0</guid>
      <description>&lt;h1&gt;
  
  
  Show Me the Diagram
&lt;/h1&gt;

&lt;p&gt;A vendor sends over an integration guide. It has three columns: a confidence value the system emits, the action you are meant to take at each level, and what each call costs. There is a cut-off at 0.8 and a note that below 0.6 you should escalate. The person who wrote it is competent and the system works.&lt;/p&gt;

&lt;p&gt;Ask one question: &lt;em&gt;show me the reliability diagram.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Not the average accuracy. Not a chart of how many calls the system makes. A plot with reported confidence on one axis, how often the system was actually right on the other, and the number of cases behind each bin. The artifact that says: when this thing said 0.9, it was right this often; when it said 0.6, this often; and here is how many times it said each.&lt;/p&gt;

&lt;p&gt;If that plot does not exist, you have not received a measurement. You have received three columns of design opinion with numbers in them. The guidance may still be good. But the one thing you were going to build a policy on — &lt;em&gt;what a 0.9 means&lt;/em&gt; — has not been established anywhere, by anyone, at any point. And you will still be switching on it in production, because the number is right there and it looks like the other numbers your systems use.&lt;/p&gt;

&lt;p&gt;That is the whole complaint, and it applies to every judge, including ours. A confidence value with no curve behind it is not a judgment. It is a feeling with a decimal point.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the number is claiming
&lt;/h2&gt;

&lt;p&gt;Read a confidence value slowly. When a system reports 0.9, it is claiming that among cases that look like this one, about 90 in 100 turn out the way it says. That is a factual claim about a frequency, and frequencies can be counted. Nothing about the interface tells you whether anyone counted.&lt;/p&gt;

&lt;p&gt;The claim is also not the same as accuracy, which is why swapping one for the other goes wrong so quietly. An average score answers "how often is this thing right across everything we tested." A calibration curve answers "what does this particular number buy me in error rate." Only the second one is usable at a cut-off, and it is almost never the one that ships with the product.&lt;/p&gt;

&lt;p&gt;Two systems can post identical accuracy and be worth very different amounts to you. One can be right 88% of the time and honest about which cases it does not know. The other can be right 88% of the time and report 0.95 on everything. The first one can be routed. The second one cannot, because a constant cannot separate cases — a gate built on it fires the same way every time, which means the gate is decoration. The score did not distinguish them. Only the curve does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calibration belongs to the population, not to the model
&lt;/h2&gt;

&lt;p&gt;Here is the part that turns "show me the diagram" from a procurement ritual into an actual question.&lt;/p&gt;

&lt;p&gt;We publish ours, and the least convenient number in it is a comparison. On JudgeBench — self-run under the official judging protocol, 620 judgments, with the 6 pairs whose first verdict failed disclosed rather than quietly retried — calls we reported at 90% confidence or above came back right &lt;strong&gt;99.6%&lt;/strong&gt; of the time, and calls in the 80–90% band came back right &lt;strong&gt;94.0%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Then the same interface, the same scoring rule, a different corpus. On ContextualJudgeBench, self-run over the full official set under the pairwise protocol where a pair counts as correct only if both presentation orders are judged correctly — which is why the random floor is 25% rather than 50% — with 12 orders excluded after repeated platform failures and that exclusion stated next to the result, our consistent accuracy is &lt;strong&gt;67.1% against the benchmark's official reference value of 65.4%&lt;/strong&gt;. The deliberately constructed near-tie splits in that benchmark sit at &lt;strong&gt;46–60%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Now put the two runs side by side. Same judge. Same confidence field. On the second corpus, the top reported band is worth dramatically less than it is on the first. That benchmark was built to contain genuinely close pairs, and a population full of close pairs is a population where "90% sure" does not buy what it bought last week.&lt;/p&gt;

&lt;p&gt;Nobody behaved badly to produce that gap. It falls out of the arithmetic. Calibration is a property of a model &lt;em&gt;and a population&lt;/em&gt;: it depends on the mix of cases you feed in, how many of them are genuinely decidable, and how the base rate sits. The number is not a fixed property of the software, the way latency or price are. It is a measurement of a distribution, and the distribution is partly yours.&lt;/p&gt;

&lt;p&gt;Which gives you the three things you actually wanted from that vendor in the first place. The curve is a claim about a corpus that is not your traffic. The curve is a claim about a corpus you can partly describe but cannot fully inspect. And the curve is only as current as the model version it was measured on — a retrained model with an unchanged interface invalidates it silently.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a diagram has to contain
&lt;/h2&gt;

&lt;p&gt;Not every chart is evidence. Three details carry most of the weight, and their absence is more informative than the headline number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The bin edges and the count in each bin.&lt;/strong&gt; A bin with four cases in it is an anecdote with axis labels. A bin with three hundred cases is a rate. Most published calibration summaries show neither, which leaves you unable to tell which of the two you are reading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happened to the failures.&lt;/strong&gt; Which cases were excluded, rerun, dropped or imputed, and how many. Ours, printed next to the results rather than in a footnote: 6 pairs on JudgeBench whose first verdict failed, 12 orders rerun-excluded on ContextualJudgeBench, out of the full 2,000-pair official set. A protocol that names its exclusions is telling you how far to trust the rows above them. A protocol with no exclusions to name has usually not looked, rather than having nothing to report.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The label set, the population and the date.&lt;/strong&gt; A calibration curve is meaningless without the corpus it was measured on and when. This is the detail that kills almost every diagram in circulation, ours included, because it is the one that makes the curve somebody's measurement of something specific instead of a general property of the product.&lt;/p&gt;

&lt;p&gt;One caution that applies to any single-number summary: a scalar can hide the only bin you care about. An aggregate error figure that averages across a corpus tells you about average performance and nothing about the band your threshold actually sits in. The bin table is the artifact; the single number is a convenience that should never be the thing a decision rests on.&lt;/p&gt;

&lt;h2&gt;
  
  
  We are asking to be measured with the same ruler
&lt;/h2&gt;

&lt;p&gt;It would be easy to read all of this as an argument against somebody else's product, so let us remove that reading.&lt;/p&gt;

&lt;p&gt;TypeSafe has not published a reliability diagram or an expected calibration error for Jev. That is a statement about what is publicly available, not about what is true internally. Their recommended usage is threshold routing — act above a line, send the rest to a person or to a slower judge — and that recommendation is sound. It simply puts the reliability question directly on the critical path of the intended deployment, and leaves the curve unpublished while it does.&lt;/p&gt;

&lt;p&gt;The reason I am comfortable writing that sentence is that the same ruler is pointed at us, and it reads worse in one specific place. The near-tie splits on our own second benchmark come back at 46–60%. That is the least flattering figure we have, it is on our own page, and it is the one that makes the other rows mean anything, because a vendor that publishes only its best row has told you which row it wants read.&lt;/p&gt;

&lt;p&gt;So the invitation is literal, and it is not a rhetorical flourish. Our bins are published with their counts. Our exclusions are numbered. The near-ties are published alongside the good rows. If you take our curve and redraw it on your own traffic, your own bin edges and your own consequences, and the bands do not hold, that is a more useful result than any of our rows — and we would rather have it than not. We have been asked for the diagram, we published it, and anyone holding the same ruler should apply it here first.&lt;/p&gt;

&lt;p&gt;The practical version, if you are the one signing: for any judge you are considering, ask for the plot, the counts, the exclusions and the date. Price the bands against your own error budget rather than someone else's average. Then write the cut-off down with the name of the person who chose it, because a threshold is a commitment and not a setting.&lt;/p&gt;

&lt;p&gt;And if the diagram turns out to exist and looks ordinary — an honest curve with a soft middle and a low band that says &lt;em&gt;this one is close&lt;/em&gt; — do not treat that as a disappointment. It is the most valuable thing a decision tool can hand you. It tells you where the machine stops and the minutes begin.&lt;/p&gt;

&lt;p&gt;If you want to point that ruler at a judge that takes one question and two candidate answers and returns a pick with a calibrated confidence and a written argument you can disagree with, the entry point is here: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-10" rel="noopener noreferrer"&gt;https://api.turingcorp.net/platform/go/decider?src=jev-10&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>When Every Gateway Ships Its Own Decision Model</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Mon, 05 Oct 2026 01:10:08 +0000</pubDate>
      <link>https://dev.to/turingcorp/when-every-gateway-ships-its-own-decision-model-3o1</link>
      <guid>https://dev.to/turingcorp/when-every-gateway-ships-its-own-decision-model-3o1</guid>
      <description>&lt;h1&gt;
  
  
  When Every Gateway Ships Its Own Decision Model
&lt;/h1&gt;

&lt;p&gt;A decision model used to be a product you called. Now it is a thing you host.&lt;/p&gt;

&lt;p&gt;On September 15, TypeSafe introduced Jev and the announcement went to the front page of Hacker News. What followed did not look like the usual adoption curve, where a new capability stays a dependency for a year before anyone builds on it. It became a genre in about two weeks.&lt;/p&gt;

&lt;p&gt;I pulled the current numbers from the Hacker News search API (&lt;code&gt;hn.algolia.com/api/v1/search?query=%22jev%22&amp;amp;tags=story&amp;amp;numericFilters=points&amp;gt;100&lt;/code&gt;) instead of quoting a snapshot, because this table keeps moving. The official announcement post is at 1,989 points and 520 comments. Then the copies and the riffs: &lt;strong&gt;Jev in 25 Lines of Python&lt;/strong&gt; (691 points, Sep 23), &lt;strong&gt;Ollaya – Ollama for open-source, Jev-style decision models&lt;/strong&gt; (614, Sep 25), &lt;strong&gt;Kev: Tiny Jev-like family of decision models&lt;/strong&gt; (462, Sep 21), &lt;strong&gt;OpenAI is well positioned to fast-follow Jev&lt;/strong&gt; (328, Sep 22), and &lt;strong&gt;Reverse-engineered Jev-like model&lt;/strong&gt; (169, Sep 16).&lt;/p&gt;

&lt;p&gt;The dates are the interesting column. &lt;code&gt;Reverse-engineered Jev-like model&lt;/code&gt; landed on September 16, one day after the announcement — a working imitation of the interface, built without the training recipe, without the weights, without any cooperation from the people who made it. By September 23 the interface fit in twenty-five lines of Python. By September 28 a small model trained at home was answering in about thirty milliseconds. The September 28 clone scored 571 points, more than the earliest family of models managed a week after it appeared. A late entrant beating early ones at the only game the front page scores, attention, is not a saturation curve.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the fast copies actually prove
&lt;/h2&gt;

&lt;p&gt;The obvious reading is that something got stolen. That reading is wrong, and being precise about why matters.&lt;/p&gt;

&lt;p&gt;What those two weeks demonstrated is that the &lt;em&gt;interface&lt;/em&gt; is now a known shape. It fits in a sentence: take a program state and typed questions, return a structured choice with a score and a confidence, write no prose. Twenty-five lines of Python is enough to re-implement that, and a weekend is enough to re-train a small open model against it. Nobody needed permission, and nobody needed to be clever.&lt;/p&gt;

&lt;p&gt;That is a genuine loss of a moat. If your business is &lt;em&gt;being the only place that offers typed judgment at low latency&lt;/em&gt;, that business just ended. It was not undercut. It was generalized.&lt;/p&gt;

&lt;p&gt;The part that is easy to miss from inside a company that builds one of these: this is mostly good. An interface many people can implement gets implemented everywhere — in gateways, routers, CI pipelines, and the classifier slot that used to hold a fine-tuned BERT and a lot of maintenance. The copies are not free-riding. They are the original becoming infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The layer that gets eaten, and why you should let it
&lt;/h2&gt;

&lt;p&gt;Every one of the clones is a &lt;strong&gt;stateless transformation with a clean boundary&lt;/strong&gt;. Text and a question go in, a structured judgment comes out, and nothing downstream depends on whether that judgment was &lt;em&gt;good&lt;/em&gt; — only on whether it was &lt;em&gt;well-formed&lt;/em&gt;. Three properties make that shape easy to commoditize. The interface is narrow, so the surface to reproduce is small. The volume is enormous, so the margin per call was always going to be competed toward the cost of compute. And errors are cheap and re-runnable: a wrong classification on a product listing is a bad row you re-run next week and never sign.&lt;/p&gt;

&lt;p&gt;When a layer has those three properties, commoditization is not a threat to defend against. It is a service. Judgment becomes a commodity the way TLS certificates and JSON parsing did: something nobody thinks about, priced near zero, available at the edge, on by default.&lt;/p&gt;

&lt;p&gt;I will be direct about the self-interest here. We sell judgment. If the only thing we sold were that stateless transformation, this article would be a eulogy. It is not, and the reason is in the next section.&lt;/p&gt;

&lt;h2&gt;
  
  
  What did not get copied
&lt;/h2&gt;

&lt;p&gt;Go back to the clones and ask a question the point totals do not answer: &lt;em&gt;how do you know whether any of them is right?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;You cannot tell from the interface, because the interface has no place to put that information. &lt;code&gt;choice&lt;/code&gt;, &lt;code&gt;score&lt;/code&gt;, &lt;code&gt;confidence&lt;/code&gt; — those are outputs, not evidence. A model that reports 0.9 on everything and a model whose 0.9 means something produce byte-identical JSON for the same input. The interface is silent on the only thing a buyer eventually needs to know.&lt;/p&gt;

&lt;p&gt;So two properties stay behind when the shape gets copied, and neither is protected by a clever trick. They are protected by being &lt;em&gt;work&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verification&lt;/strong&gt; is the first. It means a record that says: here is the data this number was computed on, here is the counting rule, here is what was excluded and why, and here is what happens when you change the order the candidates are presented in. That record cannot be reverse-engineered from an API, because it is not a property of the API. It is a property of a measurement someone had to run, on a named benchmark, under a stated protocol, and publish even where it looks bad. The thing to check is order sensitivity: does the judgment hold when you swap the two candidates? A confidence number that moves because you re-ordered the same two options is not a confidence number. It is a position in a sequence.&lt;/p&gt;

&lt;p&gt;We publish ours, which makes this a claim you can audit rather than a principle you take on faith. Self-run on JudgeBench, 620 judgments, with the 6 that failed on a first verdict disclosed rather than quietly retried, raw accuracy came out at 92.5% against 92.2% for a direct model baseline. That is a tie. We publish it as a tie, and we claim no accuracy advantage over a direct model call, or over anyone else. The bands are where the measurement earns its keep: calls reported at 90% confidence or above were right 99.6% of the time, and calls in the 80–90% band were right 94.0%. On ContextualJudgeBench, self-run under the official protocol where a pair counts as correct only if both presentation orders are judged correctly — which is why the random floor is 25% rather than 50% — consistent accuracy is 67.1% against the benchmark's official reference, with 12 orders excluded after repeated platform failures and that exclusion stated next to the result. The deliberately constructed near-tie splits sit at 46–60%.&lt;/p&gt;

&lt;p&gt;That last figure is the least flattering number we have, and it is the one that makes the others mean something. A vendor that publishes only the top row tells you which row it wants read.&lt;/p&gt;

&lt;p&gt;None of this is a moat in the patent sense. Any competitor can run the same measurement tomorrow, and if they do, the category gets more legible. What they cannot do is run it &lt;em&gt;retroactively&lt;/em&gt;, because the record is a history with dates on it. The number you can produce today is not a substitute for the number someone published before they had a reason to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Responsibility&lt;/strong&gt; is the second, and it is the one that moves the economics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the accountable end gets more expensive, not less
&lt;/h2&gt;

&lt;p&gt;Here is where this is &lt;em&gt;not&lt;/em&gt; the Jevons argument, and the difference belongs on the record. The Jevons case is about a price mechanism: make a unit of work cheaper, you buy more units, and total spend on the remaining hard cases rises even as the easy ones approach free. That mechanism is real, and it is about volume.&lt;/p&gt;

&lt;p&gt;What happens to verification and ownership is a different mechanism. It is about &lt;em&gt;scale&lt;/em&gt; — or rather the absence of it.&lt;/p&gt;

&lt;p&gt;Verification does not get cheaper when you add customers. The hundredth deployment does not make the reliability curve for the first one any easier to produce. Each materially different deployment — different data, different questions, different cost of being wrong — needs its own curve, and a curve is a measurement, which is someone's afternoon with a benchmark harness and a decision about what to exclude. No version of this batch-processes.&lt;/p&gt;

&lt;p&gt;Responsibility fails to scale at all. When a judgment runs automatically and turns out wrong, the question is not "what was the confidence." It is "who decided this was good enough to run on its own, and on what evidence." That has an answer only if a specific party put its name on a specific curve and said: &lt;em&gt;above this line, the machine acts; below it, a person does.&lt;/em&gt; The name does not get cheaper when volume goes up. A signature is a transfer of liability, and liability never had a volume discount.&lt;/p&gt;

&lt;p&gt;So the two ends move in opposite directions, and not for the reason people usually give. It is not that the high end is more intelligent, or that it has better models — we explicitly do not claim that, and the published accuracy tie is the evidence. It is that the low end's unit of value is a &lt;em&gt;transformation&lt;/em&gt;, which scales beautifully, and the high end's unit of value is a &lt;em&gt;record with a name on it&lt;/em&gt;, which scales not at all.&lt;/p&gt;

&lt;p&gt;That asymmetry has a predictable consequence for the ecosystem the gateways are building. When every gateway ships a decision model, typed judgment becomes a default — in the routing layer, the moderation layer, the triage layer, everywhere a boundary is clean and a re-run is cheap. That is a large amount of value created, and almost none of it will be captured by whoever shipped the model. The model is the floor.&lt;/p&gt;

&lt;p&gt;What gets scarce is the thing the clones left out. In a world where every system has a cheap opinion, the scarce good is not a better opinion. It is a judgment that arrives with the evidence behind it and the name in front of it — something a reviewer can read, disagree with, and hold someone to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this argument does not claim
&lt;/h2&gt;

&lt;p&gt;Two caveats. First, none of this predicts that unaccountable judgment will fail. It will work fine for a long stretch, and for most of the volume it is the correct choice — which is exactly why the commoditized layer is worth having. The claim is about where the &lt;em&gt;price&lt;/em&gt; goes, not about which layer is ethical.&lt;/p&gt;

&lt;p&gt;Second, verification is not a permanent moat. Someone can copy our protocol tomorrow, and I would rather they did. What cannot be copied is the fact that a record was published before the argument needed it, with the unflattering rows included, under a name. Publishing it is available to anyone. Doing it early, and continuing after it stops being flattering, is a policy rather than a feature.&lt;/p&gt;

&lt;p&gt;The interfaces are already copied, which is what an interface is for. What the copies cannot carry is the part where someone says: this is what we measured, this is what we left out, this is how wrong we are allowed to be, and this is who is answerable when the machine makes the call alone. That part has a name attached, and a name does not batch.&lt;/p&gt;

&lt;p&gt;If you are the one who has to put that name on the line — a hard question, two candidate answers, and a record you can be held to — that is the case we built for: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-9" rel="noopener noreferrer"&gt;Decider&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>A threshold is a policy, not a number</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Sun, 04 Oct 2026 01:10:15 +0000</pubDate>
      <link>https://dev.to/turingcorp/a-threshold-is-a-policy-not-a-number-da4</link>
      <guid>https://dev.to/turingcorp/a-threshold-is-a-policy-not-a-number-da4</guid>
      <description>&lt;h1&gt;
  
  
  A threshold is a policy, not a number
&lt;/h1&gt;

&lt;p&gt;Somewhere in a payments codebase there is a line that says: approve automatically when confidence is above 0.8. Nobody remembers the afternoon it was written. The number has three likely origins and all of them are bad — it was the first value that made the demo behave, it was copied from a vendor example, or it was chosen because 0.8 sounds strict without sounding paranoid. What it was not read off is anything that describes how this particular system behaves when it is unsure.&lt;/p&gt;

&lt;p&gt;That one value decides which refunds are executed by a machine and which ones wait for a person. Set it in the wrong place and the consequences split in two directions, easy to confuse and hard to price: some share of the volume that deserved a look gets executed at machine speed, and the rest — usually a much larger share of it — is pushed onto a human queue that was never staffed for what the gate sends it.&lt;/p&gt;

&lt;p&gt;The stack underneath is ordinary. A refund request arrives; arithmetic settles the trivial cases; a decision model answers what a predicate cannot express; a person owns the ones where both answers survive scrutiny. That shape is background. The subject is the joint: the single number that decides which of those three handles a given case, and what happens when it is set by taste instead of by measurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the joint is switching between
&lt;/h2&gt;

&lt;p&gt;The rule layer is the one everybody trusts, because it is the one everybody can read. It fails silently: the business moves, the boundary the rule was drawn around does not, and a wrong rule decision looks exactly like a correct one.&lt;/p&gt;

&lt;p&gt;The decision model is a different animal — trained rather than written, returning a distribution rather than a branch. It fails out of distribution: handed a case its training data never contained, it does not raise an error, it returns a number wearing the same face as every other number it has emitted.&lt;/p&gt;

&lt;p&gt;The third destination is not a bigger model, it is a different question — which of two defensible options should we live with. It fails by not happening. The ticket sits, the case ages, and a decision nobody made does not look like a wrong decision. It looks like a backlog.&lt;/p&gt;

&lt;p&gt;Those three paragraphs are the whole of the background, because the failure modes are not what this article is about. Each one has a price, and the latch is where you decide who pays it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the 0.8 came from
&lt;/h2&gt;

&lt;p&gt;A threshold compares two numbers: one the system emits, one you choose. The first is a claim about frequency. When a model reports 0.87, it is claiming that among cases that look like this one, roughly 87 in 100 come out the way it says. Confidence is not a feeling the model has about itself; it is a prediction about its own hit rate, and like any prediction it can be checked against outcomes.&lt;/p&gt;

&lt;p&gt;The chosen number is rarely checked against anything. It arrives from a headline accuracy figure for the whole system, from a default in a library, or from the first value that stopped the demo from embarrassing anyone. None of those is a statement about what 0.87 means.&lt;/p&gt;

&lt;p&gt;An aggregate score and a confidence value answer different questions. Accuracy tells you how often the judge is right across everything you tested. Calibration tells you what a reported 0.87 actually buys in error rate. The gate runs entirely on the second question and is almost always set with the first.&lt;/p&gt;

&lt;p&gt;This is also why calibrating a model and choosing a threshold are two separate jobs, done by two different kinds of judgment. A perfectly calibrated model still does not tell you where to cut. It tells you what each possible cut costs. Where to cut is a question about consequences, and the model has never seen your consequences.&lt;/p&gt;

&lt;h2&gt;
  
  
  Both directions are expensive; only one is invisible
&lt;/h2&gt;

&lt;p&gt;Set the cut below what the confidence numbers actually mean and cases that should have been looked at are executed instead. Those failures are irreversible: money leaves, a commitment is honored at machine speed on a judgment that never earned machine speed. They do not arrive as errors. They arrive as outcomes — from a customer, a chargeback report, or an auditor two quarters later.&lt;/p&gt;

&lt;p&gt;Set the cut above and you get the failure everyone underestimates because it looks like diligence. Everything interesting escalates. The automation runs, it costs what it costs, and the queue on the other side fills with items that did not need a person. Do the arithmetic once: at 100,000 requests a day, escalating 40% means 40,000 human reviews a day. You have not automated the work. You have moved it and paid for the tool and the worker. The failure surfaces as headcount or as aging, which is why it survives for quarters.&lt;/p&gt;

&lt;p&gt;The two mistakes are not symmetric and they are not equally visible. The first is loud but rare, and in the happy path it resembles throughput. The second is quiet, permanent, and looks like caution — and it is the one that quietly doubles the cost of the automation you just bought.&lt;/p&gt;

&lt;p&gt;Neither direction is a math error. Choosing which one you can survive is the actual decision, and it does not have a universal answer. If a wrong automated approval costs a refund, cut high. If a delayed decision costs a contract, cut low. There is no single threshold that is correct for both, which is the first sign that you are choosing a policy rather than tuning a constant.&lt;/p&gt;

&lt;h2&gt;
  
  
  The curve has to exist before the latch means anything
&lt;/h2&gt;

&lt;p&gt;The only artifact that answers the question is a reliability report: for each band of reported confidence, how often the call was correct, with the number of cases in the band.&lt;/p&gt;

&lt;p&gt;We publish ours. Self-run on JudgeBench, 620 judgments, with the 6 that failed on a first verdict disclosed rather than quietly retried: calls the system reported at 90% confidence or above were right 99.6% of the time, and calls in the 80–90% band were right 94.0%. On that same self-run, raw accuracy came out at 92.5% against 92.2% for a plain direct baseline — a tie, reported as a tie, with nothing claimed over it. The value of the bands is not the headline. It is that a threshold can now be stated as a price: cut at 90 and you are accepting a 0.4% error rate on what you automate; cut at 80 and you are accepting 6%.&lt;/p&gt;

&lt;p&gt;Read the shape of the curve, not just the top of it. A curve that flattens out in the middle, and says so honestly, is telling the stack where the third layer begins — the region where the judge's own answer is that it does not know. A curve that is high everywhere tells you nothing, and a gate built on it will either automate everything or mean nothing.&lt;/p&gt;

&lt;p&gt;One caution before you print the table into a design document: the curve is measured on a benchmark distribution, not on your traffic. It describes the judge. It says nothing about the population you are feeding it, which is why the band table is necessary and not sufficient.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three questions before you trust the latch
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Where does the error land?&lt;/strong&gt; Answer it per boundary, not for the system as a whole. A re-run and a log line mean the machine can carry the mistake; a customer, a balance sheet, or a relationship means it cannot. The boundary of what you may automate is drawn by that question and nothing else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the base rate of what you are screening for?&lt;/strong&gt; A gate that passes 97% of a rare bad case is a different object from one that passes 97% of a common one. The band table gives you a rate conditional on confidence; your base rate converts that into the number of bad outcomes per day. Two teams can read the same table and owe different amounts of money.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who signs, and have you asked them?&lt;/strong&gt; The escalation layer fails when nobody will own the call, and the cheapest way to find that out is to ask before you build the escalation. If the answer is "whoever is on call," the policy is not a policy — it is an absence of one with a queue attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  A parameter gets tuned; a policy gets defended
&lt;/h2&gt;

&lt;p&gt;Put two companies in front of the same band table. A lender reading a borderline credit file and a hospital reading a borderline discharge will compute the same numbers and should still land on different cuts, because the errors do not land on the same party. The threshold is the only place in the architecture where that asymmetry becomes executable code. Everything above it is engineering; this one line is management.&lt;/p&gt;

&lt;p&gt;That is why the number should not live only in a config file. Write it down with a date, an owner, and a trigger for revisiting it. The trigger is not a calendar reminder; it is a change in the base rate, a change in the label set, or a change in who absorbs the loss. A threshold with no owner drifts, and drift is invisible right up until someone audits the queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Close
&lt;/h2&gt;

&lt;p&gt;The line in the config file is the shortest policy document most companies have ever written, and almost nobody treats it that way. If you cannot say who chose the number, when, and what they were trading away, then it was not a policy. It was a guess with production access.&lt;/p&gt;

&lt;p&gt;The fix is not a better number. It is a name next to the number, and a curve that says what the number buys.&lt;/p&gt;

&lt;p&gt;If you are working on that joint — the confidence, the threshold, the report that is supposed to justify it — &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-8" rel="noopener noreferrer"&gt;Decider&lt;/a&gt; answers one question with two candidate options, a pick, a calibrated confidence, and a written argument you can disagree with.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The Jevons Paradox of Judgment</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Sat, 03 Oct 2026 01:12:50 +0000</pubDate>
      <link>https://dev.to/turingcorp/the-jevons-paradox-of-judgment-2nl7</link>
      <guid>https://dev.to/turingcorp/the-jevons-paradox-of-judgment-2nl7</guid>
      <description>&lt;h1&gt;
  
  
  The Jevons Paradox of Judgment
&lt;/h1&gt;

&lt;p&gt;In 1865, a British economist named William Stanley Jevons published a book called &lt;em&gt;The Coal Question&lt;/em&gt;, and in it he made an argument that his contemporaries found hard to swallow. The more efficiently a steam engine used coal, he wrote, the &lt;em&gt;more&lt;/em&gt; coal Britain would burn. He was not predicting restraint. He was predicting appetite. His colleagues thought he had the sign backwards, and for a long time the argument sat in the drawer as a curiosity — until the twentieth century kept proving him right, in electricity, in fuel, in road capacity, in paper, in storage. The pattern acquired a nickname that outlived the man: the Jevons paradox.&lt;/p&gt;

&lt;p&gt;The mechanism is not mysterious, and the sloppy version of it ("efficiency makes us consume more") is not what he said. He said the cost of a unit of useful work falls, so you perform more units. Efficiency is a price cut, and price cuts move quantity.&lt;/p&gt;

&lt;p&gt;Now put the word &lt;em&gt;judgment&lt;/em&gt; where the word &lt;em&gt;coal&lt;/em&gt; was.&lt;/p&gt;

&lt;p&gt;TypeSafe named their model Jev on purpose — the System One framing comes from Kahneman, and the name from Jevons. It is a good name, because it is an honest bet: they are wagering that once a single judgment costs a fraction of a cent, the total number of judgments being made will explode. They are almost certainly right about the explosion. The interesting question is what happens to the judgments that did not get cheaper, and that question is where most of the discussion stops.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the cheap judgments come from
&lt;/h2&gt;

&lt;p&gt;An ATM is worth an aside here, because it is the cleanest natural experiment we have. When banks started putting cash machines on the wall, the obvious forecast was fewer tellers. The United States still employed 339,200 of them in 2025, and the official projection is a 13 percent decline over the following decade — 44,700 jobs, not a collapse. The same projection expects roughly 26,800 openings a year anyway, almost all of them to replace people who leave. The headcount barely moved; the job moved. Note what the official description says tellers do: process routine transactions. The machine took the transaction. What remained was the rest — selling products, untangling the accounts the machine could not reason about, handling customers who arrived already upset.&lt;/p&gt;

&lt;p&gt;This is the part of the paradox that does not get repeated often enough. When you automate the cheap end of a category, you do not shrink the category. You &lt;em&gt;re-sort&lt;/em&gt; it.&lt;/p&gt;

&lt;p&gt;If you have ever had to sort a queue by hand, this is obvious. The items that are easy to sort are the ones you stop thinking about after the first hundred. The hard ones are hard for a reason: the answer is contested, the categories overlap, the cost of a wrong bucket is not a re-run. A filter that removes everything easy does not leave you with an easier pile. It leaves you holding exactly the cases that were never eligible for the filter.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the cheap end actually contains
&lt;/h2&gt;

&lt;p&gt;What gets automated is not a random sample of judgment. It is a specific slice with specific properties: the answer is one option from a set somebody defined; the wrong answer is recoverable at a bounded price; a downstream check would catch it anyway; the same call is made thousands of times a day. Those four properties are what make a judgment automatable. They are also what make it cheap to be wrong about.&lt;/p&gt;

&lt;p&gt;Read that list again and notice what is missing from it: the name of a person who has to live with the outcome. The judgments that a millisecond classifier can take are exactly the judgments where being wrong is paid for by a system — a retry, a log line, a customer who asks again. The judgments it cannot take are the ones where being wrong is paid for by a human being, a relationship, a reputation, or a balance sheet with somebody's name under it.&lt;/p&gt;

&lt;p&gt;So the selection effect is not just "the easy ones go first." It is sharper than that. &lt;em&gt;The boundary of the automatable set is drawn by where the consequences land.&lt;/em&gt; Automate the judgments whose cost is borne by a system, and the residual set — the part still routed to a person — is defined by accountability. That is not a temporary state of affairs waiting for a better model. It is what the residual consists of.&lt;/p&gt;

&lt;p&gt;This is the uncomfortable half of the Jevons story for anyone selling cheap judgment. The paradox says the category grows. It does not say its center of gravity stays put. Every unit of judgment that moves into the cheap layer raises the relative weight of the units that cannot, because the total pool of decisions is now larger and the expensive tail has been left untouched.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cheaper judgment buys more judgment
&lt;/h2&gt;

&lt;p&gt;The second layer is the one Jevons himself was describing, and the one that gets skipped: a price cut does not just change how existing work is done. It changes how much work is worth &lt;em&gt;initiating&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Nobody writes a rule that says "log everything," and nobody decides that every comment in a community needs a moderation verdict. Those decisions get made per-item, and per-item they are a function of price. At a fraction of a cent per call, there is no longer a reason to be selective. The system starts inserting a judgment at every point where one could possibly help — should this be retained, should this be flagged, should this be escalated, should this be sent to a person.&lt;/p&gt;

&lt;p&gt;Each of those insertions is cheap. The aggregate is not, because the aggregate is not measured in tokens. It is measured in the queue of things that need a human being to sign off, and that queue is fed by exactly these calls.&lt;/p&gt;

&lt;p&gt;Follow one judgment through a pipeline and the shape becomes visible. A cheap judge returns a verdict with a number attached. The number sits above the threshold and the verdict executes — no human involved, correctly. The number sits below the threshold and the item is escalated, which means it joins a queue. Now multiply by the volume that cheap judgment made economical in the first place, and notice that the machine that was supposed to reduce human involvement has become the most efficient generator of human work ever built.&lt;/p&gt;

&lt;p&gt;The escalations are not failures. They are the design: the model is honest about the limit of its own confidence, and the limit gets routed upward. But routing upward is a producer of responsibility, and responsibility is the one thing nobody has found a way to make cheap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things fall, one does not
&lt;/h2&gt;

&lt;p&gt;Cost curves in this industry are steep, and it is reasonable to expect the cliff to keep going. Jev runs at 70–500 milliseconds per call, priced at $0.042 per million input tokens with output free — the vendor's own figures, and by any prior standard of software economics they are extraordinary. Inference gets cheaper every quarter; none of those curves is heading toward zero cost of being wrong.&lt;/p&gt;

&lt;p&gt;The per-call cost falls. Time-to-verdict falls. The cost of generating a defensible answer falls, since the alternative is a senior person's hour. But the cost of owning the outcome does not appear on any of those curves. It is not denominated in tokens; it is denominated in consequences, and consequences are paid in the currency of the person who chose.&lt;/p&gt;

&lt;p&gt;Practical evidence, from our own published record rather than a rhetorical point. Our calibration data is self-run under named benchmarks with the failures disclosed: on JudgeBench, 620 judgments, with the 6 that failed on first verdict disclosed rather than quietly retried. Reported at 90% confidence or above, our judgments were right 99.6% of the time; in the 80–90% band, 94.0%. What matters as much is the number underneath: there is a band where the system is genuinely uncertain, and its calibration says so in public.&lt;/p&gt;

&lt;p&gt;A calibrated confidence is useful precisely because it is a map of where the machine should stop: below some line the question is not a computation but a choice somebody has to own. The cheaper judgment gets, the more precisely that line is drawn — and the more clearly the stuff above it is separated from the stuff that was never about inference at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The re-pricing, stated as an economic claim
&lt;/h2&gt;

&lt;p&gt;Put the three layers together and the claim is clean.&lt;/p&gt;

&lt;p&gt;A judgment is a bundle of two things: the inference, and the accountability. For most of history they were produced together, by the same person, which is why the bundle looked indivisible — the expensive part of deciding seemed to be the thinking. What cheap judgment reveals is that they were never the same good. Inference is a computation, and computations get cheap on a predictable schedule. Accountability is a commitment made by an identifiable party, and it has no efficiency curve, because there is nothing in it to optimize. You cannot amortize it, batch it, or cache it. It does not get faster with better hardware.&lt;/p&gt;

&lt;p&gt;So the re-pricing goes like this. The cheap half of the bundle is unbundled and commoditized — that is Jev, and Jev is very good at it. The expensive half is not merely left over; it is &lt;em&gt;sharpened&lt;/em&gt;, because once inference is nearly free, the only remaining explanation for why a decision is hard is that somebody has to own it. Judgment did not get cheap. A particular component of judgment got cheap, and the residual component got more visible, more isolated, and more expensive relative to everything around it.&lt;/p&gt;

&lt;p&gt;This is why "it is only a matter of time before the hard ones are automated too" does not follow from the trend line. The trend line describes inference. The hard ones are not hard because inference is difficult; they are hard because both candidates are defensible, the outcome is not reversible at the same price, and a person has to own the result. You can make the inference arbitrarily cheap. The signature does not become cheaper, because its cost was never computational.&lt;/p&gt;

&lt;h2&gt;
  
  
  Close
&lt;/h2&gt;

&lt;p&gt;Everything that can be made cheap will be, and the volume of judgment will rise accordingly — Jevons was right about coal, and the same argument holds here. What will not fall is the price of being the one who decided: accountability has no efficiency curve, because there is nothing in it for engineering to make efficient.&lt;/p&gt;

&lt;p&gt;If you are holding one of the questions that stayed expensive — one question, two answers that both survive scrutiny — that is the case we build for. &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-7" rel="noopener noreferrer"&gt;Decider is here&lt;/a&gt;: a pick, a calibrated confidence, and a written argument you can disagree with, for the decisions that still end with a person's name.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Tellers, Occupational Outlook Handbook, U.S. Bureau of Labor Statistics (employment 339,200 in 2025; projected change -13 percent and -44,700 jobs, 2025-35; about 26,800 openings a year; median pay $43,030 in May 2025): &lt;a href="https://www.bls.gov/ooh/office-and-administrative-support/tellers.htm" rel="noopener noreferrer"&gt;https://www.bls.gov/ooh/office-and-administrative-support/tellers.htm&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our calibration figures are self-run under the named benchmarks with the failures disclosed, and are published in full in the machine-readable file referenced in the text.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>What Jev Got Right: Judgment as an Interface, Not a Paragraph</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Fri, 02 Oct 2026 01:10:05 +0000</pubDate>
      <link>https://dev.to/turingcorp/what-jev-got-right-judgment-as-an-interface-not-a-paragraph-19pj</link>
      <guid>https://dev.to/turingcorp/what-jev-got-right-judgment-as-an-interface-not-a-paragraph-19pj</guid>
      <description>&lt;h1&gt;
  
  
  What Jev Got Right: Judgment as an Interface, Not a Paragraph
&lt;/h1&gt;

&lt;p&gt;For as long as software has been able to ask a model a question, almost every answer has come back as prose.&lt;/p&gt;

&lt;p&gt;Ask whether a support ticket is billing or account, and you get a paragraph. &lt;em&gt;"This looks like a billing issue, though it could also be an account access problem — it depends on whether the charge was a renewal."&lt;/em&gt; A person reads that and understands it in a second. A program cannot. Somebody has to turn writing back into a decision, and the machinery for that is regex, keyword lists, a second smaller model, or hope. The judgment is real. It simply arrives in a container built for humans.&lt;/p&gt;

&lt;p&gt;On September 15, 2026, TypeSafe AI shipped Jev, and the container came off. You send program state plus typed questions, and you get back choices, scores and probabilities, computed in parallel, each carrying a confidence value — and not one sentence of prose. That is a small thing to describe and a large thing to build. Before I say anything about where it stops, I want to state clearly what it got right, because it is a genuine engineering result and not a marketing claim:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It turned judgment into an interface.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The answer became a value, not a text
&lt;/h2&gt;

&lt;p&gt;The first consequence is about shape.&lt;/p&gt;

&lt;p&gt;A paragraph is an open surface. Anything plausible can be written on it, including a confident summary of a decision that was never actually made. A typed value is a closed surface: the set of things it can be was defined by the caller. You cannot improvise the shape of a value. You can only be wrong about which value it is.&lt;/p&gt;

&lt;p&gt;That moves the failure mode from &lt;em&gt;ambiguous&lt;/em&gt; to &lt;em&gt;incorrect&lt;/em&gt;, and those are different engineering problems. Ambiguity is resolved by reading harder. Incorrectness is resolved by measurement, which means you can write a test for it. Once the answer to "which of these four data sources should I read" arrives as a field with four legal values and a probability, the surrounding system stops being a pipeline with a language model in the middle and becomes a pipeline with a component in it — a component that has an input contract and an output contract, like every other component you own.&lt;/p&gt;

&lt;p&gt;Plenty of people will say this is just JSON-mode classification, and that classifiers are old. Fine. Then notice that nobody put a classifier in the path of every tool call before, because the classifier was either too weak to trust or too expensive to run per call. Which is the second thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The units of price and time changed
&lt;/h2&gt;

&lt;p&gt;TypeSafe's own published figures put Jev at 70–500 milliseconds and $0.042 per million input tokens, with output free. Those are vendor numbers, not ours, and they are the whole point.&lt;/p&gt;

&lt;p&gt;At a price measured in millionths of a dollar, judgment stops being a budget line and becomes a rounding error. At a latency measured in milliseconds, it stops being a request and becomes a step. The interesting effect is not that the same decisions got cheaper. It is that a class of decisions which was previously &lt;em&gt;never made at all&lt;/em&gt; now gets made. Nobody was going to pay a person, or a full model call, to ask whether this particular merge should wait for review. So it did not get asked. Now it does.&lt;/p&gt;

&lt;p&gt;That is a real expansion of what software is permitted to notice about itself, and it belongs to Jev.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ecosystem answered within days, and the speed is the evidence
&lt;/h2&gt;

&lt;p&gt;Adoption figures deserve suspicion, so I will label them for what they are. The company raised a $40M seed led by DCVC, and its founder, Diogo Almeida, co-authored InstructGPT and the RLHF work at OpenAI — company and public record. The launch thread on Hacker News has run to 1,989 points and 520 comments (as of October 1), which you can read straight off Hacker News' own public API. Vercel reported that within 24 hours of listing the model, nearly 13% of its paid teams had already used it, and described it as the fastest-adopted model it had seen on its own gateway. Also platform-reported.&lt;/p&gt;

&lt;p&gt;But the number I find most persuasive is not the largest one. It is that gateways, orchestration frameworks and observability tooling — organizations with no shared roadmap and no coordination committee — each wired up the same primitive inside the same short window. Interfaces get adopted at that speed when they are obvious: when the thing they expose is already the thing everyone needed and could not name. A choice, a score, a probability. The naming was part of the contribution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the interface ends
&lt;/h2&gt;

&lt;p&gt;Everything above is why I think Jev is correct, and I am not going to follow it with a "but" that takes any of it back. Cheap judgment solved the problem it was aimed at: too many small judgments, each individually too expensive to make. That problem is now solved, and it stays solved.&lt;/p&gt;

&lt;p&gt;Here is the problem standing beside it, which is a different problem rather than a criticism.&lt;/p&gt;

&lt;p&gt;Making a judgment cheap requires making it a value, and a value has a domain. The domain is the options you passed in. That closure is exactly what makes the thing computable, and it is also why one particular kind of answer has nowhere to live: &lt;em&gt;these two cannot be separated with what you have given me.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That answer is not a defect state and it is not a low probability. It is a different kind of finding, and it shows up precisely when two candidates both survive an honest reading of the question. In human affairs those are frequently the decisions that matter most: made once, not repeatable, signed by a named person, lived with for years. A per-call-priced interface has no field in which "this one deserves more of your attention than the last thousand" can be expressed. Every call costs the same, so every call looks the same size.&lt;/p&gt;

&lt;p&gt;So the two ends of the same axis look like this. One end made judgment cheap, at scale, for the decisions where being wrong costs a re-run. It did that properly, and it is right. The other end is where the two sides really are level, and where the only useful outputs are: &lt;em&gt;this is level&lt;/em&gt;; a confidence number that has actually been calibrated, so that a band means something; and an argument long enough to disagree with. An answer of that kind costs minutes. In a decision you will live with for years, the minutes are the cheapest thing in the room.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we measure, and what we refuse to claim
&lt;/h2&gt;

&lt;p&gt;We are the second kind of judge, so it is fair to ask for our numbers instead of our adjectives.&lt;/p&gt;

&lt;p&gt;They come from our own published, self-run evaluations, on named benchmarks, with failures disclosed rather than quietly retried. On JudgeBench: 620 judgments, with the 6 that failed on first verdict disclosed. Our accuracy came in at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we report it as a tie; we claim no accuracy edge over a direct model call or over anyone else. On the same run, judgments we reported at 90% confidence or above were right 99.6% of the time, and the 80–90% band was right 94.0%. The number I would most want you to carry away is the least flattering one: on ContextualJudgeBench, run by us over the full official set, with 12 orders excluded after repeated platform failures and disclosed rather than imputed, the deliberately constructed near-ties sit at 46–60%.&lt;/p&gt;

&lt;p&gt;We publish that last figure because it is the honest scale of the problem at this end of the axis. When the answer is genuinely close, the correct output is to say so — and then to spend the minutes, because that is what a decision of that weight is worth.&lt;/p&gt;

&lt;h2&gt;
  
  
  Close
&lt;/h2&gt;

&lt;p&gt;Jev showed that a judgment can be a value a program switches on, and made that value fast and nearly free — the right answer for the decisions that should never have been slow. At the other end of that same axis, where the two sides are level and a person still has to sign, we build a judge that spends the minutes and shows its work: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-6" rel="noopener noreferrer"&gt;Decider&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>When Both Answers Are Defensible, the Right Answer Is That It Is a Coin Flip</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Thu, 01 Oct 2026 01:56:54 +0000</pubDate>
      <link>https://dev.to/turingcorp/when-both-answers-are-defensible-the-right-answer-is-that-it-is-a-coin-flip-1o7j</link>
      <guid>https://dev.to/turingcorp/when-both-answers-are-defensible-the-right-answer-is-that-it-is-a-coin-flip-1o7j</guid>
      <description>&lt;h1&gt;
  
  
  When Both Answers Are Defensible, the Right Answer Is That It Is a Coin Flip
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;The scene below is a composite, written for this piece. It is not a recording of one particular engagement, and the people in it are invented. The shape of the decision is not: it is the kind of question people actually bring to a judge.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two people run a small studio. A former colleague, now running operations at a company that grew faster than its systems, offers them a year-long retainer. One client, paid monthly, roughly what their three current clients pay together for the same number of hours. The work is unglamorous and stable. The alternative is what they already have: three clients, none of whom could end their year on their own, a pipeline that has to be refilled every few months, and enough variety that neither of them has to become a specialist in something they stopped enjoying.&lt;/p&gt;

&lt;p&gt;Write the two columns out and nothing breaks. The retainer buys a floor and costs optionality: sign it, and for a year their capacity is spoken for, and the next interesting inbound email gets answered with "not now". Staying as they are buys variety and costs the floor: a bad quarter is survivable, a bad half-year is not, and the two of them carry that risk personally, in a business with no investors and no cushion. Nobody in the room is short of information. They can list every client, every rate, every hour, and what each branch would do to next year's cash. What is missing is a criterion that decides between the columns, and there is a real possibility that no such criterion exists yet.&lt;/p&gt;

&lt;p&gt;That is not a failure of preparation. It is the most common shape of a decision that changes a year: both branches survive an adversarial reading, the costs are asymmetric, and the thing that would break the tie is not a fact about the world but a preference that has not been written down yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we ask a judge to do, and why it breaks here
&lt;/h2&gt;

&lt;p&gt;A judge, human or machine, is usually asked some version of: which one is better? The question quietly assumes there is a fact of the matter. Where the gap is wide, that assumption is safe and the answer is worth having. In the scene above it is not safe. One option is better on stability, focus and sleep; the other is better on variety, upside and the freedom to say yes to something better in eight months. Those are not the same units, and nothing converts them.&lt;/p&gt;

&lt;p&gt;This is where a decision tool is most often asked to do something it cannot do. Pressed for an answer, it will give one, because giving one is what it was built to do. It will pick the retainer, or pick staying independent, and it will say so in a confident voice, and the person reading it will feel that a question has been resolved. Nothing has been resolved. A tie has been relabelled as a verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  The most useful sentence a judge can produce
&lt;/h2&gt;

&lt;p&gt;Here is what we think the useful output looks like in that situation: &lt;em&gt;Option A, 60%. The two are close.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That is not timidity, and it is not the judge refusing to work. It is the judge reporting the shape of the problem: the candidates are close enough that the pick is being decided by something outside them, which means the deciding criterion is one the person holds and has not yet articulated. Knowing that before you sign is worth as much as knowing which way a lopsided comparison went. It changes what you do next. You stop hunting for the missing fact, because there is no missing fact, and you start asking what you are actually optimising for.&lt;/p&gt;

&lt;p&gt;The opposite behaviour is easy to miss because it sounds like strength. A judge that reports high confidence on everything is not a better judge. Its number has been flattened into decoration: it separates no cases, it cannot be routed on, and a threshold rule built on it will fire the same way every time. A confidence value earns its place by moving. If it never moves, it is not information about the decision. It is a tone of voice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the number looks like when someone checks it
&lt;/h2&gt;

&lt;p&gt;Confidence is only worth reading if someone has measured what it meant. Ours is calibrated on JudgeBench, self-run under the official judging protocol over 620 pairs, with the six pairs whose first verdict failed disclosed rather than imputed. On raw accuracy that run gives 92.5% for Decider against 92.2% for a plain direct model baseline. That is a tie. We publish it as a tie, and we claim no accuracy advantage over a direct model call, or over anyone else.&lt;/p&gt;

&lt;p&gt;What we do publish is the table the number came from:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Reported confidence&lt;/th&gt;
&lt;th&gt;Judgments&lt;/th&gt;
&lt;th&gt;Share of run&lt;/th&gt;
&lt;th&gt;Actually right&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;90% or above&lt;/td&gt;
&lt;td&gt;283&lt;/td&gt;
&lt;td&gt;45.6%&lt;/td&gt;
&lt;td&gt;99.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;80–90%&lt;/td&gt;
&lt;td&gt;184&lt;/td&gt;
&lt;td&gt;29.7%&lt;/td&gt;
&lt;td&gt;94.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;70–80%&lt;/td&gt;
&lt;td&gt;82&lt;/td&gt;
&lt;td&gt;13.2%&lt;/td&gt;
&lt;td&gt;84.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Below 70%&lt;/td&gt;
&lt;td&gt;65&lt;/td&gt;
&lt;td&gt;10.5%&lt;/td&gt;
&lt;td&gt;67.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read the bottom row first, because it is the row this article is about. When the judge reported below 70% confidence, it was right 67.7% of the time. On a two-option comparison, that is a system telling you it is staring at something close to a coin flip, and being roughly honest about it. The low number is a measurement of how little separates the two candidates.&lt;/p&gt;

&lt;p&gt;Then look at the distance between the top row and the bottom row: it is a little under 32 points. That distance is the information. A judge whose confidence is always high produces a flat version of this table, where every bin reads the same and the number tells you nothing you did not already assume.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark built to make this hard
&lt;/h2&gt;

&lt;p&gt;Our second published set is ContextualJudgeBench, self-run over the full official set of 2,000 pairs under the official pairwise protocol, with 1,991 pairs completed. In our scoring, a pair counts as correct only when the judge picks the right answer in both presentation orders, which is why the random floor is 25% rather than 50%. Twelve orders, 0.3% of the total, were excluded after repeated platform failures, and that exclusion is disclosed next to the result rather than buried. On that set our consistent accuracy is 67.1% against the benchmark's official reference value of 65.4%, measured on the same judged set.&lt;/p&gt;

&lt;p&gt;The interesting part is not the top line. It is that the benchmark deliberately contains near-tie splits, and on those the consistent accuracy sits in the 46–60% range. That is the difficulty its own authors designed in, and the benchmark file says so. It also does not help, because nothing helps on a pair where both answers are defensible.&lt;/p&gt;

&lt;p&gt;The same confidence bands mean something different there. Judgments reported at 90% or above came back right 83.3% of the time across 789 judgments. Judgments reported below 70% came back right 55.4% of the time across 529 judgments. Same interface, same scoring rule, and the high band is worth more than 16 points less than it was on JudgeBench.&lt;/p&gt;

&lt;p&gt;Sit with that for a moment, because it is the reason this article does not end with "the number is reliable". Calibration is a property of a benchmark and a population, not a property that travels for free between them. On a set full of genuinely close pairs, below-70% confidence really does mean close to a coin flip, and the number is still doing its job: it is telling you that this benchmark, and possibly your problem, is full of decisions the answers do not contain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use the number as a valve, not as a verdict
&lt;/h2&gt;

&lt;p&gt;This is the practical consequence, and it is the reason a low number is worth paying for.&lt;/p&gt;

&lt;p&gt;Treat confidence as a routing signal, not as a badge. In the high band, act on the pick: the comparison was decisive and the reasoning is there if you want to check it. In the middle band, read the argument and move: the call holds up but a quick look is cheap insurance. In the low band, stop treating the judge as the decision-maker. The pick is nearly arbitrary, and the useful work has moved to you: write the criterion down, take the branch you can reverse, ask the person who pays if it goes wrong, or accept that this is a coin flip and flip it on purpose and move on.&lt;/p&gt;

&lt;p&gt;A judge built to run at volume is designed to act above a threshold and hand the rest upward; that hand-off is only as good as the number that triggers it. Ours is the branch that gets handed the low numbers, which is exactly why the low bins are the ones we publish in full.&lt;/p&gt;

&lt;p&gt;The failure this prevents is not being wrong. It is acting as if a bare conclusion were a decision. "Option A" with no number and no sense of the spread has no downstream branch: either you obey it or you ignore it. A calibrated pick has two branches built in, and the branch you take tells you what kind of problem you are actually holding. Sometimes that is a comparison a machine can settle. Sometimes it is a preference only you can supply. A judge that always sounds certain cannot tell you which one you are in.&lt;/p&gt;

&lt;p&gt;So when both answers hold up and the judge says the two are close, that is the job being done, not the job being dodged. The evidence stops; your criterion starts; and the next move is to name that criterion out loud. If, after naming it, the two options still balance, then it really is a coin flip, and the honest version of that sentence is worth more than a fabricated verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not fix
&lt;/h2&gt;

&lt;p&gt;Calibration is not correctness. On the JudgeBench run, judgments reported at 90% or above were still wrong roughly four times in a thousand. On the ContextualJudgeBench run, the same band was wrong far more often than that. A single judgment can be wrong, a well-calibrated judge can be wrong on the one case you cared about, and a 60% call is by construction the kind of call that goes the other way four times out of ten.&lt;/p&gt;

&lt;p&gt;For anything expensive, one-way, or reputationally exposed, the position written into our own benchmark file applies: set your own threshold from the accuracy column, and apply your own review policy on top of it. Our published bands describe a comparison on a named benchmark under a named protocol with a stated number of exclusions. They are not a statement about the outcome of your particular decision, and we do not promise that an output is ready to use as it lands. The signature is still yours.&lt;/p&gt;

&lt;p&gt;If you want a hard question and two candidate answers put in front of a judge that returns a pick, a calibrated confidence, and a 900–1,700 character argument for it, start here: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-5" rel="noopener noreferrer"&gt;Decider&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>You are not collecting information. You are collecting permission.</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Wed, 30 Sep 2026 01:56:53 +0000</pubDate>
      <link>https://dev.to/turingcorp/you-are-not-collecting-information-you-are-collecting-permission-1le</link>
      <guid>https://dev.to/turingcorp/you-are-not-collecting-information-you-are-collecting-permission-1le</guid>
      <description>&lt;h1&gt;
  
  
  You Are Not Collecting Information. You Are Collecting Permission.
&lt;/h1&gt;

&lt;p&gt;You have been in business with the same person for twelve years. He is not a bad partner. He is a slow one, in a way that has quietly become expensive: he has said no to the last three things that would have grown the firm, and the fourth is on your desk now. You are not angry. You are tired, and you have been tired for about a year.&lt;/p&gt;

&lt;p&gt;You have not decided anything yet. What you have done is ask.&lt;/p&gt;

&lt;p&gt;His brother-in-law, who sold his own company, over lunch. Your accountant, twice, with a spreadsheet both times. Two people from a peer group, in a thread that ran for four days. A mentor you had not called in three years, who took the call and was kind about it. Then a chatbot on a Sunday afternoon, and a different one the following week, to see whether it agreed.&lt;/p&gt;

&lt;p&gt;Every one of those eight conversations was useful in the way that a good conversation is useful. Not one of them ended the question. And here is the part that took you three weeks to see: after each one you felt better for about forty-eight hours, and then you started again.&lt;/p&gt;

&lt;p&gt;Somewhere in the middle of it, a friend says the sentence you have heard before. "Honestly? I'd wind it down." And you notice what arrives. Not clarity. Relief. Something in your chest unclenches, and for one second you are not the person who has to sign.&lt;/p&gt;

&lt;p&gt;That feeling is the story. What you have been collecting for three weeks is not information. It is permission.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question you were actually asking
&lt;/h2&gt;

&lt;p&gt;Nine times, the words that came out of your mouth were "which option is stronger". That is not the question underneath. The question underneath was "will you stand next to me while I do this". They sound like the same sentence and they are not.&lt;/p&gt;

&lt;p&gt;The first question has an answer that exists in the world. It could be a fact, a clause, a projection, another person's actual position on something. Those answers can be found, and when you find them you are done asking.&lt;/p&gt;

&lt;p&gt;The second question has no answer anywhere. It is a request for someone to carry part of the weight, and weight does not move through conversation. Advice is not liability. Your accountant can tell you what the numbers say; he cannot be the one who loses the friend.&lt;/p&gt;

&lt;p&gt;Here is the test that separates the two, and it is uncomfortable. Imagine the version of this where nobody ever learns you asked. Not your partner, not the thread, not the mentor who was kind. The answer arrives in a sealed envelope, and no one knows you needed it. Would you still want it?&lt;/p&gt;

&lt;p&gt;For most people, in most hard decisions, the appetite drops sharply. That drop is the measurement. What you wanted was not the content. It was the audience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the loop feels like work
&lt;/h2&gt;

&lt;p&gt;Asking feels like diligence, and diligence is a virtue, so the loop is easy to defend to yourself. A decision made after eight conversations feels more defensible than one made alone, and in one narrow sense it is: it is more explainable to other people afterward. Look closer and the explainability is doing something else. It is arranging, in advance, a set of people who were partially involved, in case this goes badly. "I talked to everyone" is not a decision method. It is a defence prepared before the fact.&lt;/p&gt;

&lt;p&gt;Tools feed this, and mostly without meaning to. When you put a question with two defensible answers in front of a tool that always returns a conclusion, you get a conclusion — in the same confident shape as every other conclusion it has ever given you. There is no natural place for that loop to stop, because the interface always pays out. "These two are close, and the deciding criterion is one the question never contained" reads, in the moment, like a product failing to do its job. So it gets smoothed into a pick, and you leave with an action, and actions are why you come back.&lt;/p&gt;

&lt;p&gt;The discomfort you are trying to get rid of is not a defect in the tools, or in the people you called, or in the hour of the day. It is the decision. A choice that one more opinion could settle was never hard; the fact that it has survived eight opinions is precisely what "hard" means. No tool can move the signature off your name. The ones that seem to move it are lending you the feeling of a decision while the signature stays where it was.&lt;/p&gt;

&lt;h2&gt;
  
  
  The advice tax
&lt;/h2&gt;

&lt;p&gt;Nobody invoices you for this, which is why it runs so long.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Time, and drift.&lt;/strong&gt; The firm is not paused while you consult. Three weeks is three weeks of the slow partner saying no, of good people quietly deciding their own futures, of the price and the options you actually have moving underneath you while you gather views on a situation that is no longer quite the same situation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Other people's patience.&lt;/strong&gt; Every restart costs a friend a little enthusiasm. By the sixth conversation you have become the person who asks and does not decide, and people begin answering you with a shorter version of what they said the first time. That is information too, if you are willing to read it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The quiet one.&lt;/strong&gt; Each answer buys calm, and the calm expires. Worse, the calmer you feel after an answer, the harder it becomes to admit that the answers were never the bottleneck — because stopping now would mean all that asking was not the work. Paid tools add a meter to the same loop: the eighth question costs what the first one cost, and buys less.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three questions before you ask again
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Am I asking for a fact or for an opinion?&lt;/strong&gt; Write your question in one sentence, then decide which of those two things it is. If it can be settled by something that exists in the world — a clause, a number, a date, whether the buyer would actually accept a shorter term — go and get that thing and stop asking. If the honest answer is "I want to hear what they think I should do", then you are not gathering data. You are shopping for a co-signer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. If no one ever knew I asked, would I still ask?&lt;/strong&gt; Ask the private version of the question. If the sealed envelope is worth less to you than the conversation itself, then the value was never in the answer. It was in being witnessed, and being witnessed is what permission means.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. When the paper is in front of me, whose name is on it?&lt;/strong&gt; Yours. So make the last question the signing question, not the comparing question. Not "which is stronger" but "which would I sign today, if the deadline were this afternoon?" Then the harder half: "which of these two could I explain in five years without flinching?" If you can answer either one, you already have everything the ninth conversation was going to give you. If you cannot, you have found the actual missing input — a value you hold and have not written down, which is a different piece of work and takes less time than another call.&lt;/p&gt;

&lt;p&gt;One rule makes this easier, and it has to be written before you ask, not after: name the sentence you would need to hear in order to stop. If no such sentence exists, that tells you what kind of question you have been carrying around.&lt;/p&gt;

&lt;h2&gt;
  
  
  A judge that can end the loop
&lt;/h2&gt;

&lt;p&gt;A judge can be useful here in one specific way, and it is not by being cleverer than the people you called. It is by being willing to return "these two are close" instead of a verdict, and by attaching a number to its own pick that says how likely it is to be wrong this time.&lt;/p&gt;

&lt;p&gt;For us, self-run under the named benchmarks with the failures disclosed: on JudgeBench, 620 judgments, with the 6 that failed on first verdict disclosed rather than quietly retried, we measured 92.5% against 92.2% for a plain direct baseline. That is a tie. We report it as a tie, and we claim nothing over it — not against a direct model call, not against anyone else. The column that matters more in a situation like this one is the second: judgments we reported at 90% confidence or above were right 99.6% of the time, and judgments in the 80–90% band were right 94.0%. And the least flattering number on our own page is the one I would want you to keep: on ContextualJudgeBench, self-run over the full official set with the 12 orders excluded after repeated platform failures and disclosed rather than imputed, the deliberately constructed near-tie splits sit at 46–60%. A judge that will say "this is close" in public is doing the one thing eight conversations could not do for you: telling you the decision is yours, with the record to show it looked properly first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Close
&lt;/h2&gt;

&lt;p&gt;The loop is not a character flaw. It is what a person does when every tool hands back conclusions and nobody will confirm that "close" is a legitimate result.&lt;/p&gt;

&lt;p&gt;One thing we build sits on exactly that point: a judge that takes one question and two candidate answers and returns the pick, a calibrated confidence, and a written argument you can disagree with — and that says so when the two are level. &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-4" rel="noopener noreferrer"&gt;Decider is here&lt;/a&gt; — ask it once, then sign it yourself.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The router is only 0.55 sure. Who owns that call?</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Tue, 29 Sep 2026 01:56:59 +0000</pubDate>
      <link>https://dev.to/turingcorp/the-router-is-only-055-sure-who-owns-that-call-3ibn</link>
      <guid>https://dev.to/turingcorp/the-router-is-only-055-sure-who-owns-that-call-3ibn</guid>
      <description>&lt;h1&gt;
  
  
  The router is only 0.55 sure. Who owns that call?
&lt;/h1&gt;

&lt;p&gt;A request arrives. Before any model reads a token, a router decides which model should get it.&lt;/p&gt;

&lt;p&gt;Most of the time this is genuinely easy, and that is the entire point. The request is a routine extraction, the routing decision comes back at 0.90 confidence — the number is the confidence on the Jev judgment underneath the route — the code takes the cheap path, and nobody thinks about it again. That is a System One model doing exactly what it was built to do: fast, typed, decisive, cheap enough to run on every request.&lt;/p&gt;

&lt;p&gt;Now suppose the same decision comes back at 0.55.&lt;/p&gt;

&lt;p&gt;The model it leaned toward is barely ahead of the runner-up. The code still has to do &lt;em&gt;something&lt;/em&gt;, and every option in front of it is a real engineering decision:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retry the routing call&lt;/strong&gt;, hoping the second answer is more decisive. That is buying a random number with extra latency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run both candidate models in parallel&lt;/strong&gt; and choose later. That doubles the cost of precisely the requests you were trying to make cheap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fall back to a fixed default.&lt;/strong&gt; This quietly deletes the router from your architecture on the hard cases — the only cases where it was earning its place.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Escalate.&lt;/strong&gt; Hand the choice to something that can actually reason about it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last option is where the ecosystem is visibly heading. It is also where the design gets interesting, because &lt;em&gt;escalate&lt;/em&gt; is not a function call. It is a decision — and it tends to be harder than the routing decision that triggered it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Jev Router actually is
&lt;/h2&gt;

&lt;p&gt;TypeSafe's Jev Router is listed on OpenRouter as &lt;code&gt;typesafe/jev-router&lt;/code&gt;, released on September 25, 2026 — ten days after Jev itself. &lt;a href="https://openrouter.ai/typesafe/jev-router" rel="noopener noreferrer"&gt;OpenRouter's listing&lt;/a&gt; describes it as choosing a model and a reasoning effort for each request while balancing quality, speed and cost, running on Jev, TypeSafe's first System One model, with one endpoint reaching the wider model ecosystem. The listing advertises a 1,000,000-token context, accepts text, images, audio, files and video, and states in its FAQ that the price shown is zero.&lt;/p&gt;

&lt;p&gt;That is roughly the whole public description. The listing does not say which models are eligible, how the tradeoff is weighted, or how to constrain the choice. &lt;a href="https://runtimewire.com/article/typesafe-jev-router-openrouter-launch" rel="noopener noreferrer"&gt;RuntimeWire's writeup&lt;/a&gt; noted the same, plus the reminder that the router's advertised context window and Jev's own 32,000-token window are different figures describing different things. None of this is damning — it is a new listing — but anyone building on Jev Router today is supplying the escalation policy themselves.&lt;/p&gt;

&lt;p&gt;The community has built this shape a dozen times over. &lt;a href="https://github.com/BillionsBobby/JevRouter" rel="noopener noreferrer"&gt;&lt;code&gt;BillionsBobby/JevRouter&lt;/code&gt;&lt;/a&gt; — a Jev-powered router for models, tools and subagents — picked up 230 stars within three days of its September 18 creation. &lt;a href="https://www.npmjs.com/package/jev-router-mcp" rel="noopener noreferrer"&gt;&lt;code&gt;jev-router-mcp&lt;/code&gt;&lt;/a&gt; classifies coding tasks, picks a model, and supervises tool calls behind a "Tool Gate". There are per-turn routers for Codex and Claude Code, a &lt;a href="https://github.com/yusukebe/hono-jev-router" rel="noopener noreferrer"&gt;Hono router&lt;/a&gt; that matches requests to plain-language route descriptions, and &lt;a href="https://github.com/krisitown/jev-router" rel="noopener noreferrer"&gt;&lt;code&gt;krisitown/jev-router&lt;/code&gt;&lt;/a&gt;, which hands Jev the eligible targets as Choice alternatives and skips the model call entirely when only one target is eligible. The pattern is unmistakable: put the decision model at the front door.&lt;/p&gt;

&lt;h2&gt;
  
  
  A threshold does not solve the problem. It relocates it.
&lt;/h2&gt;

&lt;p&gt;The advice everyone repeats — act on high confidence, escalate on low confidence — is correct. &lt;a href="https://docs.typesafe.ai/patterns/confidence-routing" rel="noopener noreferrer"&gt;TypeSafe's own docs&lt;/a&gt; put it well: the answer tells you &lt;em&gt;what&lt;/em&gt;, confidence tells you &lt;em&gt;whether to act&lt;/em&gt;. The same corpus names the three possible handlers for an incoming request: deterministic logic, a specialist LLM, or a human.&lt;/p&gt;

&lt;p&gt;But read the second half of that sentence carefully. Before you added the gate, you had one hard question: which model? After you add the gate, you have a second one: &lt;strong&gt;escalate to whom?&lt;/strong&gt; The gate did not answer the hard question. It moved it one level down, to a place where it is often worse.&lt;/p&gt;

&lt;p&gt;Three reasons the displaced question is harder:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Both destinations are defensible.&lt;/strong&gt; The routing decision at least had a metric — latency, price, a capability matrix. The escalation compares two reasonable ways of being right, and neither comes with a score.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The costs are asymmetric, and nobody wrote them down.&lt;/strong&gt; A frontier model costs money. A person costs attention, the scarcest resource in the system. A queue costs nothing today and everything in six weeks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;There is no ground truth to check against.&lt;/strong&gt; For a routing decision you can eventually measure whether the chosen model succeeded. For "should this have gone to a human," the honest answer is often that you will never know.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;System One models are excellent at the first question and structurally unable to answer the second, because the second is not a classification. It is a judgment about consequences.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack, layer by layer
&lt;/h2&gt;

&lt;p&gt;Three layers, with different jobs, different costs and different failure modes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1 — Rules.&lt;/strong&gt; &lt;code&gt;if&lt;/code&gt;/&lt;code&gt;else&lt;/code&gt; over fields you already have: free, auditable, deterministic. Use them when the decision is expressible as a predicate. Their failure mode is ossification — the branch nobody remembers adding, and the twenty phrasings of "I want my money back" that a regex will never cover.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2 — A decision model (System One).&lt;/strong&gt; Narrow, typed, fast, shipped with a probability distribution. Use it when the answer is one of N enumerable options and code acts on it. Its failure modes are unusually well documented — by the vendor, about its own model. &lt;a href="https://docs.typesafe.ai/model-jaggedness/jev-1.13" rel="noopener noreferrer"&gt;TypeSafe's jaggedness page&lt;/a&gt; lists nine, including literal reading ("it answers the question you wrote, not the one you meant"), unreliable counting, date comparison, degradation when the state is full of irrelevant detail, and generation — for which the guidance is, plainly, to use a generative model instead.&lt;/p&gt;

&lt;p&gt;That last one deserves a spotlight: it is where teams quietly hurt themselves. Deciding which context to keep, drop or summarize &lt;em&gt;feels&lt;/em&gt; like a routing problem; it is a writing problem. &lt;a href="https://github.com/kerpopule/hermes-jev-skills" rel="noopener noreferrer"&gt;hermes-jev-skills&lt;/a&gt;, an integration that routes model choice, skill choice and memory through Jev, measured its own handoff: one written from Jev's keep/summarize/drop digest recalled &lt;strong&gt;less&lt;/strong&gt; than one written from the plain transcript — 37.5% alone and 68.3% with one search, against 58.7% and 75.0% for shipping the whole dialogue. They shipped the whole dialogue. The instructive part is not that a decision model was used for memory; it is that the failure was silent. A digest that looks tidy while losing the record of what was already tried is how an agent ends up looping over an action it has attempted three times.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 — A signed tradeoff (System Two).&lt;/strong&gt; The end of the escalation path, for decisions where the answer is not in the state: two options that both survive scrutiny, a cost you cannot undo, a human who will have to live with the outcome. This layer is slow and expensive on purpose — and most teams have a threshold routing &lt;em&gt;into&lt;/em&gt; it and nothing inside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the escalation actually costs, measured
&lt;/h2&gt;

&lt;p&gt;There is now third-party data on the cascade. An &lt;a href="https://www.ayautomate.com/blog/jev-vs-llm-benchmark" rel="noopener noreferrer"&gt;independent benchmark from AY Automate&lt;/a&gt; — 791 labeled decisions, 3,955 calls, $1.53 of spend, run September 19, 2026 — gated Jev at 0.80 confidence and sent the rest to a frontier model. Escalation took 19.4% of the 8-way routing items and 23.0% of the 77-way items, and the cascade's accuracy matched the frontier model alone (90.0% against 89.4%, and 84.8% against 84.3%) at 26–28% of its cost and roughly half its mean latency. That is the strongest argument for confidence-gated cascades.&lt;/p&gt;

&lt;p&gt;Now read the same post's failure note. Above 0.90 confidence, Jev was still wrong 5 times out of 112 on one task and 12 out of 153 on the other. The five confident errors on the 8-way task were all the same mistake, produced by two overlapping labels the frontier model also confused. The author's conclusion is worth putting on a wall:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The confidence score cannot tell you that your label set overlaps.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the boundary of the approach. A confidence value describes how concentrated a distribution is. It does not tell you whether the question was well formed, whether the options were actually distinguishable, or whether the cost of being wrong is survivable. Those are judgments &lt;em&gt;about the decision&lt;/em&gt;, not outputs of it — and they are exactly what the escalation step is being asked to supply.&lt;/p&gt;

&lt;h2&gt;
  
  
  The end of the path needs a signature
&lt;/h2&gt;

&lt;p&gt;Follow the escalation path to its terminus. High confidence acts. Low confidence escalates. The escalation lands on something — and whatever it lands on has to do two things the millisecond layer cannot: say &lt;em&gt;which&lt;/em&gt; option is stronger when both are defensible, and state how likely it is to be wrong, in a form someone can check.&lt;/p&gt;

&lt;p&gt;That is the product we build. Decider takes one hard question and two candidate answers and returns which one is stronger, a calibrated confidence value, and a 900–1,700-character argument for the call. On accuracy we claim nothing over a direct baseline, and we say so on our own pages: on JudgeBench we measured 92.5% across 620 judgments with 6 failures disclosed, against 92.2% for a plain direct baseline — a tie. What we are willing to be measured on is the calibration and the disclosure: judgments reported at 90% confidence or higher were right 99.6% of the time, and those at 80–90% were right 94.0%, on that same named benchmark, with the method and the failures published. If you need something to catch the router's uncertain case, it is here: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-3" rel="noopener noreferrer"&gt;https://api.turingcorp.net/platform/go/decider?src=jev-3&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would actually copy into a codebase
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Give the millisecond model the decisions that are closed, small, reversible and cheap to get wrong.&lt;/strong&gt; Enumerable options, compact state, code acting on the result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the threshold in code, not in a prompt.&lt;/strong&gt; TypeSafe's own &lt;a href="https://docs.typesafe.ai/cookbooks/sde_cascade" rel="noopener noreferrer"&gt;SDE cascade cookbook&lt;/a&gt; is a clean template: extract with a cheap model, verify with typed questions, escalate to a reasoning model only when a verifier signal fires (theirs defaults to firing at 0.7). Log the confidence, the options and the threshold with every call — that log is the only way to tune the gate later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail closed but never stuck.&lt;/strong&gt; The community integrations landed on the same behavior: on a missing key, a timeout or a low-confidence route, keep the current model, drop nothing from memory, and mark the decision unknown. An outage should cost you a time budget, not a turn.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never hand the decision model a writing job.&lt;/strong&gt; Compaction, summarization and rewriting are writing jobs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide the destination before you ship the gate.&lt;/strong&gt; If the only place the uncertain case can go is the same model called twice, you have purchased latency, not judgment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For the slice where both answers survive scrutiny and the cost is irreversible, keep something in the loop that will sign its name.&lt;/strong&gt; A person, or a deliberative judge that reports its own confidence and shows its reasoning. Not because machines cannot pick between two good options, but because that pick needs an owner, and ownership is not a number.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The router's job is to decide which model runs. The escalation's job is to decide what you can live with. Those are different jobs, and only one of them can be finished in seventy milliseconds.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Order invariance is not a spot check. It is the denominator.</title>
      <dc:creator>TuringCorp</dc:creator>
      <pubDate>Mon, 28 Sep 2026 01:56:50 +0000</pubDate>
      <link>https://dev.to/turingcorp/order-invariance-is-not-a-spot-check-it-is-the-denominator-15jc</link>
      <guid>https://dev.to/turingcorp/order-invariance-is-not-a-spot-check-it-is-the-denominator-15jc</guid>
      <description>&lt;h1&gt;
  
  
  Order Invariance Is Not a Spot Check. It Is the Denominator.
&lt;/h1&gt;

&lt;p&gt;For the past two weeks, the most useful question about Jev has not been "how fast is it" or "how cheap is it." It has been: &lt;em&gt;what happens when you swap the two candidates?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The reports coming back from independent community audits say the answer is "something": changing the order of the options visibly moves the output probabilities. That is not gossip — several of these audits are public, reproducible, and some were pre-registered (see Sources). I want to be careful about how I frame it, because the interesting part is not that a model behaves this way. Pairwise judges are sensitive to presentation order in general. It is one of the oldest known failure modes in this corner of evaluation, it shows up across model families, and it is a property you either design around or you do not. It is not a character flaw.&lt;/p&gt;

&lt;p&gt;The interesting part is what a team does once it knows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two kinds of parts
&lt;/h2&gt;

&lt;p&gt;There are two things you can put in a pipeline, and they look identical from the outside.&lt;/p&gt;

&lt;p&gt;The first is a judge you &lt;em&gt;permute&lt;/em&gt;. It runs at high volume, code acts on its output, and order sensitivity is a tuning problem: you randomize the presentation, you average over a few permutations, you move on. For that job, a cheap fast judge with a wobble is still an excellent part. This is the job Jev was built for, and it is genuinely well built for it.&lt;/p&gt;

&lt;p&gt;The second is a judge you &lt;em&gt;threshold&lt;/em&gt; on. Here the &lt;code&gt;confidence&lt;/code&gt; field is not a diagnostic — it is load-bearing. Somebody downstream writes a rule against it: above this band act, in the middle band ask, below that band escalate to a human or to a slower system. The moment that rule exists, the number stops being a metric and becomes an input to policy.&lt;/p&gt;

&lt;p&gt;Those are different kinds of parts, and they demand different kinds of evidence. A benchmark accuracy score is evidence for the first. For the second, a score is close to irrelevant — what you need to know is how the confidence number itself was validated, in the exact form you are about to consume it. A number that was measured under order A and applied under order B has not been validated for your use.&lt;/p&gt;

&lt;p&gt;That is the whole argument. It is not a claim about anyone's model. It is a claim about what "confidence" has to survive before you are allowed to build a switch statement on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What our protocol actually says
&lt;/h2&gt;

&lt;p&gt;We publish a benchmark file at &lt;code&gt;api.turingcorp.net/benchmarks/latest.json&lt;/code&gt;, and the line that matters here is the protocol field for ContextualJudgeBench. Verbatim:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Self-run with the official vanilla pairwise protocol; consistent accuracy (&lt;strong&gt;both response orders judged correctly&lt;/strong&gt;); random floor &lt;strong&gt;25%&lt;/strong&gt;; failures disclosed (12 orders rerun-excluded). Reference model measured on the same judged set.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read the parenthetical twice, because it is the entire point. In our benchmark, a pair counts as correct only when the judge picks the right answer in &lt;strong&gt;both&lt;/strong&gt; presentation orders. Order invariance is not an audit we run afterward, and it is not a caveat in a footnote. It is inside the definition of "correct." A single flipped order converts a scored success into a scored failure.&lt;/p&gt;

&lt;p&gt;The consequence is visible in an unusual place: the random floor. Guess at chance on a two-option pairwise task and you get 50%. Our published floor is &lt;strong&gt;25%&lt;/strong&gt;, precisely because a coin-flip judge has to get two independent orders right, and 25% is what that costs. We think that is the honest denominator for a system whose output feeds a threshold rule, and we would rather publish a lower floor against a stricter definition than a higher number against a looser one.&lt;/p&gt;

&lt;p&gt;The same file's footnote states the rule a second time and then discloses the damage:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Self-run with the official protocol over the full 2,000 official pairs. Consistent accuracy requires the same correct pick in both presentation orders (random floor 25%). 12 orders (0.3%) were rerun-excluded after repeated platform failures; the reference model was measured on the same judged set. Previously published results measured on a subset are archived in the repository changelog.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Twelve orders were excluded rather than counted, and rather than quietly imputed. On the full 2,000-pair official set, 1,991 pairs completed, with a self-run consistent accuracy of &lt;strong&gt;67.1% against the benchmark's official reference value of 65.4%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I want to flag the least flattering number in that file before anyone else does. ContextualJudgeBench contains splits that were deliberately built as near-ties, and on those the consistent accuracy sits in the &lt;strong&gt;46–60%&lt;/strong&gt; range. That is not a bug we are hiding; it is the benchmark's own difficulty design, and order invariance does not rescue a pair where both answers are defensible. Invariance buys you the right to compare numbers across orders. It does not buy you the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part we are not allowed to skip
&lt;/h2&gt;

&lt;p&gt;The other benchmark we publish is JudgeBench: 620 pairs. Our protocol line reads:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Self-run with the official judging protocol. Both columns use first successful verdict per pair; the six pairs whose first verdict failed are disclosed rather than imputed. Reference model measured on the same judged set.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Six pairs failed on first verdict, and we disclose them instead of scoring a retry as if it were the first attempt. On that set, Decider measured &lt;strong&gt;92.5%&lt;/strong&gt;. The plain direct baseline measured &lt;strong&gt;92.2%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is a tie. We report it as a tie, and we do not make an accuracy claim over it — not against a direct model call, and not against anyone else. If you came here for a number that lets us look sharper, this is the wrong article.&lt;/p&gt;

&lt;p&gt;What we do put weight on is a different column, also from that file and also a self-run with the failures disclosed: judgments reported at &lt;strong&gt;≥90% confidence were right 99.6%&lt;/strong&gt; of the time on that benchmark, and judgments reported in the &lt;strong&gt;80–90% band were right 94.0%&lt;/strong&gt; of the time. Those are the bins a threshold rule would actually read.&lt;/p&gt;

&lt;p&gt;That is not a statement that we are sharper than a direct model call — we measured a tie, and the tie is the finding. It is a statement about what the number attached to a judgment has meant historically, on a named benchmark under a named protocol with a stated exclusion count.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is public on the other side
&lt;/h2&gt;

&lt;p&gt;TypeSafe has not published a reliability diagram or an expected calibration error for Jev. That is a statement about what is publicly available, not a statement about what is true internally — the training method is described as reinforcement learning aimed at calibrated decisions, and it may well be producing exactly what it claims. Independent audits have started filling that gap from the outside, which is a healthy sign for the ecosystem (Sources); but if you are the person writing the threshold, "not published by the vendor" is the answer you need to have.&lt;/p&gt;

&lt;p&gt;The reason this matters is not that Jev is dubious. It is that Jev's own recommended usage is &lt;em&gt;threshold routing&lt;/em&gt;: act on high confidence, and send low confidence to a human or to a stronger system. That guidance is sound, and it puts the reliability question directly on the critical path of the intended deployment. Nobody is being ambushed here; the question is simply upstream of the use case.&lt;/p&gt;

&lt;p&gt;So the two systems are not rivals. One is a decision primitive optimized for volume, and one of them — ours — spends its budget on the two things you cannot get from a cheaper judgment: a written argument, and a confidence number whose bins are published with the exclusions that produced them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three tests you can run on any confidence field
&lt;/h2&gt;

&lt;p&gt;If a decision interface hands you a confidence value and invites you to build policy on it, you can check it yourself. Three tests, in order of cost:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Permute the input.&lt;/strong&gt; Send the same question with the candidates in the opposite order. Then send it again with a different permutation. A confidence number that moves when the order moves is reporting &lt;em&gt;formatting&lt;/em&gt;, not &lt;em&gt;comparative strength&lt;/em&gt;. If your pipeline consumes that number, the movement is now inside your policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Ask for the reliability diagram.&lt;/strong&gt; Not the headline accuracy — the calibration curve. Which bins exist, how many cases fall in each, and what fraction of each bin was actually right. If the diagram does not exist, you have learned something load-bearing about the interface, and you have learned it before you shipped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Read the denominator.&lt;/strong&gt; Find out what happened to the cases that failed. Were they rerun, excluded, imputed, or dropped from the count? A protocol that names its exclusions is telling you how much to trust the number above them. A protocol that has no exclusions to name is usually not a protocol with no failures; it is a protocol that has not looked.&lt;/p&gt;

&lt;p&gt;Apply these to any judge, including ours. Everything the three tests ask for is spelled out above: the permutation requirement is inside our scoring rule, the bins and their case counts are in the published file, and the failure counts — 6 for JudgeBench, 12 rerun-excluded orders for ContextualJudgeBench — are printed next to the results rather than in a footnote nobody reads. That is deliberate. A reliability claim you cannot audit is a feeling with a decimal point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try to break the bins
&lt;/h2&gt;

&lt;p&gt;Here is the honest limit of everything above. Order invariance raises the bar for being counted correct; it does not certify that the judge is right, and it certainly does not remove the need for your own review policy on decisions that are expensive or irreversible. Our published near-tie band still runs 46–60%, and that is our own benchmark telling on us.&lt;/p&gt;

&lt;p&gt;So the useful thing you can do with this article is adversarial, not appreciative. Our confidence bins are published, our exclusions are counted, and the protocol that produces them is quoted verbatim above. If you run decisions through a judge that reports confidence and you find our bins do not hold up, that is a more valuable result than another leaderboard position, and we would like to hear it.&lt;/p&gt;

&lt;p&gt;If you want to put a hard question and two candidate answers in front of the judge whose numbers are on the record, start here: &lt;a href="https://api.turingcorp.net/platform/go/decider?src=jev-2" rel="noopener noreferrer"&gt;https://api.turingcorp.net/platform/go/decider?src=jev-2&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then run the three tests on it. That is the invitation, and it is a real one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;Everything of ours quoted above is in one file, published and machine-readable: &lt;a href="https://api.turingcorp.net/benchmarks/latest.json" rel="noopener noreferrer"&gt;https://api.turingcorp.net/benchmarks/latest.json&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The independent Jev audits referenced in the opening, all public:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Permutation-invariant reconstruction of the same interface: &lt;a href="https://github.com/TypeLLM/pijev" rel="noopener noreferrer"&gt;https://github.com/TypeLLM/pijev&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;API-only calibration audit: &lt;a href="https://github.com/jujumilk3/jev-calibration-audit" rel="noopener noreferrer"&gt;https://github.com/jujumilk3/jev-calibration-audit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Out-of-distribution calibration with an explicit ECE noise floor, reproducible for about $0.06: &lt;a href="https://github.com/scienthoon/jev-ood-calibration" rel="noopener noreferrer"&gt;https://github.com/scienthoon/jev-ood-calibration&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Decision models measured as the option order changes and the wrong answers stop being obvious: &lt;a href="https://github.com/gazelle93/decision-models-under-pressure" rel="noopener noreferrer"&gt;https://github.com/gazelle93/decision-models-under-pressure&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;An index of Jev robustness and failure-mode studies: &lt;a href="https://github.com/Yifan-Lan/awesome-jev-robustness" rel="noopener noreferrer"&gt;https://github.com/Yifan-Lan/awesome-jev-robustness&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;TypeSafe's own description of the model and its recommended threshold-routing usage: &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;https://typesafe.ai/blog/introducing-system-one-models-and-jev&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>tools</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
