<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dave</title>
    <description>The latest articles on DEV Community by Dave (@dave8172).</description>
    <link>https://dev.to/dave8172</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3987397%2F0f6b44d9-ee2f-48bc-95b1-d9433ded1ffd.png</url>
      <title>DEV Community: Dave</title>
      <link>https://dev.to/dave8172</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dave8172"/>
    <language>en</language>
    <item>
      <title>How I calibrated an LLM judge to grade like me, 25 cheaper</title>
      <dc:creator>Dave</dc:creator>
      <pubDate>Sat, 03 Oct 2026 20:06:29 +0000</pubDate>
      <link>https://dev.to/dave8172/how-i-calibrated-an-llm-judge-to-grade-like-me-25x-cheaper-f9g</link>
      <guid>https://dev.to/dave8172/how-i-calibrated-an-llm-judge-to-grade-like-me-25x-cheaper-f9g</guid>
      <description>&lt;p&gt;Businesses that sell technical products answer the same kind of question every&lt;br&gt;
day. &lt;em&gt;"What's the accuracy on this range?"&lt;/em&gt; &lt;em&gt;"Can it measure through a coating, and&lt;br&gt;
how thick?"&lt;/em&gt; &lt;em&gt;"Does it come with a calibration certificate?"&lt;/em&gt; The answers sit in&lt;br&gt;
product datasheets and manuals.&lt;/p&gt;

&lt;p&gt;I took 27 of those PDFs, three of them scans with no text layer, and wrote 47&lt;br&gt;
questions the way customers actually ask them. Then I had four AI systems answer&lt;br&gt;
every question and graded all 188 answers.&lt;/p&gt;

&lt;p&gt;Running the systems took an afternoon. Building a judge I could trust took the&lt;br&gt;
rest of the work, and that is what this post is about.&lt;/p&gt;
&lt;h2&gt;
  
  
  The result
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fph9auvk5am7xnra2i58y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fph9auvk5am7xnra2i58y.png" alt="Correct answers out of 47: my pipeline 46, Claude app 45, ChatGPT app 35, Chatbase 33" width="800" height="680"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Correct&lt;/th&gt;
&lt;th&gt;Wrong&lt;/th&gt;
&lt;th&gt;Made up&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;My pipeline&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;46&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude app&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;45&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT app&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chatbase&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My pipeline is Claude Sonnet over the API with page citations. The Claude app ran&lt;br&gt;
Opus with the documents in a project, ChatGPT ran with thinking off, and Chatbase&lt;br&gt;
was on the free plan with its default model.&lt;/p&gt;

&lt;p&gt;Every system got the same PDFs and the same instructions. One run each, so I read&lt;br&gt;
a one-question gap as a tie.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No system invented a spec value.&lt;/strong&gt; I checked every claim twice: a claim checker&lt;br&gt;
read each one against the page it came from, and I went through the ones it was&lt;br&gt;
unsure about by hand. The weaker systems failed more quietly. They said an answer&lt;br&gt;
wasn't in the documents when it was.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where the misses came from
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4wsxb6qfg59xanf4548q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4wsxb6qfg59xanf4548q.png" alt="Correct answers by question type for each system" width="800" height="730"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scanned pages.&lt;/strong&gt; Systems that only read the text layer saw blank pages. Every&lt;br&gt;
question answered only by a scan came back "not specified".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Answers that need a second look.&lt;/strong&gt; A product advertises one range on the front&lt;br&gt;
page. A different mode of the same product stops at a tenth of it, and the&lt;br&gt;
customer's question is about that mode. The same pattern showed up in accuracy&lt;br&gt;
tables split by range and in specs that change with the material.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unit conversions.&lt;/strong&gt; The customer asks in one unit. The datasheet lists the&lt;br&gt;
other unit in one row and a separate, lower limit in the customer's unit in&lt;br&gt;
another. Converting the first number gives a confident wrong answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The website and the datasheet disagreeing.&lt;/strong&gt; The product page said one value&lt;br&gt;
and the datasheet said half of it. My rule: the datasheet wins, and the reply says&lt;br&gt;
the page may be wrong. Two systems missed it.&lt;/p&gt;
&lt;h2&gt;
  
  
  Rules that live in no PDF
&lt;/h2&gt;

&lt;p&gt;Before running anything I reviewed every question. I dropped three that no&lt;br&gt;
customer would ask, corrected two gold answers, and turned the "not in the&lt;br&gt;
documents" cases into rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Datasheet beats product page.&lt;/strong&gt; Give the datasheet value and say the page may
have an error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never a bare "not in our documents".&lt;/strong&gt; Say it isn't in the published
documents, and give an email to ask.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Certifications only if the document says so.&lt;/strong&gt; Otherwise, email for
confirmation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Canned answers&lt;/strong&gt; for the two most common questions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After grading I added one more: &lt;strong&gt;answer what was asked.&lt;/strong&gt; Extra conditions&lt;br&gt;
confuse buyers and invite more questions.&lt;/p&gt;

&lt;p&gt;That short list did more for answer quality than any prompt trick.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why the judge needs a judge
&lt;/h2&gt;

&lt;p&gt;Grading 188 answers by hand takes hours, so the usual move is an LLM judge: give&lt;br&gt;
a model the question, the gold answer and the answer, and ask for correct, partial&lt;br&gt;
or wrong. You only know the judge is any good if it agrees with the person whose&lt;br&gt;
standard matters. Here, that's me.&lt;/p&gt;

&lt;p&gt;So I graded 20 answers blind: no system names, five per system, with some likely&lt;br&gt;
mistakes mixed in so the judge would be tested on errors too. Then I compared.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First try: 13 of 20.&lt;/strong&gt; Most of the gap was in my own answer key. Three gold&lt;br&gt;
answers were wrong. Grading real answers showed me that I answer from the&lt;br&gt;
datasheet &lt;em&gt;as printed&lt;/em&gt;, even when a conversion on it looks off, and flag the&lt;br&gt;
document separately. My key hadn't been written that way. Once I fixed it,&lt;br&gt;
agreement jumped.&lt;/p&gt;

&lt;p&gt;Plain agreement flatters a judge when most answers are correct. A judge that&lt;br&gt;
marks everything "correct" would agree 80% of the time here and understand&lt;br&gt;
nothing. &lt;strong&gt;Cohen's kappa&lt;/strong&gt; corrects for that.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhc3uw07pad695nzydsca.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhc3uw07pad695nzydsca.png" alt="Cohen's kappa: 67 of 100 agreement expected by luck, 28 from skill, 5 missed; kappa = 28 / 33 = 0.85" width="800" height="680"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kappa asks how far above luck the judge got, as a share of how far above luck it&lt;br&gt;
could have got. 0 means no better than random; 1 means it matched me every time.&lt;/p&gt;
&lt;h2&gt;
  
  
  Two judges
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcp23084tbe8ehxl718cg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcp23084tbe8ehxl718cg.png" alt="Agreement with my grades: Sonnet 35 of 40, hybrid 36 of 40. Cost for 188 answers: $0.80 vs $0.03" width="800" height="780"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Judge&lt;/th&gt;
&lt;th&gt;Agrees&lt;/th&gt;
&lt;th&gt;Cost, 188&lt;/th&gt;
&lt;th&gt;Speed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet alone&lt;/td&gt;
&lt;td&gt;35 / 40&lt;/td&gt;
&lt;td&gt;~$0.80&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hybrid&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;36 / 40&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$0.03&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;under 1 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The hybrid gives each part the job it does best.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foevmrd1sgz0wdbsp7kec.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foevmrd1sgz0wdbsp7kec.png" alt="The hybrid judge: code compares numbers, Jev answers typed questions, code applies the policy; Sonnet with the PDF only for unsure claims" width="800" height="860"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code&lt;/strong&gt; pulls the numbers out of the gold answer and the answer being graded,&lt;br&gt;
and records which match. It never confuses ±(1.2% + 5) with ±(0.8% + 5).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://typesafe.ai" rel="noopener noreferrer"&gt;Jev&lt;/a&gt;&lt;/strong&gt; handles meaning. It's a System One model from&lt;br&gt;
TypeSafe: it reads natural language and returns typed answers with&lt;br&gt;
probabilities, with no prose to parse. One request per answer asks four&lt;br&gt;
questions at once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"verdict"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Grade `answer` to `customer_question` against `gold` and `notes`. Use `code_checks` for which gold numbers the answer contains."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"criteria"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"correct"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"partial"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"wrong"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"brush_off"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Does `answer` brush the customer off, such as a bare 'not in our documents' with no next step?"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"unasked_extra"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Does `answer` add specs, conditions or other models that `customer_question` didn't ask about?"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rule_followed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Does `answer` do what `rule` says?"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A Choice picks one of the grades and returns a probability for each. A Noul is&lt;br&gt;
the probability that a statement holds. The judgment comes back as numbers, so&lt;br&gt;
the policy stays in code where I can see it and change it. One real answer from&lt;br&gt;
the run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;verdict        correct 0.98 · partial 0.02 · wrong 0.00
unasked_extra  0.91  → at or above 0.9: "too much information" flag
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;About 2,000 tokens per answer, under a second, and all 188 answers for about three&lt;br&gt;
cents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sonnet with the original PDF&lt;/strong&gt; only sees the few claims Jev isn't sure about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping myself honest
&lt;/h2&gt;

&lt;p&gt;I wrote the pass line down before I graded a second set of 20: &lt;strong&gt;at least 16 of&lt;br&gt;
20, kappa at least 0.6.&lt;/strong&gt; That second set got used once, as a test, and never for&lt;br&gt;
tuning.&lt;/p&gt;

&lt;p&gt;It caught a mistake. On the first 20, I had tuned a penalty: if Jev was 90% sure&lt;br&gt;
an answer added unasked details, the grade dropped to partial. It looked perfect&lt;br&gt;
there. On the fresh set it caused two of the three disagreements. The judge&lt;br&gt;
passed the line (17 of 20, kappa 0.63), but barely.&lt;/p&gt;

&lt;p&gt;Reading those two answers showed why. I grade substance and length separately.&lt;br&gt;
An answer can be right and too long. The detector itself was right: it flagged&lt;br&gt;
exactly the four answers across both sets that I had found too long, and no&lt;br&gt;
others. The penalty was the wrong policy. So it became a flag beside the grade.&lt;br&gt;
Without it the judge agrees with me 19 of 20, kappa 0.85. That change came after&lt;br&gt;
I'd seen the second set, so the next fresh set is where it has to hold.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop I'd run every time
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmrpdzmdio5siwq3y1755.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmrpdzmdio5siwq3y1755.png" alt="The eval loop: golden set, all systems answer, I grade 20 blind, the judge grades the same 20 and I fix each disagreement, a fresh 20 against a line set in advance, then grade everything" width="800" height="850"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I did it in a worse order: ran the systems, graded everything with an untested&lt;br&gt;
judge, reported numbers, then calibrated. Next time:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Build the golden set: questions as asked, reviewed gold answers.&lt;/li&gt;
&lt;li&gt;All systems answer, same documents, same instructions.&lt;/li&gt;
&lt;li&gt;I grade 20 answers blind.&lt;/li&gt;
&lt;li&gt;The judge grades the same 20. I read every disagreement and fix the cause,
which is usually the answer key.&lt;/li&gt;
&lt;li&gt;Write the pass line down, then test on a fresh 20. Use it once.&lt;/li&gt;
&lt;li&gt;Only then grade everything and report.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A golden set isn't finished until a person has checked real answers against it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it means
&lt;/h2&gt;

&lt;p&gt;On product datasheets, a frontier model with the PDFs attached is already very&lt;br&gt;
good: my pipeline and the Claude app tied. Accuracy is the starting point. The&lt;br&gt;
hard part is getting that accuracy into the place where buyers ask, across a&lt;br&gt;
catalogue too big for one prompt, while datasheets change and the website drifts&lt;br&gt;
away from them. Then proving it on the business's own questions, with a judge&lt;br&gt;
calibrated to the person who answers them. That is what I'm building now.&lt;/p&gt;

&lt;p&gt;The whole run cost about $8 in API calls. Most of that went on mistakes I won't&lt;br&gt;
repeat. Now I estimate every paid run from a one-question trial, counting every&lt;br&gt;
path that can call a paid model.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>testing</category>
    </item>
    <item>
      <title>Recurring affiliate programs: what 471 program pages say</title>
      <dc:creator>Dave</dc:creator>
      <pubDate>Thu, 24 Sep 2026 18:36:04 +0000</pubDate>
      <link>https://dev.to/dave8172/recurring-affiliate-programs-what-471-program-pages-say-om8</link>
      <guid>https://dev.to/dave8172/recurring-affiliate-programs-what-471-program-pages-say-om8</guid>
      <description>&lt;p&gt;"Recurring commission" is the phrase every affiliate page wants you to read first. I keep a directory of the affiliate terms of 471 developer and SaaS tools, and every figure in it is read off the vendor's own program page, with that URL and the date it was checked. So here is what "recurring" turns out to mean.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recurring stops after a year almost as often as it pays for life
&lt;/h2&gt;

&lt;p&gt;267 of the 471 pay recurring commission. Here is how long "recurring" lasts:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;How long it pays&lt;/th&gt;
&lt;th&gt;Programs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;For the customer's lifetime&lt;/td&gt;
&lt;td&gt;104&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12 months or less&lt;/td&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;18 months to 3 years&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The page doesn't say&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So a program that says "recurring" is about as likely to stop after a year as to pay for life. And almost one in five never tells you which.&lt;/p&gt;

&lt;p&gt;That matters more than the rate. On a $50/month product, 30% for life pays $180 a year for as long as the customer stays. 50% for 12 months pays $300 once, then nothing. By the end of year two the smaller percentage has paid more.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What to look for on the page:&lt;/strong&gt; "for life", "as long as the customer stays subscribed" or "lifetime" means lifetime. "First 12 months", "first year" or "for one year" means a cap. "Recurring" on its own means you should ask before you write anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which kinds of software pay recurring
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Pay recurring&lt;/th&gt;
&lt;th&gt;Pay once&lt;/th&gt;
&lt;th&gt;Don't say&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No-code&lt;/td&gt;
&lt;td&gt;19 of 23&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Email &amp;amp; newsletters&lt;/td&gt;
&lt;td&gt;25 of 31&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SEO&lt;/td&gt;
&lt;td&gt;16 of 21&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analytics&lt;/td&gt;
&lt;td&gt;13 of 17&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer tools&lt;/td&gt;
&lt;td&gt;27 of 36&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E-commerce&lt;/td&gt;
&lt;td&gt;28 of 41&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CRM &amp;amp; sales&lt;/td&gt;
&lt;td&gt;14 of 21&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Productivity&lt;/td&gt;
&lt;td&gt;47 of 80&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Design&lt;/td&gt;
&lt;td&gt;26 of 46&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Courses &amp;amp; communities&lt;/td&gt;
&lt;td&gt;10 of 18&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Support &amp;amp; live chat&lt;/td&gt;
&lt;td&gt;9 of 18&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HR &amp;amp; payroll&lt;/td&gt;
&lt;td&gt;5 of 14&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosting&lt;/td&gt;
&lt;td&gt;13 of 43&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CMS &amp;amp; site builders&lt;/td&gt;
&lt;td&gt;11 of 45&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Payments&lt;/td&gt;
&lt;td&gt;4 of 17&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most of the bottom is silence. Across the four lowest categories, 59 programs don't say whether they recur and 27 say they pay once. HR and payroll lean on a flat bounty per referred business instead: 9 of the 14 state a fixed amount rather than a percentage. Deel, for example, pays $500 per sales-qualified referral and $1,000 more when it becomes a paying customer.&lt;/p&gt;

&lt;p&gt;Small print on the small categories: HR has 14 programs and payments 17. Read their rows as a direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  The recurring flag is not the whole deal
&lt;/h2&gt;

&lt;p&gt;Two more terms decide whether recurring money reaches you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cookie window.&lt;/strong&gt; 244 of the 471 don't publish one. A lifetime commission is worth nothing on a buyer who converts after the cookie expired.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minimum payout.&lt;/strong&gt; 322 don't publish one. Recurring commission at low volume can sit under a threshold for a long time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The data
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://affiliateprogramterms.com/guide/#by-category" rel="noopener noreferrer"&gt;affiliateprogramterms.com/guide&lt;/a&gt;&lt;/strong&gt;: this breakdown, recomputed from the records every time the directory changes, so it stays current after this post stops being.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/dave8172/affiliate-program-terms/blob/master/data/affiliate-program-terms-free.csv" rel="noopener noreferrer"&gt;The free 50 as CSV&lt;/a&gt;&lt;/strong&gt;: the free tier, one row per program, with the source URL and check date on every row. An empty cell means the vendor doesn't state it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Terms change without notice. One program here changed both its rate and its cookie window between two checks three days apart, which is why every figure carries a date. If you find one that has gone stale, the vendor URL is enough; the correction is made against the page.&lt;/p&gt;

</description>
      <category>marketing</category>
      <category>data</category>
      <category>saas</category>
      <category>affiliate</category>
    </item>
    <item>
      <title>Ask twice: seven measurements from building with Jev</title>
      <dc:creator>Dave</dc:creator>
      <pubDate>Sun, 20 Sep 2026 11:23:30 +0000</pubDate>
      <link>https://dev.to/dave8172/ask-twice-jevs-choice-splits-its-own-vote-58kk</link>
      <guid>https://dev.to/dave8172/ask-twice-jevs-choice-splits-its-own-vote-58kk</guid>
      <description>&lt;p&gt;Jev is a System One model from TypeSafe. It reads natural language like any LLM,&lt;br&gt;
and instead of writing a reply it returns a probability distribution over options&lt;br&gt;
you define. No prose, no reasoning trace, no JSON to repair.&lt;/p&gt;

&lt;p&gt;I spent a day building waif, which reads a piece of&lt;br&gt;
text and names the feeling in it. That job has no right answer, which makes it an&lt;br&gt;
unusually honest test rig: nothing can be graded against a key, so every design&lt;br&gt;
decision has to be either argued or measured.&lt;/p&gt;

&lt;p&gt;I measured five, argued two, and the two I argued were both wrong. Sections 1&lt;br&gt;
and 7 are those, kept in place with the measurement underneath rather than&lt;br&gt;
quietly edited out.&lt;/p&gt;


&lt;h2&gt;
  
  
  1. A Choice splits its own vote between synonyms — but less than I claimed
&lt;/h2&gt;

&lt;p&gt;Here is the argument I built on, and I am leaving it in its original words&lt;br&gt;
because the correction under it is the useful part.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Annoyed&lt;/em&gt;, &lt;em&gt;irritated&lt;/em&gt; and &lt;em&gt;frustrated&lt;/em&gt; are one feeling in three wordings. A&lt;br&gt;
Choice divides probability between them, so a text the model read perfectly&lt;br&gt;
clearly comes back split three ways, and confidence collapses for a reason that&lt;br&gt;
has nothing to do with the input. The distribution is telling you about your&lt;br&gt;
option list, not about the text. Therefore: never put sixty near-synonyms in one&lt;br&gt;
Choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then I measured it, and the effect is real but far smaller than the argument&lt;br&gt;
needs.&lt;/strong&gt; One Choice over all 62 words, same glosses, 23 texts:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Mean P(chosen word)&lt;/th&gt;
&lt;th&gt;Score vs acceptable words&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Two Choices — family, then shade within it&lt;/td&gt;
&lt;td&gt;0.689&lt;/td&gt;
&lt;td&gt;43 / 46&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;One Choice over all 62 words&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.812&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;44 / 46&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The one-Choice version is &lt;em&gt;more&lt;/em&gt; certain, not less. It picks the same word 21&lt;br&gt;
times out of 23. On a plainly frustrated text it returns &lt;em&gt;frustration&lt;/em&gt; at 0.98.&lt;br&gt;
Stripping the glosses and asking with 62 bare words barely moves it: 0.79.&lt;/p&gt;

&lt;p&gt;So where does the splitting go? It shows up exactly where you would want it to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;shipped   pride 0.38  →  excitement 0.36
scan      worry 0.54  →  anxiety    0.39
sentit    guilt 0.60  →  embarrassment 0.26
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every one of those lost ground to a near-synonym — that &lt;em&gt;is&lt;/em&gt; vote splitting. But&lt;br&gt;
the two-Choice design flattens on &lt;strong&gt;the same four texts&lt;/strong&gt; (0.33, 0.53, 0.35,&lt;br&gt;
0.32). The splitting is not caused by the option list being long. It is caused&lt;br&gt;
by the text genuinely sitting between two words, and both designs report it&lt;br&gt;
because both are calibrated. &lt;em&gt;"I snapped at her in front of the kids"&lt;/em&gt; really is&lt;br&gt;
guilt and regret at once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A long option list does not flatten a clear text.&lt;/strong&gt; That is the part I had&lt;br&gt;
backwards, and the rule has to be narrower than I wrote it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A Choice splits its vote when two options are &lt;strong&gt;the same answer in different&lt;br&gt;
wordings&lt;/strong&gt;. The fix is criteria that separate them — not fewer options.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sixty-two words each carrying a gloss that distinguishes it (&lt;em&gt;annoyance: a small&lt;br&gt;
thing, quickly over&lt;/em&gt; against &lt;em&gt;resentment: an old grievance still carried&lt;/em&gt;) are&lt;br&gt;
sixty-two alternatives, not sixty-two synonyms. The count was never the problem.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Ask twice anyway — for a reason that survived
&lt;/h2&gt;

&lt;p&gt;Section 1 was the reason I split the question in two, and section 1 did not&lt;br&gt;
hold. The split stayed, and it is worth being precise about what is now holding&lt;br&gt;
it up, because "it scored the same and I had already built it" is not a reason.&lt;/p&gt;

&lt;p&gt;Families are genuinely different answers — anger, fear, sadness and shame are&lt;br&gt;
not wordings of each other. And once the family is fixed, so are its shades:&lt;br&gt;
&lt;em&gt;which shade of anger&lt;/em&gt; is a fair question, because the context has already ruled&lt;br&gt;
out the fifty-four words that were never in the running.&lt;/p&gt;

&lt;p&gt;What that buys, and one Choice over 62 words cannot, is &lt;strong&gt;two separate&lt;br&gt;
uncertainty signals&lt;/strong&gt;. Family confidence and word confidence are different&lt;br&gt;
doubts: &lt;em&gt;"I do not know whether this is sadness or affection"&lt;/em&gt; is not the same&lt;br&gt;
failure as &lt;em&gt;"it is clearly shame, but guilt or embarrassment?"&lt;/em&gt;. The page says&lt;br&gt;
different things in each case. One Choice gives you one number that cannot tell&lt;br&gt;
them apart.&lt;/p&gt;

&lt;p&gt;The cost is honest too: a wobble in the family answer corrupts the word, because&lt;br&gt;
once family says sadness, &lt;em&gt;nostalgia&lt;/em&gt; is not on the ballot. That cost me one&lt;br&gt;
text out of 23 — &lt;em&gt;"drove past the old house, the tree we planted is taller than&lt;br&gt;
the roof"&lt;/em&gt; came back &lt;strong&gt;sorrow&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So the naming became two Choices in sequence: the first picks the family, the&lt;br&gt;
second picks the shade from that family alone. TypeSafe's docs say a second&lt;br&gt;
request is warranted when an earlier answer determines the next question's&lt;br&gt;
options. This is exactly that, and here is what it bought — scored on 24 texts&lt;br&gt;
against a list of acceptable answers for each:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Design&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Nearest word in the whole vocabulary, by published valence / arousal / dominance&lt;/td&gt;
&lt;td&gt;3 / 24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nearest word within a family, by rank on the axis that separates that family&lt;/td&gt;
&lt;td&gt;9 / 24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nearest word within a family, by distance in those published ratings&lt;/td&gt;
&lt;td&gt;12 / 24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Family chosen by the model, then the shade chosen by the model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;22 / 24&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both of the two failures were the wrong &lt;em&gt;family&lt;/em&gt;. Given the family, the shade was&lt;br&gt;
right every single time.&lt;/p&gt;

&lt;p&gt;The alternative was speculative fan-out: ask all eleven within-family Choices in&lt;br&gt;
the first request, each stating its own family as a premise, and keep only the&lt;br&gt;
answer belonging to the family that won. That keeps a reading to one request, at&lt;br&gt;
roughly 3–4k input tokens against 1,250 + 444 for two. I reasoned about it and&lt;br&gt;
rejected it: two requests won on cost, and on a latency story I could explain.&lt;/p&gt;

&lt;p&gt;That paragraph was wrong when I published it. It is still here because&lt;br&gt;
section 7 is the part of this post&lt;br&gt;
I would keep if I could only keep one.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Knowing which half to give the model is a decision you can measure
&lt;/h2&gt;

&lt;p&gt;The job is: read a text, name the feeling. That job splits between the code I&lt;br&gt;
write and the model I call, and the only real decision is where the line falls —&lt;br&gt;
how much of the work do I hand over?&lt;/p&gt;

&lt;p&gt;The first three rows of that table are me keeping most of it.&lt;/p&gt;

&lt;p&gt;I had a reason. Human ratings exist for exactly the dimensions I was measuring:&lt;br&gt;
Warriner, Kuperman &amp;amp; Brysbaert scored 13,915 English words for valence, arousal&lt;br&gt;
and dominance by asking people. So the model's job shrinks to placing the text&lt;br&gt;
on those three scales, and my code finishes the job — look up the nearest word&lt;br&gt;
in the table, return it. Research-backed coordinates instead of somebody's&lt;br&gt;
guesses. It looked like the serious version.&lt;/p&gt;

&lt;p&gt;It scored &lt;strong&gt;3/24&lt;/strong&gt;. Handing the model the whole naming job scored 22/24.&lt;/p&gt;

&lt;p&gt;Three reasons, worth knowing before anyone else reaches for an emotion lexicon:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Three scales were really about two.&lt;/strong&gt; Across these emotion words, valence
and dominance move together at &lt;strong&gt;+0.87&lt;/strong&gt; — a word that reads pleasant almost
always reads in-control. The third scale is close to a copy of the first, so
it separates far less than the theory promises.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The unpleasant words are crammed into one corner.&lt;/strong&gt; Fear, frustration,
worry, terror, jealousy and embarrassment land in a ball small enough that
taking the nearest one is close to picking at random. The coordinates are
fuzzy to begin with: each is an average over about twenty raters who disagreed
by around 1.7 on a 1–9 scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rating a word on its own is not the same measurement as reading a
sentence.&lt;/strong&gt; People rate the bare word &lt;em&gt;"gratitude"&lt;/em&gt; as far more activated than
an actual grateful message reads. The two sets of numbers were never on the
same ruler.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I tried to fix that last mismatch by re-centring both sides against a sample of&lt;br&gt;
texts. It got worse, and instructively: the sample leaned negative, so the&lt;br&gt;
correction leaned negative, and a plainly warm text landed below the middle and&lt;br&gt;
got named from the sad half of the space. That is straightening a bent ruler&lt;br&gt;
with a bent ruler.&lt;/p&gt;

&lt;p&gt;The lesson is not &lt;em&gt;don't use lexicons&lt;/em&gt;. It is that &lt;strong&gt;where the line falls&lt;br&gt;
between what code owns and what the model owns is a design decision with a&lt;br&gt;
number attached&lt;/strong&gt;, and my intuition about it was wrong by a factor of seven.&lt;/p&gt;

&lt;p&gt;The norms kept the one job they are good at. A Score is not given "rate this 1&lt;br&gt;
to 5" — it is given a rubric, a written description of what each level means,&lt;br&gt;
the way a grading key spells out what a B looks like. Every level of mine now&lt;br&gt;
names words whose ratings were measured, so &lt;em&gt;"as activated as rage or panic"&lt;/em&gt; is&lt;br&gt;
a claim a reader can argue with. &lt;em&gt;"Very aroused"&lt;/em&gt; is only a word getting louder.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Confidence is peakedness — and peakedness misses a coin toss
&lt;/h2&gt;

&lt;p&gt;A Score's confidence measures how bunched together the answer is. It does not&lt;br&gt;
measure how likely the model is to be right, which is the trap everyone hits&lt;br&gt;
first.&lt;/p&gt;

&lt;p&gt;Those two sound like the same thing until a reading like this one turns up.&lt;br&gt;
Confidence came back at &lt;strong&gt;0.72&lt;/strong&gt;, comfortably above any threshold I would set,&lt;br&gt;
and here is what was underneath it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Where the answer sat&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The winning level&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;50%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The level right next to it&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The other three&lt;/td&gt;
&lt;td&gt;1%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 0.72 is not lying. Ninety-nine percent of the weight really is in two&lt;br&gt;
buckets and there is nothing anywhere else — that is about as bunched as an&lt;br&gt;
answer gets.&lt;/p&gt;

&lt;p&gt;It is also a coin toss. Choosing between the top two is 50 against 49, and my&lt;br&gt;
page announced the winner in exactly the voice it uses at 0.98.&lt;/p&gt;

&lt;p&gt;Confidence cannot catch this, because &lt;em&gt;bunched into two neighbours&lt;/em&gt; is still&lt;br&gt;
bunched. What I actually wanted was the &lt;strong&gt;gap between first and second place&lt;/strong&gt;,&lt;br&gt;
which is a different number entirely: 98 against 1 is a winner, 50 against 49 is&lt;br&gt;
a tie with a rounding error. Confidence says the same thing about both — and the&lt;br&gt;
gap is the one a reader cares about.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Audit your questions: one of mine fired on 21 of 24 inputs
&lt;/h2&gt;

&lt;p&gt;A question set grows by accretion. Each addition looks free, and none of them&lt;br&gt;
announce that they have stopped saying anything.&lt;/p&gt;

&lt;p&gt;So run them over a corpus and look at the spread. Mine had a Noul asking &lt;em&gt;is more&lt;br&gt;
than one feeling present&lt;/em&gt;. It returned ≥0.6 on &lt;strong&gt;twenty-one of twenty-four&lt;/strong&gt;&lt;br&gt;
texts. That is not a signal about the input, it is a property of writing — a&lt;br&gt;
question paying tokens to tell you something you already knew.&lt;/p&gt;

&lt;p&gt;Two more went for a different reason. They were informative, but nothing&lt;br&gt;
downstream consumed them beyond appending a line to the output. A question whose&lt;br&gt;
entire effect is an occasional footnote costs a reader more attention than it&lt;br&gt;
returns.&lt;/p&gt;

&lt;p&gt;The check is cheap: for every question, the min, max and spread of its answer&lt;br&gt;
across a representative corpus. Anything that barely moves is either a gate or a&lt;br&gt;
mistake.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Two small things that cost real time
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A rubric level must not contain a word from another axis.&lt;/strong&gt; My control rubric&lt;br&gt;
had a level reading &lt;em&gt;"Overwhelmed: struggling to keep any grip on it"&lt;/em&gt;. That&lt;br&gt;
primes the model with a feeling while asking it about agency, and it labels the&lt;br&gt;
output with a word that is not a position on a control scale at all. Every level&lt;br&gt;
of an axis has to be a point on that axis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Input is dominated by your rubrics, not by your input.&lt;/strong&gt; Criteria are sent on&lt;br&gt;
every call, so a one-line text costs almost exactly what a paragraph does. At&lt;br&gt;
this size requests are the scarce resource and tokens are not — which is the&lt;br&gt;
whole argument for batching every independent question into one call.&lt;/p&gt;

&lt;p&gt;I stopped one clause too early, and wrote that a &lt;em&gt;dependent&lt;/em&gt; second call is&lt;br&gt;
therefore a real cost rather than a rounding error. Read it again: if requests&lt;br&gt;
are scarce and tokens are not, the conclusion goes the other way.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. I reasoned where I should have measured
&lt;/h2&gt;

&lt;p&gt;Someone read section 2, noticed I had talked myself out of the fan-out, and told&lt;br&gt;
me to go and time both. It took forty minutes and a dollar's worth of nothing.&lt;/p&gt;

&lt;p&gt;Run the same 24 texts through both designs, twice each, back to back on every&lt;br&gt;
text so neither gets the warmer connection:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;th&gt;Input tokens&lt;/th&gt;
&lt;th&gt;Cost / 1,000 readings&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Two requests, the second dependent&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;762ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1,698&lt;/td&gt;
&lt;td&gt;$0.071&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One request, eleven speculative Choices&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;398ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2,803&lt;/td&gt;
&lt;td&gt;$0.118&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One request was faster on &lt;strong&gt;46 of 46&lt;/strong&gt; pairs. It named the same family &lt;strong&gt;48 out&lt;br&gt;
of 48 times&lt;/strong&gt;, and the same shade 45 of 48 — where all three misses were texts&lt;br&gt;
under 0.41 confidence that the page already reports as sitting between two&lt;br&gt;
words, and where the &lt;em&gt;sequential&lt;/em&gt; design disagreed with itself between rounds on&lt;br&gt;
one of them. The eleven extra questions barely moved the six that were already&lt;br&gt;
there: mean drift of 0.022 on a 0–4 axis score, and identical intent 48 times&lt;br&gt;
out of 48.&lt;/p&gt;

&lt;p&gt;So the fan-out is 1.9× faster for &lt;strong&gt;five hundredths of a cent&lt;/strong&gt; a thousand&lt;br&gt;
readings, and it halves the request count against the limit that actually binds.&lt;/p&gt;

&lt;p&gt;Two things I had, and did not put together:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev "ingests the state once and evaluates every question against it in&lt;br&gt;
parallel."&lt;/strong&gt; That sentence is in TypeSafe's own model card, and I had quoted the&lt;br&gt;
half of it that suited me. Question count is nearly free; a round trip is not.&lt;br&gt;
Ten wasted questions cost less than one extra wait.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev charges for input only — output tokens are free.&lt;/strong&gt; The fan-out triples the&lt;br&gt;
output, returning eleven probability distributions instead of one, and that is&lt;br&gt;
worth exactly nothing on the bill.&lt;/p&gt;

&lt;p&gt;What I actually got wrong is narrower than "I didn't measure it", and more&lt;br&gt;
useful. The docs say a second request is warranted when the first answer&lt;br&gt;
determines the second question's &lt;em&gt;options&lt;/em&gt;. It does — and I read that as&lt;br&gt;
settling the matter. But &lt;em&gt;determined&lt;/em&gt; is not &lt;em&gt;unknown&lt;/em&gt;: there were only ever&lt;br&gt;
eleven option sets, all of them written down in my own source file, so every one&lt;br&gt;
of them could be asked on spec. The rule is about what the options are. It says&lt;br&gt;
nothing about when you are allowed to ask.&lt;/p&gt;

&lt;p&gt;The tell was in my own sentence. "A latency story I could explain" is not a&lt;br&gt;
latency number. Anywhere a design note says &lt;em&gt;presumably&lt;/em&gt;, &lt;em&gt;roughly&lt;/em&gt;, or &lt;em&gt;I could&lt;br&gt;
explain&lt;/em&gt;, there is a measurement someone is about to make for you, and it is&lt;br&gt;
cheaper to make it yourself.&lt;/p&gt;




&lt;p&gt;waif is live at &lt;a href="https://dave8172-website.vercel.app/waif" rel="noopener noreferrer"&gt;dave8172-website.vercel.app/waif&lt;/a&gt;. Every number in this post is on the page, under&lt;br&gt;
&lt;em&gt;How it works&lt;/em&gt;, next to the vocabulary it was scored against.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>programming</category>
      <category>datascience</category>
    </item>
    <item>
      <title>Choice, Score and Noul: four mistakes with Jev's primitives</title>
      <dc:creator>Dave</dc:creator>
      <pubDate>Sat, 19 Sep 2026 06:25:37 +0000</pubDate>
      <link>https://dev.to/dave8172/choice-score-and-noul-four-mistakes-with-jevs-primitives-1eng</link>
      <guid>https://dev.to/dave8172/choice-score-and-noul-four-mistakes-with-jevs-primitives-1eng</guid>
      <description>&lt;p&gt;Jev is a System One model from TypeSafe. It reads natural language like any LLM,&lt;br&gt;
and instead of writing a reply it returns a probability distribution over&lt;br&gt;
options you supply. No prose, no reasoning trace, no JSON to repair — a number&lt;br&gt;
per option, summing to 1.&lt;/p&gt;

&lt;p&gt;I spent a week building a Magic 8 Ball on it, which&lt;br&gt;
sounds like a toy and turned out to be a decent test rig: every answer is a&lt;br&gt;
judgment with no ground truth, which is where calibrated probabilities either&lt;br&gt;
earn their keep or embarrass you. It's live at&lt;br&gt;
&lt;a href="https://dave8172-website.vercel.app/jevball" rel="noopener noreferrer"&gt;/jevball&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I got the primitives wrong four times. Each mistake produced working code that&lt;br&gt;
looked right, so they're worth writing down.&lt;/p&gt;
&lt;h2&gt;
  
  
  The three primitives in one paragraph each
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Choice&lt;/strong&gt; takes a map of named options and returns a probability for each, plus&lt;br&gt;
the highest-scoring one. Routing a ticket to a department, classifying a&lt;br&gt;
document, picking a handler.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Score&lt;/strong&gt; takes an ordered array of level descriptions — two to ten — and returns&lt;br&gt;
a position along them that can land between levels. Severity, frustration, skill&lt;br&gt;
level. Anything on a spectrum you can describe in words.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Noul&lt;/strong&gt; takes a yes/no statement and returns one number: the probability it's&lt;br&gt;
true. No confidence field, because the probability &lt;em&gt;is&lt;/em&gt; the answer.&lt;/p&gt;

&lt;p&gt;All three take the same two inputs: &lt;code&gt;state&lt;/code&gt;, the data being judged, and&lt;br&gt;
&lt;code&gt;instructions&lt;/code&gt;, the question asked about it. That's the whole surface.&lt;/p&gt;
&lt;h2&gt;
  
  
  Mistake 1: reading the score as a winner
&lt;/h2&gt;

&lt;p&gt;A Score returns &lt;code&gt;score&lt;/code&gt;, a position on your levels. My ball rounded it to pick a&lt;br&gt;
bucket. Someone asked it &lt;em&gt;"is black a color?"&lt;/em&gt; and got this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bars   : 0=2%  1=6%  2=19%  3=16%  4=56%
score  : 3.18
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Math.round(3.18)&lt;/code&gt; is 3 — a bucket holding &lt;strong&gt;16%&lt;/strong&gt; — while 56% of the mass sat&lt;br&gt;
on level 4 beside it.&lt;/p&gt;

&lt;p&gt;The score is a &lt;strong&gt;probability-weighted average&lt;/strong&gt;. That 19% on level 2 dragged the&lt;br&gt;
average down across a bucket boundary. An average of a skewed distribution&lt;br&gt;
points at a place the distribution isn't.&lt;/p&gt;

&lt;p&gt;TypeSafe defines a Choice's answer as the option with the highest probability.&lt;br&gt;
Score has no equivalent field, and I quietly substituted rounding for one. The&lt;br&gt;
fix is to take the argmax of &lt;code&gt;probabilities&lt;/code&gt; yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;probs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Array&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;length&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;probabilities&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;level&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;indexOf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(...&lt;/span&gt;&lt;span class="nx"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It agrees with rounding on every peaked distribution, which is why it survived&lt;br&gt;
my whole test suite. It differs exactly when the distribution is skewed, which&lt;br&gt;
is when it matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake 2: reading low confidence as doubt about the answer
&lt;/h2&gt;

&lt;p&gt;Every Choice and Score comes back with &lt;code&gt;confidence&lt;/code&gt;, 0 to 1. I built a "reply&lt;br&gt;
hazy" branch on the assumption that a vague question would produce a low one.&lt;/p&gt;

&lt;p&gt;Then I measured it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;question&lt;/th&gt;
&lt;th&gt;score&lt;/th&gt;
&lt;th&gt;confidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Should I quit my job?&lt;/td&gt;
&lt;td&gt;1.97&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.98&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Will it rain next Tuesday?&lt;/td&gt;
&lt;td&gt;1.99&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.98&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Is black a color?&lt;/td&gt;
&lt;td&gt;3.18&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.33&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two are as uncertain as a question gets, and confidence is 0.98.&lt;br&gt;
Nearly all the mass sat on the level that &lt;em&gt;means&lt;/em&gt; "could go either way" — the&lt;br&gt;
distribution was a sharp spike on a level labelled uncertainty.&lt;/p&gt;

&lt;p&gt;Confidence measures how &lt;strong&gt;peaked&lt;/strong&gt; the distribution is. It answers "did the&lt;br&gt;
options separate cleanly", and it will happily report 0.98 for a confident&lt;br&gt;
coin-flip. My hazy branch would have fired on precisely the wrong questions.&lt;/p&gt;

&lt;p&gt;Low confidence has three different causes, and they need different fixes: the&lt;br&gt;
question is genuinely contested, the options overlap each other, or &lt;code&gt;state&lt;/code&gt;&lt;br&gt;
doesn't contain enough to decide. Only the first one is the model being honest.&lt;/p&gt;

&lt;p&gt;I tried to work out the formula by fitting twelve real responses. &lt;code&gt;max(p)&lt;/code&gt; gets&lt;br&gt;
closest at 0.04 mean error and matches several exactly, but not all —&lt;br&gt;
&lt;em&gt;"will humans land on Mars before 2050?"&lt;/em&gt; reported 0.58 against a peak of 0.50,&lt;br&gt;
higher than the peak. The docs are upfront that it's a convenience statistic and&lt;br&gt;
hand you the full &lt;code&gt;probabilities&lt;/code&gt; so you can compute your own. If a decision&lt;br&gt;
rides on it, do that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake 3: offering options that split their own vote
&lt;/h2&gt;

&lt;p&gt;The classic 8 Ball has twenty answers. My first design asked Jev to pick one, as&lt;br&gt;
a Choice.&lt;/p&gt;

&lt;p&gt;Ten of those twenty mean "yes". "It is certain", "Without a doubt", "Yes&lt;br&gt;
definitely" — the same claim in different words. Here's what happens, on one&lt;br&gt;
question with four different option sets:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;options offered&lt;/th&gt;
&lt;th&gt;answer&lt;/th&gt;
&lt;th&gt;winner's share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;yes&lt;/code&gt; / &lt;code&gt;no&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;69%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;yes&lt;/code&gt; / &lt;code&gt;definitely&lt;/code&gt; / &lt;code&gt;certainly&lt;/code&gt; / &lt;code&gt;no&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;48%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The total yes-mass is identical: 70% in the second row, spread across three&lt;br&gt;
near-synonyms. Nothing changed about the question. The options ate each other,&lt;br&gt;
and confidence fell from 0.38 to 0.31 reporting a disagreement that existed only&lt;br&gt;
in my option list.&lt;/p&gt;

&lt;p&gt;So the ball asks Jev for one of &lt;strong&gt;five ordered buckets&lt;/strong&gt;, and the code picks&lt;br&gt;
which of that bucket's phrasings to show. One judgment with a right answer goes&lt;br&gt;
to the model; the theatre stays in software.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistake 4: forgetting that the options are the prompt
&lt;/h2&gt;

&lt;p&gt;Your question IDs never reach the model. The &lt;code&gt;instructions&lt;/code&gt; and every entry in&lt;br&gt;
&lt;code&gt;criteria&lt;/code&gt; do — names and descriptions both. The option list isn't a filter you&lt;br&gt;
apply to a result. It's part of what you're asking.&lt;/p&gt;

&lt;p&gt;Same question, same model, options changed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;options offered&lt;/th&gt;
&lt;th&gt;answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;yes&lt;/code&gt; / &lt;code&gt;no&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;yes&lt;/strong&gt; (69%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;yes_everyday&lt;/code&gt; / &lt;code&gt;no_physics&lt;/code&gt; / &lt;code&gt;depends&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;depends&lt;/strong&gt; (51%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Depends" won — an answer that simply didn't exist in the first row. Black is a&lt;br&gt;
colour in everyday use and the absence of light in physics, so "depends" is&lt;br&gt;
arguably the best answer available. It was unreachable until I offered it.&lt;/p&gt;

&lt;p&gt;The docs put it plainly: the model cannot choose an omitted value. Leave an&lt;br&gt;
option out and you haven't biased the result, you've made it impossible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is actually for
&lt;/h2&gt;

&lt;p&gt;Every judgment above cost about &lt;strong&gt;$0.000014&lt;/strong&gt;. Jev charges $0.042 per million&lt;br&gt;
input tokens and nothing for output, against $1.00/$5.00 for the cheapest&lt;br&gt;
frontier model I'd otherwise reach for — roughly 27× on a like-for-like&lt;br&gt;
classification, before an LLM's format instructions and reasoning tokens widen&lt;br&gt;
it further.&lt;/p&gt;

&lt;p&gt;That price only matters because of what you give up. Jev writes no prose,&lt;br&gt;
explains nothing, and can't answer anything whose answer space you can't&lt;br&gt;
enumerate. For summarising, drafting, coding or open-ended reasoning it's the&lt;br&gt;
wrong tool entirely.&lt;/p&gt;

&lt;p&gt;What it's good at is the narrow decision a pipeline makes ten thousand times a&lt;br&gt;
day, where you need a number you can threshold on. The documented pattern is to&lt;br&gt;
put it in front of the expensive work: score everything, route the confident&lt;br&gt;
cases to deterministic code, escalate the rest to a model or a person. You pay&lt;br&gt;
fourteen dollars a million to decide, and frontier prices only on the slice that&lt;br&gt;
earned it.&lt;/p&gt;

&lt;p&gt;One caveat worth keeping. Calibration is a property of &lt;em&gt;groups&lt;/em&gt; of predictions —&lt;br&gt;
across many answers, the ones marked 0.8 should be right about 80% of the time.&lt;br&gt;
It guarantees nothing about any single answer, and it's measured on TypeSafe's&lt;br&gt;
data, not yours. Validate it in your own domain before trusting a threshold.&lt;/p&gt;




&lt;p&gt;The ball is at &lt;a href="https://dave8172-website.vercel.app/jevball" rel="noopener noreferrer"&gt;/jevball&lt;/a&gt;. Ask it something you actually want to know&lt;br&gt;
and open the panel underneath — it shows the full distribution, the confidence,&lt;br&gt;
and which of the four questions produced the answer.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>typescript</category>
      <category>webdev</category>
    </item>
    <item>
      <title>252 affiliate programs, and 38% of the terms are not published</title>
      <dc:creator>Dave</dc:creator>
      <pubDate>Sun, 13 Sep 2026 06:11:58 +0000</pubDate>
      <link>https://dev.to/dave8172/252-affiliate-programs-and-38-of-the-terms-are-not-published-4fbl</link>
      <guid>https://dev.to/dave8172/252-affiliate-programs-and-38-of-the-terms-are-not-published-4fbl</guid>
      <description>&lt;p&gt;I read the affiliate page of 252 developer and SaaS tools and recorded five terms for each one: commission rate, recurring or one-time, cookie window, minimum payout, and payout method. Every field carries the URL it came from and the date it was read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;483 of those 1,260 fields — 38% — are not published by the vendor at all.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That number is the whole finding, but the shape of it is more useful than the size:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Term&lt;/th&gt;
&lt;th&gt;Share of the 252 that never state it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Minimum payout&lt;/td&gt;
&lt;td&gt;65%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cookie window&lt;/td&gt;
&lt;td&gt;52%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Payout method&lt;/td&gt;
&lt;td&gt;45%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recurring or one-time&lt;/td&gt;
&lt;td&gt;23%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Commission rate&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The term almost everyone publishes is the commission rate. The three most often missing are the three that decide whether the money reaches you at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the rate is the least useful number
&lt;/h2&gt;

&lt;p&gt;A percentage tells you nothing until you know how many times you get paid it. On a $100/month product:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;$150 one-time bounty&lt;/strong&gt; pays you $150, ever&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;60% for 12 months&lt;/strong&gt; pays about $720 in year one, then stops&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;30% for the customer's lifetime&lt;/strong&gt; pays $360 a year, indefinitely&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same customer, three different outcomes, and the lowest percentage wins by a distance. Of the 252 programs, 143 pay recurring commission, 51 pay once, and 58 don't say which.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cookie window decides whether research-heavy traffic is worth anything
&lt;/h2&gt;

&lt;p&gt;Someone clicks your link, the company drops a cookie, and if they buy after it expires you get nothing. For anything people think about before buying — which is most developer tooling — this is the term that decides whether a post you wrote is an asset or a donation.&lt;/p&gt;

&lt;p&gt;The real range: &lt;strong&gt;Jasper at 14 days against TubeBuddy at 365&lt;/strong&gt;, and two programs state that attribution never expires. Another 131 publish no window at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The minimum payout is where small affiliates actually lose
&lt;/h2&gt;

&lt;p&gt;Most programs hold your earnings until they cross a threshold. Earn $80 in a year against a $250 minimum and you have earned nothing you can spend. Thresholds here run from $5 to $450, seven state no minimum, and 164 don't say.&lt;/p&gt;

&lt;h2&gt;
  
  
  Payout method is the one to read first if you're not in the US
&lt;/h2&gt;

&lt;p&gt;Some programs pay by PayPal only. Others use Wise, Payoneer, direct bank transfer or EFT. Guides written for a US audience treat this as a footnote. If you're somewhere PayPal is awkward or expensive, it is the difference between money arriving and money technically existing. 113 of the 252 don't say how they pay.&lt;/p&gt;

&lt;h2&gt;
  
  
  One counting decision worth explaining
&lt;/h2&gt;

&lt;p&gt;152 of the 252 name no affiliate network. I deliberately left that out of the transparency count.&lt;/p&gt;

&lt;p&gt;A company that names no network may be running the program itself rather than concealing anything, and "in-house" and "undisclosed" are different facts. Inferring one from the other would mark every company that built its own system as secretive. So the network is published where a vendor states it, and counted nowhere.&lt;/p&gt;

&lt;p&gt;This is the kind of decision that makes or breaks a dataset, and it is usually invisible in a roundup post.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Not stated" is a finding, not a gap
&lt;/h2&gt;

&lt;p&gt;Where a vendor publishes nothing, the record says &lt;em&gt;not stated&lt;/em&gt; rather than filling it with an estimate. Knowing that a company won't publish its cookie window is worth knowing &lt;strong&gt;before&lt;/strong&gt; you build content around it.&lt;/p&gt;

&lt;p&gt;That rule is also why this took reading 252 pages instead of aggregating three listicles. A roundup whose author earns commission from the programs it ranks has an incentive to sort by rate rather than by accuracy. Nothing here earns a commission from any program listed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://affiliateprogramterms.com" rel="noopener noreferrer"&gt;affiliateprogramterms.com&lt;/a&gt;&lt;/strong&gt; — the full directory. 50 programs are free to read in full, each field with its source URL and check date.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/dave8172/affiliate-program-terms" rel="noopener noreferrer"&gt;github.com/dave8172/affiliate-program-terms&lt;/a&gt;&lt;/strong&gt; — the free tier as a generated README, if you'd rather read it in a repo. Every count in it is computed from the dataset rather than typed, so it can't drift from the records it describes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you find a figure that's gone stale, open an issue with the vendor URL. Terms change without notice, which is the reason every field here carries a date.&lt;/p&gt;

</description>
      <category>marketing</category>
      <category>data</category>
      <category>webdev</category>
    </item>
    <item>
      <title>doceval — eval harness for LLM document extraction pipelines</title>
      <dc:creator>Dave</dc:creator>
      <pubDate>Tue, 16 Jun 2026 12:29:37 +0000</pubDate>
      <link>https://dev.to/dave8172/show-hn-doceval-eval-harness-for-llm-document-extraction-pipelines-3gd7</link>
      <guid>https://dev.to/dave8172/show-hn-doceval-eval-harness-for-llm-document-extraction-pipelines-3gd7</guid>
      <description>&lt;p&gt;I kept seeing the same gap: people ship LLM-based document extractors (invoices, receipts, forms) with no systematic way to know how accurate they actually are. So I built doceval — point it at your extractor function + a labeled dataset and get back field-level accuracy, a failure taxonomy (missed_field / hallucination / wrong_format / wrong_value), and optional per-document cost tracking.&lt;/p&gt;

&lt;p&gt;Works with any extractor (Claude, GPT, regex, rules) and any document schema. One JSON label file per document, one Python function, one CLI command.&lt;/p&gt;

&lt;p&gt;Includes a working 20-document invoice example with a Claude Haiku extractor so you can run it immediately.&lt;/p&gt;

&lt;p&gt;GitHub: &lt;a href="https://github.com/dave8172/doceval" rel="noopener noreferrer"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>development</category>
      <category>python</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
