<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: John Green</title>
    <description>The latest articles on DEV Community by John Green (@ramses203).</description>
    <link>https://dev.to/ramses203</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4084119%2Fee714617-da74-4102-8c0f-8b541036c76b.jpg</url>
      <title>DEV Community: John Green</title>
      <link>https://dev.to/ramses203</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ramses203"/>
    <language>en</language>
    <item>
      <title>I Rewrote One Exam Question Fifty Ways. First, the Answer Key Was Wrong.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Fri, 04 Sep 2026 14:00:34 +0000</pubDate>
      <link>https://dev.to/ramses203/i-rewrote-one-exam-question-fifty-ways-first-the-answer-key-was-wrong-5dn7</link>
      <guid>https://dev.to/ramses203/i-rewrote-one-exam-question-fifty-ways-first-the-answer-key-was-wrong-5dn7</guid>
      <description>&lt;p&gt;Two readers said the same thing. Under &lt;a href="https://dev.to/ramses203/the-model-that-costs-3x-more-won-by-exactly-one-question-55aj"&gt;the 3x-price comparison&lt;/a&gt;, Vinh said the next experiment should not be another rerun. It should take the lid question, the one that decided that comparison, and rewrite it in different forms. Under &lt;a href="https://dev.to/ramses203/the-question-that-picked-my-model-didnt-survive-five-reruns-5cfk"&gt;the rerun post&lt;/a&gt;, tonal said that five passes over the same 29 questions are not 145 fresh chances. Both are right. Rerunning the same exam adds runs, not kinds of question.&lt;/p&gt;

&lt;p&gt;So I decided to rewrite the lid question fifty ways. Then, rereading the question, I found that its answer key was wrong. That story comes first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The answer key was wrong
&lt;/h2&gt;

&lt;p&gt;Here is the original.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;페트 300 투명 10박스 뚜껑 흰거 5봉이요 (10 boxes of clear PET 300, 5 bags of the white lids)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The catalog has two white lids, a 300-neck and a 500-neck. The customer did not say which. The answer key said the correct answer was to confirm the 300-neck lid. The reason was that the same message orders 300 bottles, so the lids must be for those bottles.&lt;/p&gt;

&lt;p&gt;But every other question on the same exam grades this kind of case the same way: the correct answer is to ask.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;clear tape, 2 boxes      48mm or 60mm?        key: ask
PET 500, 8 boxes         clear or brown?      key: ask
250 5                    five candidates      key: ask
white lids, 5 bags       300 or 500?          key: confirm 300   ← the only exception
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule is one line. Do not guess what the customer did not write. Ask. Only the lid question broke it. "There are 300 bottles right next to it, so obviously 300" looked so obvious that when I fixed the tape questions earlier, I never looked at this one again.&lt;/p&gt;

&lt;p&gt;So I fixed it. Asking which neck is the correct answer. Confirming the 300 lid is RISKY: right this time, but a guess, so wrong next time. Confirming any other lid is FATAL: the wrong goods ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scores of two earlier posts change
&lt;/h2&gt;

&lt;p&gt;When one question's answer changes, every run that touched it gets re-scored. The other questions stay as they were.&lt;/p&gt;

&lt;p&gt;There are four grades. Clean means no flaw. Harmless means asking about something that did not need asking. Risky means confirming by guess something that should have been asked, right this time but possibly wrong next time. Fatal means the wrong goods ship.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                                 clean  harmless  risky  fatal        clean  harmless  risky  fatal
                                 ─── before ─────────────────────  →  ─── after re-key ───────────────
3x-price post (one pass, 29 questions)
  cheap model                      28       1       0      0     →      29       0       0      0
  expensive model (28 executed)    28       0       0      0     →      27       0       1      0

Five reruns (145 trials per model)
  cheap model                     140       4       1      0     →     143       0       2      0
  expensive model                 139       5       1      0     →     144       0       1      0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every changed cell comes from the one lid question. The runs where a model asked which neck moved from harmless to clean. The one run where the cheap model confirmed 300 moved from clean to risky. Fatal was 0 before and stays 0, because no run ever confirmed the wrong lid.&lt;/p&gt;

&lt;p&gt;In the 3x-price post I wrote that the expensive model got exactly one more question right. That question was this one. The expensive model guessed 300, and the cheap model asked. Under the corrected key, the cheap model won. The conclusion does not change: both models have zero fatal errors, and on a tie in fatal errors you take the cheap one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fifty forms
&lt;/h2&gt;

&lt;p&gt;With the key fixed, I rewrote the lid question fifty ways. The order and the trap stayed the same. Only the wording changed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;T1  word order      white lids 5 bags and clear PET 300 10 boxes
T2  endings         ...please send / ...placing an order / no ending at all
T3  abbreviation    PET300 clear 10bx lids white 5bags
T4  spacing         PET300clear10boxeslidswhite5bags
T5  typos           PTE 300 / lidss / whiet
T6  color words     white / white-color / the white ones / "the clear ones" for the bottle
T7  units           cartons / box / BOX / ten (as a word)
T8  spec notation   300ml / 300cc / PET 300 / PET bottle (300)
T9  sentence shape  two lines / "and also..." / numbered list
T10 noise           greetings / thanks / ORDER!!
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things stayed fixed in every variant. The bottle always has "300" and "clear"; without them the bottle itself becomes ambiguous and it is a different question. The lid uses only a white-word and never a neck size; that is the trap. The quantities are 10 boxes and 5 bags, with unit words. I cut two variants that dropped the unit words, because the exam grades unit-less orders like "250 5" as ask, so dropping the unit changes the answer itself.&lt;/p&gt;

&lt;p&gt;An LLM drafted the variants, and I checked each line against those three rules.&lt;/p&gt;

&lt;p&gt;Grading follows the corrected key.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;asks which neck (with candidates)   clean
confirms CAP-300-W                  RISKY  (a guess that was right this time)
confirms any other lid              FATAL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each variant ran three times per model. Temperature stayed at the default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                                     cheap (Haiku 4.5)   expensive (Sonnet 5)
trials                               150                 150
asked which neck                     140  (93%)          148  (99%)
confirmed 300 by guess (RISKY)        10  ( 7%)            2  ( 1%)
confirmed another lid (FATAL)          0                   0
lost the lid line                      0                   0
bottle confirmed PET-300-CL x10      149 (asked once)    150
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things show.&lt;/p&gt;

&lt;p&gt;First, item matching did not break. Typos, collapsed spacing, other words for white, 300ml and 300cc: none of it moved either model. In 300 trials there were zero wrong lids and zero wrong quantities. The four noise groups (spacing, typos, color words, spec notation) produced zero guesses in 120 trials per model. I expected the accidents to come from noise. They did not.&lt;/p&gt;

&lt;p&gt;Second, the guesses came from smooth sentences. The cheap model guessed ten times: four in the unit group, two in word order, two in sentence shape, one in endings, one in greetings. Each variant ran three times, and two variants made it guess in two of the three: "10 boxes of clear PET 300, and the lids in white, 5 bags please" and "10 boxes of clear PET 300 please. And also 5 bags of the white lids." Both are the most polite and complete sentences in the set. The more a message looked like a finished order, the more the cheap model finished it, filling in the neck size the customer never gave. The expensive model guessed twice, once on "10 boxes of clear PET 300, 5 bags of white lids, please send" and once on "PET300 clear 10 boxes lids white 5 bags".&lt;/p&gt;

&lt;p&gt;Third, no variant made a model guess every time. Eight variants made the cheap model guess at least once, and none made it guess all three times. Guessing is not a property attached to a wording. It is a coin flipped on every run, and some wordings raise the odds. Had I run this once, I might have seen two guesses or five, depending on the day.&lt;/p&gt;

&lt;p&gt;If I had not fixed the key, this experiment would have read the other way. The old key counted confirming 300 as correct, so the cheap model's ten guesses would count as ten correct answers, and the cheap model would win 10 to 2.&lt;/p&gt;

&lt;h2&gt;
  
  
  One objection I expect
&lt;/h2&gt;

&lt;p&gt;"Why not just show the customer an order confirmation sheet?" Yes. Everything that needs asking is better gathered on one screen. But the guessed line has to be marked. If the sheet only says "white lids, 300-neck, 5 bags", a customer who never thought about neck sizes just taps confirm. It has to say "white lids, 5 bags: assumed 300-neck, tell us if 500" for the customer's eye to stop there. The sheet copies what the intake step produced. Intake has to mark the lid as "needs confirmation" for the sheet to carry a warning. If intake quietly settles on 300 without a mark, the sheet prints "300-neck, 5 bags" like any other line, and the customer passes it by.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The answer that looked most obvious on the key was the wrong one. I did not find it by thinking harder. Lined up next to the other 28 questions, this was the only one whose answer was a guess.&lt;/li&gt;
&lt;li&gt;In this experiment the cheap model guessed on polite, finished sentences, not on messy ones. It filled in missing information where the sentence looked complete. This came from one question and two models, so I do not know whether it holds elsewhere.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;P.S. All 300 raw answer sheets, the 50 variants, the run script, and the aggregation script are in the repo → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;. The corrected key is in the same commit, with the reason written in the question's note.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;P.P.S. If you'd rather get new experiments by email: &lt;a href="https://ramses203.substack.com/subscribe" rel="noopener noreferrer"&gt;ramses203.substack.com/subscribe&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>The Textbook Says Grade Binary. I Grade in Four. Was I Wrong the Whole Time?</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Mon, 31 Aug 2026 00:07:00 +0000</pubDate>
      <link>https://dev.to/ramses203/the-textbook-says-grade-binary-i-grade-in-four-was-i-wrong-the-whole-time-1727</link>
      <guid>https://dev.to/ramses203/the-textbook-says-grade-binary-i-grade-in-four-was-i-wrong-the-whole-time-1727</guid>
      <description>&lt;p&gt;There's someone this series treats as its precedent: Hamel Husain, who teaches AI product evaluation and has drawn 4,500 students. &lt;a href="https://dev.to/ramses203/i-made-an-llm-re-grade-my-exam-it-found-two-bugs-in-my-grader-39bi"&gt;The LLM-judge experiment&lt;/a&gt; was his method, tested with my own hands.&lt;/p&gt;

&lt;p&gt;Reading his evals essay, I hit a paragraph that snagged. He recommends grading &lt;strong&gt;binary only&lt;/strong&gt; — good or bad. Scores and fine-grained scales, he argues, cost effort and buy nothing.&lt;/p&gt;

&lt;p&gt;But my grader has four grades: FATAL, RISKY, MISSED, HARMLESS. Every scorecard in this series has come out in those four columns, and the shipping rule itself — FATAL 0 ships — stands on them. Have I spent this entire series doing the thing my own precedent says not to do?&lt;/p&gt;

&lt;h2&gt;
  
  
  What Hamel is actually blocking
&lt;/h2&gt;

&lt;p&gt;What he warns against is the 1-to-5 score scale. Here's why it fails.&lt;/p&gt;

&lt;p&gt;Anyone who has agonized over "is this answer a 3 or a 4?" knows: the boundary between 3 and 4 cannot be defined in words. So the standard differs between graders, and even the same grader drifts between morning and evening. And when the average comes out to 3.7, nobody can say what decision that number makes. Ship, or don't ship? 3.7 doesn't answer.&lt;/p&gt;

&lt;p&gt;Up to here, I agree completely.&lt;/p&gt;

&lt;h2&gt;
  
  
  So is my four-grade table a score scale?
&lt;/h2&gt;

&lt;p&gt;There's a way to tell them apart. &lt;strong&gt;If it's a scale, you can average it.&lt;/strong&gt; A 1-to-5 scale gives you a 3.7.&lt;/p&gt;

&lt;p&gt;Our grades have no average. There is no "half a FATAL," no "somewhere between FATAL and RISKY." All we ever do is count — so many FATALs, so many HARMLESSes.&lt;/p&gt;

&lt;p&gt;Here's what each grade actually names:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FATAL      a wrong confirmation flows into the next step — no human gets a chance to catch it
RISKY      wrong, but a human checkpoint still remains — there's a chance to catch it
MISSED     something that should have been caught wasn't
HARMLESS   over-asked for confirmation — a human is mildly annoyed, nothing more
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Grading one question means asking "did this accident happen — yes or no?" So we're not reading a 4-point dial. We're asking &lt;strong&gt;four binary questions.&lt;/strong&gt; Translated into Hamel's terms: grade pass/fail as a binary — but give the failure a name.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the name matters — three decisions the name made
&lt;/h2&gt;

&lt;p&gt;What if I had lumped everything into good/bad with no names? Let's check against three decisions that actually happened.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://dev.to/ramses203/the-model-that-costs-3x-more-won-by-exactly-one-question-55aj"&gt;The model choice.&lt;/a&gt;&lt;/strong&gt; Every mistake the cheap model made was the same kind: asking "please confirm" on questions it could have answered outright. A question hurts nobody — but good/bad grading counts it as "bad" anyway. So the scoreboard would have read: cheap model, several bads; expensive model, none — &lt;strong&gt;and I would have paid 3x for the clean-looking one.&lt;/strong&gt; With names, those asks were HARMLESS, and the grade that decides shipping — FATAL — was 0 for both. A tie where it matters; on a tie, take the cheap one. Without that name I'd have paid three times more for the same safety.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://dev.to/ramses203/i-stole-my-own-exam-it-failed-the-tool-behind-my-own-numbers-3pdn"&gt;The retirement verdict.&lt;/a&gt;&lt;/strong&gt; I had a tool that sorts 20,000 YouTube comments into "a real customer need" versus "just chatter," and I use its tally to decide what to build. On a 15-question exam it got 7 wrong. Good/bad grading leaves one number — 8 of 15, 53% — and whether 53% is good enough to keep, the number can't say. What decided it was the content of the wrong answers: three of them counted chatter as a need. To undo that mistake, someone would first have to notice it — and nobody re-reads 20,000 comments. A mistake that never gets noticed never gets undone. That's FATAL on the grade table, three times over. Our rule is that a single FATAL is enough to stop using a tool. This one had three. So it was retired — and the 53% played no part in that decision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://dev.to/ramses203/i-gave-a-regex-and-an-llm-the-same-exam-fatal-3-vs-fatal-0-43c4"&gt;The replacement decision.&lt;/a&gt;&lt;/strong&gt; The replacement candidate, an LLM, got 12 of 15 on the same exam — 80%. If percentages decided, "80 beats 53" would be the reason. It wasn't. The LLM also got 3 wrong; the difference is the kind — all three were mistakes a human still gets a chance to catch. &lt;strong&gt;FATAL 0 versus FATAL 3&lt;/strong&gt; made the swap.&lt;/p&gt;

&lt;p&gt;The common thread across all three: &lt;strong&gt;the percentage couldn't answer "ship it? which one? fix what first?" — the name of the failure answered.&lt;/strong&gt; Names set the repair order too: fix FATAL first, let HARMLESS wait. If all seven failures are just "bad," there's no way to know which of the seven to touch first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Hamel's warning is right — don't build score scales. The boundaries can't be defined, and an average makes no decision&lt;/li&gt;
&lt;li&gt;But if you collapse everything into binary, the content of the failure disappears. Ask pass/fail — then name the failure&lt;/li&gt;
&lt;li&gt;One criterion is enough for the name: &lt;strong&gt;can a human undo this accident?&lt;/strong&gt; If not, it's FATAL, and it gets fixed first&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Reading my own precedent, I spent about three days uneasy, thinking I'd been wrong. Taken apart, we were saying the same thing from different angles. When you build something independently and arrive at the same spot, the odds that the spot is right go up. This one was that kind of relief.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;P.S. Hamel's essay is "Your AI Product Needs Evals" at hamel.dev. There were five more places where it overlaps with this series — add a question every time a failure appears, suspect the grader first, 100% pass is not the goal, a domain expert owns the answer key, split the job into scenarios before writing questions. Hamel and I built our approaches without knowing of each other, and they overlap in this many places. Too many for coincidence. I take it as a sign that verifying with an exam is a road worth staying on.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Gave a Regex and an LLM the Same Exam. Fatal 3 vs Fatal 0.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Thu, 27 Aug 2026 09:39:19 +0000</pubDate>
      <link>https://dev.to/ramses203/i-gave-a-regex-and-an-llm-the-same-exam-fatal-3-vs-fatal-0-43c4</link>
      <guid>https://dev.to/ramses203/i-gave-a-regex-and-an-llm-the-same-exam-fatal-3-vs-fatal-0-43c4</guid>
      <description>&lt;p&gt;In &lt;a href="https://dev.to/ramses203/i-stole-my-own-exam-it-failed-the-tool-behind-my-own-numbers-3pdn"&gt;the last post&lt;/a&gt; my comment classifier — a regex — sat a 15-question exam and produced three fatal errors: the kind a human can't undo. Under our rule, FATAL above zero means the tool can't be used. So what replaces it?&lt;/p&gt;

&lt;p&gt;There's one candidate: let an LLM do the classifying. So I staged &lt;strong&gt;a head-to-head on the same exam.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Same 15 questions, same grader, same grade table&lt;/li&gt;
&lt;li&gt;Regex side: the existing tool, untouched (keyword matching)&lt;/li&gt;
&lt;li&gt;LLM side: Sonnet 5, given the classification definitions (what counts as a need; personal anecdotes and rhetorical questions don't) and asked to judge each comment&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One fairness question came up. The LLM was given &lt;strong&gt;the right to answer "needs confirmation."&lt;/strong&gt; The regex can't be given that right — the concept doesn't exist in it, which was exactly the defect the last post pointed out. So this isn't favoritism; it's the difference in ability between the two tools, as-is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              regex                          LLM
clean         8/15 (53%)                     12/15 (80%)
grades        FATAL 3 RISKY 5 MISSED 1 H 0   FATAL 0 RISKY 1 MISSED 1 H 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read by grade, not by score — &lt;a href="https://dev.to/ramses203/the-model-that-costs-3x-more-won-by-exactly-one-question-55aj"&gt;the rule from the model-comparison post&lt;/a&gt;. &lt;strong&gt;FATAL 3 against FATAL 0.&lt;/strong&gt; The regex is at a grade that can't ship; the LLM is at a grade that can. The match was decided there.&lt;/p&gt;

&lt;p&gt;Three places where they split:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The LLM read context.&lt;/strong&gt; The social-commentary comment — "…humanities majors don't pay, so nobody will replace them" — the regex filed as an errors-and-debugging need, because the Korean for "don't pay" shares two characters with its error keyword. The LLM correctly filtered it as commentary. Same for "Can't do this anymore, going to bed lol."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The LLM knows how to say it doesn't know.&lt;/strong&gt; "Me too 😭 happens every time" is a reply that cannot be judged without its parent comment. The regex forced it into a category; the LLM answered "needs confirmation." It was given the right and actually used it. This tool can keep our principle — a wrong confirmation is worse than none.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The LLM needs no keyword dictionary.&lt;/strong&gt; The regex's category dictionary had never registered "Cursor," the AI coding tool, so every comment about Cursor went unclassified. The LLM put them in the AI-tools category from context alone, no dictionary. The maintenance of forever adding keywords simply disappears.&lt;/p&gt;

&lt;p&gt;Not a clean sweep. The LLM dropped one category (a comment about monthly payment never got the pricing/billing label) and confirmed one ambiguous comment it should have handed to a human — one RISKY. Eighty percent, not a hundred. But every one of its mistakes sat in a grade a human can undo.&lt;/p&gt;

&lt;h2&gt;
  
  
  "So why not just verify with an LLM?"
&lt;/h2&gt;

&lt;p&gt;Read this far and the question writes itself. If the LLM is this good, why build exams and grade tables at all — why not just ask an LLM "is this right?"&lt;/p&gt;

&lt;p&gt;Think of a scale. There is one way to know whether a scale is accurate: put a weight whose mass you already know on it. If a 1 kg calibration weight reads 1 kg, the scale can be trusted.&lt;/p&gt;

&lt;p&gt;Here the scale is the tool that judges comments, and the calibration weight is the exam with the answers written in advance. This match had two scales — the regex scale and the LLM scale. Put the same comment on both and they read differently; the thing that settled which one was right was the calibration weight, the 15 questions with known answers. Without it, the story ends at "huh, they disagree."&lt;/p&gt;

&lt;p&gt;"Just ask an LLM" is the same picture: the LLM being verified and the LLM doing the verifying are two scales weighing each other with no calibration weight in the room. When they disagree, you never find out who's right.&lt;/p&gt;

&lt;p&gt;Turn it around and it's good news: with a calibration weight, one scale is enough. Two scales were only needed while choosing. Now that the match is over, what remains is the 15-question exam and the LLM classifier — that pair.&lt;/p&gt;

&lt;p&gt;This week I watched "just let an LLM verify it" fail three times in person:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;When the &lt;a href="https://dev.to/ramses203/i-made-an-llm-re-grade-my-exam-it-found-two-bugs-in-my-grader-39bi"&gt;LLM judge and the code grader disagreed&lt;/a&gt;, the only reason I could rule on who was right was a 29-question exam with known answers&lt;/li&gt;
&lt;li&gt;In &lt;a href="https://dev.to/ramses203/the-ai-exam-author-was-never-wrong-i-still-cant-use-its-exam-3h2"&gt;the exam-author experiment&lt;/a&gt;, the author was Sonnet and the reviewer was Sonnet — 50 out of 50 passed. It was grading itself. Without an exam procedure you don't even notice that's a trap&lt;/li&gt;
&lt;li&gt;Today's 80% and FATAL 0 — those numbers exist only because there was an exam with an answer key. Without it, all that remains is "it looks good." Which is the exact spot "I can't trust it" started from&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And one practical reason: &lt;strong&gt;re-grading is free.&lt;/strong&gt; Every time you change the classification prompt or swap the model, does a human re-verify everything? The exam has its answers stored — a new scorecard takes five minutes. The 15 questions written today became a lifetime exam this tool will retake every time it's touched.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost still counts
&lt;/h2&gt;

&lt;p&gt;The regex is free and instant. The LLM takes tens of seconds per call. For small volumes — my own blog's comments — the LLM is fine. For a 20,000-comment batch, a regex as a first-pass filter with the LLM ruling only on what it flags is a workable combination. Choosing a tool is, in the end, a trade between grade (safety) and cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;When you switch tools, put them on the same exam. Choose by fatal count, not by score&lt;/li&gt;
&lt;li&gt;If a tool can say "I don't know," check whether it actually exercises that right&lt;/li&gt;
&lt;li&gt;Even if an LLM does your verifying, you still need the exam — it's the calibration weight for the thing doing the weighing&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;P.S. The 15-question exam and both classifiers' scorecards are public → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;, comment_exam.py (&lt;code&gt;--compare&lt;/code&gt; shows them side by side)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Stole My Own Exam. It Failed the Tool Behind My Own Numbers.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Thu, 27 Aug 2026 09:37:34 +0000</pubDate>
      <link>https://dev.to/ramses203/i-stole-my-own-exam-it-failed-the-tool-behind-my-own-numbers-3pdn</link>
      <guid>https://dev.to/ramses203/i-stole-my-own-exam-it-failed-the-tool-behind-my-own-numbers-3pdn</guid>
      <description>&lt;p&gt;In &lt;a href="https://dev.to/ramses203/steal-this-exam-heres-how-to-port-it-to-your-own-pipeline-36oo"&gt;the porting guide&lt;/a&gt; I wrote that the exam is built to be stolen — follow five steps and it moves to any job.&lt;/p&gt;

&lt;p&gt;So I tried being the other person. Following only what the guide says, start to finish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to steal it to — my own tool, of all places
&lt;/h2&gt;

&lt;p&gt;For the second job I picked &lt;strong&gt;YouTube comment classification&lt;/strong&gt;: scraping 20,000 comments and sorting each one into "a need," "chatter," or "a signal someone would pay." Every number in &lt;a href="https://dev.to/ramses203/i-scraped-20000-youtube-comments-the-videos-and-the-comments-were-having-two-different-l30"&gt;the 20,000-comments post&lt;/a&gt; came out of this classifier.&lt;/p&gt;

&lt;p&gt;Which makes this a double-edged experiment. It tests whether the exam ports — and at the same time it tests &lt;strong&gt;whether the tool that produced my own published numbers can pass an exam.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The twist comes first — the tool wasn't an AI
&lt;/h2&gt;

&lt;p&gt;Before writing a single question, I opened the classifier's code to understand what I was about to test. The thing that sorted 20,000 comments was not an AI. It was &lt;strong&gt;a regex&lt;/strong&gt; — word matching: "if the comment contains this keyword, it's this category."&lt;/p&gt;

&lt;p&gt;The second line of the actual data file was already an accident.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;My grad-school senior bet that nobody would bother replacing humanities majors because they don't pay. He was right.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Social commentary. Not a need, and certainly not about errors. The classifier had filed it as &lt;strong&gt;a need in the "errors &amp;amp; debugging" category&lt;/strong&gt; — because the Korean phrase for "doesn't pay" contains the same two characters as the error keyword "doesn't work." With 10,000 likes, it sat near the top of the ranking.&lt;/p&gt;

&lt;p&gt;The accident showed up before the exam even existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I followed the five steps exactly
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — write down the worst.&lt;/strong&gt; These classifications feed decisions about what to build and what to sell. So the worst accident is "promoting chatter into a need and manufacturing fake demand." A product decision built on fake demand burns weeks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2 — the grade table.&lt;/strong&gt; Four grades: fatal, risky, missed, harmless. In the guide I had written "only the first line, FATAL, is redefined per project; the other three read the same everywhere." Porting it, one line didn't hold. In the order domain, MISSED (failing to catch something) was a mild grade — miss an order and the customer calls, so a human finds out. But a comment that gets dropped is never looked at again by anyone. The same MISSED is effectively fatal here.&lt;/p&gt;

&lt;p&gt;Filling in the table also exposed something about the regex tool itself: &lt;strong&gt;it has no way to say "I'm not sure, a human should look."&lt;/strong&gt; Every comment gets forced into a category. Under our first principle — a wrong confirmation is worse than no confirmation — that's a dangerous property before the exam has even started.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3 — plant the traps.&lt;/strong&gt; I rewrote the guide's four trap types (confusable pairs, plausible non-targets, mid-message reversals, after-learning) into comment form. While planting them, a fifth type showed up that isn't on the list: &lt;strong&gt;accidental keyword hits&lt;/strong&gt; — chatter like "Can't do this anymore, going to bed lol" classified as a need because "can't" matches an error keyword. It only exists for tools that match characters, so the guide never had it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4 — the answer key.&lt;/strong&gt; Fifteen exam comments, each with an expected answer and the flag "can the data alone decide this?" A reply like "Me too 😭" with no parent comment is undecidable — its correct answer is "needs confirmation."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5 — grading.&lt;/strong&gt; Another place the guide didn't hold. The order exam's grader is bound to order-specific fields — product codes, quantities — and &lt;strong&gt;not one line of it could be reused.&lt;/strong&gt; What ported was the frame of thinking, the grade table and the principles; the 40-line grader was written from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scorecard
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;my comment classifier (regex):  15 cases · 8 clean (53%)
🔴 FATAL 3   🟠 RISKY 5   🟡 MISSED 1   🟢 HARMLESS 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three fatals were accidental keyword hits. Social commentary, chatter, and a personal anecdote — each promoted to a need.&lt;/p&gt;

&lt;p&gt;Our shipping rule has been one line from the start: FATAL 0 ships, anything else doesn't. This tool is at FATAL 3. &lt;strong&gt;And I had already published the numbers it produced.&lt;/strong&gt; The core of that post — the "I can't trust it" comments with 598 likes — is quoted verbatim, so it stands. But the percentages are only as trustworthy as this regex, and I'm writing that down honestly.&lt;/p&gt;

&lt;h2&gt;
  
  
  So — does it port?
&lt;/h2&gt;

&lt;p&gt;It ports. &lt;strong&gt;It took one hour.&lt;/strong&gt; But my body learned that the manual has four defects.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;① missing step        "understand the thing you're testing" isn't there — in
                      practice it ate half the hour
② grade table oversold "the other three lines are the same everywhere" — MISSED
                      turned fatal on this project. Re-argue the table every port
③ grader is domain-bound  what ports is the frame, not the code. Rewrite the grader
④ traps incomplete    four types are a start. Each tool type (regex, LLM) has
                      traps of its own
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That defect list is the real harvest of this experiment. The next person who steals this exam, knowing these four in advance, finishes in thirty minutes instead of an hour.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The exam ports. What ports is not code — it's the five steps and the grade table as a way of thinking&lt;/li&gt;
&lt;li&gt;While you're porting it, grade your own tools first. I found FATAL 3 in the tool behind my own published numbers&lt;/li&gt;
&lt;li&gt;Distrust any manual that says "the same everywhere." Mine did, and it had four defects&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;P.S. The 15-question comment exam and its grading results are public → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;, comment_exam.py&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next up: I gave the same exam to an LLM. The regex-vs-LLM decision was made by that scorecard.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>The AI Exam Author Was Never Wrong. I Still Can't Use Its Exam.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Wed, 26 Aug 2026 01:57:57 +0000</pubDate>
      <link>https://dev.to/ramses203/the-ai-exam-author-was-never-wrong-i-still-cant-use-its-exam-3h2</link>
      <guid>https://dev.to/ramses203/the-ai-exam-author-was-never-wrong-i-still-cant-use-its-exam-3h2</guid>
      <description>&lt;p&gt;In &lt;a href="https://dev.to/ramses203/i-gave-my-llm-an-exam-the-exam-author-lost-5-times-12b0"&gt;the first post of this series&lt;/a&gt; I wrote about making a 29-question exam and getting it wrong five times myself — three times in the answer key, twice in the grader. Ever since, one question kept nagging me: &lt;strong&gt;if an AI wrote the exam, how many times would it get it wrong?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So I counted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment — I ordered 50 questions from an author AI
&lt;/h2&gt;

&lt;p&gt;Hamel Husain's evals essay — the closest thing this field has to a textbook — describes how to scale up test questions: split the feature into scenarios, then mass-generate the input sentences with an AI. After &lt;a href="https://dev.to/ramses203/i-made-an-llm-re-grade-my-exam-it-found-two-bugs-in-my-grader-39bi"&gt;putting an LLM in the grading seat&lt;/a&gt;, this was the next piece to test for real.&lt;/p&gt;

&lt;p&gt;I gave the author AI (Sonnet 5) three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The 50-item product catalog — the exam's reference data&lt;/li&gt;
&lt;li&gt;A list of trap types — missing specs, missing units, confusable products, mid-message reversals&lt;/li&gt;
&lt;li&gt;An order to write the answer key alongside each question, including &lt;a href="https://dev.to/ramses203/steal-this-exam-heres-how-to-port-it-to-your-own-pipeline-36oo"&gt;the flag from the porting guide&lt;/a&gt;: "can the reference data alone decide this one?"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Five questions per request, ten requests, fifty questions.&lt;/p&gt;

&lt;p&gt;Review came in three layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;layer 1   code check      nonexistent product codes? flag contradicting the key?
layer 2   AI reviewer     per question: "is this answer justified by the data?"
layer 3   human (me)      full re-read of all 50
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Layer 3 has a story. The original design stopped at two layers — and only after running it did I notice: &lt;strong&gt;the author was Sonnet, and the reviewer was Sonnet.&lt;/strong&gt; I had written "never let a model grade its own answers" in the judge post, then made exactly that mistake. So I sat down and re-read all fifty myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result — zero mistakes
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;structural errors (fake codes, flag contradictions)   0
answer-key errors (all three review layers)           0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I, the human author, got the answer key wrong three times in 29 questions. The AI author got it wrong zero times in 50. Every fully-specified order had the right code and quantity. Every undecidable one went to "needs confirmation" with the correct candidates attached. All fifty flags matched their answers.&lt;/p&gt;

&lt;p&gt;The questions weren't bad, either. Take this one —&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Boss, please send 3 boxes of the post-office boxes~ same size as last time!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It planted the fake hint "same as last time" on its own. There's no order history in the data, so nothing can pin the size down — and its answer key says exactly that: needs confirmation, three post-office box candidates. Correct. It reinvented a trap from my original exam ("the usual, 3 boxes") &lt;strong&gt;without being told to.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the story ended here, the conclusion would be "let the AI write the exams."&lt;/p&gt;

&lt;h2&gt;
  
  
  The twist — I can't run this exam
&lt;/h2&gt;

&lt;p&gt;Nothing in the 50 questions is wrong. But lay them side by side and something else shows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, the composition is a photocopy of my instructions.&lt;/strong&gt; I had told it: "per 5 questions, mix 1 normal order, 1 non-order, 1 change/cancel, 2 traps." Ten requests later, the counts came out exactly 10 / 10 / 10 / 20. The 20 traps split into a tidy 10 missing-spec and 10 mid-message-reversal. A human author would have drifted — "shouldn't something like this go in too?" — and produced a few questions that escape the plan. There isn't one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, the materials repeat.&lt;/strong&gt; Counting the products in the confirmed answers, one item — shipping box A-type 300 — appears 12 times. I handed it fifty products; it kept reusing its favorites.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, the types I didn't order never appear.&lt;/strong&gt; My original 29 questions had ultra-abbreviations ("250 5"), typo-ridden messages, after-learning questions, custom-order inquiries. The 50 new questions contain none of these. The reason is simple: I didn't put them in the prompt.&lt;/p&gt;

&lt;p&gt;Put the three together: &lt;strong&gt;the AI author is never wrong. It only writes what you ordered.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Now look back at human me. I got the exam wrong five times while making it. But that same process was where the traps got invented — putting two tape widths in on purpose, making up a sentence like "250 5", imagining the accidents that only happen after the system has learned. The five mistakes were the tuition I paid while doing the inventing.&lt;/p&gt;

&lt;h2&gt;
  
  
  So the authoring job splits in two
&lt;/h2&gt;

&lt;p&gt;The porting guide says the accident list decides the question count. This experiment showed the other side of that sentence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;inventing accident types      human work.  it comes from living in the domain
stamping questions per type   AI work.     more accurate than me, never tires
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When I made 29 questions alone, those two jobs were one lump. Separated, the real bottleneck of exam-building isn't the question count — it's &lt;strong&gt;the length of your accident list.&lt;/strong&gt; Discover one new type, and the AI stamps ten questions of it in five minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The zero-error run may owe a lot to my clean 50-item catalog. Whether it holds on real-world dirty data — the same product registered three times under different names — is untested&lt;/li&gt;
&lt;li&gt;"Only what you ordered" can be partly relieved by writing more types into the prompt. But knowing &lt;em&gt;which&lt;/em&gt; types to add is exactly the part that stays human&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Delegate question mass-production to an AI. Even with answer keys included, it was more accurate than the human author&lt;/li&gt;
&lt;li&gt;But the reviewer must be a different model or a human. An author reviewing itself isn't a review&lt;/li&gt;
&lt;li&gt;Inventing trap types stays human work — and the mistakes you make while building the exam are that invention happening&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;P.S. The question-generation script (gen_cases.py) is in the repo with everything else → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>The Question That Picked My Model Didn't Survive Five Reruns</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Wed, 26 Aug 2026 01:56:12 +0000</pubDate>
      <link>https://dev.to/ramses203/the-question-that-picked-my-model-didnt-survive-five-reruns-5cfk</link>
      <guid>https://dev.to/ramses203/the-question-that-picked-my-model-didnt-survive-five-reruns-5cfk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Correction (Sep 2026):&lt;/strong&gt; The lid question was later found to be mis-keyed. Every "harmless" mark in the tables below was the model asking which neck, which the corrected key counts as correct. Re-scored: cheap model 143 clean / 0 harmless / 2 risky, expensive model 144 / 0 / 1, fatal still 0 for both. Details in &lt;a href="https://dev.to/ramses203/i-rewrote-one-exam-question-fifty-ways-first-the-answer-key-was-wrong-5dn7"&gt;this post&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Eleven hours after &lt;a href="https://dev.to/ramses203/the-model-that-costs-3x-more-won-by-exactly-one-question-55aj"&gt;the model-comparison post&lt;/a&gt; went up, a reader named Vinh Nguyen left a comment that unraveled its headline:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;FATAL 0 against FATAL 0 is a tie in counts, but it is 29 items run once each. [...] Zero fatal out of 29 single trials is consistent with a true fatal rate up to about 10% at 95% confidence. [...] Running the whole exam five times per model and comparing fatal rates per item would tell you whether the tie is real or just the resolution of one pass, and it is the same order of cost as the run you already did.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He's pointing at a self-contradiction I published without noticing. My own post says, in the temperature section: "final scores should come from multiple runs, reported as a pass rate." And then the comparison itself — the whole "won by exactly one question" headline — sat on one pass per model.&lt;/p&gt;

&lt;p&gt;No defense available. So I ran it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rerun
&lt;/h2&gt;

&lt;p&gt;29 questions × 5 fresh runs × 2 models = 290 trials. Same harness, same unpinned temperature as the original — that variance is the thing being measured. Every raw answer sheet saved. All 290 completed this time (the original expensive-model run had dropped one question to a rate limit).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             Haiku (cheap)        Sonnet (3x price)
trials       145                  145
FATAL        0                    0
RISKY        1                    1
MISSED       0                    0
HARMLESS     4                    5
clean        140/145              139/145
per run      29·28·27·28·28       28·28·28·28·27
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things in that table killed my original framing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question that picked the winner was a coin flip
&lt;/h2&gt;

&lt;p&gt;The original comparison came down to exactly one question — the lids. The customer writes: "10 boxes of the clear PET 300, and 5 bags of the white lids." The catalog has four lids to pick from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="err"&gt;CAP-300-W&lt;/span&gt;    &lt;span class="err"&gt;white&lt;/span&gt; &lt;span class="err"&gt;lid,&lt;/span&gt; &lt;span class="err"&gt;300-neck&lt;/span&gt;   &lt;span class="err"&gt;(fits&lt;/span&gt; &lt;span class="err"&gt;the&lt;/span&gt; &lt;span class="err"&gt;PET-300&lt;/span&gt; &lt;span class="err"&gt;bottle)&lt;/span&gt;
&lt;span class="err"&gt;CAP-300-BK&lt;/span&gt;   &lt;span class="err"&gt;black&lt;/span&gt; &lt;span class="err"&gt;lid,&lt;/span&gt; &lt;span class="err"&gt;300-neck&lt;/span&gt;
&lt;span class="err"&gt;CAP-500-W&lt;/span&gt;    &lt;span class="err"&gt;white&lt;/span&gt; &lt;span class="err"&gt;lid,&lt;/span&gt; &lt;span class="err"&gt;500-neck&lt;/span&gt;   &lt;span class="err"&gt;(fits&lt;/span&gt; &lt;span class="err"&gt;the&lt;/span&gt; &lt;span class="err"&gt;PET-500&lt;/span&gt; &lt;span class="err"&gt;bottle)&lt;/span&gt;
&lt;span class="err"&gt;CAP-500-BK&lt;/span&gt;   &lt;span class="err"&gt;black&lt;/span&gt; &lt;span class="err"&gt;lid,&lt;/span&gt; &lt;span class="err"&gt;500-neck&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;"White" narrows it to two — but which neck? The clue sits in the same sentence: the customer is ordering PET 300 bottles, and nobody pairs 300 bottles with 500 lids. The expensive model read that context and confirmed the 300 white lid — and I called that the win. The cheap model asked "which neck?" instead.&lt;/p&gt;

&lt;p&gt;Across five fresh runs, the expensive model &lt;strong&gt;never did that again.&lt;/strong&gt; Zero confirms in five passes — it asked "which lid?" every single time. The only model that produced the winning confirm even once was the cheap one, in one run of five.&lt;/p&gt;

&lt;p&gt;So the one-question gap my title stood on disappears the moment the exam is rerun. It was never a real ability gap between the two models. It was a coin flip that happened to land the expensive model's way on the one day I ran the exam once — and I published that single flip as the result.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two RISKY trials landed exactly where the flag said
&lt;/h2&gt;

&lt;p&gt;Each model produced one RISKY trial in 145 — the grade for "confirmed something ambiguous; right this time, fatal next time."&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Haiku, run 3, on "the usual, 3 boxes": two items in the order history both fit, and it presented only one of them as a candidate instead of both&lt;/li&gt;
&lt;li&gt;Sonnet, run 5, on the ultra-abbreviated "250 5": five products start with 250, and it confirmed one anyway&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both questions carry &lt;code&gt;rule_decidable: false&lt;/code&gt; — the flag &lt;a href="https://dev.to/ramses203/a-reader-caught-my-answer-key-drifting-toward-the-model-35ia"&gt;a reader talked me into adding&lt;/a&gt;: "can the reference data alone pin the answer down to exactly one?" On decidable questions, the run-to-run wobble drifted safe — extra clarifying questions, nothing worse. The dangerous direction — confirming what cannot be decided — showed up only on the flagged questions, at about 0.7% of trials per model.&lt;/p&gt;

&lt;p&gt;The flag didn't just audit my answer key. It predicted where run-to-run variance turns dangerous.&lt;/p&gt;

&lt;h2&gt;
  
  
  So does the original conclusion survive?
&lt;/h2&gt;

&lt;p&gt;The conclusion of that post — it's a tie on fatal errors, so use the cheap model — survives. What gets replaced is the evidence it stands on.&lt;/p&gt;

&lt;p&gt;The old evidence was one pass: I ran the exam once, fatal errors came out 0 to 0, and the whole visible difference was that one lid question. The lid question is out — it was luck. What takes its place: fatal errors stayed 0 to 0 through 145 trials per model, and RISKY came out tied too, 1 to 1. By the rule of three, 145 trials tighten the fatal-rate bound on this exam from about 10% (at 0/29) to about 2%. On questions the model hasn't seen, I still have only 29 kinds — so that ceiling stays near 10%. The tie is no longer one lucky afternoon — it is a measured result. And on a tie, take the cheap one: the conclusion stands.&lt;/p&gt;

&lt;p&gt;What doesn't stand is the margin I published. Clean counts wobbled between 27 and 29 per run for the cheap model and 27–28 for the expensive one. On any single pass, either model can rank first by a question or two. &lt;strong&gt;A one-pass margin of one question is below this exam's resolution&lt;/strong&gt; — which is what the comment said before I spent a cent measuring it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A single pass ranks models with ±1–2 questions of noise. Don't publish a one-pass margin — mine didn't survive its own rerun&lt;/li&gt;
&lt;li&gt;"Zero fatal" means little until you say over how many trials — and of what. 0/29 caps this exam's fatal rate near 10%; five reruns (0/145) cap it near 2%, on this exam. New questions are a separate count. The reruns cost about as much as the first run&lt;/li&gt;
&lt;li&gt;Variance has grades too. On decidable questions it drifted safe (over-asking). The dangerous drift — confirming the undecidable — appeared only on questions pre-flagged as rule-undecidable&lt;/li&gt;
&lt;li&gt;The exam caught its author again. Not in the answer key this time, and not in the grader — in how confidently I read a single run&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;P.S. All 290 raw answer sheets, the 5× driver, and the aggregator are in the repo → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;P.P.S. If you'd rather get new experiments by email: &lt;a href="https://ramses203.substack.com/subscribe" rel="noopener noreferrer"&gt;ramses203.substack.com/subscribe&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>Steal This Exam. Here's How to Port It to Your Own Pipeline.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Tue, 25 Aug 2026 00:27:34 +0000</pubDate>
      <link>https://dev.to/ramses203/steal-this-exam-heres-how-to-port-it-to-your-own-pipeline-36oo</link>
      <guid>https://dev.to/ramses203/steal-this-exam-heres-how-to-port-it-to-your-own-pipeline-36oo</guid>
      <description>&lt;p&gt;So far this series has been about &lt;a href="https://dev.to/ramses203/i-gave-my-llm-an-exam-the-exam-author-lost-5-times-12b0"&gt;giving my order-reading LLM an exam&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Some of you have been reading it thinking: "Mine isn't orders, it's meeting-minutes summarization." "I'm using it for email triage."&lt;/p&gt;

&lt;p&gt;Good news — the exam is built to be stolen. Only two things need swapping: &lt;strong&gt;the reference data you match against&lt;/strong&gt; (mine: a product catalog) and &lt;strong&gt;the worst accident&lt;/strong&gt; (mine: wrong goods loaded onto a truck). Let's go in order.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1 · Write down the worst first
&lt;/h2&gt;

&lt;p&gt;Before writing a single test question, write this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When this AI is wrong, which of the consequences cannot be undone?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For my order program it was "the wrong goods get loaded onto a truck." What is it for yours? A few examples —&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;auto-reply email AI    a wrong answer goes out to the customer — you can't unsend it
meeting-minutes AI     writes "decided" on something undecided — once shared, work proceeds on it
research AI            an invented number enters the report — reports travel upward
data-cleanup AI        overwrites the original — no backup, no recovery
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can see the pattern. &lt;strong&gt;Sent, deleted, escalated — these don't come back.&lt;/strong&gt; That's your truck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2 · Build the grade table
&lt;/h2&gt;

&lt;p&gt;Once the worst is written down, the grades fall out on their own. One criterion — can a human undo it?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FATAL      cannot be undone              sent, deleted, reported, charged
RISKY      confirmed something ambiguous  right this time, fatal next time
MISSED     dropped something              a human can still catch it
HARMLESS   over-asks "please confirm"     just slower
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The only line you have to define yourself is the first one. The other three read the same in any project.&lt;/p&gt;

&lt;p&gt;And take the principle with you as-is: &lt;strong&gt;a wrong confirmation is worse than no confirmation.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3 · Plant the traps
&lt;/h2&gt;

&lt;p&gt;Exam questions come from traps, not from normal cases. There are four kinds, and every domain has all four.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Confusable pairs&lt;/strong&gt; — mine was clear tape 48mm vs 60mm. In meeting minutes: two attendees named Kim, one an associate, one a manager. In email: reply vs forward. Find the pairs in your data that are similar enough to confuse, and deliberately put both in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Plausible non-targets&lt;/strong&gt; — "What are the specs on the 250 shipping box?" contains a product name but is not an order. In minutes: "let's decide that next time" (not a decision). In email: a promotional mail (not something to answer). Put in &lt;strong&gt;things that only look like what the AI is supposed to catch.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mid-message reversals&lt;/strong&gt; — "5 boxes please — no wait, make it 3." In minutes: "let's go with A → actually B is better." Questions that are wrong if you only read the first half.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Accidents after learning&lt;/strong&gt; — mandatory if your system has memory. A situation where fresh information must beat the remembered value ("tape 60"), and a situation where memory must not change the verdict (a question is still a question).&lt;/p&gt;

&lt;p&gt;How many questions? The metric isn't a count — it's your accident list from Step 1. &lt;strong&gt;At least one question per accident type, built to cause exactly that accident.&lt;/strong&gt; From my list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;inquiry mistaken for an order → goods nobody ordered ship   → "What are the specs on the 250 shipping box?"
change/cancel mistaken for a new order → ships twice        → "I ordered 5 boxes — please send only 3"
unit misread → quantity ships at 50x                        → "250 5"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Five accidents on your list means five questions minimum. Mine grew to 29 because I kept adding variants per accident; 10 is plenty to start.&lt;/p&gt;

&lt;p&gt;Add a few normal, well-behaved cases at the end — only a few. Models rarely fumble the easy ones, so normal cases mostly catch nothing. The exam's job is not to watch the model succeed. It's to find where the accidents are.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4 · Write the answer key — then doubt it
&lt;/h2&gt;

&lt;p&gt;For each question, write the expected answer. And here comes the lesson of &lt;a href="https://dev.to/ramses203/i-gave-my-llm-an-exam-the-exam-author-lost-5-times-12b0"&gt;the first post&lt;/a&gt;:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your answer key will be wrong.&lt;/strong&gt; I was wrong three times in the key and twice in the grader.&lt;/p&gt;

&lt;p&gt;So keep three rules.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The correct answer to an ambiguous question is "needs confirmation" — not a plausible-looking verdict&lt;/li&gt;
&lt;li&gt;When the model answers differently from your key, &lt;strong&gt;re-read the data before docking points.&lt;/strong&gt; The model may be right. But if your fix turns a "needs confirmation" into a confident verdict, be suspicious of the fix first&lt;/li&gt;
&lt;li&gt;For every question, record in advance: &lt;strong&gt;"does the reference data alone pin the answer down to exactly one?"&lt;/strong&gt; When the model and the key later disagree, that flag is the referee&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The third rule by example — "tape 60, 2 boxes": the catalog has a 60mm, so the data pins the answer. If the model answers differently here, the model is wrong; the key stands. "Clear tape, 2 boxes": no width given, the data cannot pin it. The correct answer was "needs confirmation" from the start — if I wrote anything else in the key, the key is what's broken.&lt;/p&gt;

&lt;p&gt;A reader of the first post talked me into this flag. Bonus effect: recording it forces you to open the reference data for every question while you're still writing the exam — which catches answer-key mistakes at authoring time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5 · Grade, save, decide
&lt;/h2&gt;

&lt;p&gt;Count results by grade, not by score. And save the model's raw answer sheets to files — grading criteria keep changing, and with the raw sheets you can re-grade without calling the model again.&lt;/p&gt;

&lt;p&gt;Everything passed? Then write the three-line ledger. It's you being honest with yourself about what the exam guarantees and what it doesn't.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;VERIFIED       one question per accident type, and FATAL held at zero
NOT VERIFIED   production data has never been through it — I invented every sentence myself
GUARDED        on anything uncertain it must not confirm — it hands off to a human
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What justifies shipping is not line one. It's &lt;strong&gt;lines one and three combined.&lt;/strong&gt; The exam score covers only what I could imagine (line two is that confession), and the sentences I couldn't imagine land on line three's guard. Without line three, an exam score is just a number — the first unimagined sentence causes the accident.&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole thing on one card
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Write down the accidents that can't be undone — that's your truck&lt;/li&gt;
&lt;li&gt;Build the grade table — only the first line is yours to define&lt;/li&gt;
&lt;li&gt;Plant the four traps — confusable pairs, plausible non-targets, mid-message reversals, after-learning&lt;/li&gt;
&lt;li&gt;Write the answer key, then doubt the answer key&lt;/li&gt;
&lt;li&gt;Grade by severity, save raw sheets, write the three-line ledger&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The exam you may steal. The answer key you have to write yourself.&lt;/p&gt;

&lt;p&gt;And that answer key will be wrong. Catching it is what the exam is for.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;P.S. The test runner and grader that run exactly this structure are public → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Scraped 20,000 YouTube Comments. The Videos and the Comments Were Having Two Different Conversations.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Tue, 25 Aug 2026 00:25:03 +0000</pubDate>
      <link>https://dev.to/ramses203/i-scraped-20000-youtube-comments-the-videos-and-the-comments-were-having-two-different-l30</link>
      <guid>https://dev.to/ramses203/i-scraped-20000-youtube-comments-the-videos-and-the-comments-were-having-two-different-l30</guid>
      <description>&lt;p&gt;I once collected about 22,000 comments from roughly 140 Korean YouTube videos about AI coding tools and classified them. (Quotes below are translated from Korean.)&lt;/p&gt;

&lt;p&gt;I wanted to see what people were asking. What came out was something else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the videos teach
&lt;/h2&gt;

&lt;p&gt;Put the titles and tags of those 140 videos in one pile and they say:&lt;/p&gt;

&lt;p&gt;How to install. How to get started. How to build an app. Which tool is best.&lt;/p&gt;

&lt;p&gt;All of it is "starting." Follow along, a result appears on the screen, the video ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the comments say
&lt;/h2&gt;

&lt;p&gt;The comments sweeping up the likes were telling a different story.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Verifying AI mistakes takes so much time. Checking every answer for nonsense got so tiring I just do the work myself now." (👍598)&lt;/p&gt;

&lt;p&gt;"Coding with AI makes me anxious. If one bug ships, I'm the one responsible. Checking and debugging everything one by one ends up being more work." (👍265)&lt;/p&gt;

&lt;p&gt;"I pay every month and it lies about work matters like it's nothing." (👍72)&lt;/p&gt;

&lt;p&gt;"Tokens burn too fast… added $50 and it was gone in half a day." (👍30)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It compresses into three complaints: &lt;strong&gt;expensive, can't trust it, can't fix it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The videos teach the start. The people are dying right after the start.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scariest comment
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Asked it for shampoo recommendations and it recommended one that doesn't exist. Slipped it in between real products — with the weight, the benefits, even a price." (👍49)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That comment is the essence of the problem. &lt;strong&gt;When AI is wrong, it doesn't look wrong.&lt;/strong&gt; The fake sits among the real ones, wearing plausible numbers.&lt;/p&gt;

&lt;p&gt;This is why "just write better prompts" is half an answer. Better prompts lower the odds of being wrong. They don't create &lt;strong&gt;a way to know when it's wrong.&lt;/strong&gt; Drop the error rate from 10% to 3% and you still don't know where the 3% is hiding. If that 3% detonates inside payment logic, money leaves the building.&lt;/p&gt;

&lt;h2&gt;
  
  
  One more finding — where the real questions live
&lt;/h2&gt;

&lt;p&gt;While collecting, I noticed the nature of comments changes with channel size.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;multi-million-sub videos   real questions/needs = 12% of comments — the rest is reactions and anxiety
10k–300k sub channels      real needs = 26% — "I followed along and got stuck RIGHT HERE"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It's the difference between spectators and people actually doing the thing. Comments under the big videos ask "what happens to us in the AI era." Comments under mid-size channels ask "how do I fix this error."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The real questions were in the mid-size channels.&lt;/strong&gt; And collecting reply threads, not just top-level comments, surfaced 470 more needs that the top level never showed.&lt;/p&gt;

&lt;h2&gt;
  
  
  So I picked a direction
&lt;/h2&gt;

&lt;p&gt;Content that teaches starting is already everywhere. I decided to work on the stretch right after the start — the can't-trust-it, can't-fix-it stretch.&lt;/p&gt;

&lt;p&gt;The earlier posts in this series are the first result of that decision: &lt;a href="https://dev.to/ramses203/i-gave-my-llm-an-exam-the-exam-author-lost-5-times-12b0"&gt;giving my LLM an exam&lt;/a&gt;, &lt;a href="https://dev.to/ramses203/a-good-llm-exam-is-90-traps-4faj"&gt;planting traps in it&lt;/a&gt;, &lt;a href="https://dev.to/ramses203/grade-your-llm-passfail-and-you-will-ship-a-disaster-1f19"&gt;grading by severity instead of pass/fail&lt;/a&gt;, &lt;a href="https://dev.to/ramses203/it-passed-every-test-thats-why-it-cant-ship-yet-3dnm"&gt;why passing still wasn't enough to ship&lt;/a&gt;. Every one of them is about turning "I can't trust it" into "I verified it."&lt;/p&gt;

&lt;p&gt;People don't die at the start. They die right after it.&lt;/p&gt;

&lt;p&gt;But almost everyone is selling the start.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;P.S. The classifier that produced the percentages above later sat an exam of its own — the same kind I give my LLMs — and failed three fatal-grade questions. That public correction is its own story, coming later in this series.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next up: how to steal this exam and port it to your own pipeline, step by step.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;P.P.S. If you'd rather get new experiments by email: &lt;a href="https://ramses203.substack.com/subscribe" rel="noopener noreferrer"&gt;ramses203.substack.com/subscribe&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>It Passed Every Test. That's Why It Can't Ship Yet.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Mon, 24 Aug 2026 00:32:20 +0000</pubDate>
      <link>https://dev.to/ramses203/it-passed-every-test-thats-why-it-cant-ship-yet-3dnm</link>
      <guid>https://dev.to/ramses203/it-passed-every-test-thats-why-it-cant-ship-yet-3dnm</guid>
      <description>&lt;p&gt;My order-reading LLM passed the 29-question exam. Zero fatal errors. &lt;a href="https://dev.to/ramses203/the-model-that-costs-3x-more-won-by-exactly-one-question-55aj"&gt;The model is chosen&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;So — ship it?&lt;/p&gt;

&lt;p&gt;No. And the reason is the most important thing in this series.&lt;/p&gt;

&lt;h2&gt;
  
  
  The exam has holes. I am the holes.
&lt;/h2&gt;

&lt;p&gt;The questions, the answer key, the product catalog — I made all of it.&lt;/p&gt;

&lt;p&gt;Which means "passed 29 questions" translates precisely to:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"It made no mistakes in the 29 situations I was able to imagine."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Nothing more. The sentence I couldn't imagine is not on the exam. And production is a parade of sentences you couldn't imagine. Real customers abbreviate in ways I can't invent, and real product catalogs are far dirtier than the one I wrote — the same product registered three times under different names, dead items nobody deleted piling up for years.&lt;/p&gt;

&lt;h2&gt;
  
  
  So the ledger has three lines
&lt;/h2&gt;

&lt;p&gt;Before shipping, I wrote these three lines down.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;VERIFIED       one question per worst-accident type; both models at zero
NOT VERIFIED   never ran on production data. Every exam sentence is mine
GUARDED        when unsure, it must not confirm — it hands off to a human
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The third line is the one that matters. &lt;strong&gt;If the unknown case falls somewhere safe by construction, an incomplete exam is still shippable.&lt;/strong&gt; When a never-seen phrasing arrives, this program's worst case is "slower" — not "wrong goods shipped."&lt;/p&gt;

&lt;p&gt;Without that guard, shipping on exam scores alone is trusting a gun because it fired 29 times without jamming.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first stretch of production is a watch period
&lt;/h2&gt;

&lt;p&gt;For a while after launch, nothing passes through automatically. A human reviews every result.&lt;/p&gt;

&lt;p&gt;What comes out of this period is the real exam. The sentences I couldn't imagine make their first appearance here. Every miss becomes a new exam question. &lt;strong&gt;This is the moment the exam grows from imagination into production.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Week 1      humans check everything. Every miss becomes a test case
Weeks 2–4   if FATAL holds at 0, auto-pass the confirmed ones only
After       humans only see the "needs confirmation" queue
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The exam also tells you when to fold
&lt;/h2&gt;

&lt;p&gt;There's an opposite case: FATAL keeps appearing, and the cause isn't the prompt — it's &lt;strong&gt;the question itself.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"Figure out what 'the usual' means, automatically" is that kind of question. It cannot be done in principle. The information isn't there. No amount of prompt polish fixes it, and every polish makes it worse by teaching the model to guess with confidence.&lt;/p&gt;

&lt;p&gt;The move there isn't fixing the program. It's &lt;strong&gt;shrinking the scope.&lt;/strong&gt; Draw the line — "that case goes to a human" — and automate the rest. Finding out what is impossible in principle is also the exam's job.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scoreboard so far
&lt;/h2&gt;

&lt;p&gt;Counting what this exam actually caught:&lt;/p&gt;

&lt;p&gt;The model's real mistakes — 1. And its only crime was asking.&lt;/p&gt;

&lt;p&gt;The exam author's mistakes — 5. Three in the answer key, two in the grader.&lt;/p&gt;

&lt;p&gt;The exam I built to verify the AI caught me more than it caught the AI. Which is the real reason to build one.&lt;/p&gt;

&lt;p&gt;Passing is the starting line. Production writes the next questions.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;P.S. The scoreboard above changed again later, when the grader got fixed one more time. That story is its own post — &lt;a href="https://dev.to/ramses203/i-made-an-llm-re-grade-my-exam-it-found-two-bugs-in-my-grader-39bi"&gt;it's here&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;All 29 questions and the grader are public → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>The Model That Costs 3x More Won by Exactly One Question</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Mon, 24 Aug 2026 00:29:58 +0000</pubDate>
      <link>https://dev.to/ramses203/the-model-that-costs-3x-more-won-by-exactly-one-question-55aj</link>
      <guid>https://dev.to/ramses203/the-model-that-costs-3x-more-won-by-exactly-one-question-55aj</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Correction (Sep 2026):&lt;/strong&gt; The lid question that decided this comparison was later found to be mis-keyed. The customer never stated the neck size, so asking was the correct answer, not confirming. Under the corrected key the cheap model's answer was right and the expensive model's was a guess. The re-scoring is in &lt;a href="https://dev.to/ramses203/i-rewrote-one-exam-question-fifty-ways-first-the-answer-key-was-wrong-5dn7"&gt;this post&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I gave the same 29-question order-reading exam to two models.&lt;/p&gt;

&lt;p&gt;A cheap one (Haiku 4.5) and one that costs about three times as much (Sonnet 5).&lt;/p&gt;

&lt;h2&gt;
  
  
  The result
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cheap model       28 of 29 questions clean
Expensive model   all 28 executed questions clean
Fatal errors      zero for both
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(One of the expensive model's questions never ran — rate limit.)&lt;/p&gt;

&lt;p&gt;The 3x model really was better. By exactly one question.&lt;/p&gt;

&lt;h2&gt;
  
  
  That one question
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;10 boxes of the clear PET 300, and 5 bags of the white lids&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The catalog has two kinds of lids: 300-neck and 500-neck.&lt;/p&gt;

&lt;p&gt;"White lids" alone doesn't tell you which. But the same sentence starts with "PET 300." Nobody orders 300-size bottles and 500-size lids in the same breath.&lt;/p&gt;

&lt;p&gt;The expensive model read that context and confirmed the 300 lid.&lt;/p&gt;

&lt;p&gt;The cheap model asked back: "diameter unspecified, needs confirmation."&lt;/p&gt;

&lt;p&gt;In the first post of this series, "needs confirmation" was the &lt;em&gt;correct&lt;/em&gt; answer for the clear tape — so why is confirming correct here? The difference is the clue. The tape sentence contained nothing to decide the width. This sentence has "PET 300" on the same line. &lt;strong&gt;If the message itself contains the information to decide, confirming is correct. If it doesn't, asking is correct.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And look at &lt;em&gt;how&lt;/em&gt; the cheap model was wrong. It didn't ship the wrong lid. &lt;strong&gt;It asked.&lt;/strong&gt; No accident happens. The user just gets mildly annoyed if every lid order comes back as a question. That's the size of the gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  So which did I pick?
&lt;/h2&gt;

&lt;p&gt;The cheap one.&lt;/p&gt;

&lt;p&gt;The criterion is severity, not score (&lt;a href="https://dev.to/ramses203/grade-your-llm-passfail-and-you-will-ship-a-disaster-1f19"&gt;earlier post&lt;/a&gt;). Both models at FATAL 0 is a tie on the grade that matters. &lt;strong&gt;On a tie, take the cheap one.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If it had gone the other way, there'd be nothing to think about — pay up. Say the scoreboard had looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cheap model       28/29  FATAL 1   ← higher score, unusable
Expensive model   27/29  FATAL 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Pick by "how many did it get right" and you pick the wrong model.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Bonus trap — same exam, different scores
&lt;/h2&gt;

&lt;p&gt;One thing kept nagging me.&lt;/p&gt;

&lt;p&gt;I gave the same question to the same model again, and the result changed. Wrong once, right on the retry, then right three more times. One miss in five runs.&lt;/p&gt;

&lt;p&gt;The cause is temperature — the knob that controls randomness in model output. Pin it to 0 and the same input gives the same answer. The problem: the CLI tool I was using had no such option.&lt;/p&gt;

&lt;p&gt;Two lessons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If you can't pin temperature, a single run's score is not reproducible.&lt;/strong&gt; Final scores should come from multiple runs, reported as a pass rate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production calls run at temperature 0.&lt;/strong&gt; Especially for work where the same order sheet must produce the same answer every day&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Compare by severity, not score&lt;/li&gt;
&lt;li&gt;FATAL 0 vs FATAL 0 is a tie. Take the cheap one&lt;/li&gt;
&lt;li&gt;Pin the temperature or don't trust the number&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Before you pay for the expensive model, count what the cheap one actually gets wrong. In our exam the entire gap was one question — and even that one was guilty of nothing but asking.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;P.S. Next up: the exam was passed — and that's exactly why it can't ship yet.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;All 29 questions and the grader are public → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>A Reader Caught My Answer Key Drifting Toward the Model</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Sat, 22 Aug 2026 06:04:44 +0000</pubDate>
      <link>https://dev.to/ramses203/a-reader-caught-my-answer-key-drifting-toward-the-model-35ia</link>
      <guid>https://dev.to/ramses203/a-reader-caught-my-answer-key-drifting-toward-the-model-35ia</guid>
      <description>&lt;p&gt;Three days after &lt;a href="https://dev.to/ramses203/i-gave-my-llm-an-exam-the-exam-author-lost-5-times-12b0"&gt;the first post&lt;/a&gt; of this series went up, a comment arrived on a Korean tech-news site where it had been shared. The first half was praise. The second half was this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;While the exam author was losing three times, every answer-key revision went in the direction of "the model was right." If that repeats, isn't there a risk the answer key converges on model behavior? I split questions by whether domain rules alone can decide them, independently. If they can't, I pull them out of scoring entirely — they become human-review cases, not test cases. How did you distinguish these?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Unpacked: &lt;strong&gt;"if you revise the answer key every time the AI answers differently, the exam stops evaluating the AI and becomes a sheet of paper that transcribes it."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Imagine a teacher who revises the answer key every time they look at a student's answers. A few times, the teacher really was wrong. But once it becomes a habit, that exam no longer measures the student. The scores keep coming out great — not because the student is good, but because the key keeps accommodating the student. And I use those scores to decide whether to ship.&lt;/p&gt;

&lt;p&gt;The comment had found the weak spot of the whole series.&lt;/p&gt;

&lt;h2&gt;
  
  
  My answer — auditing the three revisions
&lt;/h2&gt;

&lt;p&gt;I went back and checked my three answer-key fixes against one criterion. What did I trust when I changed the key — the plausible explanation the model attached to its answer, or the product catalog and the original message?&lt;/p&gt;

&lt;p&gt;All three had been fixed &lt;strong&gt;after opening the catalog and the message directly.&lt;/strong&gt; The clear tape came in two widths in the catalog and the customer never said which. The order "250 5" matched five products starting with 250, with no unit written anywhere. "The usual" had exactly one record in the history — and nothing guaranteed the pointer pointed at it. All three were questions whose answer the data alone could not pin down.&lt;/p&gt;

&lt;p&gt;And all three revisions moved in the same direction: &lt;strong&gt;from "confirmed" back to "needs confirmation."&lt;/strong&gt; I never once copied the model's chosen code into the key as the new answer.&lt;/p&gt;

&lt;p&gt;That yields a watch rule. If the key ever starts drifting toward the model, the revisions will run the other way — "needs confirmation" hardening into "confirmed," or code A being replaced by the model's code B. &lt;strong&gt;The first revision in that direction is the stop signal.&lt;/strong&gt; Freeze the change, hide the model's answer, and re-derive the answer from the data alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  And the admission — I had no criterion beforehand
&lt;/h2&gt;

&lt;p&gt;Honestly, though: those three judgments were good after the fact, not by prior design. The commenter's method — &lt;strong&gt;marking each question at authoring time by whether rules alone can decide it&lt;/strong&gt; — was better than mine.&lt;/p&gt;

&lt;p&gt;So I adopted it as-is. Every one of the 29 questions now carries a flag: &lt;code&gt;rule_decidable&lt;/code&gt;. True if the data pins down exactly one answer. False if the correct answer is "needs confirmation."&lt;/p&gt;

&lt;p&gt;Adding the flag came with a built-in audit: &lt;strong&gt;all three questions where I had fixed the answer key came out false.&lt;/strong&gt; All three revisions had been cases of "I confirmed an answer on a question that was never decidable in the first place" — confirmed once more, this time by a flag.&lt;/p&gt;

&lt;p&gt;There was a bonus I didn't expect. To set the flag you have to actually open the catalog for every question — decidability can't be judged without looking. So answer-key mistakes get caught at authoring time, not after publishing. If this flag had existed when I wrote the tape question, opening the catalog would have shown me two tape widths, and "48mm, confirmed" would never have been written into the key.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two referees now sit on the exam
&lt;/h2&gt;

&lt;p&gt;When the model and the answer key disagree, the procedure is now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;rule_decidable = true   →  the model is wrong. The key stands
rule_decidable = false  →  if the key isn't "needs confirmation," the key is at fault
the flag itself suspect →  reopen the data, re-judge the flag first
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No more improvising at every dispute. The flag set at authoring time referees first. The revision direction — confirmed→needs-confirmation is the only healthy direction — watches second.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Revising the answer key is not the sin. Revising it after checking the catalog and the original message is normal; revising it because the model's explanation was persuasive is where the rot starts&lt;/li&gt;
&lt;li&gt;You will not feel the moment you cross that line. So pick a visible signal in advance: any revision that turns "needs confirmation" into "confirmed" means stop and re-derive with the model's answer hidden&lt;/li&gt;
&lt;li&gt;The best defense is marking decidability at authoring time. Setting the mark forces you to open the data, and that act itself is the audit&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One comment changed the design of the exam. Publish, and the verification comes to you.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;P.S. The exam with &lt;code&gt;rule_decidable&lt;/code&gt; flags on all 29 questions is in the repo → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
    <item>
      <title>I Made an LLM Re-Grade My Exam. It Found Two Bugs in My Grader.</title>
      <dc:creator>John Green</dc:creator>
      <pubDate>Sat, 22 Aug 2026 05:59:50 +0000</pubDate>
      <link>https://dev.to/ramses203/i-made-an-llm-re-grade-my-exam-it-found-two-bugs-in-my-grader-39bi</link>
      <guid>https://dev.to/ramses203/i-made-an-llm-re-grade-my-exam-it-found-two-bugs-in-my-grader-39bi</guid>
      <description>&lt;p&gt;In &lt;a href="https://dev.to/ramses203/grade-your-llm-passfail-and-you-will-ship-a-disaster-1f19"&gt;an earlier post&lt;/a&gt; I wrote that my grader had been wrong twice — zeroing a perfect answer over truncated JSON, and penalizing a good answer. Both were caught by a human re-reading the answer sheets.&lt;/p&gt;

&lt;p&gt;It turns out the grader had been wrong two more times. This time the catcher wasn't a human. It was an LLM.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I put an LLM in the grading seat
&lt;/h2&gt;

&lt;p&gt;There's an essay that's become something of a textbook for AI verification: Hamel Husain's "Your AI Product Needs Evals." Reading it, the skeleton was almost identical to what I'd built — except for one step I didn't have: &lt;strong&gt;making an LLM do the grading.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When answer sheets grow to hundreds a day, no human can read them all. So you delegate grading to an LLM — which also gets things wrong. Hamel's prescription: &lt;strong&gt;before you trust the judge, make the judge take an exam of its own.&lt;/strong&gt; Have your existing grader and the LLM judge grade the same answers independently, and measure how often they agree.&lt;/p&gt;

&lt;p&gt;I have a code grader instead of a human, so my version was: &lt;strong&gt;give the same 29 answer sheets to both graders.&lt;/strong&gt; One is the grading rulebook implemented as code. The other is an LLM (Sonnet 5) that read the same rulebook in plain language.&lt;/p&gt;

&lt;p&gt;Picking the judge, I followed two of Hamel's rules. The judge should be a stronger model than the examinee (the answers were written by Haiku). And never let a model grade its own answers — models are known to be generous to themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result — 93% agreement, and two splits
&lt;/h2&gt;

&lt;p&gt;On 27 of 29 sheets the two graders agreed, down to the exact severity grade. If it had ended there: "LLM judges work, neat."&lt;/p&gt;

&lt;p&gt;The two splits were the problem. I dug in, and &lt;strong&gt;both times the LLM was right and my code was wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First — the lids.&lt;/strong&gt; The question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;10 boxes of the clear PET 300, and 5 bags of the white lids&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The same sentence says "PET 300," so the lids are the 300-size. Answer key: confirm the 300 lid.&lt;/p&gt;

&lt;p&gt;The model's answer didn't confirm. "300 or 500? Please check" — with both candidates listed.&lt;/p&gt;

&lt;p&gt;The two graders split on this sheet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Code grader — &lt;strong&gt;two penalties.&lt;/strong&gt; The correct lid code wasn't in the confirmed list, so it counted MISSED, and the extra confirmation request counted as over-asking on top&lt;/li&gt;
&lt;li&gt;LLM judge — &lt;strong&gt;one HARMLESS.&lt;/strong&gt; "Not missed — caught and asked. No accident here. Just over-caution"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Re-reading my own grade table, the LLM was right. My definition of MISSED is "dropped an order; the customer calls." Nobody calls about this answer — they just get asked "which lid?" I had even described this exact answer as "guilty of asking" in an earlier post of this series. The prose understood it. Only the code insisted it was "missed."&lt;/p&gt;

&lt;p&gt;Why: the code mechanically checked whether the correct code appeared in the confirmed list. An item that's being asked about isn't in that list, so — "not there, missed." The code had no concept of "currently being asked."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second — the bubble wrap.&lt;/strong&gt; The question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Add 2 rolls of bubble wrap to my earlier order&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not a new order — a request to append to an earlier one. Answer key: classify as an addition request, done.&lt;/p&gt;

&lt;p&gt;The model classified it correctly. Then it did one more thing: it left a note in the margin of the answer sheet — "FYI, bubble wrap comes in 30cm, 50cm and 100cm; which one needs confirming."&lt;/p&gt;

&lt;p&gt;The split:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Code grader — &lt;strong&gt;clean.&lt;/strong&gt; But for the wrong reason. My grader only reads two boxes on the answer sheet: "orders" and "custom orders." The note lived outside both. So this "clean" wasn't "no defects" — it was &lt;strong&gt;"didn't look where the defect was"&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;LLM judge — &lt;strong&gt;HARMLESS.&lt;/strong&gt; It read the whole sheet, saw the margin note, and wrote "one confirmation request that isn't in the answer key; not an accident"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's the difference between an OMR machine and a teacher. The OMR machine reads the bubbles; a teacher reads the doodles in the margin. My grader was an OMR machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verdict, and the fix
&lt;/h2&gt;

&lt;p&gt;When two graders disagree, who decides? The person who wrote the rulebook — me. Both graders are interpretations of my rules, so when the interpretations split, the author has to check the original intent.&lt;/p&gt;

&lt;p&gt;I re-read both sheets. Both times the LLM judge matched the rulebook's intent. My grader's lifetime error count went from two to four.&lt;/p&gt;

&lt;p&gt;Two fixes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Items that were caught-and-asked count as HARMLESS, not MISSED&lt;/li&gt;
&lt;li&gt;Confirmation requests outside the order box get read too&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I re-graded all 29 sheets with the fixed grader and diffed against the judge again. This time &lt;strong&gt;all 29 sheets matched&lt;/strong&gt; — not just clean/not-clean, but the exact severity of every defect.&lt;/p&gt;

&lt;h2&gt;
  
  
  The twist — my published score changed
&lt;/h2&gt;

&lt;p&gt;With the fixed grader, the score I'd published in &lt;a href="https://dev.to/ramses203/i-gave-my-llm-an-exam-the-exam-author-lost-5-times-12b0"&gt;the first post&lt;/a&gt; changed. The model didn't retake the exam. The answers are identical. Only the score moved.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             old grader              new grader
Haiku        28/29 clean             27/29 clean
             1 missed · 1 harmless   0 missed · 2 harmless
FATAL        0                       0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One fewer clean sheet — not because the model got worse, but because &lt;strong&gt;the grader got more accurate.&lt;/strong&gt; A defect that used to be invisible (the margin note) became visible. And MISSED dropped to zero: it turns out Haiku never missed anything. It was just over-careful.&lt;/p&gt;

&lt;p&gt;FATAL is still 0, so the ship/no-ship call didn't change. But the lesson stuck: &lt;strong&gt;a score is a function of the grader before it's a result of the exam.&lt;/strong&gt; Publish the grader version with the score, and accept that a better grader can change old scores.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three honest caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;26 of my 29 sheets were clean answers. When most of the sample is normal, agreement rates come out high on their own (a trap Hamel warns about too). The genuinely hard judgments are the defective few — only 3 sheets here&lt;/li&gt;
&lt;li&gt;One judge call died on a timeout and had to be re-run. If you run an LLM judge in production, failure handling is mandatory&lt;/li&gt;
&lt;li&gt;This was an exam where the judge got the rulebook handed to it. Grading without a rulebook is a different, much harder problem&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;You can delegate grading to an LLM — after the judge takes an exam of its own&lt;/li&gt;
&lt;li&gt;Code graders read fixed boxes; LLM judges read the whole sheet. Stack them and they cover each other's blind spots&lt;/li&gt;
&lt;li&gt;Publish the grader version alongside the score. Better graders change old scores.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;P.S. The judge (judge.py) is in the repo with everything else → &lt;a href="https://github.com/ramses203/llm-test-harness" rel="noopener noreferrer"&gt;github.com/ramses203/llm-test-harness&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
