Two readers said the same thing. Under the 3x-price comparison, Vinh said the next experiment should not be another rerun. It should take the lid question, the one that decided that comparison, and rewrite it in different forms. Under the rerun post, tonal said that five passes over the same 29 questions are not 145 fresh chances. Both are right. Rerunning the same exam adds runs, not kinds of question.
So I decided to rewrite the lid question fifty ways. Then, rereading the question, I found that its answer key was wrong. That story comes first.
The answer key was wrong
Here is the original.
페트 300 투명 10박스 뚜껑 흰거 5봉이요 (10 boxes of clear PET 300, 5 bags of the white lids)
The catalog has two white lids, a 300-neck and a 500-neck. The customer did not say which. The answer key said the correct answer was to confirm the 300-neck lid. The reason was that the same message orders 300 bottles, so the lids must be for those bottles.
But every other question on the same exam grades this kind of case the same way: the correct answer is to ask.
clear tape, 2 boxes 48mm or 60mm? key: ask
PET 500, 8 boxes clear or brown? key: ask
250 5 five candidates key: ask
white lids, 5 bags 300 or 500? key: confirm 300 ← the only exception
The rule is one line. Do not guess what the customer did not write. Ask. Only the lid question broke it. "There are 300 bottles right next to it, so obviously 300" looked so obvious that when I fixed the tape questions earlier, I never looked at this one again.
So I fixed it. Asking which neck is the correct answer. Confirming the 300 lid is RISKY: right this time, but a guess, so wrong next time. Confirming any other lid is FATAL: the wrong goods ship.
The scores of two earlier posts change
When one question's answer changes, every run that touched it gets re-scored. The other questions stay as they were.
There are four grades. Clean means no flaw. Harmless means asking about something that did not need asking. Risky means confirming by guess something that should have been asked, right this time but possibly wrong next time. Fatal means the wrong goods ship.
clean harmless risky fatal clean harmless risky fatal
─── before ───────────────────── → ─── after re-key ───────────────
3x-price post (one pass, 29 questions)
cheap model 28 1 0 0 → 29 0 0 0
expensive model (28 executed) 28 0 0 0 → 27 0 1 0
Five reruns (145 trials per model)
cheap model 140 4 1 0 → 143 0 2 0
expensive model 139 5 1 0 → 144 0 1 0
Every changed cell comes from the one lid question. The runs where a model asked which neck moved from harmless to clean. The one run where the cheap model confirmed 300 moved from clean to risky. Fatal was 0 before and stays 0, because no run ever confirmed the wrong lid.
In the 3x-price post I wrote that the expensive model got exactly one more question right. That question was this one. The expensive model guessed 300, and the cheap model asked. Under the corrected key, the cheap model won. The conclusion does not change: both models have zero fatal errors, and on a tie in fatal errors you take the cheap one.
Fifty forms
With the key fixed, I rewrote the lid question fifty ways. The order and the trap stayed the same. Only the wording changed.
T1 word order white lids 5 bags and clear PET 300 10 boxes
T2 endings ...please send / ...placing an order / no ending at all
T3 abbreviation PET300 clear 10bx lids white 5bags
T4 spacing PET300clear10boxeslidswhite5bags
T5 typos PTE 300 / lidss / whiet
T6 color words white / white-color / the white ones / "the clear ones" for the bottle
T7 units cartons / box / BOX / ten (as a word)
T8 spec notation 300ml / 300cc / PET 300 / PET bottle (300)
T9 sentence shape two lines / "and also..." / numbered list
T10 noise greetings / thanks / ORDER!!
Three things stayed fixed in every variant. The bottle always has "300" and "clear"; without them the bottle itself becomes ambiguous and it is a different question. The lid uses only a white-word and never a neck size; that is the trap. The quantities are 10 boxes and 5 bags, with unit words. I cut two variants that dropped the unit words, because the exam grades unit-less orders like "250 5" as ask, so dropping the unit changes the answer itself.
An LLM drafted the variants, and I checked each line against those three rules.
Grading follows the corrected key.
asks which neck (with candidates) clean
confirms CAP-300-W RISKY (a guess that was right this time)
confirms any other lid FATAL
Each variant ran three times per model. Temperature stayed at the default.
Results
cheap (Haiku 4.5) expensive (Sonnet 5)
trials 150 150
asked which neck 140 (93%) 148 (99%)
confirmed 300 by guess (RISKY) 10 ( 7%) 2 ( 1%)
confirmed another lid (FATAL) 0 0
lost the lid line 0 0
bottle confirmed PET-300-CL x10 149 (asked once) 150
Three things show.
First, item matching did not break. Typos, collapsed spacing, other words for white, 300ml and 300cc: none of it moved either model. In 300 trials there were zero wrong lids and zero wrong quantities. The four noise groups (spacing, typos, color words, spec notation) produced zero guesses in 120 trials per model. I expected the accidents to come from noise. They did not.
Second, the guesses came from smooth sentences. The cheap model guessed ten times: four in the unit group, two in word order, two in sentence shape, one in endings, one in greetings. Each variant ran three times, and two variants made it guess in two of the three: "10 boxes of clear PET 300, and the lids in white, 5 bags please" and "10 boxes of clear PET 300 please. And also 5 bags of the white lids." Both are the most polite and complete sentences in the set. The more a message looked like a finished order, the more the cheap model finished it, filling in the neck size the customer never gave. The expensive model guessed twice, once on "10 boxes of clear PET 300, 5 bags of white lids, please send" and once on "PET300 clear 10 boxes lids white 5 bags".
Third, no variant made a model guess every time. Eight variants made the cheap model guess at least once, and none made it guess all three times. Guessing is not a property attached to a wording. It is a coin flipped on every run, and some wordings raise the odds. Had I run this once, I might have seen two guesses or five, depending on the day.
If I had not fixed the key, this experiment would have read the other way. The old key counted confirming 300 as correct, so the cheap model's ten guesses would count as ten correct answers, and the cheap model would win 10 to 2.
One objection I expect
"Why not just show the customer an order confirmation sheet?" Yes. Everything that needs asking is better gathered on one screen. But the guessed line has to be marked. If the sheet only says "white lids, 300-neck, 5 bags", a customer who never thought about neck sizes just taps confirm. It has to say "white lids, 5 bags: assumed 300-neck, tell us if 500" for the customer's eye to stop there. The sheet copies what the intake step produced. Intake has to mark the lid as "needs confirmation" for the sheet to carry a warning. If intake quietly settles on 300 without a mark, the sheet prints "300-neck, 5 bags" like any other line, and the customer passes it by.
Takeaways
- The answer that looked most obvious on the key was the wrong one. I did not find it by thinking harder. Lined up next to the other 28 questions, this was the only one whose answer was a guess.
- In this experiment the cheap model guessed on polite, finished sentences, not on messy ones. It filled in missing information where the sentence looked complete. This came from one question and two models, so I do not know whether it holds elsewhere.
P.S. All 300 raw answer sheets, the 50 variants, the run script, and the aggregation script are in the repo → github.com/ramses203/llm-test-harness. The corrected key is in the same commit, with the reason written in the question's note.
P.P.S. If you'd rather get new experiments by email: ramses203.substack.com/subscribe
Top comments (0)