Here is a subtotal from a Japanese invoice rendered at 300 dpi, and the answer azure/gpt-5.6-sol@low gave for it. The printed value is ¥1,237,500, in the same body type this series has been testing all along. The answer was ¥1,235,000 — and it was ¥1,235,000 on 118 of the model's 120 reads, including every material where the field sits in plain sight.
That number appears on no line of the document. It is not a misread — no digit is wrong the way OCR gets digits wrong. It is the invoice's total, ¥1,358,500, divided by 1.1, exactly. The wrong answer has a name, and the name is a formula.
In the last article I wrote that the 5.5/5.6-generation @low variants read exactly four things on this invoice: the title and the three boxed money figures. A follow-up run has since reached back into that sentence — its postscript already says so — and this article is the follow-up. Two of the four fields were never read.
Give the wrong answer a name
The stamp benchmark ended on an honest gap. Models were answering a destroyed subtotal correctly, and two arithmetic routes could explain it — total − tax, or total ÷ 1.1 — but the tax was a round 10%, so both routes produced the same number and the value couldn't tell them apart. The article went as far as the evidence did: the materials for subtraction were present; nobody watched them used.
The v2 material changes that without touching the harness. The tax is now a mixed 8%/10% — a per-rate breakdown box on the page, 適格請求書 style — while the printed subtotal keeps the exact v1 value, so the occlusion ladder stays geometrically identical. And the arithmetic is built so that total ÷ 1.1 = 1,235,000, an exact integer that appears on no printed line. Every flat-10% algebra collapses onto it: total − total/11 gives the same number. Subtraction of the printed tax, or summing the printed per-rate bases, still gives the true 1,237,500.
That is the design principle of this whole series, finally stated in both directions. Fictional ground truth makes a right answer prove reading. v2 adds the complement: make each wrong route produce an answer you can name. Unguessable truths, nameable wrongs. The same run swapped in a fictional seal text and a fictional branch name — v1's other two stated gaps — and both sections are below.
What the fingerprint caught
Deep coverage first, since that was the original question. With the subtotal's pixels destroyed under the opaque seal, 71 of 270 answers across the catalog are exactly 1,235,000, and the senders are the seven 5.5/5.6 @low variants, at nine and ten out of ten. The route is settled: not subtraction. The flat-10% prior.
But the fingerprint's real catch is at the other end of the ladder. At L0 — zero occlusion, the subtotal in plain sight — the same seven variants answer 1,235,000 in 69 of their 70 reads.
And the tax field, which no stamp touches on any material in the benchmark, comes back as 123,500 — total/11, the flat-10% shadow — 817 times: at every level, on every instance, including the ones where the seal is on the far side of the page, sitting on the issuer's name. Two of the seven variants emit the pair on literally all 120 of their reads. Per model, the subtotal count and the tax count match almost exactly — the two manufactured numbers travel together, in the same responses.
I can find only one reading of this. These models are not falling back to arithmetic when a field is destroyed. They never read the field. They read the total — the one number in a 16 pt box — and manufacture the subtotal and the tax outward from it, on every document, every time. The last article's "four fields" were two fields and two shadows: the title, and the total.
And v1's scores were the collision. On v1's invoice the tax was a round 10%, so the manufactured subtotal equaled the printed subtotal, the manufactured tax equaled the printed tax, and the seven variants banked 272 "correct" money reads out of 280 that the scorer had no way to doubt. The last article has a section about a branch-name cell that lied by being right — six cells, caught only by co-occurrence evidence — and calls it the edge of the fictional-ground-truth doctrine. The money columns were the same failure, spread across two full columns of the table, hidden by an arithmetic coincidence instead of a common phrase.
The seal, read
The v1 seal read 検収済印 — a real, common inspection stamp — so the one model that could read it couldn't be separated from a model remembering it. The v2 seal reads 納検済印, a phrase that does not exist: web-zero, and therefore training-zero.
claude-fable-5 read it 111 times out of 120 — 93%, against 94% on the real phrase in v1. Question closed: it was reading. And it is not a family trait — claude-sonnet-5 read it 0 times in 120, claude-opus-4-8 three. Whatever fable-5 does with a red square, it does alone.
The rest of the catalog answered a question I hadn't thought to ask. The most common wrong answer — 966 times, 30% of every stamp read in the benchmark — is 済納印検: all four characters, correctly recognized, in the wrong order. The seal's traditional layout reads right column first, top to bottom. The models walk it in horizontal rows, left to right — the default order of modern text. They see the glyphs; they don't know the traversal. The trap I actually set — a model completing from memory should emit the real phrase 検収済印 — fired nine times in 3,240. The criminal I caught was not memory. It was reading order, and I know of no document benchmark that tests it. (Another 190 answers were some version of the issuer's company name: the model answering the wrong box entirely.)
The branch, cured
The v1 branch was 本店営業部 — the most common branch name in Japan — and six @low cells scored "correct" on it while fabricating the bank in the same response. The v2 branch is 月芝支店, which exists at no bank (芝支店 does; 月隈支店 does; this one doesn't). The six cells did not survive the change: zero "correct" branch cells remain. They were coin-flips, and now the coin is gone.
What else moved in six days
- The gate structure got more binary, not less: 124 of 130 model×field cells are exactly 0 or exactly 20, up from 113 — the money columns joined the zeros.
- The invention rate on unread fields rose from 84% to 86%, the literal account number 1234567 from 88 to 114 of 140 reads, and the new generation's blank ratio from one blank per 5.2 inventions to one per 6.3. The confident author got more confident.
- The year shift survived a date change. The document now says 2026-10-30; the most-invented dates are 2025-10-30 and 2025-11-30 — month and day preserved, year decremented. "Last year" is a transformation, not a memorized date; the fabrication pattern from the start of this series holds on fresh input.
-
gemini-3.5-flash@lowremains the counter-example, intact: reads all thirteen fields, blanks at destruction, zero manufactured numbers, 16 credits a page.
One honest confound, and the autumn leg
The mixed rate required printing a per-rate breakdown, and that box is itself a derivation path v1 didn't offer: the two tax-exclusive bases sum to the subtotal. Stated plainly: my "read or derived" class cannot distinguish reading the subtotal from summing the printed bases from subtracting the printed tax — all three produce the true value. That ambiguity doesn't touch the @low verdict (a tier that cannot read 10.5 pt body text cannot read the 9 pt box either, and its answer isn't the true value anyway). But it does confound the strong models' movement at deep coverage: two of the three Anthropic models went from blanking thirty times out of thirty in v1 to deriving half their deep answers, and the OpenAI 5.6 @high variants to seven-to-ten out of ten. Maybe the box unlocked them; maybe behavior moved in six days. A v3 would A/B the box. And one scope line: the flat-10% verdict is proven on the v2 document; its reach back into v1's scores is an inference — same models, same week, same layout family, but an inference.
The harness is frozen and runs again in autumn on whatever generation ships by then, unchanged to the byte. It carries four questions. Whether the new cheap tiers still step on 1,235,000. Whether anything still reads the seal once fable-5 leaves the catalog. Whether the traversal trap closes. And whether a 2027 document gets 2026 written into it — whether "last year" is really always last year.
Blank, fiction, or arithmetic
The last article's closing word for this tier was authored. This run adds the mechanism: authored around one anchor. The model reads a single large number and writes an internally consistent document outward from it — names from the category's statistics, dates from last year, arithmetic from a formula — and it all reconciles because reconciliation was the generating rule, not the check.
The subtotal on this invoice was never occluded. It was never read, either.
Method notes: leg 1 of the v2 benchmark — 3,240 jobs run 2026-07-18 on the same 27-variant catalog as the stamp benchmark's run six days earlier. Materials, ground truth, scorer, the route classifier, and the recorded results are public at ldxhub-io/examples › analyzedoc/hanko-benchmark-v2. The route values are derived from the ground truth, never hardcoded — fed v1's output, the classifier reports the routes as inseparable, which is the point. Provider vision pipelines change — re-run before trusting any of this for anything current.

Top comments (0)