This is a submission for the Kaggle Benchmarking Challenge.
A Dutch reader asks where, in the back office, they set the colors of their app's interface, role by role:
Kunt u me het menupad in de back-office geven waarmee ik het kleurenpalet van de interface per rol instel?
The model gets the English edition of the article that answers it. Fourteen models on Kaggle got that exact message, and all fourteen gave the same path, thirteen of them in a Dutch sentence:
Het menupad is My App > App Style > Colors.
On a Dutch screen, that menu reads Mijn App > App-Stijl > Kleuren. The fact is right, the language is right, and the menu is the one from another edition. Dutch is close enough here to find your way. A Portuguese reader handed the German edition of another article gets less help: ten of the fourteen models told them, in Portuguese, to choose „Mit KI erstellen", the German edition's label.
I co-founded GoodBarber, an app platform, and I run its engineering. Our blog comes in seven languages, and much of it names a back-office menu or an app's button. A retrieve-then-answer loop over these pages doesn't always return one in the reader's language. I wanted to know what happens to the label then.
MTM-Bench already measures how the content's language leaks into replies, and another entry in this challenge checks Italian UI strings. This one scores the label the reader acts on, against their own edition, with controls where English is right.
TL;DR
- The models copy the label from the page: 58.8% of 5,266 answers on localized labels named only another edition's label, all but two the page's own. Only 1.9% came back in another language.
- English pages are the worst case: 0.14 on average, against 0.22 to 0.35 for the other page languages.
- GPT-6 Astra (0.689), Claude Opus 5 (0.592), GPT-5.5 (0.499) and the open-weight Qwen3 235B (0.483, $0.10 a run) clearly beat the always-English baseline (0.32). The other ten score 0.286 to 0.405.
- One sentence in the system prompt, asking for the reader's labels, takes two Gemini Flash models from about 0.37 to 0.77 and 0.86, past the best model without it.
What I Benchmarked
The material. Eleven items from ten articles of our blog, each something a reader acts on: four back-office paths or options, five strings an app shows, and two controls whose English name is right in every edition (AI Extension Builder, a Lost & Found board).
| Item | English edition | French edition | German edition |
|---|---|---|---|
| Colors by role | My App > App Style > Colors | Mon app > Style de l'app > Couleurs | Meine App > App-Stil > Farben |
| Start a new game | Play again | Rejouer | Nochmal |
| Podcast episode already started | 12 min left | Il reste 12 min | noch 12 Min. |
I checked every back-office label on screenshots of the real back office in the seven languages; the label on the reader's screen also gets full credit.
The minimal pair. The reader asks in their language, Q. A simulated retrieval step puts two blog sections in the message, in language C: the one that holds the answer and one on a nearby subject. Everything else is fixed; only C changes. Seven reader languages times seven page languages make 49 cells per item. In the seven cells where C equals Q, the page already holds the reader's label: these same-language cells check the setup and aren't scored. That leaves 462 scored cells per model.
The system prompt says nothing about the language of the answer:
You are the reader help assistant for a software company's blog. Answer the reader's question using only the retrieved documents provided in the message. If the documents do not answer the question, say so. End your answer with the id of the document you relied on.
"Using only the retrieved documents" bounds the facts, not the wording: the models already write their answer in the reader's language, and a label they translate themselves gets full credit, so nothing beyond the page is needed to score 1.
The scoring. No model judges the answers. The scorer is code, with 39 tests, and three gates:
- The fact is there: a label from any edition, or the item's key concepts. Otherwise 0.
- The answer is in the reader's language, once quotes, URLs and ids are stripped. Otherwise 0.
- The label:
| Outcome | What the answer contains | Score |
|---|---|---|
| Reader's label | the reader's label, no other edition's | 1 |
| Both | the reader's label and another edition's | 0.5 |
| Own translation | its own translation, no other edition's label | 1 |
| Other edition's label | another edition's label, not the reader's | 0 |
| Control kept | the English name, on a control | 1 |
| Control translated | a translated name, on a control | 0 |
Pasting a label from an edition the reader doesn't see costs points, which is why both labels score half: the reader still has to work out which one is theirs.
The second row loses only the controls: translating everything doesn't score 1 either.Baselines, run through the same scorer
Baseline
Score
Oracle (the reader's edition, word for word)
1
The reader's label everywhere, controls translated too
0.84
Always answer with the English label
0.32
Copy the label from the retrieved page
0.19
Vague answer ("go to the settings")
0
Refusal
0
Models Tested
Fourteen models, one run each on Kaggle's model proxy, from September 30 to October 2. The lineup is drawn from the roster Kaggle offers for benchmarks, not from what we run in our products: a large and a small model from each vendor, the latest generation where I could, plus open-weight models, to see whether size or vendor changes what happens to the label. Reasoning was set to low, with at most 4,096 output tokens and four attempts per cell. A run that loses more than 10% of its cells to proxy errors posts no score. Three aren't measured: gpt-oss-120b lost too many cells to rate limits in three runs at different hours, and every request to Gemma 4 26B and Grok 4.6 came back with an error (400 and 404).
Findings
| Model | Score | Reader's label | Both | Own translation | Other edition's label only | Wrong language | Run cost |
|---|---|---|---|---|---|---|---|
| GPT-6 Astra | 0.689 | 146 | 91 | 43 | 98 | 0 | $3.82 |
| Claude Opus 5 | 0.592 | 80 | 193 | 13 | 92 | 0 | $5.00 |
| GPT-5.5 | 0.499 | 89 | 77 | 19 | 192 | 1 | $2.35 |
| Qwen3 235B | 0.483* | 73 | 30 | 42 | 201 | 2 | $0.10 |
| Claude Sonnet 5 | 0.405 | 44 | 96 | 12 | 222 | 4 | $1.49 |
| Claude Haiku 4.5 | 0.379 | 30 | 102 | 11 | 205 | 29 | $1.36 |
| Gemini 3.7 Flash | 0.377 | 9 | 140 | 11 | 218 | 0 | $0.44 |
| Gemini 3.8 Flash | 0.374 | 11 | 132 | 12 | 223 | 0 | $0.31 |
| Gemini 3.1 Pro | 0.345 | 20 | 81 | 15 | 261 | 0 | $3.42 |
| GLM-5 | 0.337 | 17 | 77 | 17 | 265 | 2 | $1.30 |
| DeepSeek-R1 | 0.330 | 37 | 73 | 7 | 223 | 37 | $1.71 |
| GPT-5.4 nano | 0.317 | 23 | 41 | 21 | 278 | 14 | $0.11 |
| Gemini 3.1 Flash-Lite | 0.303 | 24 | 32 | 17 | 303 | 2 | $0.21 |
| GPT-5.4 mini | 0.286 | 29 | 8 | 15 | 315 | 7 | $0.45 |
The score covers the 462 scored cells; the counts, the 378 cells of the nine localized items. *Qwen3 235B: 431 scored cells, 352 localized.
The models copy the label from the page. On the localized items, 3,096 of the 5,266 answers (58.8%) named only another edition's label, and 3,094 of those named the page's own. Copying is right in two places, and the models get both: on the same-language cells and the controls, every model but DeepSeek-R1 scores at least 0.97. Everywhere else it's wrong, yet the most common outcome. The language itself rarely leaks: 1.9% of answers came back in another one. DeepSeek-R1 has a different problem: it answered in English 58 times across its 539 cells, 10 of them where the page was already in the reader's language.
English pages are the worst case. With an English page, the localized items average 0.14, against 0.22 to 0.35 for the other languages, and 78% of answers named only the English label. Even GPT-6 Astra drops to 0.278, from 0.50 or more elsewhere. Asked by a French reader about a game's results screen, with the English page, all fourteen models named Play again; two added (Rejouer).
The leaders get there differently. GPT-6 Astra most often gives the reader's label alone (146 cells) or its own translation (43). Claude Opus 5 most often gives both (193), which scores half. Tied for third, GPT-5.5 and Qwen3 235B copy nearly as often as Claude Sonnet 5 (51% and 57% of localized cells, against 59%) but commit to one label far more often (108 and 115 cells, against 56).
12 min left is the one label the models often translate themselves: 29% of its answers, when no other item goes above 5%. It's also the only label with a number in it, a lead from a single item. It scores 0.592, Play again 0.077.
One sentence in the prompt changes most of it. I reran the grid on three models with one more sentence in the system prompt: "Name menus, buttons and other interface labels as the reader's own screen shows them, in the reader's language, not as the documents write them."
| Model | Without | With | Other edition's label only |
|---|---|---|---|
| Gemini 3.7 Flash | 0.377 | 0.855 | 218 → 37 |
| Gemini 3.8 Flash | 0.374 | 0.768 | 223 → 63 |
| GPT-5.4 nano | 0.317 | 0.405 | 278 → 235 |
With it, both Flash models beat GPT-6 Astra without it. Nano gains much less, and 3.8 Flash translated 4 of the 84 controls.
My Benchmark
- Kaggle task: wrong-screen-answers
- Dataset: wrong-screen-answers-data (frozen sections, questions, labels, scorer and prompt builder; blog text © GoodBarber, reproduced with permission for evaluation)
The task builds every prompt from the dataset, scorer included, so a new model gets the exact bytes these did.
What it doesn't measure:
- Run-to-run variation. A second run moved Gemini 3.7 Flash by 0.004 and GPT-5.4 nano by 0.010, so gaps of a few hundredths are ties.
- The instruction on bigger models. It ran on three small ones; both of Qwen3 235B's runs lost too many cells to rate limits.
- The scorer's blind spots. A glossed « encore 12 min » / « nog 12 min » counts as a leak; an invented hybrid, "Com KI criar", as an own translation.
- The retrieval. The sections are chosen in advance, the answer comes first, with neutral ids. This measures what a model does with the context it's given, not a retrieval stack.
What I take from it. The label in the answer comes from the page, not from the reader, and the best model still gives only the page's label in a quarter of its answers. So in a retrieve-then-answer loop over a multilingual corpus, the first thing I'd check is which edition comes back, before choosing which model reads it.
How this was made
I built this with Claude Code over two days: harness, scorer, tests, Kaggle task and analysis. I chose the angle and reviewed the questions. The questions were drafted with Claude and back-translated by GPT-6 Astra, from another family, since Claude models are in the lineup.

Top comments (0)