This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Companies ship Italian that no Italian would write. In my own QA audits of software websites this year I found a transcription tool whose "Tools" menu said Utensileria (a hardware store), a homepage bragging about 5 Mld+ Altoparlanti nativi (five billion native loudspeakers), and more than one Fidato da, a word-for-word "Trusted by" that reads as machine output.
Models are now the reviewers many teams rely on, so I wanted to know: shown a string on its own, the way a reviewer sees it, can a model tell a shipped defect from its native fix?
The benchmark has 29 real pairs: the string as it shipped, and the version a native speaker would ship. Brand names are masked. The model sees each half separately, with the product type, where the text appears and the English source when there is one, and answers ok or error.
The scoring is the part I care about. A pair counts only when the model flags the shipped string and leaves the fix alone. A model that flags everything scores zero, and so does one that approves everything. Reviewers that cry wolf are useless in practice, and a plain accuracy number would hide them.
The defects cover ten types: 8 untranslated leftovers, 5 calques, 4 agreement errors, 3 number formats, 2 each of wrong word sense, misspelling and broken grammar, and one each of a broken encoding, an unidiomatic phrase and English-style capitalisation.
Models Tested
I ran the task on the 18 models Kaggle Benchmarks offered me, from Google, Anthropic, OpenAI, xAI, Alibaba, DeepSeek, Zhipu and the open Gemma and gpt-oss families, big and small, because a localisation team picks a reviewer by cost as much as by quality.
Not all of them produced answers. Seven returned API errors on most calls (permission denied, rate limits, one model not found), so I report them as not measured rather than as zero. I ran the task twice, and a run counts only when at most 2 of its 58 answers were unreadable. That leaves nine models with clean runs.
Findings
Pairs fully right, bugs caught and fixes left alone, out of 29. Where a model has two clean runs, both are shown.
| Model | Pairs fully right | Bugs caught | Fixes left alone |
|---|---|---|---|
| Gemini 3.8 Flash | 26 · 25 | 27 · 26 | 28 · 27 |
| Gemini 3.7 Flash | 26 · 23 | 27 · 27 | 28 · 25 |
| Gemini 3.5 Flash-Lite | 26 · 25 | 26 · 27 | 29 · 27 |
| Gemma 4 31B | 23 · 24 | 25 · 27 | 26 · 26 |
| Gemini 3.1 Pro (preview) | 23 · 24 | 27 · 27 | 25 · 26 |
| Gemini 2.5 Flash | 21 | 24 | 26 |
| GPT-5.4 mini | 19 · 19 | 24 · 24 | 23 · 23 |
| GLM-5 | 19 | 26 | 22 |
| GPT-5.4 nano | 18 · 16 | 26 · 26 | 20 · 19 |
Catching the bug is the easy half. Every clean model caught 24 or more of the 29 shipped defects. What separated them was the other half: leaving a correct native string alone. GPT-5.4 nano catches as many bugs as the leaders and then flags a third of the native fixes too. A reviewer like that sends a translator to "fix" good Italian, which is worse than no reviewer.
Numbers are the blind spot. Across the 16 clean runs, models caught every untranslated leftover, calque and broken encoding, and 63 of 64 agreement errors. They caught the three number-format errors only 18 times in 48. All three are the English "+" glued to a number: 6+ piattaforme, 15.000+ computer, and oltre 100+ siti, which says "more than" twice. Italian writes oltre 6. The strings look like numbers, so models wave them through.
The small model tied the big ones. Gemini 3.5 Flash-Lite, the cheapest Gemini on the list, left all 29 native fixes alone in one run and tied the top score. Gemini 3.1 Pro did not beat the Flash models.
One run is not a result. The same model moved by up to three pairs between runs (Gemini 3.7 Flash: 26, then 23, and 25 in my first notebook run). On a 29-pair benchmark, a one- or two-pair gap between models is noise; I would not rank models on it.
What I'd measure next: more pairs per defect type, especially numbers and idiom, so each type gets its own reliable score, and the same test with a style guide in the prompt, to see whether rules fix the number blind spot or only make models flag more.
My Benchmark
Benchmark on Kaggle: https://www.kaggle.com/benchmarks/giuseppecastelluccio/italian-ui-strings-shipped-vs-fixed
Task (public): https://www.kaggle.com/benchmarks/tasks/giuseppecastelluccio/italian-ui-strings-shipped-vs-fixed
How this was made
The strings and their corrections come from my QA audits, made before the challenge; I checked every correction again as a native speaker for this benchmark. The task code was written during the challenge with an AI coding assistant (Claude Code): I set the scoring rule and the checks, it drafted the code, and I reviewed and ran it. A stub model that always answers right scores 1.0 and one that flags everything scores 0.0.
Top comments (0)