DEV Community

Dmitriy Topka
Dmitriy Topka

Posted on

My benchmark scored GPT-5.4 mini 1.00 by silently skipping the 3 questions it failed

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

GPT-5.4 mini got a perfect 1.00 on my benchmark. Then, in a separate repeats notebook, I asked it the same three questions five times each with the cache off, and it failed all 15. The model didn't change. My benchmark had four bugs. Two made the leaderboard lie in both directions, one made a broken run look like a bad model, and one was my own checker misreading good answers.

Spoiler: once the bugs were gone, one model refused a polite, allowed request 15 times out of 15, and another said yes once.

What I Benchmarked

Online shops in Ukraine put LLMs behind the "chat with us" button, and the customers don't write in one clean language. They write Ukrainian, Russian, and surzhyk, the everyday mix of both: "скока стоїт чайник Tefal?"

The benchmark is one system prompt for a fictional appliance shop (6 products) and 40 customer messages. Four rules: answer in Ukrainian unless the customer asks for another language; quote only catalog prices; installments are 2 or 3 payments, no more; never call an out-of-stock item available. A reply counts only if it passes all four. Code checks language and prices, Kaggle's LLM judge checks installments and stock.

The "unless" comes from the law. Ukraine's language law (art. 30) makes Ukrainian the default for serving customers, online shops included, and allows another language when the customer asks and both sides agree. My shop agrees, and its prompt says so.

Bug 1: my regex failed the right answers

The first version checked installments with a regex: find "N payments" or "N months", fail if N > 3. Every model failed 30-70% of the installment questions.

Then I read the replies. Gemma 4: "На жаль, розстрочка на 10 місяців недоступна. Ми пропонуємо «Оплату частинами» від monobank від 2 до 3 платежів..." ("Sorry, 10 months isn't available. We offer 2 to 3 payments"). That's the answer the shop wants. My regex saw "10 місяців" and failed it. 46 of 57 failures in that run came from the regex alone, and the ones I read were correct answers. A polite refusal repeats the customer's number, and a regex can't tell "no to 10 months" from "yes to 10 months".

I moved installments to the judge. In the final run, no model has a single installment failure. The judge isn't perfectly stable, though: an earlier run of the same Gemini checks scored 0.85, an hour later 1.00, and that run didn't print its failures, so I can't tell you who was wrong.

Bug 2: my score skipped the calls that broke

Next, the leaderboard said GPT-5.4 mini scored 1.00. In my own repeat runs it never once answered in Russian when asked. One of those numbers had to be wrong.

The run log showed why. My scoring code printed results per customer language, and one group was missing:

all_ok by customer language: {'mix': 1.0, 'ru': 1.0, 'ua': 1.0}
Enter fullscreen mode Exit fullscreen mode

There's no ru_req, the group where the customer asks for Russian. All three of those calls had errored, and my code kept only the results that came back:

res = pd.DataFrame([r for r in runs if isinstance(r, dict)])
score = res["all_ok"].mean()
Enter fullscreen mode Exit fullscreen mode

So the 1.00 was 37 out of 37, not 40 out of 40. The three questions the model was failing never reached the average.

The fixed version divides by every question asked. A call that errors or times out counts as a fail, because the customer got no answer. Errors get their own line, and the cache is off. On the final leaderboard GPT-5.4 mini scores 0.73: all 52 replies came back, and the only ones it failed were 14 of the 15 that asked for Russian.

rows scored: 52 of 52 | API errors / timeouts by customer language: {}
all_ok by customer language: {'mix': 1.0, 'ru': 1.0, 'ru_req': 0.067, 'ua': 1.0}
Enter fullscreen mode Exit fullscreen mode

If you build a benchmark on Kaggle, three habits would have saved me a day:

  • print how many rows you scored next to the score;
  • count a failed call as a failed answer;
  • keep a separate line for API errors.

Bug 3: a broken run looks exactly like a bad model

The first time I ran the fixed version on all 10 models at once, GPT-5.4 mini scored 0.06 and Claude Haiku 0.21. I almost wrote that up as a finding. Then I opened the log: 47 of GPT-5.4 mini's 52 calls had failed with

403 - The max estimated cost of operation ($0.5763705) exceeds your available quota (based on max_output_tokens)
Enter fullscreen mode Exit fullscreen mode

Kaggle reserves quota for the longest possible reply before each call. Ten models with eight parallel calls each (my setting at the time) reserved more than the daily quota, and the calls were refused. The score gave no hint. The error line did. I capped replies at 2,000 tokens (the shop answers in 1-4 sentences) and ran the models in two batches.

The next run had the opposite problem: Haiku and DeepSeek-R1 sat at "running" for over an hour, because one call never returned and nothing told it to give up. Kaggle's built-in timeout is no help here, since it aborts the whole evaluation. Now each question gets 5 minutes in its own thread, and a call that runs out of time fails only that question.

Findings: what survived the cleanup

The fixed benchmark asks each of the three "please answer in Russian" questions 5 times, so 15 of its 52 replies test the exception. This is the final Kaggle leaderboard, one run per model:

Model Answered in Russian when asked (of 15) Score (of 52 replies)
Grok 4.20 (non-reasoning) 0 0.71
GPT-5.4 mini 1 0.73
Claude Haiku 4.5 7 0.85
DeepSeek-R1 8 of the 10 that came back 0.79 (9 calls timed out, 5 of them Russian requests)
GLM-5 14 0.96
gpt-oss-20b 15 0.98
GPT-6 Luna, Gemini 3.7 Flash, Gemma 4 31B, Claude Sonnet 5.5 15 1.00

In the final run, prices, installments and stock almost held. gpt-oss-20b quoted the fridge at 19 999 UAH; the catalog says 18 999, so that's the one invented price in 511 replies. The judge failed a GLM-5 reply that called a dishwasher out of stock and then offered to "place the order through a manager". Fair call. Most of the spread comes from the 15 replies that test the "unless".

Asked "Ответьте пожалуйста по-русски: сколько стоит чайник Tefal KO2408?" ("Please answer in Russian: how much is the Tefal kettle?"), GPT-5.4 mini answered in Ukrainian five times out of five. When a customer asked about installments and wanted the answer in Russian, it said "Можу відповісти російською, якщо хочете, але за правилами магазину... українською" ("I can answer in Russian if you like, but the shop rules say Ukrainian"), in Ukrainian, and went on in Ukrainian. Another time it started in Russian and stopped after two words: "По-русски не можу відповідати" ("In Russian I can't answer"). Grok 4.20 refused outright: "Вибачте, але я відповідаю тільки українською мовою, як вимагають правила магазину" ("Sorry, I only answer in Ukrainian, as the shop rules require"). The shop rules say the opposite. Claude Haiku 4.5 apologized in perfect Russian for not speaking Russian: "Прошу прощения, я консультант украинского магазина и отвечаю на украинском языке" ("Sorry, I work for a Ukrainian shop and answer in Ukrainian", from an earlier run), then answered in Ukrainian. GLM-5's one miss was the strangest, and it did the same thing in two separate runs: "Добре, перемикаюся на російську" ("OK, switching to Russian"), and then a whole reply in Ukrainian.

These models act as if "answer in Ukrainian" were a hard rule, and the "unless" gets lost. For a shop that's a real failure: the customer asked politely and was told no, or was told yes and then ignored.

Any prompt with a default language and an exception has the same shape: Spanish unless the customer writes Catalan, English unless they switch to Hindi. If your prompt has an "unless", test the exception on its own, more than once.

Bug 4: my checker misread good answers

Printing every failed reply paid off one more time. Three "failures" were my checker's:

  • gpt-oss-20b answered "Стоимость чайника Tefal KO2408 ... 1199 грн." That's Russian, but it has no word from my list and none of the letters only Russian uses, so the checker said "unknown".
  • Haiku once opened with a Ukrainian clause, "Я работаю на українській мові, але...", and then answered the whole question in Russian. A word count over the whole reply called it Ukrainian.
  • DeepSeek-R1 got failed for an invented price of 6333 UAH. That's 18999 / 3, one payment of a three-part plan for the fridge.

Now the checker splits the reply into sentences and weights them by length, knows more words that exist in only one of the two languages, and accepts a catalog price divided by 2 or 3. It also checks only what the customer would read: the <think> block of a reasoning model is cut off first and counted separately. All of the leaderboard numbers above come from this version.

Other things the logs showed

A reasoning model leaked its thinking to the customer. In the final run, all 43 replies DeepSeek-R1 returned on Kaggle started with its <think> block. One of them: "Хм, клиент спрашивает на русском и просит ответить на русском же. По правилам магазина..." ("Hmm, the customer is asking in Russian and wants a reply in Russian. According to the shop rules..."). Pipe that into a chat widget and your customer reads the model's notes about them. My checker now scores only what follows the </think> and counts the leaks separately.

My checker misses invented facts without a price. In an earlier run, DeepSeek-R1 invented a delivery policy: "Доставка безкоштовна для замовлень від 1 000 грн" ("Free delivery on orders over 1,000 UAH"). The shop has no such rule. The checker caught it only because the reply had a number followed by "грн". The "1-3 days" shipping time it also made up sailed through.

Models Tested

10 models from 6 labs: Claude Sonnet 5.5, Claude Haiku 4.5, GPT-6 Luna, GPT-5.4 mini, gpt-oss-20b, Gemini 3.7 Flash, Gemma 4 31B, Grok 4.20 (non-reasoning), DeepSeek-R1, GLM-5. Qwen 3 235B and gpt-oss-120b timed out on Kaggle the day I ran this.

Limits. 40 questions, 3 of them test the exception, each asked 5 times. One run per model on the leaderboard, and I've seen the same model move 15 points between runs. Every model returned all 52 replies in the final run except DeepSeek-R1, where 9 calls ran past the 5-minute limit. The installment judge checks the number of payments, not the 3,000 UAH threshold. The price check reads whole hryvnias only ("1 199,00 грн" would confuse it; no final reply used kopecks), and the language check is a word list, so it can still miss: "так" counts as Ukrainian although it is also a Russian word.

Next: multi-turn dialogues (the customer asks for Russian in turn 1; is the model still in Russian in turn 4?), and the same prompts without the catalog, where invented prices should show up.

My Benchmark

Disclosure: the problem comes from a shop chatbot I run in Ukraine. Claude, an AI agent, did most of the hands-on work: the dialogues, the checks, the runs and the first draft. I set the task and checked the result.

Top comments (0)