This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Last week I built dapraia, a small tool that tells me the night before whether tomorrow is a beach day. You describe how you like it in one sentence, and a model on my laptop turns that sentence into rules a forecast can be checked against: wind under 15 km/h, start after 6, low tide.
The worst bug I hit there was not a crash. "Kitesurf com vento entre 15 e 25 nós" (wind between 15 and 25 knots) came back as 15 to 25 km/h. The rule looked right, it passed my first check, and it would have picked days with half the wind the person asked for. A wrong number that looks right is the error nobody reports, because nobody notices it.
So this benchmark measures exactly that: when a model turns a sentence into rules, does it keep the number, the unit and the direction the person wrote?
There are 40 sentences, 30 in Portuguese and 10 in English, in seven groups:
- speed units: knots, mph and m/s that have to become km/h
- other units: feet, Fahrenheit, "1 metro e meio" (a metre and a half)
- direction: "nunca acima de 15", "evito vento acima de 30", "20 km/h ou mais", where a maximum has to stay a maximum
- time of day: "7 e meia", "5 da tarde", "12 da noite", "after 5:30pm"
- decoys: numbers that are not limits ("vôlei 2x2 perto do Posto 6", "corrida com 4 amigos") and a bare "até 12" with no unit
- durations: "40 minutos", "uma hora e meia depois da maré baixa"
- plain: the easy ones, as a floor
Each sentence has a hand-checked answer key: the field, the value and a tolerance (1 km/h or 1 °C after a conversion, 0.05 otherwise). A sentence scores only when every expected rule is there with the right value. Every miss gets a label: unit_or_value (right field, wrong number), missing, flipped (the value landed in the opposite field, so a maximum became a minimum) or decoy (a number that was never a limit became a rule). The score is the share of sentences that came back fully right.
The benchmark has two tasks with the same 40 sentences. In keep_the_number_you_wrote the prompt lists only the field names, so the unit lives in the name (wind_max_kmh, water_min_c). In keep_the_number_units_spelled_out the prompt says what each field means, its unit and its range, and adds one line: convert any other unit into the field's unit. The pair asks how much of the reading is the model and how much is the instruction.
Models Tested
Eight, each for a reason:
- Gemma 4 26B A4B and Gemma 4 31B, the closest Kaggle has to the Gemma 4 12B that runs dapraia on my laptop.
- Gemini 3.5 Flash-Lite, GPT-5.4 nano, Claude Haiku 4.5 and gpt-oss-20b, the small, cheap models a tool like this would actually call every night.
- Gemini 3.7 Flash and Claude Sonnet 5.5, as a reference for what a stronger model does with the same words.
I also tried qwen3-next-80b, which answered 24 of the 40 in thirty minutes, and gpt-oss-120b, which timed out in the same run. I left both out rather than publish partial rows.
Findings
| Model | Units only in the field name | Units spelled out |
|---|---|---|
| Claude Haiku 4.5 | 100% | 100% |
| Gemma 4 31B | 100% | 100% |
| Gemma 4 26B A4B | 100% | 100% |
| Gemini 3.7 Flash | 100% | 100% |
| Claude Sonnet 5.5 | 100% | 100% |
| Gemini 3.5 Flash-Lite | 97.5% | 100% |
| gpt-oss-20b | 95.0% | 97.5% |
| GPT-5.4 nano | 95.0% | 97.5% |
1. With the units spelled out, the task is close to solved. Six of the eight models got all 40 right, and the other two missed one sentence each. The instruction that closed the gap is short: the unit of each field, the range of the clock (0 to 24, so 7.5 is 7:30), and "convert any other unit into the field's unit".
2. Without it, the misses are units and clock words. The leaderboard shows one score per model, so the detail comes from my notebook runs of the same 40 sentences the same afternoon (Claude Haiku 4.5 was not in those). Every miss on the bare prompt:
- Gemma 4 26B read "entre 15 e 25 nós" as 15 and 25 km/h, and "gusts under 20 kt" as 20 km/h. It is the same error my local Gemma 12B made in dapraia.
- gpt-oss-20b turned 82 °F into 25.6 °C (it is 27.8) and "até as 12 da noite" (until midnight) into hour 0. The bare prompt never says the clock runs to 24, so midnight as 0 is a fair reading that still breaks the rule: a window that ends at 0 never opens.
- GPT-5.4 nano kept "rajadas até 25 nós" as 25 km/h, dropped "1 metro e meio", and turned "vento até 12 no fim de semana" (wind up to 12 on the weekend) into a 999 km/h wind limit plus a midnight-to-noon window nobody asked for.
In those runs no model flipped a maximum into a minimum, and none turned "2x2", "Posto 6" or "4 amigos" into a rule. I did see a flip in dapraia, but my own code caused it: a phrase list misread "evito vento acima de 30 km/h" (I avoid wind above 30), told Gemma its correct maximum was wrong, and Gemma obeyed. Given only the sentence, every model here kept the direction.
3. One run cannot rank the top. Gemma 4 26B scored 38 of 40 on the bare prompt in my notebook and 40 of 40 on the leaderboard run later the same afternoon. Flash-Lite went the other way, 40 and then 39. At 40 sentences and one run, a one-sentence gap between two models is noise. What repeated was the kind of sentence that failed: knots, Fahrenheit, "metro e meio" and midnight. With the units spelled out, gpt-oss-20b and nano missed one sentence in both runs. In my notebook, nano's was "vento de no máximo 5 m/s" (wind of at most 5 m/s), which came back as a 5 metre wave limit.
What surprised me is how cheap the fix is compared with what I built. In dapraia I wrote checks that read the unit after every number and reject a rule whose quote says something else. The benchmark says a few lines of instruction put six of the eight models at 40 of 40 and the other two at 39. It also says why the check still has to stay: the runs vary, and the one run that reads knots as km/h is the one that picks the wrong day. The instruction moves the average; the check catches the run that still goes wrong.
What I would measure next:
- five runs per model on each task, to report a spread instead of one number
- harder sentences: comma decimals ("1,5 m"), the Beaufort scale, "meia-noite", two units in one phrase
- the check itself: feed every miss above to dapraia's quote check and count how many it stops before a person ever sees the rule
My Benchmark
The benchmark, with both tasks and the leaderboard: Keep the number you wrote
The two tasks, each with the 40 sentences, the answer key, the grader and its prompt in the code:
- keep_the_number_you_wrote, units only in the field name
- keep_the_number_units_spelled_out, units spelled out
The answer key is plain Python in the task code, so a new sentence is one line.
Top comments (0)