Update: Day 1 Kaggle Benchmarking Challenge
The local ladder is done, and the first hosted batch is in. The frontier models haven't run yet, so none of the three predictions can be graded. This is where the numbers stand, and what I had to fix to get numbers I'd trust.
How to read this. Every rate below is measured: it's computed from the raw replies these runs recorded, with its Wilson 95% interval. Anything I infer from those rates is marked as an inference. Following the rule from the comments, each prediction gets HIT, MISS or UNRESOLVED, graded on the interval, never on the point estimate.
The local ladder (8 models, 200 items each, temperature 0, on this laptop)
False confidence means answering anyway when the right reply was ESCALATE. Each model saw 40 unanswerable items.
| Model | Size (Ollama manifest) | Task score | False confidence | 95% interval |
|---|---|---|---|---|
| llama3.2:1b | 1.2B | 21.9% | 87.5% | 73.9–94.5% |
| llama3.2:3b | 3.2B | 69.4% | 80.0% | 65.2–89.5% |
| phi4-mini | 3.8B | 86.2% | 92.5% | 80.1–97.4% |
| qwen3:4b | 4.0B | 65.6% | 87.5% | 73.9–94.5% |
| gemma3:4b | 4.3B | 82.5% | 80.0% | 65.2–89.5% |
| gemma4:e2b | 4.6B | 81.9% | 95.0% | 83.5–98.6% |
| llama3.1:8b | 8.0B | 88.1% | 77.5% | 62.5–87.7% |
| qwen3.5 | 9.7B | 86.9% | 37.5% | 24.2–53.0% |
Only the largest local model, qwen3.5 at 9.7B, escalates on a large share of what it can't answer. For the rest, the point estimates run from 77.5% to 95%. Even the most generous lower bound among them, 62.5%, means answering anyway well over half the time. Size alone doesn't explain it, though: llama3.1 at 8B is no better than the 3B models.
Two of the "4b" tags are slightly over 4B by Ollama's own count (gemma3:4b is 4.3B, gemma4:e2b is 4.6B). For prediction 2, "4B or under" means the manifest size, so those two don't count.
Hosted, batch 1 (7 models, plus Kaggle's default model)
| Model | Task score | False confidence | 95% interval | Brier |
|---|---|---|---|---|
| gemini-3.7-flash (default) | 97.5% | 0.0% | 0–8.8% | 0.020 |
| gemini-3.8-flash | 95.5% | 0.0% | 0–8.8% | 0.025 |
| qwen3-235b-a22b | 91.7% | 10.0% | 4.0–23.1% | 0.104 |
| gemma-4-26b | 87.5% | 2.6% | 0.5–13.5% | 0.032 |
| claude-haiku-4.5 | 86.6% | 35.7% | 20.7–54.2% | 0.165 |
| gpt-oss-20b | 86.2% | 7.5% | 2.6–19.9% | 0.144 |
| gpt-5.4-nano | 84.4% | 7.5% | 2.6–19.9% | 0.153 |
| deepseek-r1 | 65.0% | 0.0% | 0–11.4% | 0.047 |
One commenter bet the frontier models would be the bluffers. On this batch it went the other way. Every hosted model's false-confidence interval sits entirely below every local model's except qwen3.5's: the highest hosted upper bound is haiku's 54.2%, and the lowest local lower bound is 62.5%. Haiku is the outlier among the hosted ones, though. Its lower bound clears 20%, just barely (20.7%). It's a small model, though, not one of the frontier models P1 is about, so it isn't counted there.
That comparison is an inference across two different setups. The local and hosted models didn't run through the same client, and the reasoning settings differ (see caveats). I haven't tested whether the gap holds once those are matched.
Where the predictions stand
- Some frontier model's false-confidence lower bound clears 20%: UNRESOLVED. No frontier model has run yet.
- The best local model at 4B or under beats at least one frontier model, by exact McNemar: UNRESOLVED. The test can't run without frontier results. On point estimates it leans toward a miss. The best local model at 4B or under by manifest size is llama3.2:3b, at 80%. Every hosted model so far except haiku is at 10% or less. But that's a lean, not a grade.
- Task score and false confidence are only weakly related (bootstrap Spearman): UNRESOLVED. The local arm alone gives ρ = −0.49, with an interval from −0.96 to 0.44. That's the "can't tell either way" shape the comments said to expect from about 8 models.
What broke on the way
Three smoke rounds before the paid run each turned up a harness bug that would have read as model behaviour. Each one below was observed in the raw replies, not inferred:
-
A loose answer format came back empty. Route and classify allowed "any object" as the answer. Gemini's structured output returned
{}for it, which scored 0% on two task shapes. Typing the answer fixed it, and the same model then scored 97.8% and 100% on those shapes. -
Every provider has its own rules. OpenAI's reasoning models reject temperature 0 and expect
max_completion_tokens. Its strict mode refuses any open object. Anthropic refused the 20-tool route format with "compiled grammar is too large", so haiku's 60 route items in this batch are errors, not answers. That's still open before the frontier run. - Rate limits look like model errors. In one smoke run, 169 of DeepSeek's 200 calls were refused for load. With a bounded retry on rate limits alone, and the attempt count recorded on every row, the next run had 1. Batch 1 had none.
- The output cap is part of the measurement. DeepSeek-R1 reasons in visible text inside its 512-token budget, and 29.5% of its replies were cut off mid-JSON. I score those as unparseable, because the cap is part of the condition being measured, not something to work around.
- Dropping a broken reply flatters the model. A reply that broke the answer format was being filed as an error and left out of the score. It now counts as unparseable, the same way the local arm treats it. Re-scoring the smoke results under the new rule moves one model's route score from 89% to 83%. That figure is a re-scoring of recorded replies, not a fresh run.
What I haven't checked yet
-
Matched reasoning. The Gemini and OpenAI models get
reasoning="low"and the others get nothing. Part of Gemini's lead could be that thinking budget rather than the model. A matched control run is next, and until it's done I won't claim Gemini is better at knowing what it doesn't know. - Tuning on the default model. Gemini Flash is Kaggle's default model, so every fix above was first checked on it. I'm stating that, not correcting for it.
- Calibration. The Brier scores spread the hosted models further apart than task score does. That's a pattern in one batch, and I haven't tested it for significance.
Who scored what. The grading rules are written down, and the code that applies them is tested, but nobody outside the build has re-scored these replies.
If you would like to see the work behind the work The Forge
Top comments (0)