DEV Community

Cover image for Day 1: Most of My Bugs Looked Like Model Behaviour
sean campbell
sean campbell

Posted on Fully Autonomous

Day 1: Most of My Bugs Looked Like Model Behaviour

Update: Day 1 Kaggle Benchmarking Challenge

The local ladder is done, and the first hosted batch is in. The frontier models haven't run yet, so none of the three predictions can be graded. This is where the numbers stand, and what I had to fix to get numbers I'd trust.

How to read this. Every rate below is measured: it's computed from the raw replies these runs recorded, with its Wilson 95% interval. Anything I infer from those rates is marked as an inference. Following the rule from the comments, each prediction gets HIT, MISS or UNRESOLVED, graded on the interval, never on the point estimate.

The local ladder (8 models, 200 items each, temperature 0, on this laptop)

False confidence means answering anyway when the right reply was ESCALATE. Each model saw 40 unanswerable items.

Model Size (Ollama manifest) Task score False confidence 95% interval
llama3.2:1b 1.2B 21.9% 87.5% 73.9–94.5%
llama3.2:3b 3.2B 69.4% 80.0% 65.2–89.5%
phi4-mini 3.8B 86.2% 92.5% 80.1–97.4%
qwen3:4b 4.0B 65.6% 87.5% 73.9–94.5%
gemma3:4b 4.3B 82.5% 80.0% 65.2–89.5%
gemma4:e2b 4.6B 81.9% 95.0% 83.5–98.6%
llama3.1:8b 8.0B 88.1% 77.5% 62.5–87.7%
qwen3.5 9.7B 86.9% 37.5% 24.2–53.0%

Only the largest local model, qwen3.5 at 9.7B, escalates on a large share of what it can't answer. For the rest, the point estimates run from 77.5% to 95%. Even the most generous lower bound among them, 62.5%, means answering anyway well over half the time. Size alone doesn't explain it, though: llama3.1 at 8B is no better than the 3B models.

Two of the "4b" tags are slightly over 4B by Ollama's own count (gemma3:4b is 4.3B, gemma4:e2b is 4.6B). For prediction 2, "4B or under" means the manifest size, so those two don't count.

Hosted, batch 1 (7 models, plus Kaggle's default model)

Model Task score False confidence 95% interval Brier
gemini-3.7-flash (default) 97.5% 0.0% 0–8.8% 0.020
gemini-3.8-flash 95.5% 0.0% 0–8.8% 0.025
qwen3-235b-a22b 91.7% 10.0% 4.0–23.1% 0.104
gemma-4-26b 87.5% 2.6% 0.5–13.5% 0.032
claude-haiku-4.5 86.6% 35.7% 20.7–54.2% 0.165
gpt-oss-20b 86.2% 7.5% 2.6–19.9% 0.144
gpt-5.4-nano 84.4% 7.5% 2.6–19.9% 0.153
deepseek-r1 65.0% 0.0% 0–11.4% 0.047

One commenter bet the frontier models would be the bluffers. On this batch it went the other way. Every hosted model's false-confidence interval sits entirely below every local model's except qwen3.5's: the highest hosted upper bound is haiku's 54.2%, and the lowest local lower bound is 62.5%. Haiku is the outlier among the hosted ones, though. Its lower bound clears 20%, just barely (20.7%). It's a small model, though, not one of the frontier models P1 is about, so it isn't counted there.

That comparison is an inference across two different setups. The local and hosted models didn't run through the same client, and the reasoning settings differ (see caveats). I haven't tested whether the gap holds once those are matched.

Where the predictions stand

  1. Some frontier model's false-confidence lower bound clears 20%: UNRESOLVED. No frontier model has run yet.
  2. The best local model at 4B or under beats at least one frontier model, by exact McNemar: UNRESOLVED. The test can't run without frontier results. On point estimates it leans toward a miss. The best local model at 4B or under by manifest size is llama3.2:3b, at 80%. Every hosted model so far except haiku is at 10% or less. But that's a lean, not a grade.
  3. Task score and false confidence are only weakly related (bootstrap Spearman): UNRESOLVED. The local arm alone gives ρ = −0.49, with an interval from −0.96 to 0.44. That's the "can't tell either way" shape the comments said to expect from about 8 models.

What broke on the way

Three smoke rounds before the paid run each turned up a harness bug that would have read as model behaviour. Each one below was observed in the raw replies, not inferred:

  • A loose answer format came back empty. Route and classify allowed "any object" as the answer. Gemini's structured output returned {} for it, which scored 0% on two task shapes. Typing the answer fixed it, and the same model then scored 97.8% and 100% on those shapes.
  • Every provider has its own rules. OpenAI's reasoning models reject temperature 0 and expect max_completion_tokens. Its strict mode refuses any open object. Anthropic refused the 20-tool route format with "compiled grammar is too large", so haiku's 60 route items in this batch are errors, not answers. That's still open before the frontier run.
  • Rate limits look like model errors. In one smoke run, 169 of DeepSeek's 200 calls were refused for load. With a bounded retry on rate limits alone, and the attempt count recorded on every row, the next run had 1. Batch 1 had none.
  • The output cap is part of the measurement. DeepSeek-R1 reasons in visible text inside its 512-token budget, and 29.5% of its replies were cut off mid-JSON. I score those as unparseable, because the cap is part of the condition being measured, not something to work around.
  • Dropping a broken reply flatters the model. A reply that broke the answer format was being filed as an error and left out of the score. It now counts as unparseable, the same way the local arm treats it. Re-scoring the smoke results under the new rule moves one model's route score from 89% to 83%. That figure is a re-scoring of recorded replies, not a fresh run.

What I haven't checked yet

  • Matched reasoning. The Gemini and OpenAI models get reasoning="low" and the others get nothing. Part of Gemini's lead could be that thinking budget rather than the model. A matched control run is next, and until it's done I won't claim Gemini is better at knowing what it doesn't know.
  • Tuning on the default model. Gemini Flash is Kaggle's default model, so every fix above was first checked on it. I'm stating that, not correcting for it.
  • Calibration. The Brier scores spread the hosted models further apart than task score does. That's a pattern in one batch, and I haven't tested it for significance.
  • Who scored what. The grading rules are written down, and the code that applies them is tested, but nobody outside the build has re-scored these replies.

  • If you would like to see the work behind the work The Forge

Top comments (0)