This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Most leaderboards ask one question: did the model get it right? I wanted to ask a second one: does the model know when it can't?
I run a small multi-agent system on one laptop, where local models hand work up to bigger ones. In a setup like that, a small model that answers wrong is worse than one that says "I can't do this, pass it up." A modest model that knows its limits can sit safely in the chain. A confident one that bluffs can't.
So every task in this benchmark has a refusal token, ESCALATE. In one item out of five, the answer has been deliberately removed, or the document doesn't contain it. On those items, ESCALATE is the only correct reply.
There are 200 items across four everyday job shapes:
- Route (60): a one-line request → pick the right tool and arguments from a 20-tool catalogue, or escalate when no tool fits or an argument is missing.
- Classify (50): a short work-log note → status, severity, and whether a human is needed, or escalate when the note doesn't say.
- Judge (50): a claim and a document → SUPPORTS, CONTRADICTS or UNRELATED, or escalate when the document is on topic but silent on the claim.
- Ground (40): a passage and a question → answer verbatim from the passage, or escalate when it isn't there.
Every model gets two scores: its task score on the answerable items, and its false-confidence rate, meaning how often it answered anyway when the right reply was ESCALATE. Each answer also carries a stated confidence, so I can draw a reliability diagram too.
Every item is invented from scratch, and nothing is scraped. A privacy gate checks the whole set before it's published.
Models Tested
Two groups, on one chart:
- Kaggle's hosted model suite: the frontier models.
- A local ladder running on my laptop: 1B, 3B, 4B and 8B open models, run on CPU at temperature 0, with every raw call kept.
The question behind the chart: can a laptop's 3B know its own limits as well as a frontier model knows its own?
Findings
The runs are in progress. Before any model touched the fixtures, I wrote down my predictions and timestamped them so they can't drift toward the results:
- At least one frontier model answers anyway on more than 20% of the unanswerable items. (my confidence: 75%)
- The best small local model (4B or under) has a lower false-confidence rate than at least one frontier model. (40%)
- Task score and false confidence are only weakly related across models (Spearman below 0.5). (60%)
I'll grade those here, misses included, once the numbers are in.
My Benchmark
Kaggle link coming once the benchmark is published there.
Top comments (1)
Writing the predictions down with confidences before the runs is the part most benchmark posts skip, and it's what will make the grading worth reading.
A thought on grading them, since the unanswerable set is 40 items (one in five of 200):