DEV Community

sean campbell
sean campbell

Posted on AI-assisted

Does your model know when it doesn't know? A benchmark for the ESCALATE answer

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

Most leaderboards ask one question: did the model get it right? I wanted to ask a second one: does the model know when it can't?

I run a small multi-agent system on one laptop, where local models hand work up to bigger ones. In a setup like that, a small model that answers wrong is worse than one that says "I can't do this, pass it up." A modest model that knows its limits can sit safely in the chain. A confident one that bluffs can't.

So every task in this benchmark has a refusal token, ESCALATE. In one item out of five, the answer has been deliberately removed, or the document doesn't contain it. On those items, ESCALATE is the only correct reply.

There are 200 items across four everyday job shapes:

  • Route (60): a one-line request → pick the right tool and arguments from a 20-tool catalogue, or escalate when no tool fits or an argument is missing.
  • Classify (50): a short work-log note → status, severity, and whether a human is needed, or escalate when the note doesn't say.
  • Judge (50): a claim and a document → SUPPORTS, CONTRADICTS or UNRELATED, or escalate when the document is on topic but silent on the claim.
  • Ground (40): a passage and a question → answer verbatim from the passage, or escalate when it isn't there.

Every model gets two scores: its task score on the answerable items, and its false-confidence rate, meaning how often it answered anyway when the right reply was ESCALATE. Each answer also carries a stated confidence, so I can draw a reliability diagram too.

Every item is invented from scratch, and nothing is scraped. A privacy gate checks the whole set before it's published.

Models Tested

Two groups, on one chart:

  • Kaggle's hosted model suite: the frontier models.
  • A local ladder running on my laptop: 1B, 3B, 4B and 8B open models, run on CPU at temperature 0, with every raw call kept.

The question behind the chart: can a laptop's 3B know its own limits as well as a frontier model knows its own?

Findings

The runs are in progress. Before any model touched the fixtures, I wrote down my predictions and timestamped them so they can't drift toward the results:

  1. At least one frontier model answers anyway on more than 20% of the unanswerable items. (my confidence: 75%)
  2. The best small local model (4B or under) has a lower false-confidence rate than at least one frontier model. (40%)
  3. Task score and false confidence are only weakly related across models (Spearman below 0.5). (60%)

I'll grade those here, misses included, once the numbers are in.

My Benchmark

Kaggle link coming once the benchmark is published there.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

Writing the predictions down with confidences before the runs is the part most benchmark posts skip, and it's what will make the grading worth reading.

A thought on grading them, since the unanswerable set is 40 items (one in five of 200):

  • A false-confidence rate from 40 items has a wide interval. 8 of 40 (20%) gives a 95% interval of roughly 10% to 35%, so for prediction 1 a frontier model at 9 of 40 hasn't clearly crossed 20%. Printing the interval next to each rate, and saying before the numbers arrive whether a prediction is graded on the point estimate, keeps the grading honest either way.
  • Prediction 2 compares two models on the same 40 items, so a paired test (McNemar, on the items where exactly one of the two escalates) is much sharper than comparing their two intervals.
  • For prediction 3, with around 8 models even a true correlation of zero produces Spearman values out to about ±0.7 by chance, so "below 0.5" will be hard to tell apart from either no relation or a moderate one. A bootstrap interval over models next to the value would show how much the ranking can actually say.