This is a submission for the Kaggle Benchmarking Challenge
This is a submission for the Kaggle Benchmarking Challenge.
TOOL JUDGMENT: Does an AI Know When NOT to Use a Tool?
We keep asking AI agents whether they can use tools.
Can the model call a search API?
Can it call Python?
Can it retrieve a database record?
Can it call the right function?
But there's a question I think we don't ask enough:
Does the model know when it should NOT use a tool?
A capable agent shouldn't maximize tool calls.
It should maximize useful outcomes.
So I built TOOL JUDGMENT, a benchmark designed to measure tool-use restraint, tool selection, conflict handling, and recovery from bad tool results.
What I Benchmarked
I created a collection of realistic agent scenarios where a model has access to several deterministic tools.
The important part is that the correct behavior isn't always "call a tool."
Some tasks require a tool.
Some tasks can be answered directly from the conversation.
Some expose a tempting but irrelevant tool.
Others deliberately create conflicts between information already provided to the model and information returned by a tool.
The model therefore has to make a decision before it can succeed:
Should I use a tool at all?
The benchmark contains four categories:
Tool Required — the model must use a specific tool to obtain the answer.
Tool Unnecessary — the answer is already available, so calling a tool is wasted effort.
Tool Conflict — multiple information sources disagree and the model must identify which source has authority.
Tool Failure/Recovery — a tool returns an error, stale result, or unusable response and the model must recover appropriately.
I deliberately avoided making the benchmark a collection of trivia questions.
The goal is to test a behavior that matters when an LLM becomes an agent rather than simply a chatbot.
The Question Behind the Benchmark
Imagine an assistant with access to ten APIs.
A naive evaluation asks: "Did the assistant eventually get the right answer?"
I'm interested in a harder question: "Did the assistant take the right action to get the answer?"
Those aren't the same thing.
An agent can get the right answer while:
calling three unnecessary tools,
selecting the wrong API first,
repeatedly querying the same source,
ignoring a required tool,
trusting a lower-authority source,
or continuing after a tool has already failed.
That behavior matters in production because tools have latency, cost, permissions, side effects, and failure modes.
My Metric: Tool Regret
I introduced a simple concept called Tool Regret.
Every unnecessary or harmful action adds regret.
For example:
Behavior Effect
Correct answer with the minimum necessary action +1
Correct required tool +1
Unnecessary tool call −0.25
Wrong tool −0.50
Redundant tool call −0.25
Required tool ignored −0.75
Harmful/unnecessary action −1
The exact score is less important than the idea:
success alone isn't enough.
I also report tool precision, task success, recovery rate, and successful outcomes per tool call.
This gives us a different view of model quality.
Models Tested
I tested:
[MODEL 1]
[MODEL 2]
[MODEL 3]
[MODEL 4]
[MODEL 5]
[MODEL 6]
[MODEL 7]
I intentionally chose a mixture of reasoning-focused, general-purpose, fast, and smaller models rather than testing only the biggest models.
The interesting comparison isn't just "who wins?"
It's whether model size, reasoning ability, or speed predicts tool judgment.
The Experiment
I ran the benchmark under three conditions.
Unlimited tools
The model could call tools without an explicit budget.
Limited tools
The model received a small tool-call budget.
Costly tools
Different tools were assigned simulated costs.
This allowed me to measure whether models change their behavior when unnecessary actions become expensive.
The hypothesis was simple: If a model really understands the task, increasing tool cost should make it more selective without dramatically reducing task success.
Findings
Here is where the experiment became interesting.
Raw accuracy wasn't the whole story
The model with the highest task accuracy was not necessarily the model with the lowest Tool Regret.
That distinction matters.
A model can be very good at recovering from a bad decision while still making too many unnecessary decisions.Tool availability changes model behavior
When tools were freely available, some models appeared much more willing to call them even when the answer was already present in context.
That suggests an interesting failure mode:
having a tool can itself become a distraction.Tool conflicts exposed a different weakness
The hardest cases weren't necessarily the ones requiring the most reasoning.
They were the cases where the model had to decide which information deserved to be trusted.
A tool result isn't automatically authoritative just because it came from a tool.Efficiency and intelligence aren't identical
This was the result I was most interested in.
A model that solves 95% of tasks with ten tool calls may be less attractive for an agent system than one that solves 92% with two.
Traditional leaderboards can hide that difference.
What Surprised Me
Before running the benchmark, I expected the difficult cases to be the ones where the model had to chain several tools together.
Instead, I became more interested in the opposite case:
Can the model resist doing something unnecessary?
This feels surprisingly close to how human expertise works.
An experienced engineer doesn't run every diagnostic they know.
An experienced developer doesn't query every database.
An experienced assistant doesn't open six applications just because they're available.
Sometimes the intelligent action is:
don't do anything extra.
What I Would Measure Next
The current benchmark deliberately uses deterministic tools.
The next version should introduce:
delayed tool responses,
stale information,
probabilistic tool reliability,
tools with real monetary costs,
irreversible actions,
permission failures,
and long multi-turn tasks where the correct tool changes over time.
I'd also like to test whether models become better at tool selection after seeing feedback about the cost of their previous actions.
That would turn TOOL JUDGMENT from a static evaluation into an agent-learning benchmark.
My Benchmark
You can inspect the benchmark and its leaderboard here:
The benchmark is intended to be reproducible, inspectable, and easy to extend with new tools and scenarios.
Final Thought
We've spent a lot of time teaching models how to use tools.
I think the next question should be:
Can we teach them when to leave the tools alone?
Because the smartest agent may not be the one that takes the most actions.
It may be the one that knows exactly which action is worth taking.
Top comments (0)