An agent picks a tool by matching the current task to each tool's name and description, then it writes the arguments for that call. Chip Huyen names the two ways that goes wrong: an invalid tool, or the right tool with the wrong parameter. Write each description so the right choice and the valid inputs are obvious, score that selection on its own, and make the agent ask before it runs anything that can change or delete data.
This post draws on two Chain of Thought episodes. Chip Huyen, then at Tép Studio (formerly NVIDIA), and Vivienne Zhang of NVIDIA, discussed agent evaluation in episode 5 (Dec 2024). Angie Jones, then VP of Engineering for AI Tools and Enablement at Block, described deploying Goose with MCP in episode 48 (Jan 2026).
How does an agent choose a tool at each step?
It reads the task against the tools it has been given, picks one (or answers with no tool), and generates the call. Chip Huyen treats an agent as tools plus a plan. Web search, a calculator, and a retriever are tools. Given a task, a constraint, and that set, the agent is supposed to produce a strategy, then carry it out one call at a time. After a call returns, the next step uses that result.
The call itself is where selection goes wrong. A weather function is not a bare call weather. The model has to pull out a location and a time, and those values have to be valid. A search string, a JSON payload, and function arguments fail the same way: the tool name can be right while the fields are not.
"The other is whether it can call the right tool. So a lot of time, we'll try to call like an invalid tool. Or it will try to use the correct tool, but not with the right parameter."
Chip Huyen, then at Tép Studio (formerly NVIDIA), on Chain of Thought ep 5
So a description has two jobs. It says when this tool is the right choice, and it says which arguments the call needs. Protocols like MCP, which Angie Jones uses with Goose, put many systems behind one client. That only helps if each server's tool cards stay distinct. One shared interface does not fix two cards that sound the same.
Why do overlapping descriptions break selection?
The model picks the wrong tool, or it picks the right one and still sends bad arguments. Either way the rest of the task follows a bad step. Chip Huyen's method is to map how the application will fail, then put a metric on each of those points. For agents she uses three checks:
- Can it finish the task inside the constraints? Her example is a two-week trip on a $5,000 budget.
- Does it call the right tool, and are the parameters valid?
- Is the sequence of actions short enough to stop, rather than looping on and on?
Keep those scores separate. A single "accuracy" number hides a wrong tool behind a lucky final answer, and it hides a correct tool that was called with garbage arguments. If two cards could both match the same sentence in the user request, rewrite them until a reader can point at one card and say why the other does not apply. State the data the tool touches, the case it is for, and what it returns.
How do you keep a destructive tool from running unchecked?
Mark it on the tool, and require a person to approve that call. Angie Jones does this in Goose at the tool level, not per model. The annotation tells the agent the tool can change, edit, or delete things. A separate toggle lets the agent run tools that are not destructive, and forces a question when it hits one that is.
"annotating tools that say like this tool is this word is strong, destructive. But that just means, uh, it can change things. It can edit or it can delete or something like that. So helping your agent understand that this is a destructive tool, right?"
Angie Jones, then VP of Engineering for AI Tools and Enablement at Block, on Chain of Thought ep 48
The same episode covers a second gate: an allowed list of MCP servers. If a server is not on the list, it will not install. In her words, the agent will say nope. Most of the servers on Block's list are built in-house, and those still go through security review, because not every problem is intentionally malicious. Pair that with the destructive flag. The list decides which MCP servers will install. The annotation identifies destructive tools, and the permission toggle requires approval before they run.
"agent, you're allowed to run any tools that are not destructive. But if you run into one, that is, you need to ask my permission first. Like, give me some insight into what this is, what you're trying to do, and I'll. I'll say yes or no as the human in the loop."
Angie Jones, then VP of Engineering for AI Tools and Enablement at Block, on Chain of Thought ep 48
A minimal sketch to adapt, not a drop-in library.
def agent_step(task, allowed_tools):
if not allowed_tools:
return no_tool(task)
scored = [(t, match_score(task, t.desc)) for t in allowed_tools]
applicable = [pair for pair in scored if pair[1] > 0]
if not applicable:
return no_tool(task)
applicable.sort(key=lambda pair: pair[1], reverse=True)
if len(applicable) > 1 and applicable[0][1] == applicable[1][1]:
return flag_overlap(task, applicable[0][0], applicable[1][0])
best = applicable[0][0]
args = generate_args(task, best.desc)
if not valid(best, args):
return no_tool(task)
if best.flag == "destructive" and not ask_permission(best.name, args):
return no_tool(task)
return run_tool(best, args)
Read that as a decision shape. If the allowed list is empty or no description matches, answer with no tool. If two descriptions tie for the top score, flag the pair for a rewrite instead of guessing. Otherwise take the best match, generate and check the arguments, and run a tool that can change data only after approval.
What should you measure once the agent is in production?
Selection quality on its own, then the same failure points again on live traffic. A benchmark you curate before launch is a start. Vivienne Zhang's point is that production queries are a different set: more complex, full of edge cases. What she has seen work is a loop. Capture user feedback and usage data, find what tripped the LLM, and change the system. In her NVIDIA copilot example, the early failures were missing answers in the data, and chunks cut in the wrong places, so the fix was the data and the chunking, not the model.
"I think what I've seen working is that you simply have to keep iterating. You have to capture the user feedback, the usage data, and systematically troubleshoot what have tripped up your LLM and then improve the system."
Vivienne Zhang of NVIDIA, on Chain of Thought ep 5
For tools, log the choice and the arguments, not only the final reply. When users correct a run, check whether the description was ambiguous, whether a parameter was invalid, or whether the plan never stopped. Vivienne Zhang also describes a familiar handoff: a support bot keeps the query until it cannot handle the complexity, then sends it to a person. Treat that handoff as one more failure point, and run the same evals Chip described on it. Fix the card or the constraint, ship again, and read the next batch of traces.
FAQ
What is the fastest fix for wrong agent tool calls?
Measure whether the right tool was chosen and the parameters were valid (Chip Huyen), then rewrite the descriptions behind the misses.
Should destructive tools run without approval?
No. Annotate tools that can change, edit, or delete things, and use a toggle so the agent asks a person before those calls run (Angie Jones).
How do I start measuring agent reliability without a big benchmark?
List the likely breaks (wrong tool, bad parameters, a plan that never stops) and put one metric on each. After release, use feedback and usage data to see which of those still fire (Chip Huyen; Vivienne Zhang).
Takeaway
- Write each tool card with the case it covers, the arguments it needs, and what it returns.
- Compare cards side by side and rewrite any pair that could match the same request.
- Annotate tools that can change or delete data, and require a yes before they run.
- Score task completion, tool-and-parameter validity, and plan length as separate checks.
- After release, log choices and arguments, then fix the descriptions that real queries trip.
The longer explainer, with the episode clips, is on Chain of Thought. It draws on this episode.
Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.
Drafted with AI assistance from the episode transcripts.
Top comments (2)
The "right tool, wrong parameter" case is the one that bit us most, and better descriptions only got us partway. What actually helped was the tool's response when the args were bad: instead of a stack trace or a generic 400, return "rejected" plus one plain sentence like "date must be ISO, got 'next friday'". The model fixes its own call on the next step most of the time. Do you score that recovery separately from first-try selection, or does it fold into the sequence-length check?
The "two cards that could both match the same sentence" test is a good one. A trick for near-twins like search vs fetch vs headless browser: put the escalation path into the description of the tool the agent reaches for first. Something like "Returns clean markdown for one URL. If the result is nearly empty or asks you to enable JavaScript, use render_page instead." The model sees the hint exactly when it needs it, instead of having to compare every card up front.
On @anciwasim's point about "rejected" responses for bad args, I'd score that recovery separately too. An agent that always gets the call right on the second try looks fine on final-answer accuracy and still doubles your tool calls.