DEV Community

Cover image for Tool names and parameter expectations
Charles Solar for Favur

Posted on

Tool names and parameter expectations

If your agent's tool is called read_file, its parameter had better be called path

Every model you can put behind an agent has seen read_file more times than you will ever write it, and almost every time it was called with a parameter named path. That habit comes with the name. So if you build a read_file tool whose parameter is file_to_read, you have set the model up to guess wrong. It will reach for path, your validation will reject the call, and you pay for the retry. On a bad day it half-follows your schema and half follows its habit, and you get a call that is valid and wrong.

When you design tools for an agent, you have two good options and one bad one. Keep a common convention exactly, name and parameters both. Or break it on purpose, with a name the model will not confuse with the common one. The bad option is the middle ground, a familiar name with unfamiliar parameters, and it is where a surprising number of tool calls go wrong.

Why a name carries its parameters

A model does not read your tool list the way a person reads an API reference. It sees a name, and everything it has absorbed about that name comes along for free. That includes the parameter names that usually follow it, the shape of the values, and even the order they tend to appear in. Agent frameworks, open-source harnesses and the vendors' own examples have repeated the same few tool names so often that those names now act like prompts in their own right.

That is mostly a gift. A model handed a tool called read_file with a path parameter needs almost no description to use it correctly, because the name and the schema agree with everything it already knows. The trouble starts when they disagree. Your schema says one thing, the name says another, and the model has to decide which to believe on every single call. Sometimes it picks your schema. Sometimes it picks its habit. You rarely get to choose which.

Common names come with a schema attached

read_file, write_file, bash, apply_patch, execute_command - names like these show up across agent frameworks and in the vendors' own tool definitions, and models have strong priors about what goes into each one.

Some vendors make it explicit. Anthropic's bash tool has to be named bash, and it takes one string called command. You do not even send a schema for it, because the docs say the schema "is built into Claude's model and can't be modified." Its text editor tool, str_replace_based_edit_tool, is fixed the same way. A command field picks view, str_replace, create or insert, and path, old_str and new_str carry the edit. OpenAI's GPT-4.1 prompting guide publishes an apply_patch tool that takes the whole patch as one string, and says the model "has been extensively trained" on its diff format.

The vendors also disagree with each other. OpenAI's shell tools take the command as a list of arguments, while Anthropic's bash takes one string for the whole line. So a tool called "shell" or "command" means one thing to one model and something else to the next. If your harness runs models from more than one vendor, as ours does, the only safe conventions are the ones they all share.

These are the shapes models expect from the most common names.

Name What models expect
read_file path
write_file path, content
bash command, one string for the whole line
text editor (str_replace) path, old_str, new_str
apply_patch the whole patch as one string

How the middle ground goes wrong

A familiar name with unfamiliar parameters fails in a few recognizable ways, and it helps to know them because they do not all look like errors.

The first is the loud one. The model sends path to a tool that wants file_to_read, validation rejects it, and the model tries again. Each retry is a full round trip, so a naming choice you made once gets paid for on every run, by every agent, forever.

The second is quieter. The parameter names line up but the meaning does not. If your read_file takes a path that is relative to some project root the model has never been told about, or your line numbers start at zero when the model expects one, every call passes validation and a fraction of them read the wrong thing.

The third is shape. A model that has written bash("pytest -v tests/") thousands of times will try to hand that same string to any tool whose name suggests running a command, whatever your schema says. If your command tool wants its arguments split up, the name is working against you before the first call.

The fourth is extras. Models add the parameters they expect even when your schema does not have them. Depending on how strict your validation is, those get rejected, which costs a retry, or silently dropped, which means the model believes it asked for something it did not get.

If you use the name, use the parameters

When your tool does the same job as a common one, keep the name and keep the parameter names that go with it. read_file takes path. write_file takes path and content. bash takes command. The model gets these right without being taught, which is the cheapest accuracy you will ever buy.

That is how Favur's file tools work. read_file and write_file both take path. The file-editing tool goes further and accepts the *** Begin Patch format from OpenAI's guide as one of its five input formats, so a model that already knows that format can just use it.

Keeping the convention does not mean your tool has to be basic. It means the extras go around the familiar core instead of replacing it. Favur's read_file has optional start and end lines for large files, and it returns far more than the file. The response carries line numbers, a symbol outline, lint results and a hash of the file. That hash can be passed back to write_file so a stale write gets caught. None of that required renaming anything. The model calls read_file(path=...) exactly as it expects to, and the richness arrives in the response, where it costs the model nothing to understand.

The rule of thumb is simple. Required parameters should match the convention exactly, and anything new should be optional, with a name that explains itself.

If your tool is different, break the convention entirely

inline shape

Sometimes your tool genuinely needs a different shape, and then the worst thing you can do is keep the familiar name. The model sees the name, recalls the usual shape, and fights your schema on every call.

Favur hit exactly this with commands. Every command runs through a security layer that checks each argument before anything executes, and a single shell string makes that hard. The parser has to split the line itself, so it must refuse unclosed quotes and tokens that could be read two ways, and it cannot safely split a command it has no profile for at all. So the main command tool takes the program, its boolean switches, its key-value options and its paths as separate fields.

Calling that tool execute_command, or worse, bash, would have invited the model to hand it a shell string. It is called execute_structured_command instead. The name is close enough to be obvious and different enough that the model does not mistake it for the tool it has seen a thousand times. That is the kind of break you want. The meaning stays in the same family, the word is clearly different, and the model reads the schema instead of assuming it.

A new name does not erase every habit, so the tool also coaches. Our guide for it warns specifically against putting an option like --tb=short into the switches field, because that is exactly the shortcut a model trained on command lines reaches for. The parser rejects a path that lands in the switches field and says which field it belongs in. And the old string parser, when it hits a token it cannot classify, answers with a message pointing the model at execute_structured_command. A clean break in the name, plus errors that teach the new shape, gets you most of the way.

The same pytest command sent to execute_command as one string, and to execute_structured_command as program, switches, options and paths

Keep the familiar name as a quiet fallback

Breaking the convention still leaves the question of what happens when a model reaches for the old name anyway. Our answer is to separate what the model is shown from what it can call.

The original string version, execute_command, still exists. It is no longer in the list of tools the model is shown, but it stays registered and callable, so a model that reaches for the old name out of habit gets a command that runs, through the same security checks, instead of an unknown-tool error. Hiding it took one flag on the tool class, and any other tool can use the same flag. The model is steered toward the new tool by the list, and forgiven by the registry.

Models really do reach for names they were not given. We traced a model calling write_document in a turn where that tool was not on its list, because it had seen write_document calls earlier in the same conversation. Nothing handled that call, the error was logged as a warning, and the model never found out why nothing happened. The fix we wrote up for it is simple. When a model calls a tool it does not have, answer with a short message saying that tool is not available here and naming the one to use instead, and let it try again. An unknown tool call should be a correction, not a silence.

How to tell if your names are costing you

You do not have to guess whether a name is fighting the model. Count validation failures per tool and look at what the model actually sent. We split every tool failure into categories, and "the model called the tool wrong" is kept apart from "the tool refused on purpose" and "something broke", because only the first one says anything about your naming.

If one tool's malformed calls keep showing the same wrong parameter name, the model is telling you what it expects the tool to look like. Either give it that, or change the name so it stops expecting it.

What two weeks of our own runs show

We ran that count across every tool our agents called between September 14 and 27, about 440,000 calls. We left out three internal workflow-control tools, which would otherwise dominate the count. We counted only structural errors, meaning calls where the model's arguments failed the tool's schema or were not valid JSON. Anything that failed for another reason, like a file that does not exist, is left out.

Structural argument errors per 1,000 calls, grouped by how common the tool's name is

The two tools with common names and common parameters, read_file and list_dir, failed 45 times in 144,655 calls, or 0.3 per thousand. All 45 were read_file, across 129,951 calls, and only 14 of them were a missing path. list_dir had none in 14,704 calls.

The middle bar is our own middle-ground case. attempt_completion is a name models know from other coding agents, where it carries the result of the task. Ours adds a category for why the task ended, plus a list of missing capabilities that becomes required for one of those categories. That tool failed 13.2 times per thousand calls, and most of those failures were the category rule.

The tools with names models have not seen come to 5.9 per thousand taken together, and most of that comes from four of them, window_observe, add_agent_notes, delegate_to_agent and execute_structured_command. The typical uncommon tool called 1,000 or more times gets well under one per thousand. A name the model has never seen gives it no habit to fall back on, so it reads the schema.

Structural argument errors per 1,000 calls for every tool called 1,000 or more times

The uncommon names run from zero to 78 per thousand, so a new name is not a guarantee, and a few of our own tools clearly need work. Three of them err more often than attempt_completion does. Still, the pattern matches the argument of this article. The familiar name with familiar parameters is the cheapest spot on the chart, and our one familiar name with its own parameters errs at more than twice the rate of the typical-to-heavy mix of names the model has never seen. That is the middle ground, measured on a single tool of ours, so treat it as one data point rather than a law.

A short checklist for naming agent tools

Keep read_file with path, break to execute_structured_command, avoid read_file with file_to_read

Before you name a tool, look at what the common version of that name takes, in the vendor docs and in the popular agent frameworks.

If you keep a common name, keep its required parameters exactly (path, command, content). Add new parameters as optional ones, but never rename the ones the model already expects.

If your parameters are genuinely different, break the convention entirely. Pick a name that is clearly not the common one and says how it differs, such as structured, batch or safe.

When you break a convention, make your error messages teach the new shape, and name the right field or the right tool in the message.

If you run models from more than one vendor, assume the common names mean slightly different things to each of them, and stick to the conventions they share.

If you retire a familiar name, leave it callable but unlisted, and send whatever arrives through the same checks as everything else.

When a model calls a tool it does not have, tell it so, and name the tool it should use instead.

We wrote earlier about what your tools should say back to the model, the other half of designing a tool for a model. You can see how Favur runs its agent team on our site.

Name your tools the way the model already expects, or nothing like it. The middle is where the retries live.

Top comments (0)