Summary
I built an MCP tool that assigns multiple independent tasks to subagents and has the subagents execute them in parallel. Below, I call this MCP tool the parallel-execution MCP tool. MCP stands for Model Context Protocol, a standard for AI agents to find and call external tools. A subagent is a separate AI agent that an AI agent launches in order to delegate a task to it. However, the three AI coding agents OpenHands, OpenCode, and Qwen Code did not call the parallel-execution MCP tool on their own.
To find the conditions under which it gets called, I tried three things. First, I increased the number of short memos I had the AI agent summarize from 6 to 30. Second, I replaced the task I gave the AI agent with 8 independent tasks that require reading, fixing, and verification. Third, in the rule file that the AI agent loads at startup, I wrote the name of the parallel-execution MCP tool and the condition that it is to be used when the items are independent of each other and there are 5 or more of them. With these three, there was not a single run in which subagents completed the tasks.
The AI agents called the parallel-execution MCP tool only when I added to the MCP tool description one sentence stating when to use the tool. The MCP tool description is the text written in the description field of an MCP tool definition. The conditions are the 6 combinations of the three AI agents and whether a rule file is present. I ran each of the 6 conditions twice with the sentence in the description and twice without it, and compared 24 runs in total. With the sentence, the AI agent called the parallel-execution MCP tool in 7 of the 12 runs. In 4 of those 7 runs, subagents completed the tasks. Without the sentence, every one of the 12 runs had 0 calls.
When I tested this difference with Fisher's exact test, the p-value was 0.0046 with the run as the unit and 0.0152 with the condition as the unit. However, 6 of the 12 runs with the sentence were made before I aligned two points other than the description. The two points are how numbers were written in the prompt and the recording of the time limit given to the AI agent. When I compared only the runs in which the two points were aligned, the runs with a call were 3 of 6 with the sentence and 0 of 6 without it. The p-value of this comparison is 0.182, which is larger than 0.05.
For comparison, I also measured a consultation MCP tool. The consultation MCP tool is an off-the-shelf MCP tool with which an AI agent asks another model for an opinion just once before deciding on a design. Whether the AI agent called the consultation MCP tool depended on whether I had written a provision about consultation in the rule file. With the provision, there was a call in all 6 runs. Without the provision, there were 0 calls in all 6 runs.
A call to the parallel-execution MCP tool is heavy processing that launches multiple subagents. A call to the consultation MCP tool is light processing that asks another model for an opinion just once. I present the following explanation as a hypothesis: the factor that decides whether the AI agent calls a tool changes with the weight of the processing. The reason is that the consultation MCP tool is an off-the-shelf tool whose description I cannot change.
In the body, I first explain the results of the 24 runs and my decision to add the sentence to the default description. Next, I explain the settings I kept the same across the four measurements and the method of judgment. The four measurements are the 6-memo measurement, the 30-memo measurement, the 8-task measurement, and the measurement comparing the description with and without the sentence. After that, I show the results of the four measurements and the results of the statistical tests. Then I explain the three cases I measured again because of defects in the measurement program. Finally, I explain the measurement of the consultation MCP tool and the four factors I have not separated. I also explain the design of the run records that lets me later tell apart the reasons for 0 calls, as well as what cannot be said from the results of this experiment. In the last section, I have gathered the bibliographic details of the materials I referred to.
What you can take away
For those who build their own MCP servers and are having trouble because AI agents do not call their MCP tools, I explain the following three things.
- When an MCP tool is not called even though you wrote its name and the conditions for using it in the rule file, you will be able to check whether the cause lies in the MCP tool description. In the 12 runs without the sentence, including the 6 runs in which I had the AI agent load the rule file, there were 0 runs in which the AI agent called the MCP tool. In the 12 runs with the sentence, there was a call in 7
- You will be able to build, in your own environment, an experiment that changes one point at a time and measures whether an AI agent calls an MCP tool. I explain the settings I kept the same, the points I changed, and a method of judgment that does not rely on the AI agent's reports
- You will be able to keep the records of runs with 0 calls in a form that lets you tell two cases apart later. The two cases are the case where the AI agent did not call the tool and the case where it called the tool but the call did not meet the criteria for counting
All three are based on the four measurements I ran and the three cases I measured again because of defects in the measurement program.
This article is a discussion based on primary records from work I am pursuing. In the body I name the three AI coding agents I compared. The content does not represent the views of any of their providers.
The MCP tool that was called only when I added one sentence to its description
I measured the conditions under which AI agents call, on their own, an MCP tool I built. In this article, an AI agent means an AI coding agent. MCP stands for Model Context Protocol, a standard for AI agents to find and call external tools.
I implemented the parallel-execution MCP tool as an MCP server and made it the target of the measurement. The parallel-execution MCP tool assigns multiple independent tasks to subagents and has the subagents execute them in parallel. A subagent is a separate AI agent that an AI agent launches in order to delegate a task to it.
In this article, I count two numbers separately. The first is the number of runs in which the AI agent called the parallel-execution MCP tool. The second is the number of runs in which subagents completed the tasks.
The conclusion is as follows. The AI agents called the parallel-execution MCP tool only when I added to the MCP tool description one sentence stating when to use the tool. The MCP tool description is the text written in the description field of an MCP tool definition. An AI agent reads MCP tool descriptions when it chooses which tool to call.
There are 6 conditions. For each of the three AI agents, OpenHands, OpenCode, and Qwen Code, I prepared a case in which I had it load a rule file and a case in which I did not. A rule file is a file, loaded by the AI agent at startup, that describes how to proceed with the work.
For each of the 6 conditions, I ran twice with the sentence stating when to use the tool in the MCP tool description and twice without it. There were 24 runs in total, and the results are as shown in the following table.
| The sentence in the description | Number of runs | Runs in which the AI agent called the parallel-execution MCP tool | Runs in which subagents completed the tasks |
|---|---|---|---|
| Present | 12 runs | 7 runs | 4 runs |
| Absent | 12 runs | 0 runs | 0 runs |
Of the 12 runs without the sentence stating when to use the tool in the MCP tool description, 6 were runs in which I had the AI agent load the rule file. In these 6 runs as well, there were 0 runs in which the AI agent called the parallel-execution MCP tool.
On the basis of this result, I added the sentence stating when to use the tool to the default description of the parallel-execution MCP tool. However, so that I can compare again in the future, I kept a mechanism for replacing the description through an environment variable. I also pinned down with a unit test that the default description contains this sentence.
Because I added the sentence stating when to use the tool to the default description of the parallel-execution MCP tool, what can be compared has changed. After I added the sentence, when to use the parallel-execution MCP tool is written in both the MCP tool description and the rule file. So from now on, even if I compare having the AI agent load the rule file with not having it load the file, the comparison for calls to the parallel-execution MCP tool becomes a different one. It is a comparison between having the statement of when to use the tool in two places and having it in one place. The two places are the MCP tool description and the rule file, and the one place is the MCP tool description.
Even so, I added the sentence stating when to use the tool to the default description of the parallel-execution MCP tool. The reason is that among the 12 runs without the sentence, there were 0 runs in which subagents completed the tasks. I gave priority to the parallel-execution MCP tool being used over the rigor of the comparison.
I wrote about how I think about which tasks can be delegated to AI in another article, The difference between work that broke when I delegated it to AI and work that didn't. This time, I measured whether an AI agent delegates tasks to subagents.
The settings I kept the same across the four measurements, and the method of judgment
With four measurements, I investigated the conditions under which AI agents call, on their own, the parallel-execution MCP tool I built. The parallel-execution MCP tool is an MCP tool that assigns multiple independent tasks to subagents and has the subagents execute them in parallel. The four measurements are the 6-memo measurement, the 30-memo measurement, the 8-task measurement, and the measurement comparing the description with and without the sentence. In all four measurements, I kept four things the same: the model, the inference endpoint, the order of runs, and the method of judgment.
I compared three AI agents: OpenHands, OpenCode, and Qwen Code. On 2026-08-17, I installed them on macOS and recorded their versions at that time. The versions are OpenHands CLI 1.16.0, OpenCode 1.18.18, and Qwen Code 0.21.13.
In every run, I used the same single model. The three AI agents call a preview version of a model for code through the shared endpoint of an inference service. In the run records that the measurement program writes for each run, the model field has the same value on every line.
The shared endpoint of the inference service has a limit on how many runs can execute at the same time, so I ran them one at a time, in order.
To check whether the AI agent chooses a tool by itself, I did not write the name of the MCP tool in the prompt. The prompt is the text of the instructions given to the AI agent.
To judge whether the AI agent called a tool, I did not use the reports the AI agent wrote in its output. For the judgment, I used the following records, decided for each type of tool.
- For the MCP tool that launches subagents, I used the call records written by that MCP tool.
- For the tool that sends requests to an AI agent running in another session, I used the file for receiving requests.
- For the consultation MCP tool, I used the records of requests left on the relay server. The consultation MCP tool is an MCP tool with which an AI agent consults another model. The relay server relays requests from the consultation MCP tool to the model being consulted.
I judged whether each task passed or failed by the exit code of the tests alone. I placed the original tests where the AI agent could not rewrite them, and when judging pass or fail, I ran those originals. I also detected, with a separate mechanism, whether the AI agent had rewritten the tests.
The points I changed across the four measurements are as shown in the following table.
| Name of the measurement | Task given to the AI agent | Point I changed |
|---|---|---|
| 6-memo measurement | Summarizing 6 short memos | The measurement used as the baseline |
| 30-memo measurement | Summarizing 30 short memos | Increased only the number of items |
| 8-task measurement | 8 independent tasks that require reading, fixing, and verification | Increased the amount of reasoning one task requires |
| Measurement comparing the description with and without the sentence | 8 independent tasks that require reading, fixing, and verification | Changed only whether the MCP tool description has the sentence |
In all four measurements, the conditions are the 6 combinations of the three AI agents and the 2 options of having or not having a rule file. A rule file is a file, loaded by the AI agent at startup, that describes how to proceed with the work.
I developed the method of measurement that compares three AI agents on the same model in another article, I compared three AI agents on the same AI model, and the only differences that came out were elapsed time and the number of retries. This time I used that method and changed what I measured from the quality of the deliverables to calls to the MCP tool.
The MCP tool that was not called even when I increased the number of items and the amount of reasoning
Before I added the sentence stating when to use the tool to the description of the parallel-execution MCP tool, I ran three measurements. The parallel-execution MCP tool is an MCP tool I built that assigns multiple independent tasks to subagents and has the subagents execute them in parallel. The purpose of the measurements was to find in what cases AI agents call the parallel-execution MCP tool on their own.
In the three measurements, I changed the number of items and the amount of reasoning one task requires. In all three measurements, the conditions are the 6 combinations of the three AI agents, OpenHands, OpenCode, and Qwen Code, and whether a rule file is present. A rule file is a file, loaded by the AI agent at startup, that describes how to proceed with the work.
The results of the three measurements are as shown in the following table.
| Name of the measurement | Task given to the AI agent | Conditions in which the AI agent called the parallel-execution MCP tool | Conditions in which subagents completed the tasks | Time until the AI agent finished processing |
|---|---|---|---|---|
| 6-memo measurement | Summarizing 6 short memos | 0 of 6 | 0 | 22 to 143 seconds |
| 30-memo measurement | Summarizing 30 short memos | 1 of 6 | 0 | 42 to 101 seconds |
| 8-task measurement | 8 independent tasks | 1 of 6 | 0 | 39 to 264 seconds |
The times in the table are the time of the shortest run and the time of the longest run among the runs for the 6 conditions.
In the 6-memo measurement, I had the AI agent summarize each of 6 short memos in one line. In none of the 6 conditions did the AI agent call the parallel-execution MCP tool, and it processed the 6 summaries itself through to the end.
On the same day as the 6-memo measurement, I also measured the request-sending tool. The request-sending tool is a tool that sends requests to an AI agent running in another session. In all 6 conditions, the AI agent called the request-sending tool. In the file on the side that receives the requests, lines containing the marker string remained. The marker string is a string I had put into the text of the task. Whether a rule file was present had no bearing on the results for the request-sending tool.
I thought the cause of the different results between the parallel-execution MCP tool and the request-sending tool lay in the descriptions of the two tools. The description of the request-sending tool stated the purpose of the tool and when to use it.
In the 30-memo measurement, I increased the short memos to be summarized from 6 to 30. In 1 of the 6 conditions, the AI agent called the parallel-execution MCP tool. However, as far as the remaining records show, not a single subagent had been launched even in this 1 condition. In the run for this 1 condition, there were 2 call records written by the parallel-execution MCP tool. The measurement program at the time saved only the numbers from the first record. The measurement program is the program that launches the AI agent, writes the run records, and judges whether each task passed or failed. In the first record, the number of subagents launched was 0. The content of the second record does not remain.
I interpreted the result of the 30-memo measurement as follows. When the processing of each item is a repetition of the same steps, the AI agent can process all 30 items together by itself, even when there are 30 of them. So just increasing the number of items did not give the AI agent a reason to delegate tasks to subagents.
In the 8-task measurement, I switched the task given to the AI agent from summarizing memos to 8 independent tasks. Each of the 8 tasks requires reading, fixing, and verification, and each task has its own dedicated explanatory text and tests. However, I did not give the AI agent the correct answers.
In the 8-task measurement as well, the AI agent called the parallel-execution MCP tool in only 1 of the 6 conditions. However, in this 1 condition, the AI agent incorrectly wrote the script that launches the subagents. Because of the error in the script, not a single subagent was launched. So the parent agent processed the 8 tasks itself. The parent agent is the AI agent that tried to launch the subagents.
Combining the 6 conditions of the 30-memo measurement and the 6 conditions of the 8-task measurement gives 12 conditions. For the 8-task measurement, I counted the runs made after I fixed a defect in the measurement program. The AI agent called the parallel-execution MCP tool in only 2 of the 12 conditions. However, subagents were not launched in either of the 2 conditions. So the conditions in which subagents completed the tasks were 0 of 12.
At the time of the 8-task measurement, the rule file contained the name of the parallel-execution MCP tool. As for when to use it, the rule file also said "use it when the items are independent of each other and there are 5 or more of them." However, even with these two things written in the rule file, there was not a single condition in which subagents completed the tasks.
The sentence stating when to use the tool, added to the MCP tool description
In this section, I explain the measurement comparing the description with and without the sentence. The target of this measurement is the parallel-execution MCP tool I built. The parallel-execution MCP tool is an MCP tool that assigns multiple independent tasks to subagents and has the subagents execute them in parallel.
For this measurement, I used the same 8 tasks as in the 8-task measurement. The 8-task measurement is the measurement in which I had the AI agent solve 8 independent tasks that require reading, fixing, and verification. I also did not change the steps the AI agent needs in order to call the parallel-execution MCP tool. However, the one thing I did change was whether to put the sentence stating when to use the tool at the beginning of the description of the parallel-execution MCP tool. The MCP tool description is the text written in the description field of an MCP tool definition. I switched between putting the sentence in and leaving it out with an environment variable.
Before the measurement, I checked the provisions on MCP tool descriptions in the 2026-07-28 revision of the MCP specification. According to the specification, MCP tools are designed to be model-controlled. The "model" in the specification refers to the language model that the AI agent uses when it chooses tools. There is also an explanation that a language model can discover and call tools automatically, based on its understanding of the context and the user's prompts. The description field is defined as a human-readable description of functionality. The specification has no provision requiring the description field to state when to use a tool. However, the section of the specification on tools that hold state contains the following explanation. If the retention policy is written in the description of the tool responsible for creation, the model can read that policy. So the premise that the model reads descriptions and makes judgments is written in the specification.
In the sentence stating when to use the tool, I wrote the following two things. The first is the situation in which to use the parallel-execution MCP tool. That situation is one where there are multiple mutually independent items and you want to process every item without missing any. In the sentence, I wrote that in this situation, instead of processing the items in order itself, the parent agent has subagents process them in parallel with this MCP tool. The second is the behavior of the parallel-execution MCP tool. The number of subagents launched and the number of subagents that completed are left in the records. The two things above are the gist of the sentence, and I do not write the actual wording of the sentence in this article.
I took the way of writing the sentence stating when to use the tool from the description of the request-sending tool. The request-sending tool is a tool that sends requests to an AI agent running in another session. The 6-memo measurement is the measurement in which I had the AI agent summarize each of 6 short memos in one line. In the measurement on the same day as the 6-memo measurement, the AI agent called the request-sending tool in all 6 conditions. The conditions are the combinations of the three AI agents and whether a rule file is present. A rule file is a file, loaded by the AI agent at startup, that describes how to proceed with the work.
Next, I explain how I came to measure again. In the first round, the AI agent called the parallel-execution MCP tool in 4 of the 6 runs with the sentence stating when to use the tool. However, the review that followed found two points, other than the description, that were not aligned. The first was that how numbers were written in the prompt differed between the case with the sentence and the case without it. The second was that the time limit given to the AI agent had not been recorded. I aligned the two points and measured again. In the second round, there were calls in 0 of 6 runs without the sentence and in 3 of 6 runs with the sentence.
I requested another review and received 1 finding. The finding was as follows. If the AI agent ignores an environment variable whose value is empty, then even when a run uses the setting that leaves out the sentence stating when to use the tool, the MCP server uses the default description. In that case, the runs measured as runs without the sentence become runs that used the default description. So I added to the MCP server a process that writes the description actually used out to the records. After adding it, as a third round, I ran the case without the sentence 6 more times. Of the 6 runs, there were 0 runs in which the AI agent called the parallel-execution MCP tool. In all 6 runs, the description written out to the records did not contain the sentence.
Adding up the runs so far, there were 12 runs each with and without the sentence stating when to use the tool. The results split by whether a rule file was present are as shown in the following table.
| The sentence in the description | Rule file | Number of runs | Runs in which the AI agent called the parallel-execution MCP tool |
|---|---|---|---|
| Present | Present | 6 runs | 5 runs |
| Present | Absent | 6 runs | 2 runs |
| Absent | Present | 6 runs | 0 runs |
| Absent | Absent | 6 runs | 0 runs |
With the sentence stating when to use the tool, there were more runs in which the AI agent called the parallel-execution MCP tool when I had it load the rule file. However, there are only 1 or 2 runs that share the same type of AI agent, the same presence or absence of the rule file, and the same presence or absence of the sentence, and the results also vary widely. There are cases in which, under the same condition, the first run had a call and the second run did not. So this difference can be read only as a tendency, but it cannot be said that the rule file has no effect.
What I can say comes down to the following two points. Without the sentence stating when to use the tool, even when I had the AI agent load the rule file, the AI agent called the parallel-execution MCP tool in 0 of 6 runs. With the sentence, there were more runs with a call when I had the AI agent load the rule file.
Whether subagents completed the tasks differed by the type of AI agent.
Qwen Code called the parallel-execution MCP tool in 3 runs and launched 8 subagents in every one of them. In 2 of the 3 runs, all 8 subagents completed. In the remaining 1 run, 7 of the 8 completed. The run in which 7 completed was a run in which I did not have the AI agent load the rule file. In this run, 7 of the 8 subagents completed their tasks even without the rule file.
In the first-round run in which I had it load the rule file, OpenCode fixed on its own the error in the script that launches the subagents. After that, the 8 subagents completed. This run took 935 seconds.
OpenHands called the parallel-execution MCP tool in 3 runs. Across the 3 runs, there were 4 call records written by this MCP tool, and 32 subagents were launched. However, 0 of the 32 subagents completed. I wrote the cause of this in the section that explains the defects in the measurement program.
In every run, all 8 tasks passed their tests, and there were 0 cases in which the AI agent rewrote the tests.
External research also reports that results change with the quality of MCP tool descriptions. The first study is MCP Tool Descriptions Are Smelly!. This study covers the descriptions of MCP tools provided by many MCP servers. According to this study, many descriptions have defects, and filling in the elements missing from a description raises the task success rate. The second study is Learning to Rewrite Tool Descriptions. This study rewrites tool descriptions to make tool calls more reliable. This study also reports that rewriting descriptions raises the success rate.
What the two studies measured is the task success rate, not the rate at which AI agents call tools. So I do not use the two studies as support for the results of my measurements. However, from the two studies I took it that the observation that results change with the quality of MCP tool descriptions is not an observation of mine alone. I wrote the scale of the two studies' investigations and their figures in the section on the materials I referred to.
Fisher's exact test comparing the numbers of runs in which an MCP tool was called
I tested the differences in the number of runs in which the AI agent called an MCP tool with Fisher's exact test. There are two measurements I tested.
The first measurement is the measurement of the parallel-execution MCP tool. The parallel-execution MCP tool is an MCP tool that assigns multiple independent tasks to subagents and has the subagents execute them in parallel, and I built it. In this measurement, I compared the case where the description of the parallel-execution MCP tool has the sentence stating when to use the tool with the case where it does not.
The second measurement is the measurement of the consultation MCP tool. The consultation MCP tool is an off-the-shelf MCP tool with which an AI agent asks another model for an opinion just once before deciding on a design. In this measurement, I compared the case where the rule file has a provision about consultation with the case where it does not. A rule file is a file that describes how to proceed with the work, and the AI agent loads it at startup.
The records of the experiment did not include results of Fisher's exact test. So I calculated the test results on 2026-09-10. Using only the Python standard library, I calculated two-sided p-values directly from the hypergeometric distribution.
I counted the cases in which the AI agent called the MCP tool, using three kinds of units.
- In the method with the run as the unit, I counted one run as one item.
- In the method with the condition as the unit, I counted one condition as one item. The conditions are the combinations of the three AI agents and whether a rule file is present, and there are 6 of them. The three AI agents are OpenHands, OpenCode, and Qwen Code. Each condition has 2 runs. In this method, I counted the conditions in which the AI agent called the MCP tool in at least 1 of the 2 runs.
- In the method with the type of AI agent as the unit, I counted one type of AI agent as one item. Each type of AI agent has 2 runs. In this method, I counted the AI agents that called the MCP tool in at least 1 of the 2 runs. I used this method for the measurement of the consultation MCP tool.
The results of the test are as shown in the following table.
| What I compared | Values counted | p-value |
|---|---|---|
| Parallel-execution MCP tool. The run as the unit | With the sentence, 7 of 12 runs. Without the sentence, 0 of 12 runs | 0.0046 |
| Parallel-execution MCP tool. The condition as the unit | With the sentence, 5 of 6 conditions. Without the sentence, 0 of 6 conditions | 0.0152 |
| Parallel-execution MCP tool. Only the runs in which the two points other than the description were aligned | With the sentence, 3 of 6 runs. Without the sentence, 0 of 6 runs | 0.182 |
| Parallel-execution MCP tool. Whether a rule file was present, with the sentence | With a rule file, 5 of 6 runs. Without one, 2 of 6 runs | 0.242 |
| Consultation MCP tool. The run as the unit | With the provision, 6 of 6 runs. Without the provision, 0 of 6 runs | 0.0022 |
| Consultation MCP tool. The type of AI agent as the unit | With the provision, 3 of 3 types. Without the provision, 0 of 3 types | 0.10 |
There are four cautions for reading the p-values in the table.
The first caution is as follows. The 24 runs of the parallel-execution MCP tool are not independent of each other. That is because they are runs in which I repeated the 6 conditions twice each, with and without the sentence stating when to use the tool. So the p-value with the run as the unit tends to come out smaller than it really is.
The second caution is as follows. In statistics, a way of counting under which it is harder to judge that there is a difference is called a conservative way of counting. Counting with the condition as the unit is more conservative than counting with the run as the unit. In the measurement of the parallel-execution MCP tool, the p-value of the comparison with the condition as the unit is 0.0152, which is larger than 0.0046, the p-value of the comparison with the run as the unit. However, counting with the condition as the unit is not as conservative as counting that compares only the runs in which the two points other than the description of the parallel-execution MCP tool were aligned. Comparing only the runs in which the two points other than the description were aligned, the p-value is 0.182.
The third caution is as follows. The 12 runs with the sentence stating when to use the tool are the 6 runs of the first round and the 6 runs of the second round. The 6 runs of the first round were made before I aligned the two points other than the description. In 4 of these 6 runs, the AI agent called the parallel-execution MCP tool. The two points are how numbers were written in the prompt and the recording of the time limit given to the AI agent. The 6 runs of the second round were made after I aligned the two points. In 3 of these 6 runs, the AI agent called the parallel-execution MCP tool. The 12 runs without the sentence were all made after I aligned the two points. Comparing only the runs in which the two points were aligned, the numbers of runs in which the AI agent called the parallel-execution MCP tool are as follows. With the sentence, it was 3 of 6 runs, and without the sentence, 0 of 6 runs. The p-value of this comparison is 0.182, which is larger than 0.05.
The fourth caution is as follows. In the comparison for the consultation MCP tool, taking the type of AI agent as the unit gives a p-value of 0.10. This value is also larger than 0.05.
Three cases I measured again because of defects in the measurement program
The four measurements are measurements in which I investigated in what cases AI agents call, on their own, the parallel-execution MCP tool I built. The four measurements are the 6-memo measurement, the 30-memo measurement, the 8-task measurement, and the measurement comparing the description with and without the sentence. The parallel-execution MCP tool is an MCP tool that assigns multiple independent tasks to subagents and has the subagents execute them in parallel.
In the four measurements, three defects in the measurement program were found. The measurement program is the program that launches the AI agent, writes the run records, and judges whether each task passed or failed. I fixed all three defects and measured again. In this section, I explain the three defects in the order they were found.
The first defect was found in the first round of the 8-task measurement. The 8-task measurement is the measurement in which I had the AI agent solve 8 independent tasks. The measurement program had no process for detecting that the AI agent had rewritten the tests. In a review, it was pointed out to me that this process was missing. I added to the measurement program a process that detects rewrites of the tests and a process that restores the tests to their original state, and measured again.
However, I have not deleted the records of the measurement that had the defect. I attached to those records a flag indicating a defect in the measurement program and excluded them from the aggregation.
In the 8-task measurement after I fixed the measurement program, all 8 tasks passed their tests in all 6 conditions. The conditions are the combinations of the three AI agents, OpenHands, OpenCode, and Qwen Code, and whether a rule file is present. A rule file is a file, loaded by the AI agent at startup, that describes how to proceed with the work. There were 0 cases in which the AI agent rewrote the tests and 0 cases in which it stopped the work partway.
The second defect was found in the first round of the measurement comparing the description with and without the sentence. The measurement comparing the description with and without the sentence is a measurement that changes only whether the description of the parallel-execution MCP tool has the sentence stating when to use the tool. However, in the first round, there were two points other than the description that were not aligned between the runs with the sentence and the runs without it. The first was that how numbers were written in the prompt differed. The second was that the time limit given to the AI agent had not been recorded. I aligned the two points and measured again.
The third defect was found in the runs of OpenHands. In the OpenHands runs, not a single subagent completed. Using a program that reproduces the problem, I extracted the errors from when the subagents failed. All 8 subagents had terminated about 1.2 seconds after they were launched, with an error saying that no authentication method had been selected. This error occurred because three causes overlapped.
The first cause was that the parent process's environment variables were not passed to the MCP server. With the method the measurement program uses to launch the AI agent, OpenHands CLI 1.16.0 was the only case in which the environment variables were not passed. In the case of OpenCode and Qwen Code, the environment variables were passed with the same method. However, I have not checked whether the parent process's environment variables are passed to the MCP server if the OpenHands configuration is written differently.
The second cause was that there was no environment configuration file in the place the parallel-execution MCP tool reads from. Instead of environment variables, the parallel-execution MCP tool reads the environment configuration file located where this MCP tool is placed. Because I ran the measurement in an isolated working directory, there was no environment configuration file in that place.
The third cause was that the measurement program did not pass the authentication key used by the subagents through the environment variable entry of the MCP configuration.
I fixed the measurement program and stopped using a method that assumes the parent process's environment variables are inherited. The fixed measurement program writes the authentication key used by the subagents into the environment variable entry of the MCP configuration and passes it that way. After the fix, I measured again with OpenHands.
In the run in which I had it load the rule file, OpenHands called the parallel-execution MCP tool on its own and launched 8 subagents. All 8 subagents completed, and all 8 tasks passed their tests. This run was the first in which all of the OpenHands subagents completed.
In the run in which I did not have it load the rule file, OpenHands also called the parallel-execution MCP tool. There were 3 call records, and the subagents launched in this run totaled 16. Of the 16 subagents, 1 completed.
So in the measurements before the fix, it was because of the defect in the measurement program that not a single OpenHands subagent completed. A difference in the capability of the AI agents was not the cause, so this result does not show which of the three AI agents is better or worse. However, in the run in which I did not have it load the rule file, 15 of the 16 subagents launched did not complete, for reasons other than the authentication key. I have not yet investigated why these 15 subagents did not complete.
The 2 runs of OpenHands after I fixed the measurement program are additional runs to investigate why not a single subagent completed in the measurements before the fix. So I recorded these 2 runs separately from the aggregation of the 24 runs of the measurement comparing the description with and without the sentence.
The consultation MCP tool that was called only when I wrote a provision in the rule file
To compare with the parallel-execution MCP tool, I also measured the consultation MCP tool. The parallel-execution MCP tool is an MCP tool that assigns multiple independent tasks to subagents and has the subagents execute them in parallel, and I built it. A call to the parallel-execution MCP tool is heavy processing that launches multiple subagents.
The consultation MCP tool is a tool included in an off-the-shelf MCP server. Before deciding on a design, the AI agent uses the consultation MCP tool to ask another model for an opinion just once. A call to the consultation MCP tool is light processing.
In the prompt, I wrote neither the name of the consultation MCP tool nor the model being consulted. The prompt is the text of the instructions given to the AI agent. In this measurement, I changed only whether to write a provision about consultation in the rule file. A rule file is a file that describes how to proceed with the work, and the AI agent loads it at startup.
The provision about consultation says two things. The first is that the AI agent asks another model for an opinion before deciding on a design, and decides after taking into account the main points of the opinion that comes back. The second is that the AI agent must not decide on its own without asking another model for an opinion. I also wrote the name of the consultation MCP tool in the provision.
This measurement had 12 runs, which are the combinations of the following: the 2 options of having or not having the provision about consultation; the three AI agents, OpenHands, OpenCode, and Qwen Code; and 2 repetitions. With the provision, the AI agent called the consultation MCP tool in all 6 runs. Without the provision, the AI agent did not call the consultation MCP tool even once in any of the 6 runs. In this measurement, there were 0 defects in the measurement program. The measurement program launches the AI agent, writes the run records, and judges whether each task passed or failed.
I judged whether the AI agent called the consultation MCP tool from the records of the relay server. The relay server relays the requests that the consultation MCP tool sends to the model being consulted. In this judgment, of the requests in the relay server's records, I counted only the requests that met all three of the following.
- The request is addressed to the model being consulted.
- The status code of the response is 200.
- The body of the response is not empty.
I set an invalid key in the consultation MCP tool. The valid key exists only on the relay server, and the relay server replaces the invalid key with the valid key when it forwards a request. So a request that does not go through the relay server fails authentication, and a consultation that is not left in the relay server's records does not succeed.
However, this measurement has one constraint. Because the description of the consultation MCP tool is included in the off-the-shelf MCP server, I could not change this description. The measurement of the consultation MCP tool is a comparison that changes only whether the rule file has the provision.
The hypothesis that the factor changes with the weight of the processing, and four factors I have not separated
In this section, I explain the extent of what I can say from the results of the two measurements. The first is the measurement of the parallel-execution MCP tool. The parallel-execution MCP tool is an MCP tool I built that assigns multiple independent tasks to subagents and has the subagents execute them in parallel. The second is the measurement of the consultation MCP tool. The consultation MCP tool is an off-the-shelf MCP tool with which an AI agent asks another model for an opinion just once before deciding on a design.
The results of the two measurements are as follows. The AI agent called the parallel-execution MCP tool only when the description of the parallel-execution MCP tool had the sentence stating when to use the tool. Also, the AI agent called the consultation MCP tool only when the rule file had a provision about consultation. A rule file is a file, loaded by the AI agent at startup, that describes how to proceed with the work.
The conclusion first written from the results of the two measurements was as follows. The factor that decides whether an AI agent calls an MCP tool on its own changes with the weight of the processing being called. This conclusion divided heavy and light processing as follows. The processing of calling the parallel-execution MCP tool is heavy processing that launches multiple subagents. The processing of calling the consultation MCP tool is light processing that asks another model for an opinion just once. The first conclusion explained that for heavy processing the MCP tool description is the factor, and for light processing the rule file is the factor.
I requested a review of the first conclusion and received a serious finding. There were two reasons for the finding.
The first reason was that in the measurement of the parallel-execution MCP tool as well, there were more runs in which the AI agent called the parallel-execution MCP tool when I had it load the rule file. When the sentence stating when to use the tool was present and I had the AI agent load the rule file, the AI agent called the parallel-execution MCP tool in 5 of 6 runs. When the sentence was present and I did not have it load the rule file, it called the parallel-execution MCP tool in 2 of 6 runs. So the explanation that the factor switches with the weight of the processing does not match the results of the measurement.
The second reason was that four points differ between the two measurements. The four points are the MCP tool used, the quality of the MCP tool description, the task given to the AI agent, and how specifically the MCP tool was written into the provision of the rule file. So I have not been able to separate which difference is the cause of the results.
I accepted the review's finding and withdrew the assertion in the first conclusion. So I now write the results of the two measurements separately. I write the explanation that the factor changes with the weight of the processing as a hypothesis, and I also write that I have not separated the factors. Turning this explanation back into an assertion would mean writing again, as it is, the claim that was sent back in the review, so I do not turn it back into an assertion.
There are four factors I have not separated. In statistics, such factors are called confounding factors. The four factors are as follows.
- The MCP tool used differs between the two measurements.
- The quality of the MCP tool description differs between the two MCP tools.
- The task given to the AI agent differs between the two measurements.
- How specifically the MCP tool was written into the provision of the rule file differs between the two measurements.
Because the consultation MCP tool is an off-the-shelf MCP tool, I cannot change the description of the consultation MCP tool. To separate the four factors, it is necessary to take a single MCP tool as the target and measure all combinations of the presence or absence of the sentence stating when to use the tool and the presence or absence of the provision in the rule file. However, I have not yet run this measurement.
The design of the run records for telling apart the reasons for 0 calls
In this section, I explain the design of the run records. A run record is the record that the measurement program writes for each run. The measurement program is the program that launches the AI agent, writes the run records, and judges whether each task passed or failed. With the measurement program, I measured whether AI agents call MCP tools on their own.
When a run record says there were 0 calls to the MCP tool, there are two possible reasons. The first is that the AI agent did not call the MCP tool. The second is that the AI agent called the MCP tool, but the call did not meet the criteria for counting. I set the criteria for counting for each measurement. For example, in the measurement of the MCP tool that consults another model, I counted only requests whose response had status code 200 and a body that was not empty. So that these two reasons can be told apart later, I divided the roles of two fields in the run records as follows.
- The call field records whether the MCP tool was called, the total number of calls, and the breakdown of that number.
- The run status field records whether the measurement program worked correctly and the run can be counted as a result of the measurement. This field does not record whether the AI agent succeeded at the task.
On the line for a run in which the measurement program had a defect, I set the run status field to failure and attach a flag indicating the defect. So as not to record values that could not be measured, I do not write the call field on that line.
I have been using the recording method explained so far since 2026-08-27. The lines of the run records from before 2026-08-27 remain as recorded with the old method, and I have not rewritten them. That is because there is a provision that lines are only appended to the run records. I also have not deleted the records of measurements that became invalid because of defects in the measurement program; I keep those records with a flag attached.
I received a review of the conclusions of the measurements. During this review, I split three processes into separate functions. The three processes are the process that counts the number of calls, the process that judges whether the measurement program worked correctly, and the process that assembles the run records. I pinned down the behavior of the three functions with 10 unit tests.
Previously, I copied part of the MCP tool's output into the notes field of the run records, but I stopped using this method. Now I record the total number of calls and the breakdown of that number as structured data.
What cannot be said from the results of this experiment
In this section, I explain what cannot be said from the results of this experiment, and what can. In this experiment, I measured the conditions under which the three AI agents, OpenHands, OpenCode, and Qwen Code, call, on their own, the parallel-execution MCP tool I built. The parallel-execution MCP tool is an MCP tool that assigns multiple independent tasks to subagents and has the subagents execute them in parallel.
The following six things cannot be said from the results of this experiment.
- Only one model was used in this experiment, so I cannot say that other models would give the same results.
- There are only 1 or 2 runs under the same condition. A condition is a combination of the type of AI agent and whether a rule file is present. A rule file is a file, loaded by the AI agent at startup, that describes how to proceed with the work. Also, because I measured the same conditions repeatedly, the p-values of Fisher's exact test calculated with the run as the unit tend to come out smaller than they really are.
- The measurement of the consultation MCP tool is a comparison that changes only whether to write a provision about consultation in the rule file. The consultation MCP tool is an off-the-shelf MCP tool with which an AI agent asks another model for an opinion just once. Because it is off-the-shelf, I could not change the description of the consultation MCP tool. Also, between the measurement of the parallel-execution MCP tool and the measurement of the consultation MCP tool, there are four factors I have not separated. The four factors are the MCP tool used, the quality of the MCP tool description, the task given to the AI agent, and the specificity of the provision in the rule file.
- In the OpenHands runs, there are 15 subagents that did not complete even after I fixed the defect in the measurement program. The defect I fixed is the one in which the authentication key did not reach the subagents. I have not yet investigated why these 15 did not complete.
- As of 2026-09-27, the version of OpenCode has gone up from 1.18.18, which I used for the measurements, to 1.18.32. The version of Qwen Code has also gone up from 0.21.13, which I used for the measurements, to 0.24.6. So measuring the same way with the current versions may give different results. However, the version of the OpenHands CLI is 1.16.0, the same as the version I used for the measurements.
- The values measured in the external research on the quality of MCP tool descriptions are task success rates. They are not the rates at which AI agents call MCP tools. So I cannot use the external research as support for the results of this experiment.
What can be said from the results of this experiment is as follows. With the three AI agents, even when I wrote in the rule file the name of the parallel-execution MCP tool and in what cases to use it, there was not a single run in which subagents completed the tasks.
Even when I increased the memos I had the AI agent summarize to 30, and even when I increased the amount of reasoning one task requires, the result was the same. Of the 12 conditions, which combine the 6 conditions of the 30-memo measurement and the 6 conditions of the 8-task measurement, the AI agent called the parallel-execution MCP tool in only 2. However, in neither of the 2 conditions was a single subagent launched.
However, only when I added the sentence stating when to use the tool to the description of the parallel-execution MCP tool were there runs in which subagents completed the tasks.
Still, I have not yet separated which is the factor that decides whether an AI agent calls an MCP tool: the MCP tool description, the rule file, or the weight of the processing executed after the call.
Materials I referred to
In the body of this article, I referred to three public materials. I checked the original text of all three on 2026-09-10 and checked it again on 2026-09-27. For each material, I write below its content and the range I quoted from it in this article.
- The tools section of the Model Context Protocol specification (revision 2026-07-28) https://modelcontextprotocol.io/specification/2026-07-28/server/tools
This material is the specification that defines the fields of an MCP tool definition. This material also contains the provision that MCP tools are designed to be model-controlled.
In the body of this article, I quoted the following three points from this material. The description field of an MCP tool definition is defined as a human-readable description of functionality. This material has no provision requiring the description field to state when to use a tool. The section on tools that hold state contains the premise that the model reads MCP tool descriptions and makes judgments.
- Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions (Mohammed Mehedi Hasan, Hao Li, Gopi Krishnan Rajbahadur, Bram Adams, Ahmed E. Hassan, arXiv:2602.14878, submitted 2026-02-16, revised 2026-05-31) https://arxiv.org/abs/2602.14878
This material is a study that examined 856 MCP tools provided by 103 MCP servers. It reports that 97.1% of the descriptions of the MCP tools examined had at least one defect, and that 56% did not state their purpose explicitly. It reports that filling in the elements of a description improves the task success rate by a median of 5.85 percentage points. However, it reports that the number of execution steps increased by 67.46%, and that performance dropped in 16.67% of the cases.
In the body of this article, I quoted from this material only the point that results change with the quality of MCP tool descriptions. I wrote the scale of the investigation and the figures in this section.
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use (Ruocheng Guo, Kaiwen Dong, Xiang Gao, Kamalika Das, arXiv:2602.20426, submitted 2026-02-23, revised 2026-04-29) https://arxiv.org/abs/2602.20426
This material is a study that rewrites tool descriptions to make tool calls more reliable, and it addresses the following problem. Tool descriptions are often written for human developers and contain ambiguity. Especially when the number of candidate tools grows, AI agents cannot resolve this ambiguity. It reports that in an experiment that increased the number of candidate tools to 150 or more, the per-query success rate improved by an average of 60.89% compared with the original descriptions.
In the body of this article, I quoted from this material only the point that rewriting tool descriptions raises the success rate. I wrote the figures of this study in this section.
Originally published at The Future of Humans, AI, and the Web, a site where my research and development is recorded and analyzed by a human and an AI.
Top comments (1)
7 of 12 against 0 of 12 says the description outweighs the rule file for a heavy tool and the publisher owns that sentence re read every session.
Did rewording it, rather than removing it, move the call rate?
That's the case I'd worry about.. an upgrade that changes one sentence and quietly changes which tool the agent picks.