DEV Community

Guilherme Dalla Rosa
Guilherme Dalla Rosa

Posted on

An agent with hundreds of tools

Tools are what turned language models into agents, and the Model Context Protocol (MCP) made plugging them in cheap enough that nobody stops at three. Every system the agent has to touch adds its own set, each with dozens of endpoints, so an agent that started with three tools ends up with far more in front of it than any task needs. The trouble is that every one of those definitions is loaded into the context window up front. A good chunk of the context is gone before the first message arrives. And with so many tools that look alike, the model picks the wrong one more often than you'd expect. Nothing fails outright, but the agent takes longer, costs more per turn and, every so often, does the wrong thing without a single error in the logs.

Connecting that many tools isn't a wiring problem, it's a context problem. What the model sees and when it sees it decides whether the agent works. That decision belongs to the harness, the loop and the tooling around the model, and it's yours whichever model sits inside it. In the pattern that has settled for this, called progressive disclosure, the model never sees the full catalogue.

Where all those tools come from

Most of those tools come from a habit that goes back to the first MCP servers, when the usual way to build one was to map every endpoint of an API to its own tool. That's how a single integration turns into a few dozen tools before anyone has asked whether the agent needs them. Neon's recent video MCP Just Got a Whole Lot Better walks through how that habit came about and what Neon changed on their own server as a result, so it's a good place to start if you want the server-side view.

The tool definitions are billed on every turn, at a discount if prompt caching is on. The MCP client best practices draw the difference between loading them all and fetching them on demand:

Loading every tool up front against progressive discovery

Source: MCP client best practices, from the Model Context Protocol documentation.

 

The wrong tool is the part of the bill you can't see, and Anthropic's tool search documentation says Claude's ability to pick the right tool degrades once you exceed 30 to 50 available tools, a line that two integrations of a few dozen tools each already cross. The post that introduced advanced tool use measured the difference tool search makes on Anthropic's own MCP evaluations:

Model Without tool search With tool search
Opus 4 49% 74%
Opus 4.5 79.5% 88.1%
Source: Introducing advanced tool use on the Claude Developer Platform, Anthropic.

 

The costs don't stop once the agent starts working, because every tool result lands in the context and the model reads past pages of output on every turn. The model gets less reliable as its input grows, even on simple tasks, which Chroma's Context Rot report measured across 18 models. In an agent, that's the rule you agreed ten turns earlier being forgotten, and I've seen it happen with a much smaller set of tools.

Three tools instead of all of them

Progressive disclosure replaces the catalogue with three meta-tools:

  • search takes a plain description of what the model needs and returns the matching names, each with a one-line description
  • get_definition returns the full schema of a single tool
  • invoke calls it and returns the result

Only the definitions the model asked for ever enter the context.

The three layers of progressive disclosure

The MCP guide calls the pattern progressive discovery, a name worth knowing before you search for it, and recommends it for clients that connect to many servers.

Progressive disclosure only pays off once the tool definitions take a share of the context window worth recovering. Below that, or when every tool is used on every request, plain tool calling is the better fit. The guide gives 1% to 5% of the window as an example of where to put the line, while Anthropic draws its own at 10 tools or 10,000 tokens of definitions.

Once the model asks for something, somebody has to decide what a match is, and the guide lists four ways to do it:

  • Keyword matching, BM25 or a regex, which is simple and works well when tool names and descriptions are descriptive
  • Embeddings over the tool descriptions, which handle synonyms and phrasing the keywords miss
  • A small, fast model as a sub-agent that picks the tools for the task, which usually works well but can cost more
  • A hybrid, scoring across keyword and embedding rankings, or switching strategy by use case

The four above are what you build when your provider doesn't ship a tool search of its own, or when you need your own ranking.

One caveat: a search can come back empty, and nothing in the pattern tells the agent what to do about it. Anthropic's tool search returns an empty list rather than an error when nothing matches, so the model does what models do and tries again with a different query, and then again. The stop has to come from the harness, which in Strands means the limits you pass with each invocation to cap the number of turns and the total tokens a run can use. What happens when the budget runs out is your decision, whether that's a partial answer, a plain "I can't do that" or a hand-off to a person.

Where the search can live

These three tools can run on the server, on a gateway in front of every server or on the client that runs the loop, and what changes between them is simply who holds the index.

On the server

A server can trade one tool per endpoint for those three. Neon's video calls this a layered tool call and describes how a lot of MCP servers moved to it. AWS built its own AWS API MCP Server this way, and it covers the whole API with two tools, suggest_aws_commands to search and call_aws to run the command it found:

from fastmcp import Context, FastMCP
from typing import Any

server = FastMCP(name='AWS-API-MCP')

@server.tool(name='suggest_aws_commands')  
async def suggest_aws_commands(query: str, ctx: Context) -> dict[str, Any]:
    """Suggest AWS CLI commands based on the provided query."""
    with get_requests_session() as session:
        response = session.post(
            ENDPOINT_SUGGEST_AWS_COMMANDS,
            json={'query': query},
            timeout=30,
        )
        response.raise_for_status()
        return response.json()

@server.tool(name='call_aws') 
async def call_aws(cli_command: str, ctx: Context) -> CallAWSResponse:
    """Call AWS with the given CLI command and return the result as a dictionary."""
    # call_aws_helper validates the command and applies the security policy
    response = await call_aws_helper(cli_command, ctx, None, None)
    return CallAWSResponse(cli_command=cli_command, response=response)

if __name__ == '__main__':
    main()  # loads the read-only operations index, then server.run()
Enter fullscreen mode Exit fullscreen mode
Adapted from AWS API MCP Server, AWS Labs.

 

The search itself isn't in the code, because suggest_aws_commands sends the query to a service AWS hosts, which returns the most likely CLI commands with their descriptions and parameters, so the search and the definition arrive in one step. AWS has since marked the server as entering end of development and points users to its managed AWS MCP Server instead.

The limit shows up when there are ten servers built the same way, because the agent is then choosing between ten indexes and ten sets of meta-tools. To be fair, Neon's own counter, which Andre Landgraf's post Give your agent Neon tools walks through, is that a tool bundling three calls into one workflow saves the agent from making the same mistakes every time, the way an SDK wraps raw endpoints in convenience methods. Those workflow tools still belong on the server.

On the gateway

A gateway sits between the agent and every server, so there is one connection and one index whatever is behind it. Solo Enterprise for agentgateway does it with two meta-tools, get_tool and invoke_tool, in the setup that Michael Levan, AI Architect at Solo.io, walks through in his post MCP Progressive Disclosure: Save Tokens, Retrieve Schemas. On AWS, AgentCore Gateway ships the search as a built-in tool, x_amz_bedrock_agentcore_search, which has to be enabled when the gateway is created and can't be turned on afterwards. In Strands, AWS ships a plugin for that search in its bedrock-agentcore package:

from mcp_proxy_for_aws.client import aws_iam_streamablehttp_client
from strands import Agent
from strands.tools.mcp import MCPClient
from bedrock_agentcore.gateway.integrations.strands.plugins import AgentCoreToolSearchPlugin

mcp_client = MCPClient(lambda: aws_iam_streamablehttp_client(
    endpoint="https://<gateway-id>.gateway.bedrock-agentcore.<region>.amazonaws.com/mcp",
    aws_region="us-east-1",
    aws_service="bedrock-agentcore",
))

with mcp_client:
    agent = Agent(plugins=[AgentCoreToolSearchPlugin(mcp_client=mcp_client)])
    agent("Find me afternoon flights to New York")
Enter fullscreen mode Exit fullscreen mode
From Amazon AgentCore Tool Search, Strands Agents docs.

Before each request, the plugin has the model sum up in one line what the user wants, searches the gateway with that line and loads only the tools that come back, so the model never searches for itself. The price is an extra model call per request and one more hop on every tool call.

On the client

The guide itself is written for the client, where the host lists the tools from every server once, keeps the definitions on its side and exposes the search, so the servers stay as they are. Strands 1.57 added a coarser version of the pattern as a vended tool, the MCP router:

from strands import Agent
from strands.vended_tools import make_mcp_router

mcp_router = make_mcp_router(
    servers={
        "orders": {"url": "https://orders.example.com/mcp"},
        "billing": {"url": "https://billing.example.com/mcp"},
    },
    max_connections=5,
)
# Nothing connects yet. The model sees one tool, mcp_router, whose
# description ends "Permitted server names: 'billing', 'orders'."
agent = Agent(tools=[mcp_router])
agent("Which orders shipped late this week?")
Enter fullscreen mode Exit fullscreen mode
Adapted from MCP Router, Strands Agents docs.

 

The model connects to a server by name, lists its tools and calls one, all through that same tool, so list_tools does the job of get_definition for a whole server and call_tool is invoke. The router has no search, which leaves the model picking a server by its name alone and reading every definition on it. A search_tools tool and deferred loading are on the Strands roadmap, in its tool selection design.

Anthropic's tool search and OpenAI's are the native version of the full pattern, where you add the provider's search tool to the tools array and mark the others defer_loading. On Bedrock, Anthropic's version works through the InvokeModel API but not Converse, which Strands uses, so there the search on top of the router is still yours to build.

The catch on the client is the prompt cache, because most providers cache the prompt prefix with the tools array in it, so adding a definition mid-conversation invalidates it and the miss can cost more than the definitions you saved. Of the guide's two fixes, the router takes the one that sends every call through a single stable tool so the array never changes, and the native versions take the other, appending what the search finds after the cached prefix.

Code mode

Search cuts the definitions the model sees, but each call still makes a round trip through the model. In code mode, the model writes a script that chains the calls, the script runs in a sandbox and only what it prints comes back. Kenton Varda and Sunil Pai at Cloudflare named it Code Mode, and the guide calls it programmatic tool calling. Their colleague Matt Carey's follow-up pairs it with search to put more than 2,500 Cloudflare endpoints behind two tools.

In Strands, the strands-harness package ships it as a built-in tool, programmatic_tool_caller, which turns every other tool the agent has into a function the script can call. With the gateway search plugin loading the tools first, the model calls them from one script:

from strands_harness import create_harness
from bedrock_agentcore.gateway.integrations.strands.plugins import AgentCoreToolSearchPlugin

# mcp_client: an MCPClient for the AgentCore Gateway
with mcp_client:
    agent = create_harness(
        plugins=[AgentCoreToolSearchPlugin(mcp_client=mcp_client)],
        builtin_tools={"*": False, "programmatic_tool_caller": {"timeout": 60}},
    )
    agent("Which team members exceeded their Q3 travel budget?")

# The model sends one script, such as:
#   import asyncio, json
#   team = await hr___get_team_members(department="engineering")
#   expenses = await asyncio.gather(*[
#       hr___get_expenses(user_id=m["id"], quarter="Q3") for m in team
#   ])
#   ...  # compare each total with its budget, inside the sandbox
#   print(json.dumps(exceeded))
Enter fullscreen mode Exit fullscreen mode
Adapted from Add tools and instructions, Strands Agents docs, with the model's script from Introducing advanced tool use on the Claude Developer Platform, Anthropic.

 

The script runs in Monty, Pydantic's Python interpreter written in Rust, with no access to the filesystem, the network or the environment, so the credentials stay with the host and the gateway. On the server side, the managed AWS MCP Server offers the same idea as aws___run_script, where the script reaches only the AWS APIs and IAM checks each call.

Code mode opens a code execution surface that's now yours to run. The guide is explicit that approving a script doesn't approve every call it makes, so in strands-harness each call still goes through the agent's hooks and policies, then through the gateway's own authorisation. A call that needs a person's approval fails, though, because the package can't pause a running script to ask for it.

Where to put it

In my opinion, the search belongs in the harness you own, whether that's the client or the gateway. Jiquan Ngiam, co-founder and CEO of MintMCP, which sells an MCP gateway, is blunter about the servers, writing that "progressive tool disclosure should be done by the agent, not as an MCP server tool".

A small API with a handful of endpoints is still fine with one tool per endpoint. To be fair, a single API as large as Cloudflare's or AWS's is where Matt's case for putting code mode on the server makes sense. As always, it depends on your requirements. If you want to go further, the Strands docs on programmatic tool calling explain how code mode runs in the harness you own.

Top comments (0)