DEV Community

Toine
Toine

Posted on

I gave 360 browser tools an MCP server. Here is what broke.

I run ToolForte. It is a site with 360 small tools: merge a PDF, validate an IBAN, parse a cron line, diff two texts. Most of them run in the browser, so the file never leaves your machine. I built it for people who do not want an account for a five second job.

Then I started using Claude and Cursor for real work and noticed I was copying results out of my own site into a chat. That is silly. The tools are deterministic. An agent should be able to call them.

So I gave the site an MCP server. This is what I learned, including the parts that went wrong.

The first version was too big to be useful

The obvious move is to expose everything. I did. The server listed 111 tools, later 180. It worked, technically. In practice a list that size eats a large slice of the context window before the agent has done anything. And the agent picks the wrong tool more often, because twelve of them sound alike.

The fix was progressive disclosure. The server now shows 14 tools by default. Those cover most jobs. A search tool finds the other 166 by name or by what you need. A registry or a power user can ask for the full list with ?tools=all on the endpoint. Same server, two views.

Registries score you on metadata, not on what the tools do

I listed the server on Smithery and got a quality score of 43 out of 100. The tools worked. The score was about what the listing said about them. No titles. Thin parameter descriptions. No output schema. No annotations saying which tools are read only.

I went through all of them. Every tool got a title, a description per parameter, an output schema with structured content, and annotations. The score went to 96. Nothing about the tools changed. If you publish an MCP server, do this first. It is what an agent reads to decide whether to call you.

Version numbers drift when nothing ties them together

My registry manifest said 1.10.0. The server told every client it was 1.8.0. The number was typed by hand in the handler and nothing checked it against the manifest that gets published. Two releases went out like that.

Now the handler reads its version from the manifest file. The publish script asks the live server for its version and refuses to publish while the two differ. Boring, but it is the kind of thing a registry scan flags and you do not notice yourself.

What is in it now

  • 150 deterministic utilities. IBAN and VAT validation, cron parsing, regex testing, date maths, unit conversion, hashing, JSON and text tools.
  • 14 render tools. HTML or a URL in, a hosted PDF or screenshot out, as a signed link that lives for 24 hours.
  • Workflows. Chain tools into one job, save it, and let an agent run it by id.
  • Memory that survives between sessions.
  • 6 AI tools for the jobs that are not deterministic.

Free to try without a key: 3 renders a day. Pro is 19 euro a month. That buys a monthly balance of credits, and the API, the MCP server and the AI tools all draw from the same one.

Adding it takes about thirty seconds

Remote server, streamable HTTP:

https://toolforte.com/api/mcp
Enter fullscreen mode Exit fullscreen mode

Claude Desktop, Cursor, ChatGPT and the others each have a config screen for this; the exact steps per client are on toolforte.com/mcp. If you prefer stdio, npx toolforte-mcp wraps the same endpoint. The server is in the official registry as io.github.Toinedotcom/toolforte.

The same 150 tools are also a plain REST API, documented at toolforte.com/developers. Everything is searchable in one box at toolforte.com/everything.

What I want to know from you

Which tool is missing? I add deterministic ones quickly. If your agent calls the server and something comes back wrong, tell me the tool and the input. I will look at it the same day.

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli • • Edited

Your version-drift story has a cousin we hit today: the Claude Desktop manifest of our MCP server still listed 18 tools after the server had 20, because nothing compared the two lists. A test now registers every tool and checks the names against the manifest, which is worth doing for every copy of the tool list (manifest, registry server.json, README tables). On the context cost of big lists, one measurement that helps: in our agent benchmark on gpt-5.4-mini, 48% to 90% of prompt tokens were served from the provider's cache depending on the run, because the tool definitions were byte-identical every turn. So progressive disclosure is cheapest when the default list stays fixed for the whole session and search returns tool names as results, rather than adding tools mid-session, which changes the list and breaks the cache.

Collapse
 
nikolas_dimitroulakis_d23 profile image
Nikolas Dimitroulakis •

Did you pick the 14 defaults by hand or from call logs? On the ApyHub MCP (I co-founded it) we went one step further: no defaults at all, just four generic tools (catalog, search, spec, call) over 1,500+ endpoints, and teams curate which endpoints their agents can see. The jump from 43 to 96 without touching a single tool is the part I'd show every API team.