<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: MCPulse</title>
    <description>The latest articles on DEV Community by MCPulse (@getmcpulse).</description>
    <link>https://dev.to/getmcpulse</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4101660%2Fa645fc21-6709-43be-9a49-f98f4b99b52a.png</url>
      <title>DEV Community: MCPulse</title>
      <link>https://dev.to/getmcpulse</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/getmcpulse"/>
    <language>en</language>
    <item>
      <title>Your MCP tool schemas are billed on every session, used or not</title>
      <dc:creator>MCPulse</dc:creator>
      <pubDate>Mon, 28 Sep 2026 10:09:57 +0000</pubDate>
      <link>https://dev.to/getmcpulse/your-mcp-tool-schemas-are-billed-on-every-session-used-or-not-2g3m</link>
      <guid>https://dev.to/getmcpulse/your-mcp-tool-schemas-are-billed-on-every-session-used-or-not-2g3m</guid>
      <description>&lt;p&gt;When a client connects to your MCP server it calls &lt;code&gt;tools/list&lt;/code&gt;. Every tool you registered comes back — name, description, and the full JSON Schema for its arguments — and all of it goes into the model's context window.&lt;/p&gt;

&lt;p&gt;That happens before the user has typed anything. It happens on every session. It happens for tools nobody will call.&lt;/p&gt;

&lt;p&gt;This is the cost of your server that nobody measures, and for most servers it's the largest one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Doing the arithmetic
&lt;/h2&gt;

&lt;p&gt;Take a plausible tool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"search_orders"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Search orders by customer, status, or date range. Returns up to 50 matching orders with their line items."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"inputSchema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"object"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"properties"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"customer_email"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Customer's email address"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"enum"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"pending"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"shipped"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"delivered"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cancelled"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"from"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"format"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"date"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"to"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"format"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"date"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Serialised, that's about 480 bytes — roughly 120 tokens.&lt;/p&gt;

&lt;p&gt;Twelve tools of that size is around 1,400 tokens per session. Verbose descriptions and nested schemas push it much higher; servers in the 4,000–6,000 token range are common, and it isn't hard to find worse.&lt;/p&gt;

&lt;p&gt;For context from a survey of 4,749 public MCP servers: the median server ships about 1,251 tokens of schema, the 90th percentile about 7,992, and the largest one measured is on the order of 216,000 — which doesn't fit a 200K context window at all, before anyone has asked a question.&lt;/p&gt;

&lt;p&gt;At $3 per million input tokens the money is trivial. That isn't the cost that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost that matters is attention
&lt;/h2&gt;

&lt;p&gt;The real price is paid in the context window. Every token of schema is a token not available for the conversation, and — more importantly — one more thing the model has to read past to find the tool it needs.&lt;/p&gt;

&lt;p&gt;This is the effect people underestimate. A model choosing between four sharply described tools picks well. The same model choosing between eighteen, several of which overlap, picks worse. You'll see it in your first-call success rate long before you see it on a bill.&lt;/p&gt;

&lt;p&gt;So the question for every tool isn't "does this work" but &lt;strong&gt;"does this earn its place in every conversation"&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which is why dead tools are worse than they sound
&lt;/h2&gt;

&lt;p&gt;Schema size comes from the startup payload — the serialised length of each tool's schema, sent once when the server boots. Put that next to call counts and you get a specific, slightly uncomfortable number:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;export_report&lt;/code&gt; has never been called, but costs 480 tokens of schema every session.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A tool with no calls isn't neutral. It's a fixed tax on every conversation, paid in the scarcest resource the model has, in exchange for nothing at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trimming without losing accuracy
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cut tools before cutting words.&lt;/strong&gt; Removing one unused tool saves more than tightening the prose in five.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Descriptions are for disambiguation, not documentation.&lt;/strong&gt; The model needs to know when to pick this tool over its neighbours. It doesn't need your changelog, your rate limits, or a worked example — those cost tokens on every session to serve a case that arises rarely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enums are cheap and worth it.&lt;/strong&gt; Four values in an &lt;code&gt;enum&lt;/code&gt; cost a handful of tokens and remove an entire class of argument-validation failures. This is one of the few places where more schema pays for itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't describe self-evident parameters.&lt;/strong&gt; &lt;code&gt;"customer_email": { "type": "string", "description": "Customer's email address" }&lt;/code&gt; — the name already said it. Delete the description and lose nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Watch for duplication across tools.&lt;/strong&gt; Five tools that each explain your pagination convention are paying for that explanation five times per session. One server in the corpus appends the same 51-word context block to all 275 of its tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  The measurement that settles arguments
&lt;/h2&gt;

&lt;p&gt;Sort your tools by schema bytes, then look at calls next to them. The tools at the top of the first list and the bottom of the second are your answer.&lt;/p&gt;

&lt;p&gt;You're not looking for a small saving spread across everything. You're looking for the two tools costing 800 tokens a session between them and getting called once a week — because deleting those is free, and it makes every remaining tool easier for the model to choose correctly.&lt;/p&gt;




&lt;p&gt;If you want your own number without installing anything, the &lt;a href="https://getmcpulse.com/check" rel="noopener noreferrer"&gt;schema checker&lt;/a&gt; reports your total &lt;code&gt;tools/list&lt;/code&gt; size and per-tool breakdown from a paste of your JSON, scored against servers your own size. Runs in the browser, nothing uploaded.&lt;/p&gt;

&lt;p&gt;Call counts need instrumentation — that's &lt;a href="https://getmcpulse.com" rel="noopener noreferrer"&gt;MCPulse&lt;/a&gt;, two lines inside your own process, never your arguments or your results.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://getmcpulse.com/blog/schema-tokens-every-session" rel="noopener noreferrer"&gt;getmcpulse.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>mpc</category>
      <category>devtools</category>
    </item>
    <item>
      <title>How to test whether a model can tell your MCP tools apart</title>
      <dc:creator>MCPulse</dc:creator>
      <pubDate>Fri, 25 Sep 2026 10:00:53 +0000</pubDate>
      <link>https://dev.to/getmcpulse/how-to-test-whether-a-model-can-tell-your-mcp-tools-apart-4p5n</link>
      <guid>https://dev.to/getmcpulse/how-to-test-whether-a-model-can-tell-your-mcp-tools-apart-4p5n</guid>
      <description>&lt;p&gt;You can read your own tool descriptions all day and not know whether a model can tell them apart. I spent a few weeks measuring 82,549 public MCP tool descriptions, and the most useful thing I learned is that the measurement has a ceiling.&lt;/p&gt;

&lt;p&gt;Here are two tools from a real server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;score_resume             "Score a resume for ATS compatibility."
analyze_job_description  "Extract what a job posting actually screens on."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They share no content word. Not one, even after stemming. Every lexical metric I have scores them as perfectly distinct.&lt;/p&gt;

&lt;p&gt;Their author says they're confusable, and he's right — because "check my resume for this job" names both objects at once. The collision doesn't live between the two descriptions. It lives between a sentence and a set of tools, and nothing computed from descriptions alone will ever see it.&lt;/p&gt;

&lt;p&gt;So here's the test that does. It takes an afternoon, needs no traffic, and produces evidence you can put in a pull request.&lt;/p&gt;

&lt;p&gt;The method came out of a comment thread on a previous post — credit where it's due, I'd been circling something much more expensive.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of it
&lt;/h2&gt;

&lt;p&gt;For each pair of tools you suspect might be confusable, write &lt;strong&gt;two sentences&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;control&lt;/strong&gt; — a request that should route to exactly one of them, unambiguously.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;probe&lt;/strong&gt; — a request that could fairly route to either.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run each through a model with the full tool list, at fixed temperature, a handful of times. Record which tool gets picked.&lt;/p&gt;

&lt;p&gt;That's the whole method. The interesting part is reading the results.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the control
&lt;/h2&gt;

&lt;p&gt;The control runs first, and it's a gate rather than a finding.&lt;/p&gt;

&lt;p&gt;If the control doesn't route correctly, &lt;strong&gt;stop&lt;/strong&gt; — you don't have an ambiguity problem between this pair. You have one description that's broken on its own terms, and that's a different fix. Rewriting it in relation to its sibling would be solving the wrong problem.&lt;/p&gt;

&lt;p&gt;Most of the time the control passes and you move on. When it fails, it's usually worth more than anything the probe would have told you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the probe
&lt;/h2&gt;

&lt;p&gt;This is where the instinct misleads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If the probe routes to the same tool every time&lt;/strong&gt;, that's a &lt;strong&gt;pass&lt;/strong&gt;, even though the sentence was genuinely ambiguous to you. It means the model has found a distinction you can't see in the words. Maybe it's the parameter names. Maybe it's the tool name. Maybe it's something about the phrasing you didn't consciously encode. Doesn't matter — it works.&lt;/p&gt;

&lt;p&gt;My instinct would have been to flag any non-uniform distribution as a collision. That's wrong, and acting on it would mean rewriting descriptions that are already doing their job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If the probe splits&lt;/strong&gt;, that's a real collision, and you now have the sentence that proves it. Not a suspicion. Not a metric above a threshold. An actual user request that lands in two different places depending on the roll.&lt;/p&gt;

&lt;p&gt;That distinction — stable versus split — is the whole reason this is worth running. And a handful of runs at fixed temperature is enough to tell them apart, which is a far smaller budget than I'd assumed before someone spelled it out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Picking the pairs
&lt;/h2&gt;

&lt;p&gt;You can't do this pairwise on a large server. 191 tools is 18,145 pairs, and at even a handful of model calls each, that's not an afternoon.&lt;/p&gt;

&lt;p&gt;This is where static analysis earns its place. It stops being the answer and becomes the filter that produces candidates.&lt;/p&gt;

&lt;p&gt;Rank pairs by &lt;strong&gt;input overlap&lt;/strong&gt; — do these two tools act on the same object? Compare the nouns in their descriptions plus their parameter names. Tools that share their input are the ones where a single sentence can route to either, which is exactly the failure the probe tests for.&lt;/p&gt;

&lt;p&gt;Then gate that on &lt;strong&gt;output distinguishability&lt;/strong&gt;. If two tools take the same input but their descriptions clearly state different returns — "flat details for many records" versus "a bounded evidence bundle" — that's good design, not a problem. Don't spend probes on it.&lt;/p&gt;

&lt;p&gt;Rank, take the top fifty, probe those. The static metric was never going to find the collision; it's very good at telling you where to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with a split
&lt;/h2&gt;

&lt;p&gt;Three outcomes, and they're genuinely different kinds of work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rewrite for contrast.&lt;/strong&gt; The cheapest fix and it changes no behaviour. The rule that works: every tool in a cluster shares its input by definition — that's what makes them a cluster — so describing the input describes what they have in common. Name the &lt;strong&gt;output&lt;/strong&gt; instead. "Search reports" describes ten tools. "Search reports by client, including archived ones, returning IDs and titles only" describes one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;State the trigger condition.&lt;/strong&gt; Call this one before acting, that one after a decision is made. No lexical measure sees this and it's frequently the real distinction. It's also the part a description template flattens, which is why collision rates climb sharply past about thirty tools — that's where people stop writing descriptions individually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Merge into one tool with a mode parameter.&lt;/strong&gt; Sometimes the split isn't telling you your wording is bad. It's telling you the sentence names an intent your tools have partitioned and the user hasn't. "Check my resume for this job" is not a failure of either description — it's a request that legitimately spans both. If two tools can ever be correct for the same sentence, you may not have two tools. You may have one tool with a parameter.&lt;/p&gt;

&lt;p&gt;That last one is the conclusion I've now reached from two completely different directions: as a design rule about interfaces, and empirically from a probe that splits. The second version is the one you can put in a pull request without arguing, because it comes with the sentence attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the failures
&lt;/h2&gt;

&lt;p&gt;Every split is a regression case: the prompt, the expected tool, the tool actually picked.&lt;/p&gt;

&lt;p&gt;Store them. Then a description rewrite stops being hopeful and becomes testable — you re-run the probe and see whether the split closed. And the next person who "improves" the wording six months from now finds out immediately if they've broken the distinction.&lt;/p&gt;

&lt;p&gt;This is the thing missing from every schema-linting approach, mine included. A linter tells you a description looks risky. A regression suite tells you whether the fix worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The limits
&lt;/h2&gt;

&lt;p&gt;Fixed temperature and a handful of runs gives you a signal, not a rate. If you want "this pair splits 60/40", that needs many more runs and it's probably not worth the budget — the useful information is binary.&lt;/p&gt;

&lt;p&gt;Different models may split differently, which is itself worth knowing. If a pair is stable in one client and splits in another, the problem is more specific than your descriptions.&lt;/p&gt;

&lt;p&gt;And a probe that doesn't split tells you about that sentence, not about every sentence. Someone will eventually phrase something you didn't think of. That's an argument for keeping the regression suite and adding to it, not for distrusting the method.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is the interesting layer
&lt;/h2&gt;

&lt;p&gt;I've published a fair amount of static analysis on MCP schemas: distinctive share, input overlap, name distinctiveness, parameter coverage. All of it measures properties of text.&lt;/p&gt;

&lt;p&gt;None of it observes a model choosing a tool. Which means every number reports how often a pattern &lt;em&gt;exists&lt;/em&gt;, not how often it costs you a call — and the gap between those two is where all the value is.&lt;/p&gt;

&lt;p&gt;The probe closes that gap for a pair at a time, cheaply, before you have any production traffic at all. If you maintain an MCP server and you've been meaning to audit your descriptions, this is the thing I'd do first.&lt;/p&gt;




&lt;p&gt;I build &lt;a href="https://getmcpulse.com" rel="noopener noreferrer"&gt;MCPulse&lt;/a&gt;, which reports what models actually do with your tools once real traffic arrives — retries, empty results, first-call success, schema cost. There's also a &lt;a href="https://getmcpulse.com/check" rel="noopener noreferrer"&gt;free schema checker&lt;/a&gt; that ranks your tool pairs by input overlap, which is the prefilter step above.&lt;/p&gt;

&lt;p&gt;The control-and-probe design came from readers of a previous post. Best methodology suggestion I've had, and it came from people who run servers rather than measure them.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>llm</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Readers took my MCP schema study apart. Here's what they found.</title>
      <dc:creator>MCPulse</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:19:16 +0000</pubDate>
      <link>https://dev.to/getmcpulse/readers-took-my-mcp-schema-study-apart-heres-what-they-found-d40</link>
      <guid>https://dev.to/getmcpulse/readers-took-my-mcp-schema-study-apart-heres-what-they-found-d40</guid>
      <description>&lt;p&gt;I published a survey of 4,749 public MCP server schemas a few weeks ago. The headline was that 17.7% of tools carry a description containing no word that distinguishes them from a sibling tool on the same server.&lt;/p&gt;

&lt;p&gt;Then people who actually run MCP servers showed up in the comments, and over about ten days they took the measurement apart.&lt;/p&gt;

&lt;p&gt;One found a bug. Several found a design flaw. Two of them found opposite design flaws that cancel each other out, which was the most useful thing that happened. And one found a ceiling the whole approach can't get past.&lt;/p&gt;

&lt;p&gt;Here's all of it, because the corrections are worth more than the original number.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug: no stemming
&lt;/h2&gt;

&lt;p&gt;The metric counts content words in a tool's description that appear in no other tool's description on the same server. It does exact string matching.&lt;/p&gt;

&lt;p&gt;Someone pointed at two descriptions from his own server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;analyze_job_description   "Extract what a job posting actually screens on."
optimize_resume           "Rewrite a resume so it passes ATS screening."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;screens&lt;/code&gt; and &lt;code&gt;screening&lt;/code&gt;. Same word, two forms, and my counter treats them as unrelated — so these two tools score as distinctive on the one word that actually links them.&lt;/p&gt;

&lt;p&gt;That's not a nuance. It's wrong, and it's wrong across all 82,549 tools. Stemming before counting is a small change with an unknown effect on the headline figure, which I'll report when I've re-run it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The design flaw: lexical distinctness isn't ambiguity
&lt;/h2&gt;

&lt;p&gt;The same author gave me his full ten-tool list and said three of them were confusable in practice. So I ran the metric on it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; 50.0%  score_resume              "Score a resume for ATS compatibility."
 60.0%  analyze_job_description   "Extract what a job posting actually screens on."
 60.0%  optimize_resume           "Rewrite a resume so it passes ATS screening."
 75.0%  search_jobs               "Return the user's job matches."
100.0%  update_job_preferences    "Set the roles and locations it hunts for."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;His ambiguous cluster — score, analyze, optimize — comes out at 50–60%. My corpus calls that healthy.&lt;/p&gt;

&lt;p&gt;My first fix was to compute distinctiveness on verbs alone, on the theory that verbs carry the action and nouns are shared boilerplate.&lt;/p&gt;

&lt;p&gt;That fails harder. His ten verbs are score, extract, rewrite, produce, write, translate, return, set and build. All distinct. Verb-only distinctiveness scores every tool at 100%, including all three ambiguous ones.&lt;/p&gt;

&lt;p&gt;The problem isn't granularity. Lexical distinctness and semantic distinctness are different properties, and no word-counting metric at any resolution closes the gap.&lt;/p&gt;

&lt;p&gt;What does show the cluster is the shared input nouns — &lt;code&gt;resume&lt;/code&gt; in four descriptions, &lt;code&gt;job&lt;/code&gt; in three, &lt;code&gt;ats&lt;/code&gt; in two. Those three tools are ambiguous because they act on the same object, not because they're worded alike.&lt;/p&gt;

&lt;p&gt;His rule, better than my number: &lt;strong&gt;if two tools can ever be correct for the same sentence, you don't have two tools. You have one tool with a parameter.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So the fix isn't a better distinctive share. It's a second signal, measuring something else entirely. Distinctive share asks whether two descriptions look alike. Input overlap asks whether two tools could both be right.&lt;/p&gt;

&lt;h2&gt;
  
  
  The correction to the correction: overlap alone flags good design
&lt;/h2&gt;

&lt;p&gt;Someone running a six-tool ticket server pushed back, with a pair he'd deliberately kept:&lt;/p&gt;

&lt;p&gt;Batch retrieval and an analysis bundle. Identical input shape. Different intended output — flat details for more tickets, versus bounded evidence with optional comments and source anchors.&lt;/p&gt;

&lt;p&gt;Input overlap would flag that pair. He'd be right to dismiss it, because the boundary is stated: a model reading those two descriptions has something to discriminate on even though the inputs are the same.&lt;/p&gt;

&lt;p&gt;So overlap needs a partner. Overlapping inputs with distinguishable outputs is fine design. Overlapping inputs with indistinguishable outputs is the failure. Flagging the first is exactly the false positive that gets a signal switched off — and on a large server, a 10% false positive rate buries everything.&lt;/p&gt;

&lt;p&gt;His test is better than any metric I can compute: &lt;strong&gt;can a model infer the intended output and boundary from the name, description and parameter contract alone, without knowing the repo or the author's intent?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The two corrections that fight each other
&lt;/h2&gt;

&lt;p&gt;This is the pair I'd have got wrong silently.&lt;/p&gt;

&lt;p&gt;One reader pointed out that the metric penalises focused servers. A memory server says "memory" in every description by design. A ticket server says "ticket". Those words count as shared, but they aren't a collision — they're the name of the server, repeated. Drop words appearing in more than 80% of a server's tools before counting, and small focused servers score higher.&lt;/p&gt;

&lt;p&gt;Correct, and I was about to implement it globally.&lt;/p&gt;

&lt;p&gt;Then the ten-tool author showed why that breaks the &lt;em&gt;other&lt;/em&gt; signal. On a focused server, the collision &lt;strong&gt;is&lt;/strong&gt; the domain noun. Strip &lt;code&gt;resume&lt;/code&gt; from his descriptions and you delete the exact words that make score, optimize and analyze confusable. I'd have shipped a sharper distinctive-share number that quietly destroyed input overlap.&lt;/p&gt;

&lt;p&gt;The resolution is two readings with different preprocessing, not one corrected score:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Domain-stripped distinctiveness&lt;/strong&gt; — do these descriptions look alike?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Domain-intact input overlap&lt;/strong&gt; — could two tools be right for the same sentence?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the gap between them is the finding. A server that scores high on the first and badly on the second is focused, acts on one object throughout, and never says what its tools return. That's a common shape and it deserves its own line rather than being averaged away.&lt;/p&gt;

&lt;p&gt;A third reader then improved the threshold out of existence. Frequency isn't a dial to tune — it's a distinction the same count already makes. A noun in nearly every tool names the server. A noun in a &lt;em&gt;subset&lt;/em&gt; of tools names a cluster, and that subset is your candidate collision list.&lt;/p&gt;

&lt;p&gt;Run it on the ten-tool server and the tightest group is &lt;code&gt;ats&lt;/code&gt;, shared by score and optimize and nothing else. Which is the pair he flagged first. Same counter, no threshold, no tuning.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ceiling
&lt;/h2&gt;

&lt;p&gt;The same reader then found the thing none of this reaches.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;analyze_job_description&lt;/code&gt; and &lt;code&gt;score_resume&lt;/code&gt; share no word at all, stemmed or not. And they still collide, because "check my resume for this job" names both objects at once. The tools partition an intent that the sentence doesn't.&lt;/p&gt;

&lt;p&gt;That collision exists between a request and a set of tools, not between two descriptions. Nothing computed on descriptions alone will ever see it.&lt;/p&gt;

&lt;p&gt;That's a hard ceiling on static analysis, not a gap to close with a cleverer metric. The only artefact that catches it is a small labelled set of real requests — a handful of plausible sentences and which tool each should route to. Which is the one thing a corpus can't generate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I hadn't measured at all
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Names.&lt;/strong&gt; Two people independently said that when descriptions collide, models fall back to pattern-matching on tool names. One put it: "the model stops reading descriptions and pattern-matches on names, and then your error rate is really a naming problem."&lt;/p&gt;

&lt;p&gt;That changes what a collision means. &lt;code&gt;Gmail_DeleteDraftEmail&lt;/code&gt; and &lt;code&gt;Gmail_SendDraftEmail&lt;/code&gt; have zero-distinctive descriptions — that pair was my headline example — but their names are clear. The names are carrying the load. A hypothetical pair with the same descriptions and names like &lt;code&gt;delete_item&lt;/code&gt; and &lt;code&gt;remove_item&lt;/code&gt; scores identically and is in far worse shape.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Annotations.&lt;/strong&gt; MCP has &lt;code&gt;readOnlyHint&lt;/code&gt;, &lt;code&gt;destructiveHint&lt;/code&gt;, &lt;code&gt;idempotentHint&lt;/code&gt;, &lt;code&gt;openWorldHint&lt;/code&gt;. Two readers raised the same point: authors set them &lt;em&gt;and&lt;/em&gt; write "Read-only" and "REPLACES" in the prose by hand, because the annotations are read by the client deciding whether to prompt for confirmation, while the description is read by the model deciding whether to call the thing. Two readers, two channels, and nothing requires them to agree.&lt;/p&gt;

&lt;p&gt;Which makes a mechanical check available: descriptions mentioning constraints with no matching annotation set, and annotations set with nothing in the prose. The first means the client can't protect the user. The second means the model doesn't know a tool is destructive at the moment it's choosing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd take from this
&lt;/h2&gt;

&lt;p&gt;Every correction came from someone who knew what the text was &lt;em&gt;for&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I measured a property of the text and treated it as a proxy for ambiguity. Lexical distinctness looked like a proxy until someone showed me ten tools where it isn't. Populated descriptions looked like described parameters until someone pointed at one that says nothing. Descriptions looked like the whole interface until two people said names do the work when descriptions fail.&lt;/p&gt;

&lt;p&gt;The original study was careful about one thing: it drew a hard line between what a schema shows and what models do with it, and refused to claim the second. That line held. What didn't hold was the assumption that measuring the text well is the same as measuring the thing that matters.&lt;/p&gt;

&lt;p&gt;The published numbers stand as published — they measure what they say they measure, and they're all floors. The more useful metrics are the ones nobody had asked for yet.&lt;/p&gt;




&lt;p&gt;Original study, data and analysis scripts: &lt;a href="https://github.com/getmcpulse/mcp-schema-study" rel="noopener noreferrer"&gt;https://github.com/getmcpulse/mcp-schema-study&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://getmcpulse.com/check" rel="noopener noreferrer"&gt;schema checker&lt;/a&gt; runs these measurements on your own &lt;code&gt;tools/list&lt;/code&gt; in the browser. The corrections above are being added to it.&lt;/p&gt;

&lt;p&gt;I build &lt;a href="https://getmcpulse.com" rel="noopener noreferrer"&gt;MCPulse&lt;/a&gt;, an SDK that reports what models actually do with your tools under real traffic. If something here matches what you've seen on your own server, I'd like to hear it — every improvement above came from exactly that.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>llm</category>
      <category>devtools</category>
    </item>
    <item>
      <title>How deep do MCP input schemas nest?</title>
      <dc:creator>MCPulse</dc:creator>
      <pubDate>Fri, 18 Sep 2026 07:18:09 +0000</pubDate>
      <link>https://dev.to/getmcpulse/how-deep-do-mcp-input-schemas-nest-3lb0</link>
      <guid>https://dev.to/getmcpulse/how-deep-do-mcp-input-schemas-nest-3lb0</guid>
      <description>&lt;p&gt;A model filling in a tool call has to construct whatever shape your &lt;code&gt;inputSchema&lt;/code&gt; describes. A flat object of strings is one thing; an array of objects each containing an object is another.&lt;/p&gt;

&lt;p&gt;So how nested are real MCP schemas? We measured depth across 74,666 tools that take parameters, following &lt;code&gt;properties&lt;/code&gt;, &lt;code&gt;items&lt;/code&gt;, and &lt;code&gt;anyOf&lt;/code&gt;/&lt;code&gt;oneOf&lt;/code&gt;/&lt;code&gt;allOf&lt;/code&gt; branches.&lt;/p&gt;

&lt;p&gt;Then we split the undescribed-parameter rate by depth, expecting deep schemas to be worse. They're better, and the reason is more interesting than the number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Almost everything is flat
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Depth&lt;/th&gt;
&lt;th&gt;Tools&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;62,054&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;83.1%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;9,279&lt;/td&gt;
&lt;td&gt;12.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;2,333&lt;/td&gt;
&lt;td&gt;3.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;599&lt;/td&gt;
&lt;td&gt;0.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;216&lt;/td&gt;
&lt;td&gt;0.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;77&lt;/td&gt;
&lt;td&gt;0.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7–13&lt;/td&gt;
&lt;td&gt;~108&lt;/td&gt;
&lt;td&gt;&amp;lt;0.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Five tools in six are a flat object of scalars.&lt;/strong&gt; No nesting at all — &lt;code&gt;{"query": string, "limit": integer}&lt;/code&gt; and nothing more.&lt;/p&gt;

&lt;p&gt;Add depth 2 and you have 95.5% of the corpus. Depth 2 is the array-of-strings and the single options object: &lt;code&gt;{"tags": string[]}&lt;/code&gt;, &lt;code&gt;{"filter": {"from": string}}&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The deepest schema in the corpus is thirteen levels. There are five of those.&lt;/p&gt;

&lt;h2&gt;
  
  
  The finding that reverses the expectation
&lt;/h2&gt;

&lt;p&gt;We expected deep schemas to be worse documented — more structure, more fields, more places to skip a description.&lt;/p&gt;

&lt;p&gt;The opposite:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tools at depth 3 or more: &lt;strong&gt;18.2%&lt;/strong&gt; of top-level parameters undescribed&lt;/li&gt;
&lt;li&gt;Tools below depth 3: &lt;strong&gt;21.8%&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Deep schemas are better documented than flat ones, by 3.6 points.&lt;/p&gt;

&lt;p&gt;Our reading is that it's about who writes them. A depth-5 schema isn't something you arrive at casually. It's either generated from a typed source — an OpenAPI spec, a Zod schema, a protobuf definition — where descriptions come along for the ride, or hand-built by someone modelling a genuinely structured input and paying attention.&lt;/p&gt;

&lt;p&gt;Flat schemas are where the two-line afterthought lives. The single-parameter tool is the worst-described shape in the whole corpus, at 27.0%.&lt;/p&gt;

&lt;p&gt;So depth isn't a warning sign. It correlates with care.&lt;/p&gt;

&lt;h2&gt;
  
  
  The caveat that matters more than the finding
&lt;/h2&gt;

&lt;p&gt;That comparison counts &lt;strong&gt;top-level&lt;/strong&gt; &lt;code&gt;inputSchema.properties&lt;/code&gt; only. A field three levels down isn't in the denominator.&lt;/p&gt;

&lt;p&gt;That was the right call for the survey — per-tool parameter counts had to mean one consistent thing — but it means the corpus number understates the real gap. A tool whose top-level parameters are all described and whose nested object is completely bare scores clean.&lt;/p&gt;

&lt;p&gt;And an undescribed field at depth 4 is exactly as invisible to a model as one at depth 1. It has a name, a type, and nothing else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means practically
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Don't flatten a schema to look simpler.&lt;/strong&gt; Depth isn't the problem the data points at. If your input genuinely has structure — a filter object, a list of line items — modelling it honestly beats a dozen underscore-joined top-level fields. The corpus suggests people who do it describe their fields better anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do describe the nested fields.&lt;/strong&gt; They're the ones most likely to be missed, because most tooling shows you the top level. This is the one place where our own methodology would let you off and a model wouldn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Past about four levels, ask what the model is meant to construct.&lt;/strong&gt; 0.4% of the corpus goes deeper than four. A model generating a five-level nested object has many more ways to get the shape wrong than right, and there's usually a flatter representation of the same request. Not always — but at that depth it's worth checking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Watch &lt;code&gt;anyOf&lt;/code&gt; and &lt;code&gt;oneOf&lt;/code&gt; specifically.&lt;/strong&gt; They're depth without looking like depth. A parameter that is "either a string or an object with three fields" is two shapes the model has to choose between, and the choice is rarely documented. A &lt;code&gt;description&lt;/code&gt; on the branch point is worth more than one on either branch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this can't tell you
&lt;/h2&gt;

&lt;p&gt;We measured depth, not difficulty. A depth-3 schema of well-named described fields is easier for a model than a flat one with two bare &lt;code&gt;id&lt;/code&gt; parameters, and nothing here captures that.&lt;/p&gt;

&lt;p&gt;The correlation between depth and better documentation is a correlation. The generated-from-typed-source explanation is our reading of it, not something the data shows — we can't see how a schema was produced.&lt;/p&gt;

&lt;p&gt;And as always: no model was run against any of these servers. Whether nesting depth actually causes malformed arguments is unmeasured, and it's one of the more testable things on our list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring your own
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://getmcpulse.com/check" rel="noopener noreferrer"&gt;schema checker&lt;/a&gt; walks your schema to every depth — through &lt;code&gt;properties&lt;/code&gt;, &lt;code&gt;items&lt;/code&gt; and union branches — and names undescribed fields by their full path: &lt;code&gt;filter.from&lt;/code&gt;, &lt;code&gt;tags[].label&lt;/code&gt;. It reports the top-level count separately, so the number scored against the corpus stays comparable while the list of things to fix doesn't leave anything out.&lt;/p&gt;

&lt;p&gt;Runs in your browser, nothing uploaded. Data in &lt;a href="https://github.com/getmcpulse/mcp-schema-study" rel="noopener noreferrer"&gt;getmcpulse/mcp-schema-study&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://getmcpulse.com/blog/how-deep-mcp-schemas-nest" rel="noopener noreferrer"&gt;getmcpulse.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>mcp</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Why the model won't call your tool</title>
      <dc:creator>MCPulse</dc:creator>
      <pubDate>Wed, 16 Sep 2026 08:26:29 +0000</pubDate>
      <link>https://dev.to/getmcpulse/why-the-model-wont-call-your-tool-4k43</link>
      <guid>https://dev.to/getmcpulse/why-the-model-wont-call-your-tool-4k43</guid>
      <description>&lt;p&gt;Your tool is registered. It appears in &lt;code&gt;tools/list&lt;/code&gt;. The model never calls it, or calls a different one instead.&lt;/p&gt;

&lt;p&gt;There are six reasons this happens. They are not equally likely, and the instinct — assume the description needs to be better — is right about a third of the time.&lt;/p&gt;

&lt;p&gt;Here they are in order of how often they turn up across 4,749 public MCP servers.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Another tool is winning a competition you cannot see
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Most likely.&lt;/strong&gt; 17.7% of public tools have a description containing no word that distinguishes them from a sibling. On servers with more than sixty tools it is 32.4%.&lt;/p&gt;

&lt;p&gt;The tell is that a &lt;em&gt;neighbouring&lt;/em&gt; tool gets called instead of yours, consistently. Not randomly — the same substitution every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to check.&lt;/strong&gt; Print your whole tool list and read your tool's description next to the one that keeps getting called. If you can swap the two descriptions and both still read correctly, the model has nothing to choose on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix.&lt;/strong&gt; Lead with the difference, not the category. "Search orders" describes four of your tools. "Search orders by customer, including cancelled ones" describes one. Naming the return shape is the single most distinguishing sentence available and almost nobody writes it.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The parameters are unfillable
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Second most likely&lt;/strong&gt;, and the one people check last. 21.5% of parameters across the corpus have no description at all.&lt;/p&gt;

&lt;p&gt;A model that cannot work out what to put in a required field will avoid the tool rather than guess. This looks identical to a description problem from outside, and it is not.&lt;/p&gt;

&lt;p&gt;The clearest version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;fetch    id, document_id      (neither described)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A model has to pick one and nothing tells it which, or whether they differ. So it picks a different tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to check.&lt;/strong&gt; Cover the parameter &lt;em&gt;names&lt;/em&gt; and look only at types and descriptions. If you could not supply a value, neither can the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Watch for the internal identifier specifically.&lt;/strong&gt; A tool requiring a &lt;code&gt;warehouse_id&lt;/code&gt; is unreachable from a conversation where nobody has ever seen a warehouse ID. That is not a description problem — it is a tool that needs a lookup step before it, or a name-based alternative.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The name is doing all the work and losing
&lt;/h2&gt;

&lt;p&gt;When descriptions collide, the name is what decides — in 89.6% of collision cases. So a colliding description plus a generic name is the case where nothing is choosing at all.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;get&lt;/code&gt; leads &lt;strong&gt;15.6%&lt;/strong&gt; of all public tool names. If your tool is &lt;code&gt;get_thing&lt;/code&gt; and a sibling is &lt;code&gt;fetch_thing&lt;/code&gt;, those are synonyms in English and the model is guessing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix.&lt;/strong&gt; One verb per operation, across the whole server. Never ship two words from the same row:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;get&lt;/code&gt; / &lt;code&gt;fetch&lt;/code&gt; / &lt;code&gt;read&lt;/code&gt; / &lt;code&gt;retrieve&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;list&lt;/code&gt; / &lt;code&gt;getAll&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;search&lt;/code&gt; / &lt;code&gt;query&lt;/code&gt; / &lt;code&gt;find&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;delete&lt;/code&gt; / &lt;code&gt;remove&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Your tool list is too long to hold
&lt;/h2&gt;

&lt;p&gt;Schema goes into the context window on every connection. The median server spends about 1,251 tokens; the 90th percentile spends 7,992. The largest in the corpus spends around 216,000 — which does not fit a 200K window at all.&lt;/p&gt;

&lt;p&gt;The failure here is not that your tool is badly described. It is that it is the fortieth candidate in a list the model is skimming, and attention is finite.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to check.&lt;/strong&gt; Count your tools. Under fifteen, this is not your problem. Past thirty, it is probably your main one.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. A closed set was written in prose
&lt;/h2&gt;

&lt;p&gt;The model sends &lt;code&gt;"Direct"&lt;/code&gt; where you accept &lt;code&gt;"direct"&lt;/code&gt;, your server rejects it, and from the model's side that reads as a broken tool rather than a fixable mistake — so it stops trying.&lt;/p&gt;

&lt;p&gt;857 parameters in the corpus name their valid values in the description and leave the schema as an open string. Moving them into an &lt;code&gt;enum&lt;/code&gt; is one line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tell for this one is distinctive:&lt;/strong&gt; the tool &lt;em&gt;does&lt;/em&gt; get called, once, and then never again in that session.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. It is actually unreachable
&lt;/h2&gt;

&lt;p&gt;Two rarer causes worth ruling out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A duplicate tool name.&lt;/strong&gt; Only 0.19% of servers — nine in the whole corpus — but when it happens, one of the two tools cannot be called at all, and the symptom points somewhere else entirely. Almost always a generated tool list. One assertion in CI catches it forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tool is not in the response you think it is.&lt;/strong&gt; Servers that build their list conditionally can emit something different from what the code appears to register. Read the real &lt;code&gt;tools/list&lt;/code&gt; output rather than the registration code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The order to actually work through
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Read your whole tool list as one block. Which tool would &lt;em&gt;you&lt;/em&gt; pick? — catches causes 1 and 3.&lt;/li&gt;
&lt;li&gt;Cover the parameter names, read only types and descriptions — catches cause 2.&lt;/li&gt;
&lt;li&gt;Count your tools — catches cause 4.&lt;/li&gt;
&lt;li&gt;Grep descriptions for &lt;code&gt;valid values&lt;/code&gt; / &lt;code&gt;options:&lt;/code&gt; — catches cause 5.&lt;/li&gt;
&lt;li&gt;Assert unique names in CI — catches cause 6.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is fifteen minutes and it covers all six.&lt;/p&gt;

&lt;h2&gt;
  
  
  What none of this can tell you
&lt;/h2&gt;

&lt;p&gt;Every cause above is a property of your schema, and this is the honest limit: we have never observed a model choosing a tool on any of these servers. The rates say how often each pattern &lt;em&gt;exists&lt;/em&gt;, not how often it costs you a call.&lt;/p&gt;

&lt;p&gt;It is entirely possible your tool is skipped for a reason no schema shows — the client truncates long lists, the model was mid-task, the user phrased something unusually. Distinguishing those from the six above needs data from inside your server: which tools get called, which get retried, which return empty. A schema check is what you can do before you have any of that.&lt;/p&gt;




&lt;p&gt;The &lt;a href="https://getmcpulse.com/check" rel="noopener noreferrer"&gt;schema checker&lt;/a&gt; tests causes 1, 2, 3, 5 and 6 on a paste of your &lt;code&gt;tools/list&lt;/code&gt;, scored against servers your own size. It names the sibling each colliding tool is hardest to tell apart from — which for cause 1 is usually the whole answer. Runs in your browser, nothing uploaded.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://getmcpulse.com/blog/why-the-model-wont-call-your-tool" rel="noopener noreferrer"&gt;getmcpulse.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>devtools</category>
      <category>llm</category>
    </item>
    <item>
      <title>How many tools should an MCP server have?</title>
      <dc:creator>MCPulse</dc:creator>
      <pubDate>Mon, 14 Sep 2026 06:52:52 +0000</pubDate>
      <link>https://dev.to/getmcpulse/how-many-tools-should-an-mcp-server-have-2ang</link>
      <guid>https://dev.to/getmcpulse/how-many-tools-should-an-mcp-server-have-2ang</guid>
      <description>&lt;p&gt;If you maintain an MCP server, at some point you ask this. You've got twenty tools, you're about to add five more, and something feels wrong about it — but you can't say what, and there's no guidance anywhere.&lt;/p&gt;

&lt;p&gt;I read the tool schemas of 4,951 public MCP servers to answer it. 87,146 tools, 270,487 parameters.&lt;/p&gt;

&lt;p&gt;The short answer: around thirty. Past that, the thing that breaks isn't your server. It's whether a model can tell your tools apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  The measurement
&lt;/h2&gt;

&lt;p&gt;For every tool, I took the content words in its description and counted how many appear in no other tool's description on the same server. Call it the tool's distinctive share.&lt;/p&gt;

&lt;p&gt;If that number is zero, every word in the description is a word its siblings also use. A model choosing between your tools has nothing in the descriptions to choose on — the names are doing all the work.&lt;/p&gt;

&lt;p&gt;Then I split the corpus by how many tools each server publishes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tools on the server&lt;/th&gt;
&lt;th&gt;Servers&lt;/th&gt;
&lt;th&gt;Zero-distinctive tools&lt;/th&gt;
&lt;th&gt;Parameters with no description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1–3&lt;/td&gt;
&lt;td&gt;1,256&lt;/td&gt;
&lt;td&gt;0.5%&lt;/td&gt;
&lt;td&gt;14.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4–7&lt;/td&gt;
&lt;td&gt;1,292&lt;/td&gt;
&lt;td&gt;1.6%&lt;/td&gt;
&lt;td&gt;22.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8–15&lt;/td&gt;
&lt;td&gt;1,044&lt;/td&gt;
&lt;td&gt;4.3%&lt;/td&gt;
&lt;td&gt;21.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16–30&lt;/td&gt;
&lt;td&gt;767&lt;/td&gt;
&lt;td&gt;7.7%&lt;/td&gt;
&lt;td&gt;24.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31–60&lt;/td&gt;
&lt;td&gt;377&lt;/td&gt;
&lt;td&gt;16.3%&lt;/td&gt;
&lt;td&gt;22.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;61+&lt;/td&gt;
&lt;td&gt;215&lt;/td&gt;
&lt;td&gt;31.3%&lt;/td&gt;
&lt;td&gt;20.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A factor of sixty, rising monotonically. On servers with more than sixty tools, nearly one tool in three has no distinguishing word at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it gets worse
&lt;/h2&gt;

&lt;p&gt;Part of it is arithmetic. More tools means more chances that two of them collide, and that would happen even if every author wrote carefully.&lt;/p&gt;

&lt;p&gt;But the curve steepens around thirty, and arithmetic alone doesn't explain that. What happens around thirty is that authors stop writing descriptions one at a time and start generating them from a pattern. A template is a machine for producing tools that read alike.&lt;/p&gt;

&lt;p&gt;The most extreme case in the corpus: one server appends the same 51-word context block to all 275 of its tools.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;3land_createCollection   "Create a new NFT collection on 3.Land marketplace.
                          SAP MCP context: Protocol 3land; operation class
                          write. Use for 3.Land NFT collection, minting,
                          listing, cancellation, and purchase flows…"

3land_buyNFT             "Purchase an NFT from a 3.Land listing.
                          SAP MCP context: Protocol 3land; operation class
                          write. Use for 3.Land NFT collection, minting,
                          listing, cancellation, and purchase flows…"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Creating a collection and buying one are different operations, and the opening sentence says so — in 8 words out of 59. The other 51 are identical across both, and across all 275.&lt;/p&gt;

&lt;p&gt;That block was added deliberately, to help.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zero doesn't mean badly written
&lt;/h2&gt;

&lt;p&gt;This is the part worth internalising, because it's counterintuitive.&lt;/p&gt;

&lt;p&gt;Four tools from a widely-installed Gmail server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gmail_DeleteDraftEmail   "Delete a draft email using the Gmail API."
Gmail_SendDraftEmail     "Send a draft email using the Gmail API."
Gmail_ListLabels         "List all the labels in the user's mailbox."
Gmail_SearchThreads      "Search for threads in the user's mailbox."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every one of those is clear, correct English. No reviewer would flag them. Every one is also built entirely out of words the other tools use — delete, draft, email, gmail, api, list, search, threads, mailbox all recur across the set.&lt;/p&gt;

&lt;p&gt;The description tells you what the tool does. It doesn't tell you what &lt;em&gt;this&lt;/em&gt; tool does and the others don't. That second thing is what a model needs at the moment it's choosing, and it's a different question from "is this description good."&lt;/p&gt;

&lt;h2&gt;
  
  
  The column that argues against splitting
&lt;/h2&gt;

&lt;p&gt;Look at the right-hand column again. Parameters with no description at all sit between 20% and 25% at every size above the smallest bucket. It doesn't improve as servers get smaller.&lt;/p&gt;

&lt;p&gt;A two-tool server has the habit about as much as a two-hundred-tool one.&lt;/p&gt;

&lt;p&gt;So the two failures are independent. Splitting a large server reduces your description collisions and does nothing whatsoever for your undescribed parameters. They need separate fixes, and the split only buys you one of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  So should you split?
&lt;/h2&gt;

&lt;p&gt;Split when the collisions are real, not because you crossed a number.&lt;/p&gt;

&lt;p&gt;The threshold in the data is around thirty, but that's a population average and your server isn't the population. A server with forty tools that all do genuinely different things to genuinely different objects may be fine. A server with twelve tools where four of them are variations on "search" is not.&lt;/p&gt;

&lt;p&gt;The test that actually tells you: read your tool list as one block, the way a model receives it. Nothing else. No README, no repo, no memory of what you meant. Then ask which tool you'd pick for a request that could plausibly go to two of them.&lt;/p&gt;

&lt;p&gt;Three options when the answer is "I can't tell":&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rewrite for contrast rather than clarity.&lt;/strong&gt; Not "is this description clear" but "is it clear which of my tools this is." If you have a shared preamble on every tool, it's costing more than it's buying — the distinguishing sentence shouldn't be a seventh of the text.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Collapse near-identical tools into one with a mode parameter.&lt;/strong&gt; Four search variants become one &lt;code&gt;search&lt;/code&gt; with a &lt;code&gt;scope&lt;/code&gt; enum. Fewer things to choose between, and the choice the model has to make moves from "which tool" to "which value," which an enum can constrain and a description can't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Split the server.&lt;/strong&gt; Fewer tools per connection, more servers to maintain. Worth it when the tools genuinely belong to different domains, less so when you're just cutting an arbitrary list in half.&lt;/p&gt;

&lt;h2&gt;
  
  
  The limits of this
&lt;/h2&gt;

&lt;p&gt;This is static analysis. I never ran a model against any of these servers, so I can't tell you how often collisions actually cost anything. A zero-distinctive description might be harmless when the tool name is unambiguous, and expensive when it isn't. Ranking these signals by how well they predict a real mistake needs a model in the loop, which is the next study.&lt;/p&gt;

&lt;p&gt;One more hole, found by a server author after I published: the undescribed-parameter figure measures &lt;em&gt;absence&lt;/em&gt; only. A parameter described as &lt;code&gt;"query: The query"&lt;/code&gt; counts as described and passes. Restating the parameter name is arguably the more common failure and it passes every linter, so 21.8% is a floor.&lt;/p&gt;




&lt;p&gt;Data and analysis scripts: &lt;a href="https://github.com/getmcpulse/mcp-schema-study" rel="noopener noreferrer"&gt;https://github.com/getmcpulse/mcp-schema-study&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you want your own numbers rather than the corpus averages, there's a free checker at &lt;a href="https://getmcpulse.com/check" rel="noopener noreferrer"&gt;https://getmcpulse.com/check&lt;/a&gt; — paste your &lt;code&gt;tools/list&lt;/code&gt; JSON and it scores your distinctive share, undescribed parameters, and token cost against all 4,951 servers. Browser only, nothing uploaded.&lt;/p&gt;

&lt;p&gt;I'm building &lt;a href="https://getmcpulse.com" rel="noopener noreferrer"&gt;MCPulse&lt;/a&gt;, an SDK that reports what models actually do with your tools once real traffic arrives. Everything above came from outside the server, which is exactly its limit: a schema can tell you a model has nothing to choose on, but only traffic tells you whether it chose wrong.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://getmcpulse.com/blog/should-you-split-a-large-mcp-server" rel="noopener noreferrer"&gt;getmcpulse.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>llm</category>
      <category>devtools</category>
    </item>
    <item>
      <title>We read the schemas of 4,951 public MCP servers</title>
      <dc:creator>MCPulse</dc:creator>
      <pubDate>Wed, 09 Sep 2026 09:14:53 +0000</pubDate>
      <link>https://dev.to/getmcpulse/we-read-the-schemas-of-4951-public-mcp-servers-537o</link>
      <guid>https://dev.to/getmcpulse/we-read-the-schemas-of-4951-public-mcp-servers-537o</guid>
      <description>&lt;p&gt;Your MCP server's tool schema is the entire interface a model has. No README, no repo, no idea what you meant — just the JSON from &lt;code&gt;tools/list&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;We read that JSON for 4,951 public servers: 87,146 tools and 270,487 parameters. Here's what's in it.&lt;/p&gt;

&lt;p&gt;The short version: one tool in six carries a description containing no word that distinguishes it from a sibling tool on the same server. On servers with more than sixty tools, it's nearly one in three.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this is, and what it isn't
&lt;/h2&gt;

&lt;p&gt;This post is about what models are given. It is not about what they do with it.&lt;/p&gt;

&lt;p&gt;We did not run a model against these servers. We never observed a tool selection, an argument, or a retry, so nothing here can tell you how often models actually get it wrong. Every number below is a property of a schema sitting still.&lt;/p&gt;

&lt;p&gt;We're drawing that line hard because it's the line the interesting claim sits on. "A model has nothing to discriminate on" is a fact about a schema. "A model therefore picks wrong 30% of the time" is a fact about traffic, and we don't have it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How we did it
&lt;/h2&gt;

&lt;p&gt;We took the tool schemas from the Smithery registry, which stores the &lt;code&gt;tools/list&lt;/code&gt; response for every server it has scanned — &lt;code&gt;inputSchema&lt;/code&gt; included. That's byte-for-byte the JSON a model receives.&lt;/p&gt;

&lt;p&gt;The obvious alternative was to boot each server in a container and call &lt;code&gt;tools/list&lt;/code&gt; ourselves. We didn't, and the reason is the finding underneath the method: a large share of public MCP servers won't start without real credentials. The set that boots cleanly on a machine with no API keys isn't a random subset of anything.&lt;/p&gt;

&lt;p&gt;Of 5,123 servers in the frame, 154 detail requests failed, 5 had never been scanned, and 13 published no tools at all. That leaves 4,951 servers with schemas.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check that nearly ended the study
&lt;/h2&gt;

&lt;p&gt;Before any of this counts for anything, one question has to be answered: does the registry hand back the schema the server actually serves?&lt;/p&gt;

&lt;p&gt;We had reason to think it might not. Across the first 1,377 tools we collected, not one carried a top-level &lt;code&gt;required&lt;/code&gt; array. Real MCP servers mark parameters required constantly.&lt;/p&gt;

&lt;p&gt;So we booted the four official reference servers locally — they need no credentials and no network — took their real &lt;code&gt;tools/list&lt;/code&gt; over stdio, and diffed field by field.&lt;/p&gt;

&lt;p&gt;Everything survives except &lt;code&gt;required&lt;/code&gt;, which is stripped. On one server that's 27 of 37 tools with a required array on the real server, and 0 via the API.&lt;/p&gt;

&lt;p&gt;So this study measures nothing about required-versus-optional parameters. If you're doing your own analysis on registry data, that field is not there, and it does not announce itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. One parameter in five has no description at all
&lt;/h2&gt;

&lt;p&gt;59,038 of 270,487 parameters — 21.8% — ship with no description. They appear on 33.7% of servers.&lt;/p&gt;

&lt;p&gt;A parameter with no description is a parameter the model guesses at. It has the name, it has the type, and that's the entire brief. Sometimes the name carries it: &lt;code&gt;query&lt;/code&gt; on a search tool isn't mysterious. Often it doesn't — we found plenty of bare &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;type&lt;/code&gt;, &lt;code&gt;mode&lt;/code&gt; and &lt;code&gt;filter&lt;/code&gt; parameters with nothing to say which of several plausible things they meant.&lt;/p&gt;

&lt;p&gt;The striking part is the contrast with tool descriptions. Only 0.4% of tools have no description. Authors describe the tool and forget the arguments — which is understandable, because the tool is the thing you're thinking about when you write it, and the arguments are the thing the model has to fill in.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. One tool in six has nothing to tell it apart from its neighbour
&lt;/h2&gt;

&lt;p&gt;For every tool, we took the content words in its description and asked how many appear in no other tool's description on the same server. Call it the tool's distinctive share.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The median tool's description is 26% distinctive. Three-quarters of the words it spends are words its siblings also use.&lt;/li&gt;
&lt;li&gt;26.7% of tools are under 10% distinctive.&lt;/li&gt;
&lt;li&gt;17.4% are exactly zero. 23.1% of servers have at least one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Zero doesn't mean the description is bad. Here are four from one widely-installed Gmail server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gmail_DeleteDraftEmail   "Delete a draft email using the Gmail API."
Gmail_SendDraftEmail     "Send a draft email using the Gmail API."
Gmail_ListLabels         "List all the labels in the user's mailbox."
Gmail_SearchThreads      "Search for threads in the user's mailbox."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every one of those is clear, correct English. Every one is also built entirely from words the other tools use — delete, draft, email, gmail, api, list, search, threads, mailbox all recur across the set. The description tells you what the tool does. It doesn't tell you what &lt;em&gt;this&lt;/em&gt; tool does and the others don't, and that second thing is the one a model needs when it's choosing.&lt;/p&gt;

&lt;p&gt;The pattern that produces the most extreme cases is shared boilerplate. One server appends the same 51-word context block to all 275 of its tools. "Create a new NFT collection" and "Purchase an NFT from a listing" differ by 8 words out of 59. That block was added deliberately, to help.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. It gets worse the more tools you ship
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tools on the server&lt;/th&gt;
&lt;th&gt;Servers&lt;/th&gt;
&lt;th&gt;No description&lt;/th&gt;
&lt;th&gt;Zero-distinctive tools&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1–3&lt;/td&gt;
&lt;td&gt;1,256&lt;/td&gt;
&lt;td&gt;14.8%&lt;/td&gt;
&lt;td&gt;0.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4–7&lt;/td&gt;
&lt;td&gt;1,292&lt;/td&gt;
&lt;td&gt;22.4%&lt;/td&gt;
&lt;td&gt;1.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8–15&lt;/td&gt;
&lt;td&gt;1,044&lt;/td&gt;
&lt;td&gt;21.9%&lt;/td&gt;
&lt;td&gt;4.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;16–30&lt;/td&gt;
&lt;td&gt;767&lt;/td&gt;
&lt;td&gt;24.8%&lt;/td&gt;
&lt;td&gt;7.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;31–60&lt;/td&gt;
&lt;td&gt;377&lt;/td&gt;
&lt;td&gt;22.2%&lt;/td&gt;
&lt;td&gt;16.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;61+&lt;/td&gt;
&lt;td&gt;215&lt;/td&gt;
&lt;td&gt;20.5%&lt;/td&gt;
&lt;td&gt;31.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Description collision rises monotonically and by a factor of sixty. Some of that is arithmetic — more tools means more chances for two to collide — but it's also the point at which authors start generating descriptions from a template, and a template is a machine for producing tools that read alike.&lt;/p&gt;

&lt;p&gt;And notice the column that doesn't move. Missing parameter descriptions sit between 20% and 25% at every size above the smallest bucket. It's not a scale problem; it's a habit. The two failures are independent, which means shipping fewer tools won't fix your undescribed parameters and writing better descriptions won't fix your collisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The median tool list costs about 1,250 tokens before anyone asks a question
&lt;/h2&gt;

&lt;p&gt;Tool schemas are sent on every connection, whether or not a single tool gets called.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Median server: 4,991 bytes, roughly 1,250 tokens&lt;/li&gt;
&lt;li&gt;90th percentile: 32,636 bytes, roughly 8,200 tokens&lt;/li&gt;
&lt;li&gt;Largest in the corpus: 1,145,575 bytes — on the order of 280,000 tokens of schema, from one server, before the conversation starts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Median tools per server is 7 and the mean is 17.6. One server publishes 2,530 tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Some parameters name their valid values and then don't enforce them
&lt;/h2&gt;

&lt;p&gt;8.0% of all parameters carry an &lt;code&gt;enum&lt;/code&gt;. What you can find from outside is the case where the author wrote the values down in prose and left the schema as an open string:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;outcome_attribution   "Attribution type for the outcomes.
                       Valid values: "direct", "influenced",
                       "unattributed", "total"."

commitment            "Optional processed|confirmed|finalized commitment"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;895 parameters, on 4.6% of servers. That's a floor rather than an estimate: it only catches authors who documented the set.&lt;/p&gt;

&lt;p&gt;Worth reporting how we got there. Our first version of this measurement said 108 hits in a 40-server sample. It was matching text like &lt;code&gt;Filter by line (e.g. "1", "A", "F")&lt;/code&gt;, which is an illustration, not a closed set. Requiring explicit closed-set language took that sample from 108 to 9, and all nine were real.&lt;/p&gt;

&lt;p&gt;One signal we expected to find and didn't: duplicate tool names within a server, on 0.2% of servers. Effectively nobody does this. If it's on your review checklist, take it off.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we couldn't see
&lt;/h2&gt;

&lt;p&gt;This is static analysis, and we never observed a single real request to any of these servers.&lt;/p&gt;

&lt;p&gt;That means we missed everything that only shows up under traffic. The tool that works in isolation but gets called in the wrong order. The parameter that's fine until someone phrases a request unusually. The retry loop that only triggers on a specific error path.&lt;/p&gt;

&lt;p&gt;We also can't tell you the thing you most want to know, which is how much of this matters. A zero-distinctive description might cost nothing when the tool name is unambiguous, and a great deal when it isn't.&lt;/p&gt;

&lt;p&gt;One more hole, found by a server author after publication: the 21.8% measures &lt;em&gt;absence&lt;/em&gt; only. A parameter described as &lt;code&gt;"query: The query"&lt;/code&gt; counts as described and passes. Restating the parameter name is the more common failure in his experience, and it passes every linter. So that number is a floor too.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you maintain a server
&lt;/h2&gt;

&lt;p&gt;Three things, ordered by how common the problem is in the data. None of them changes your server's behaviour — they're all changes to the text a model reads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Describe every parameter.&lt;/strong&gt; The most common gap by a distance, and it doesn't get better at any size. If you do one thing, do this one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make each description say what the others don't.&lt;/strong&gt; Not "is it clear" — is it clear &lt;em&gt;which of my tools this is&lt;/em&gt;. Read your tool list as one block and ask what a reader with only that block would use to choose.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you have more than about thirty tools, audit for collisions specifically.&lt;/strong&gt; Above that size, roughly one tool in six has no distinguishing word, rising to one in three past sixty.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data
&lt;/h2&gt;

&lt;p&gt;Analysis scripts and raw data: &lt;a href="https://github.com/getmcpulse/mcp-schema-study" rel="noopener noreferrer"&gt;https://github.com/getmcpulse/mcp-schema-study&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every figure is in &lt;code&gt;analysis.json&lt;/code&gt;, with the servers and tools behind each one in &lt;code&gt;examples.json&lt;/code&gt;. The collector rebuilds the whole corpus from a public API in about twenty minutes.&lt;/p&gt;

&lt;p&gt;There's also a browser-based checker that runs these measurements on your own &lt;code&gt;tools/list&lt;/code&gt; — paste the JSON, get your numbers scored against the corpus: &lt;a href="https://getmcpulse.com/check" rel="noopener noreferrer"&gt;https://getmcpulse.com/check&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We ran this because we're building &lt;a href="https://getmcpulse.com" rel="noopener noreferrer"&gt;MCPulse&lt;/a&gt;, an SDK that reports what actually happens under real traffic. Everything in this post came from outside the server, which is exactly its limit. A schema can tell you a model has nothing to choose on. Only traffic can tell you whether it chose wrong.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://getmcpulse.com/blog/reading-5000-mcp-schemas" rel="noopener noreferrer"&gt;getmcpulse.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>llm</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
