<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hiroki Nomura</title>
    <description>The latest articles on DEV Community by Hiroki Nomura (@nomunomu0504).</description>
    <link>https://dev.to/nomunomu0504</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4170064%2F2801507a-070b-4222-ac24-f5edd64501db.png</url>
      <title>DEV Community: Hiroki Nomura</title>
      <link>https://dev.to/nomunomu0504</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nomunomu0504"/>
    <language>en</language>
    <item>
      <title>I open-sourced a linter for MCP tool descriptions</title>
      <dc:creator>Hiroki Nomura</dc:creator>
      <pubDate>Thu, 08 Oct 2026 03:54:27 +0000</pubDate>
      <link>https://dev.to/nomunomu0504/i-open-sourced-a-linter-for-mcp-tool-descriptions-12nj</link>
      <guid>https://dev.to/nomunomu0504/i-open-sourced-a-linter-for-mcp-tool-descriptions-12nj</guid>
      <description>&lt;p&gt;AI agents choose an MCP tool by reading three things: its name, its description and its input schema. Many servers write those descriptions like docstrings, for a human who already knows why they opened the file. An agent with thirty tools in its context doesn't know that. It needs to know when to use a tool, what comes back, and which lookalike to use instead.&lt;/p&gt;

&lt;p&gt;Forecall's linter checks exactly that, and its code is now public under Apache-2.0: &lt;a href="https://github.com/forecall/forecall-cli" rel="noopener noreferrer"&gt;github.com/forecall/forecall-cli&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it on a real server
&lt;/h2&gt;

&lt;p&gt;Two commands. The first starts a server and saves its tools/list; the second scores it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx forecall dump &lt;span class="nt"&gt;-o&lt;/span&gt; memory.json &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @modelcontextprotocol/server-memory
npx forecall lint memory.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here is part of the report for the official memory server (0.6.3):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Average score     43.8 / 100
Tools             9
Confusable pairs  1

Across the server
  Major     confusable_pair  delete_entities and delete_relations are easy to mix up.
                             Description similarity 0.62 · Name similarity 0.33

  39 / 100  delete_relations
    Purpose 6/20 · When to use 0/20 · Arguments 18/25 · Return value 8/15 · Constraints and side effects 7/10 · Examples 0/10
    Major     restates_name     The description only restates the name (7 words).
    Major     too_short         The description is too short (7 words).
    Major     no_usage_context  Says nothing about when to use it, or when not to.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two delete tools with similar descriptions, and neither says when to use it. That's the kind of pair where an agent deletes the wrong thing. I've suggested rewrites for them in &lt;a href="https://github.com/modelcontextprotocol/servers/issues/5071" rel="noopener noreferrer"&gt;an issue on the servers repo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;For a remote server, pass its URL instead, with &lt;code&gt;--header "Authorization: Bearer …"&lt;/code&gt; if it needs a key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx forecall dump https://example.com/mcp &lt;span class="nt"&gt;-o&lt;/span&gt; tools.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What it checks
&lt;/h2&gt;

&lt;p&gt;Each tool gets a score out of 100, in six parts:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Part&lt;/th&gt;
&lt;th&gt;Points&lt;/th&gt;
&lt;th&gt;What it looks for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Purpose&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;What the tool does, beyond restating its name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When to use&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;When to use it, and which similar tool to use instead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arguments&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;Each argument described, typed and marked required&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Return value&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;What comes back, in an output schema or the description&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Constraints and side effects&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Preconditions, side effects, auth, rate limits, annotations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Examples&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;An example of a call&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Across the server, it flags too many tools, tools that are easy to mix up, and descriptions that share a long paragraph. Every issue has a code and a severity, so you can work through them one by one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use it in CI
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;--fail-under 60&lt;/code&gt; exits with 1 when the average is below 60, and &lt;code&gt;--json&lt;/code&gt; prints the report as JSON.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx forecall dump -o tools.json -- node dist/server.js&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx forecall lint tools.json --fail-under &lt;/span&gt;&lt;span class="m"&gt;60&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Nothing leaves your machine
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;lint&lt;/code&gt; works offline. The scoring lives in &lt;code&gt;@forecall/lint&lt;/code&gt;, a package with no dependencies and no network code. &lt;code&gt;dump&lt;/code&gt; is the only command that talks to the network, and only to the server you name; the JSON it writes records where the tools came from, never headers or environment values.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it can't tell you
&lt;/h2&gt;

&lt;p&gt;The score comes from rules that read the text and the schema. A short description like &lt;code&gt;browser_close&lt;/code&gt; can score low even when every model uses it correctly, and a high score doesn't prove that a model will pick the right tool. Use the score to find where to write more. (Forecall's hosted Evaluations run real models against your tools for that. They're paid, and not in this repository.)&lt;/p&gt;

&lt;h2&gt;
  
  
  What we found with it
&lt;/h2&gt;

&lt;p&gt;We ran it over the 46 most-used servers on Smithery: 769 of 1,118 tools (69%) never say when to use them. The full write-up is &lt;a href="https://dev.to/nomunomu0504/we-checked-46-popular-mcp-servers-on-smithery-7-in-10-tools-never-say-when-to-use-them-4gb9"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Feedback welcome
&lt;/h2&gt;

&lt;p&gt;The rules are at v1, and I'd like to hear where they're wrong. A reader of that write-up already suggested one: when two tools are flagged as lookalikes, check that each description names the other. It's now &lt;a href="https://github.com/forecall/forecall-cli/issues/7" rel="noopener noreferrer"&gt;issue #7&lt;/a&gt;. Issues and pull requests are welcome.&lt;/p&gt;

&lt;p&gt;I'm the author of Forecall.&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>mcp</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
    <item>
      <title>We checked 46 popular MCP servers on Smithery: 7 in 10 tools never say when to use them</title>
      <dc:creator>Hiroki Nomura</dc:creator>
      <pubDate>Thu, 08 Oct 2026 02:55:15 +0000</pubDate>
      <link>https://dev.to/nomunomu0504/we-checked-46-popular-mcp-servers-on-smithery-7-in-10-tools-never-say-when-to-use-them-4gb9</link>
      <guid>https://dev.to/nomunomu0504/we-checked-46-popular-mcp-servers-on-smithery-7-in-10-tools-never-say-when-to-use-them-4gb9</guid>
      <description>&lt;p&gt;We took 46 of the most-used MCP servers on Smithery and checked the descriptions of all 1,118 of their tools. &lt;strong&gt;769 tools (69%) never say when to use them.&lt;/strong&gt; 507 (45%) never say in their description what they return. And 15 servers have tools whose descriptions are so alike that an agent can mix them up: 90 such pairs in all.&lt;/p&gt;

&lt;p&gt;The 69% is exactly what we found last time, when we checked 8 public servers with 126 tools and 87 of them lacked it. Nine times the sample, the same share.&lt;/p&gt;

&lt;p&gt;The checks were done with Forecall, a static linter for MCP tool descriptions that I build.&lt;/p&gt;

&lt;h2&gt;
  
  
  How we checked: the top 50 servers on Smithery, judged by their text only
&lt;/h2&gt;

&lt;p&gt;On October 8, 2026, we took the first 500 servers in Smithery's public registry (in the order its API returns), removed duplicates, picked the 50 with the most uses, and collected their tool definitions. We scored them with Forecall's CLI (&lt;code&gt;forecall&lt;/code&gt; 0.3.2). Four of the 50 have more than 200 tools, the CLI's limit, so we left them out: 46 servers and 1,118 tools remain.&lt;/p&gt;

&lt;p&gt;One caveat. The tool definitions Smithery's API returns leave out the input schemas' &lt;code&gt;required&lt;/code&gt; lists and any output schemas, so Forecall's overall scores come out lower than they should. &lt;strong&gt;This article does not use the overall scores. It counts only the findings that depend on the description text alone&lt;/strong&gt;: whether a tool says when to use it, whether its description says what it returns, whether the description is too short, and pairs of tools whose descriptions are alike.&lt;/p&gt;

&lt;h2&gt;
  
  
  7 in 10 tools never say when to use them
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Finding&lt;/th&gt;
&lt;th&gt;Tools&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No "when to use"&lt;/td&gt;
&lt;td&gt;769&lt;/td&gt;
&lt;td&gt;69%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Description doesn't say what it returns&lt;/td&gt;
&lt;td&gt;507&lt;/td&gt;
&lt;td&gt;45%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Too short, or restates the name&lt;/td&gt;
&lt;td&gt;101&lt;/td&gt;
&lt;td&gt;9%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;By server it is starker. In 16 of the 46 servers, &lt;strong&gt;not one&lt;/strong&gt; tool says when to use it. Count the servers where 80% or more lack it, and you have half of them: 23.&lt;/p&gt;

&lt;p&gt;"When to use" is what decides between lookalike tools. Without it, a model picks by the tools' names and the feel of their descriptions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same notice on every tool makes the tools indistinguishable
&lt;/h2&gt;

&lt;p&gt;PubMed has the most lookalike pairs. It has only 7 tools, yet &lt;strong&gt;all 21 pairs&lt;/strong&gt; are flagged as easy to mix up. Among the 12 most similar pairs, the descriptions share 73% to 94% of their words.&lt;/p&gt;

&lt;p&gt;The cause is a long notice pasted into all 7 descriptions:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;IMPORTANT - PubMed Database Scope: This server provides access to PubMed, which ONLY indexes biomedical and life sciences literature including: …&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;What sets each tool apart, its first sentence ("search articles", "get metadata", "find related articles" and so on), is buried in the shared paragraph.&lt;/p&gt;

&lt;p&gt;A note that applies to the whole server belongs once in the server's instructions, which the client receives at initialization, not in every tool's description. A tool's description should say what only that tool does.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Send a reply" and "draft a reply" are written almost the same way
&lt;/h2&gt;

&lt;p&gt;Gmail has 7 lookalike pairs. The closest two share 92% of their words:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Start of its description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gmail_ReplyToEmail&lt;/td&gt;
&lt;td&gt;Send a reply to an email message, optionally with one or more file attachments.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gmail_WriteDraftReplyEmail&lt;/td&gt;
&lt;td&gt;Compose a draft reply to an email message, optionally with one or more file attachments.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;After that, both go on with the same paragraph about attaching files. The difference is a few words: "Send" against "Compose a draft".&lt;/p&gt;

&lt;p&gt;Mix these two up, and an agent meant to write a draft sends the email. Tools that do something you can't undo, above all, should say in one sentence how to choose: "To prepare a reply without sending it, use Gmail_WriteDraftReplyEmail."&lt;/p&gt;

&lt;h2&gt;
  
  
  A good example: Brave Search says how its tools connect
&lt;/h2&gt;

&lt;p&gt;Brave Search's descriptions, by contrast, spell out how the tools relate. Here is brave_web_search:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Performs web searches using the Brave Search API and returns comprehensive search results with rich metadata. To chain into local-POI enrichment, pass &lt;code&gt;result_filter=locations&lt;/code&gt; and feed the resulting &lt;code&gt;locations.results[].id&lt;/code&gt; values into &lt;code&gt;brave_local_search&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It says what comes back and which value to pass to which tool next, so an agent can use several tools in sequence. All 8 of Brave Search's tools say in their descriptions what they return, and it has no lookalike pairs.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to fix it: one sentence each for three things
&lt;/h2&gt;

&lt;p&gt;From these results, this is the order to fix descriptions in:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Say when to use the tool.&lt;/strong&gt; "Use this when …", and if there is a lookalike, "for …, use X instead."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move shared notes to the server's instructions.&lt;/strong&gt; The same paragraph on every tool buries what tells them apart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the difference between lookalikes into words.&lt;/strong&gt; Send or draft? Delete or remove? Say it plainly, especially for actions you can't undo.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;To check your own server, run these two commands. The first gets the tools/list from your server, the second scores it. Both run on your machine and send the results nowhere.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx forecall dump &lt;span class="nt"&gt;-o&lt;/span&gt; tools.json &lt;span class="nt"&gt;--&lt;/span&gt; &amp;lt;&lt;span class="nb"&gt;command &lt;/span&gt;that starts your server&amp;gt;
npx forecall lint tools.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a Streamable HTTP server, pass its URL instead: &lt;code&gt;npx forecall dump https://example.com/mcp &amp;gt; tools.json&lt;/code&gt;. The CLI is open source: &lt;a href="https://github.com/forecall/forecall-cli" rel="noopener noreferrer"&gt;https://github.com/forecall/forecall-cli&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The weakness doesn't shrink with scale
&lt;/h2&gt;

&lt;p&gt;With 8 servers or 46, 7 in 10 tools never said when to use them, popular servers included. The fix is small: one sentence on when to use the tool and one on how it differs from its lookalikes leave an agent far less room to pick the wrong one.&lt;/p&gt;

&lt;p&gt;We'll keep this study going with new data. Next, we'll score the remote servers in the official MCP Registry, schemas included.&lt;/p&gt;

&lt;p&gt;The same study is also on Qiita, in Japanese: &lt;a href="https://qiita.com/nomunomu0504/items/b2eab1960a7aad38b59d" rel="noopener noreferrer"&gt;https://qiita.com/nomunomu0504/items/b2eab1960a7aad38b59d&lt;/a&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>llm</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
