DEV Community

Hiroki Nomura
Hiroki Nomura

Posted on Fully Autonomous

I open-sourced a linter for MCP tool descriptions

AI agents choose an MCP tool by reading three things: its name, its description and its input schema. Many servers write those descriptions like docstrings, for a human who already knows why they opened the file. An agent with thirty tools in its context doesn't know that. It needs to know when to use a tool, what comes back, and which lookalike to use instead.

Forecall's linter checks exactly that, and its code is now public under Apache-2.0: github.com/forecall/forecall-cli.

Try it on a real server

Two commands. The first starts a server and saves its tools/list; the second scores it.

npx forecall dump -o memory.json -- npx -y @modelcontextprotocol/server-memory
npx forecall lint memory.json
Enter fullscreen mode Exit fullscreen mode

Here is part of the report for the official memory server (0.6.3):

Average score     43.8 / 100
Tools             9
Confusable pairs  1

Across the server
  Major     confusable_pair  delete_entities and delete_relations are easy to mix up.
                             Description similarity 0.62 · Name similarity 0.33

  39 / 100  delete_relations
    Purpose 6/20 · When to use 0/20 · Arguments 18/25 · Return value 8/15 · Constraints and side effects 7/10 · Examples 0/10
    Major     restates_name     The description only restates the name (7 words).
    Major     too_short         The description is too short (7 words).
    Major     no_usage_context  Says nothing about when to use it, or when not to.
Enter fullscreen mode Exit fullscreen mode

Two delete tools with similar descriptions, and neither says when to use it. That's the kind of pair where an agent deletes the wrong thing. I've suggested rewrites for them in an issue on the servers repo.

For a remote server, pass its URL instead, with --header "Authorization: Bearer …" if it needs a key:

npx forecall dump https://example.com/mcp -o tools.json
Enter fullscreen mode Exit fullscreen mode

What it checks

Each tool gets a score out of 100, in six parts:

Part Points What it looks for
Purpose 20 What the tool does, beyond restating its name
When to use 20 When to use it, and which similar tool to use instead
Arguments 25 Each argument described, typed and marked required
Return value 15 What comes back, in an output schema or the description
Constraints and side effects 10 Preconditions, side effects, auth, rate limits, annotations
Examples 10 An example of a call

Across the server, it flags too many tools, tools that are easy to mix up, and descriptions that share a long paragraph. Every issue has a code and a severity, so you can work through them one by one.

Use it in CI

--fail-under 60 exits with 1 when the average is below 60, and --json prints the report as JSON.

- run: npx forecall dump -o tools.json -- node dist/server.js
- run: npx forecall lint tools.json --fail-under 60
Enter fullscreen mode Exit fullscreen mode

Nothing leaves your machine

lint works offline. The scoring lives in @forecall/lint, a package with no dependencies and no network code. dump is the only command that talks to the network, and only to the server you name; the JSON it writes records where the tools came from, never headers or environment values.

What it can't tell you

The score comes from rules that read the text and the schema. A short description like browser_close can score low even when every model uses it correctly, and a high score doesn't prove that a model will pick the right tool. Use the score to find where to write more. (Forecall's hosted Evaluations run real models against your tools for that. They're paid, and not in this repository.)

What we found with it

We ran it over the 46 most-used servers on Smithery: 769 of 1,118 tools (69%) never say when to use them. The full write-up is here.

Feedback welcome

The rules are at v1, and I'd like to hear where they're wrong. A reader of that write-up already suggested one: when two tools are flagged as lookalikes, check that each description names the other. It's now issue #7. Issues and pull requests are welcome.

I'm the author of Forecall.

Top comments (3)

Collapse
 
launchgatecheck profile image
Launch Gate •

The average-only CI threshold could hide a weak destructive tool: nine tools scoring 90 and one delete tool scoring 10 still average 82. Adding well-described read tools would let that server pass --fail-under 60 without improving the delete description.

Would you consider a per-tool floor or a fail-on-severity option alongside the average? A useful regression fixture would keep the bad delete tool unchanged, add several high-scoring read tools, and assert that its blocking finding still fails CI. That tests the aggregation rule without pretending a description score proves safe model behavior. I haven't run Forecall; this is based on the CI rule shown here.

Collapse
 
nomunomu0504 profile image
Hiroki Nomura •

You're right: --fail-under only looks at the average today, so strong read tools can carry a weak delete tool. Until there's a flag for it, the JSON report lets you gate per tool or by severity:

npx forecall lint tools.json --json > report.json
jq -e '[.tools[] | select(.score < 40)] | length == 0' report.json
jq -e '[.tools[].issues[], .serverIssues[] | select(.severity == "critical" or .severity == "major")] | length == 0' report.json
Enter fullscreen mode Exit fullscreen mode

jq -e exits non-zero when the result is false, so either line fails the job. For your delete-tool case the per-tool floor is the direct fix: a missing destructiveHint is a major finding, not a critical one. I've filed both options as github.com/forecall/forecall-cli/i..., with your fixture as the test case.

Collapse
 
nomunomu0504 profile image
Hiroki Nomura •

Follow-up: both are in forecall 0.5.0. --min-tool-score 40 fails the job when any tool scores below 40, and --fail-on major fails on any major or critical issue. The failed gate and the tools it caught are named on stderr, and --json still prints the report. Your fixture is now the regression test: a weak delete_record plus seven well-described read tools passes --fail-under 60 but fails --min-tool-score 40. Docs: forecall.dev/en/docs/cli#ci