AI agents choose an MCP tool by reading three things: its name, its description and its input schema. Many servers write those descriptions like docstrings, for a human who already knows why they opened the file. An agent with thirty tools in its context doesn't know that. It needs to know when to use a tool, what comes back, and which lookalike to use instead.
Forecall's linter checks exactly that, and its code is now public under Apache-2.0: github.com/forecall/forecall-cli.
Try it on a real server
Two commands. The first starts a server and saves its tools/list; the second scores it.
npx forecall dump -o memory.json -- npx -y @modelcontextprotocol/server-memory
npx forecall lint memory.json
Here is part of the report for the official memory server (0.6.3):
Average score 43.8 / 100
Tools 9
Confusable pairs 1
Across the server
Major confusable_pair delete_entities and delete_relations are easy to mix up.
Description similarity 0.62 · Name similarity 0.33
39 / 100 delete_relations
Purpose 6/20 · When to use 0/20 · Arguments 18/25 · Return value 8/15 · Constraints and side effects 7/10 · Examples 0/10
Major restates_name The description only restates the name (7 words).
Major too_short The description is too short (7 words).
Major no_usage_context Says nothing about when to use it, or when not to.
Two delete tools with similar descriptions, and neither says when to use it. That's the kind of pair where an agent deletes the wrong thing. I've suggested rewrites for them in an issue on the servers repo.
For a remote server, pass its URL instead, with --header "Authorization: Bearer …" if it needs a key:
npx forecall dump https://example.com/mcp -o tools.json
What it checks
Each tool gets a score out of 100, in six parts:
| Part | Points | What it looks for |
|---|---|---|
| Purpose | 20 | What the tool does, beyond restating its name |
| When to use | 20 | When to use it, and which similar tool to use instead |
| Arguments | 25 | Each argument described, typed and marked required |
| Return value | 15 | What comes back, in an output schema or the description |
| Constraints and side effects | 10 | Preconditions, side effects, auth, rate limits, annotations |
| Examples | 10 | An example of a call |
Across the server, it flags too many tools, tools that are easy to mix up, and descriptions that share a long paragraph. Every issue has a code and a severity, so you can work through them one by one.
Use it in CI
--fail-under 60 exits with 1 when the average is below 60, and --json prints the report as JSON.
- run: npx forecall dump -o tools.json -- node dist/server.js
- run: npx forecall lint tools.json --fail-under 60
Nothing leaves your machine
lint works offline. The scoring lives in @forecall/lint, a package with no dependencies and no network code. dump is the only command that talks to the network, and only to the server you name; the JSON it writes records where the tools came from, never headers or environment values.
What it can't tell you
The score comes from rules that read the text and the schema. A short description like browser_close can score low even when every model uses it correctly, and a high score doesn't prove that a model will pick the right tool. Use the score to find where to write more. (Forecall's hosted Evaluations run real models against your tools for that. They're paid, and not in this repository.)
What we found with it
We ran it over the 46 most-used servers on Smithery: 769 of 1,118 tools (69%) never say when to use them. The full write-up is here.
Feedback welcome
The rules are at v1, and I'd like to hear where they're wrong. A reader of that write-up already suggested one: when two tools are flagged as lookalikes, check that each description names the other. It's now issue #7. Issues and pull requests are welcome.
I'm the author of Forecall.
Top comments (3)
The average-only CI threshold could hide a weak destructive tool: nine tools scoring 90 and one delete tool scoring 10 still average 82. Adding well-described read tools would let that server pass --fail-under 60 without improving the delete description.
Would you consider a per-tool floor or a fail-on-severity option alongside the average? A useful regression fixture would keep the bad delete tool unchanged, add several high-scoring read tools, and assert that its blocking finding still fails CI. That tests the aggregation rule without pretending a description score proves safe model behavior. I haven't run Forecall; this is based on the CI rule shown here.
You're right: --fail-under only looks at the average today, so strong read tools can carry a weak delete tool. Until there's a flag for it, the JSON report lets you gate per tool or by severity:
jq -eexits non-zero when the result is false, so either line fails the job. For your delete-tool case the per-tool floor is the direct fix: a missing destructiveHint is a major finding, not a critical one. I've filed both options as github.com/forecall/forecall-cli/i..., with your fixture as the test case.Follow-up: both are in forecall 0.5.0.
--min-tool-score 40fails the job when any tool scores below 40, and--fail-on majorfails on any major or critical issue. The failed gate and the tools it caught are named on stderr, and--jsonstill prints the report. Your fixture is now the regression test: a weak delete_record plus seven well-described read tools passes--fail-under 60but fails--min-tool-score 40. Docs: forecall.dev/en/docs/cli#ci