DEV Community

ScriptMasterLabs
ScriptMasterLabs

Posted on Originally published at scriptmasterlabs.com

I Built the Tool That Scores MCP Listings, Then Ran It on My Own API — It Gave Me a B-

I built a tool that scores MCP server listings for AI agent discoverability. Then I did the only honest thing you can do with a tool like that: I ran it on myself first.

The tool

SqueezeRank MSO scores any MCP server listing across six axes — Purpose Clarity (25%), Usage Guidelines (20%), Behavioral Transparency (20%), Parameter Semantics (15%), Conciseness & Structure (10%), Contextual Completeness (10%). Per-tool grades A+ to F, a weighted final score, and a prioritized fix list. It's a deterministic scorer, not a vibe check: same manifest in, same score out, every time.

It sells itself pay-per-call: $0.002–$0.010 USDC per tool on Base via x402. A full optimization cycle — score, rewrite the description, audit every tool, probe the registries, build the fix plan — costs $0.023. Two cents, with a receipt for every step.

The dogfood run

Before launch, I pointed it at its own listing. Result: A (4.55/5.0). Good — the product can grade its own homework.

Then I pointed it at SqueezeOS — the live MCP server the whole thing runs on. Thirteen real tools, production traffic, the works.

Result: B- (2.94/5.0).

The report was blunt. None of the 13 tools told an agent when to use them. The server description ran over 20 words. Most parameters had no descriptions at all. My own flagship listing had the exact disease my product exists to cure.

The fix

So I fixed all of it. Every tool got when-to-use guidance. Every parameter got a description. The server description got trimmed to 14 words. Return formats, error states, and truthful behavioral terms (read-only, idempotent, no-side-effects) documented per tool. No behavior changes — descriptions only, because the code was never the problem. The listing was.

Re-score of the rewritten manifest: A- (4.14/5.0), every one of the 13 tools individually grading A.

And here's the part I didn't expect: 4.14 is the mathematical ceiling. I worked the rubric's own caps — Purpose Clarity tops out at 4.0, Conciseness at 4.0 given the fixed tool names, coherence is dragged by non-verb-first names I can't change without breaking the API. A "perfect" 4.5+ is unreachable for this manifest without renaming tools, which would break every agent already calling them. Chasing the last half-point would mean gaming the rubric — the exact dishonest optimization this product exists to punish. So 4.14 stands, honestly.

One more honest note: the rewritten manifest hit a deployment snag on the hosting side (the deploy pipeline froze on a strings-only commit — no code change), so the live listing still serves the B- version while that's sorted. The rewrite is preserved and ships the moment deploys cooperate. The scores above are computed from the rewritten manifest, not claimed from the live one.

Why this matters beyond one listing

AI agents pick tools from their descriptions. In 2026, registries like Glama, Smithery, and mcp.so list over 100,000 MCP servers, and the vast majority are functionally invisible — not because the software is bad, but because the listing doesn't start with a capability verb, doesn't explain when to use it, and wastes the character limit.

Every one of those listings is a $0.023 fix. That's the whole product: the grading layer for the MCP ecosystem, sold per-call, settled on-chain.

The scores in this post are computed, not claimed. Run the scorer yourself: POST /v1/mso/score — $0.003 USDC, and it'll tell you exactly what's wrong with your listing too.

https://scriptmasterlabs.com/squeezerank-mso

Top comments (0)