Most AI agent directories answer one question: who pitched best? If an agent is going to act for you (book things, read your files, call your tools), that is the wrong question. The right one is: what does the code actually do, and what do we not know?
That is why I built Metal Mantra, a registry of public reports on AI agents.
What a report is
You paste a public GitHub repository and an official domain. Deterministic rules read the source and produce a public report across five weighted pillars:
| Pillar | Weight |
|---|---|
| Security & privacy | 25% |
| Guardrails & loop control | 25% |
| Schema & tooling | 20% |
| Token & cost efficiency | 15% |
| Entity trust | 15% |
Every deduction points to a rule and a file. There is no AI in the scanner, so the same repository gives the same result. (AI appears in exactly one place: an optional box that ranks agents against a buyer's description. It never touches a score.)
The rule I will not break
A grade is an automated signal about public source. It is not a security certification, and each report carries a plain "what this does not establish" section: runtime safety, real-world capability, fit for your use case.
That sounds like a disclaimer. It is the product. A registry that quietly turns a heuristic into a seal is worse than no registry, because people stop reading.
What nine real reports looked like
At launch the registry holds nine reports of well-known open source agents. Scores ranged from 69 to 84 out of 100. Nobody scored near 100 and nobody scored near zero, which is what you would expect from static analysis of serious projects: it finds missing guardrails, loose schemas and unbounded loops, not catastrophes.
More useful than the scores was what the scanner admitted it did not know. Each report shows how many files were eligible and how many were actually read. A "partial selection" is labelled as such.
The bug that taught us the most
For our first scans we read the first 80 eligible files in archive order. For a monorepo like an agents SDK, those 80 files were examples and docs. The scanner was effectively grading the framework by its demos, and it failed the "is this an agent?" check on one of the best-known agent SDKs in existence.
The fix was boring and important: rank files before capping them (product source first, then unclassified files, then examples, docs and tooling), and widen the agent-detection rules to recognise common SDK imports. The lesson generalises: in an evidence-led system, the sampling strategy is part of the claim. If you cannot say what you read, you cannot say what you found.
Private code
Commercial agents rarely publish source. For those, owners run our scanner inside their own GitHub Actions job. It sends only rule ids and counts, authenticated by GitHub's OIDC token. The server rebuilds each finding from its own rule catalogue, so a caller can omit findings but cannot dress them up. The report is labelled "CI-attested", visibly weaker than a scan we ran ourselves.
Try it
The first scan of any public repo is free: metalmantra.io/scan. If you maintain an agent, tell me which rules feel wrong. The method is public, and the best bug reports so far came from people who disagreed with a deduction.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.