Secret scanners have a noise problem. Point one at a big, healthy codebase and you get hundreds of alerts: test fixtures, docs examples, long identifiers that happen to look random. People stop reading them, and then the real leak slips through.
That gets worse with AI coding agents, which write files faster than anyone reviews them. So I built Klarion, a scanner that asks an LLM before it raises an alert, and then benchmarked it against the usual tools to see whether that actually helps.
The test
Five scanners, four large open-source repos in four languages, all pinned to a commit:
- spring-boot (Java), terraform (Go/HCL), next.js (TypeScript), symfony (PHP)
- 61,514 files in total
- Control set: flask and rails (5,054 files)
Versions: klarion 0.4.3 (with Claude Haiku), gitleaks 8.30.1, trufflehog 3.97.5, detect-secrets 1.5.0, ripsecrets 0.1.11.
These repos have no live credentials, so every alert is something a human would have to check and throw away.
| Tool | Alerts on the 4 repos | Per 1k files | Control (flask + rails) |
|---|---|---|---|
| Klarion + model | 11 | 0.18 | 1 |
| ripsecrets | 299 | 4.86 | 61 |
| trufflehog | 344 | 5.59 | 9 |
| gitleaks | 361 | 5.87 | 33 |
| detect-secrets | 12,073 | 196.26 | 247 |
That's 33x fewer alerts than gitleaks and over 1,000x fewer than detect-secrets on the same code. Each of Klarion's 11 comes with the model's verdict and a one-line reason, so checking them takes minutes, not an afternoon.
Why entropy scanners blow up on typed languages
Entropy-based detection, including Klarion's own first stage, raises 3–4x more alerts per file on Java, Go, TypeScript and PHP than on Python and Ruby: trufflehog jumps 3.1x and detect-secrets 4.0x. gitleaks and ripsecrets, which lean on rules, stayed flat.
The reason: Java, Go, TypeScript and PHP are full of long identifiers like SseCustomerKeySHA256AttrName that look high-entropy. Entropy alone can't tell them apart from a credential. Reading the surrounding code can, and that's what the LLM step does. After it, Klarion's rate is 0.18 per 1k files on the typed repos and 0.20 on the control, so the language effect is gone.
Does it miss real secrets?
Fewer alerts are useless if real leaks go missing, so I also used Samsung's CredData: labelled lines from 337 real repositories.
On code outside test folders:
| Recall | Precision | |
|---|---|---|
| Klarion + model | 0.35 | 0.89 |
| gitleaks 8.30 | 0.21 | 0.91 |
| detect-secrets 1.5 | 0.35 | 0.34 |
That's about 1.7x the real secrets gitleaks finds, at similar precision.
How it works
- Detect: keyword gate, 81 regex rules, normalized Rényi entropy.
- Judge: an LLM reads each candidate with its surrounding code and returns secret / false positive / uncertain, with a reason.
It needs a model: Anthropic, OpenAI-compatible, a local Ollama model (nothing leaves your machine), or your Claude Code login.
Where it runs
- Claude Code: a plugin hook blocks the write before the file exists (file edits and Bash)
- Cursor, Cline or any MCP agent: through its MCP server
- CI: a GitHub Action that scans only what a PR adds; GitLab CI works too
-
Git hooks:
klarion protector the pre-commit framework -
Locally:
klarion scan .
Try it, and tell me where it's wrong
Free and open source (MIT): https://github.com/0x1Adi/Klarion
The full method, pinned commits and every finding are in benchmark/REPORT.md.

Top comments (0)