DEV Community

Onur Kesim
Onur Kesim

Posted on AI-assisted

I benchmarked my own security tool against 3 others — and wrote down what it lost

I built gedik, a security-audit skill for Claude Code. Before telling anyone it was good, I wanted a number. So on 30 September 2026 I ran it against three other tools on two targets, with the answer key kept outside every run directory. Then I put the results in the README, including the parts where it lost.

This post is about those parts. If you build tools on top of an LLM, the losses are more useful to you than the wins.

The setup

Four tools. The three LLM-based ones ran on the same model (claude-sonnet-5-5), each in its own isolated claude -p process, with the same prompt:

  • gedik v2.6.0
  • cloudflare/security-audit-skill
  • Anthropic's built-in /security-review
  • Semgrep 1.178.0, free registry rules only (no Pro, no login)

A fifth candidate, Anthropic's claude-security plugin, was left out: its licence restricts use with non-Anthropic products, including competing ones, and I read that as covering this test.

Two targets:

  • Target 1: NodeGoat (commit c5cb68a), an intentionally vulnerable Node app. I deleted the tutorials, the tests and every JavaScript comment so the tools couldn't read the answers. That left 17 known issues in the copy.
  • Target 2: a small Supabase + LLM-tool repo I built myself. 11 seeded issues, 5 decoys.

I couldn't restrict file-system access, so I checked afterwards instead: every session and sub-agent log searched for the key's path, zero matches.

One person scored everything: me.

The numbers

Target 2, my repo, 11 seeded issues (runs of 30 Sep 2026):

gedik cloudflare /security-review Semgrep
found 11/11 11/11 (9 confirmed) 9/11 1/11
false positives 0 0 0 0

Target 1, NodeGoat, 17 known issues (30 Sep and 1 Oct 2026; competitors one run each):

gedik cloudflare /security-review Semgrep
detection, when the report arrived 16–17/17 (5 delivered runs) 16/17 5/17 6/17
report delivered 2 of 6 runs (30 Sep, evening and night) · 3 of 3 (1 Oct) yes yes yes

That second row is the story.

Loss 1: the user got an error message

On 30 September, in the first scored run on NodeGoat, gedik's report never reached me. Twice. Every sub-agent finished its work, and then, while the final report was streaming, the session stopped with stop_reason: refusal and a [cyber] safeguard tag. One run had written 6,876 characters and announced 27 findings before it was cut. The other stopped at 3,114 characters. What I received was the error text.

So whatever it had found, the honest score for that first run was 0/17 delivered.

The other three tools audited the same code and were never stopped. Their reports carried little or no exploit detail. gedik's design rule is "no proof, no finding": every finding comes with a working proof of concept. My guess, and it is only a guess because I never isolated the trigger, is that a report full of working code-execution and injection proofs is the kind of text a real-time cyber safeguard interrupts.

I tried to engineer my way out the same night. I moved the proof inputs out of the Markdown report into a JSON field and tightened the rules for harmless markers. It didn't fix it. One patched run delivered. One was refused during the research phase, before any report existed. In a third, after being refused, the agent started splitting its file write into smaller pieces. That is behaviour I don't want in a security tool. I stopped that run, didn't count it, and reverted the patch.

Across the six unpatched runs on that target, 4 were cut and 2 delivered.

Loss 2: the fix was not mine

What changed delivery was an account-level change, not code. My account was approved into Anthropic's Cyber Verification Program, and on 1 October 2026 I re-ran the unmodified v2.6.0 on the same target: 3 of 3 reports delivered, 17/17 each (in one run, 16 confirmed and 1 marked as suspected), 0 false positives.

That sounds like a happy ending. It is weaker than it looks:

  • 3 of 3 after, 2 of 6 before. One-sided Fisher p ≈ 0.12. Consistent, not proven.
  • Time, run-to-run model variation and the reverted patch are all confounders.
  • A user without that approval should still expect the report to be cut on exploit-heavy code like this: 4 runs in 6 on this target on 30 Sep, not re-measured since.

So the README says detection was at or above the best tool here, and delivery depends on your account. It does not say "better".

Loss 3: slower and more expensive

On Target 2 (30 Sep 2026) gedik took 176 seconds; /security-review took 40 and the cloudflare skill 96. The client-reported total_cost_usd field showed $1.02 for gedik against $0.35 and $0.74. I haven't verified what that field means on a subscription account, so read it as relative, not as a bill. Semgrep finished in seconds without a model.

gedik does more per run: a triage step, a refutation step in most runs, mutation testing of the project's own tests. But "does more" is a cost the user pays every time.

Loss 4: my favourite target proves the least

Target 2 is where gedik looks best, and it is the weakest evidence in this post. I built it, I seeded it, and I picked issues close to gedik's strongest areas: Supabase row-level security and LLM tool permissions. In the 30 Sep run the cloudflare skill also found all 11, two of them marked "needs validation". On a target designed by the person who wrote the tool, a tie is the most a fair reader should take from it.

NodeGoat has its own problem. It is almost certainly in the model's training data. Deleting the comments didn't hide it: gedik and the cloudflare skill both called the target "NodeGoat-style" on their own. Memorisation may inflate every LLM-based score in that table, gedik's included.

Loss 5: the machine-readable output didn't show up

gedik v2.6 is supposed to write a findings file in JSON, checked against a schema. In the three runs of 1 October it wrote prose reports and no JSON file. My prompt said "put the findings in your final reply", and the skill read that as "don't write files". The cloudflare skill already ships a schema'd findings.json with tested validators. That gap is real and current.

What it won

One thing the others didn't do at all. On Target 2 (30 Sep 2026) gedik mutated the project's own code in a scratch copy and ran its test suite: three of four mutants that removed an access check or an input check survived with green tests. These were library-level mutants, not every route-level mutant in my key. The other tools said nothing about the tests.

That check, "would your tests notice if this guard disappeared?", is the reason I wrote gedik, and it's the part I'd defend hardest.

The caveats, in one place

  • One annotator: me, the person who wrote the tool. No second reading.
  • Target 2 was built by me and seeded toward my tool's strengths.
  • The cleaned NodeGoat copy is weaker than the original, and probably memorised.
  • One run per tool and target, except gedik on NodeGoat: six runs on 30 Sep before the account change, three on 1 Oct after.
  • The cloudflare agent said it did not run its full multi-phase workflow, so I measured "the skill under this prompt", not the skill's complete pipeline.
  • Semgrep ran free rules only.
  • Everything ran on Windows.

What I'd like from you

The weakest part of all this is that only I scored it. The method and the run records are in KRITER.md (written in Turkish) in the repository, github.com/onur-kesim/gedik; the answer key and the raw reports aren't published yet. If you'd be willing to be the second reader, say so in the comments.

And if you build on an LLM: measure delivery, not just detection. My tool found everything and, for a day, delivered nothing.

This post was written with AI assistance. The numbers come from my run records (gedik commit f34aaf4, 1 October 2026).

Top comments (0)