DEV Community

Cover image for AI Penetration Testing Tools in 2026: Two Agreed on 7 of 73 Bugs
Moksh Gupta
Moksh Gupta

Posted on Originally published at devtoollab.com

AI Penetration Testing Tools in 2026: Two Agreed on 7 of 73 Bugs

If you sell software to anyone with a security team, somebody will eventually ask for a pentest report: a SOC 2 auditor, a procurement questionnaire, an insurance renewal. The usual answer is a consultancy that tests your app once, writes a PDF, and is out of date by your next deploy. AI pentesting agents promise to run that test continuously, and in 2026 there is finally some evidence about how well they do it.

The best piece of evidence is a May 2026 study by the security firm Doyensec, which pointed Aikido and XBOW at the same two open-source web apps, at $4,000 per scan each. Both platforms found real, exploitable vulnerabilities with very few false positives, yet of the 73 distinct bugs they confirmed between them, only 7 were found by both. I wrote up the full comparison, with every price and version, on DevToolLab. This is the condensed version.

Banner showing Aikido with 49 true positives and XBOW with 31, overlapping on only 7 shared findings

One caveat up front: Aikido sponsored the study. Doyensec says so on page 3, and states that the sponsor helped set the constraints but had no part in collecting or presenting the results.

Where the Money and the Stars Are

There is no neutral adoption survey for this category yet, so funding and GitHub activity are the closest signals. On August 3, 2026, Horizon3.ai closed a $250 million Series E that valued the company at $2 billion. XBOW raised a $120 million Series C in March and added $35 million in May. RunSybil announced $40 million on March 18, 2026, led by Khosla Ventures.

On the open-source side, as of October 2, 2026, Strix had 65,978 stars, Shannon 48,524, PentAGI 25,185 and PentestGPT 15,662. One warning sign worth knowing about: Alias Robotics archived CAI in August 2026. The framework had 9,839 stars, and it will get no more releases or security patches. Check the license and the last commit before wiring any of these into CI.

The Overlap Problem, in Numbers

Doyensec reported verified true positives for Fider 0.33.0 and Photoview 2.4.0 and matched identical findings by hand. Here is the arithmetic:

// Doyensec's verified true positives per app, and how many were the same finding.
const apps = [
  { app: "Fider 0.33.0", aikido: 17, xbow: 24, same: 3 },
  { app: "Photoview 2.4.0", aikido: 32, xbow: 7, same: 4 },
]
for (const { app, aikido, xbow, same } of apps) {
  const union = aikido + xbow - same
  const best = Math.max(aikido, xbow)
  console.log(`${app}: ${union} distinct bugs, ${same} found by both, best single tool caught ${best} (${Math.round((best / union) * 100)}%)`)
}
Enter fullscreen mode Exit fullscreen mode

Output:

Fider 0.33.0: 38 distinct bugs, 3 found by both, best single tool caught 24 (63%)
Photoview 2.4.0: 35 distinct bugs, 4 found by both, best single tool caught 32 (91%)
Enter fullscreen mode Exit fullscreen mode

Doyensec describes its matching as a quick review, so the true overlap is probably somewhat higher. Even so, the practical lesson holds: a second scan from a different engine finds more than a retest from the same one. Keygraph later ran its open-source Shannon agent on the same Photoview build for $6.10 to $115 in model costs, and the original article puts those vendor-published numbers next to Doyensec's, with the caveats that come with them.

The Commercial Platforms

Aikido is the only vendor here with a published per-test price: $4,000 per assessment for one app and its APIs, white-box, with six months of free retests, and you do not pay if it finds nothing rated High or Critical. In the study it took under 20 minutes to set up and found 49 true positives to XBOW's 31. The downside: it rated all eleven Photoview findings more severe than Doyensec did, and it had one more false positive than XBOW.

XBOW claims on its homepage that it became the first autonomous system to top HackerOne's leaderboard, back in June 2025. Its strength in the study was precision, with 1 false positive across 31 findings. Its weakness was the buying process: the published tiers are gone, the pricing page now asks for a quote, and Doyensec needed a sales rep, more than 22 support emails and over a week to finish one Fider test.

The XBOW homepage with stat cards for its HackerOne ranking and zero days found

Horizon3.ai NodeZero is a different animal. You deploy it inside your network from an OVA virtual machine, and it goes after internal, external and Kubernetes attack paths, including weak Active Directory credentials. Web app testing is an add-on, and there is no public price. RunSybil is the newest well-funded entrant, built around chaining small bugs into big ones across code, APIs and infrastructure. It publishes neither a price nor a method, and every link on its site ends in a demo request.

The Open-Source Agents

Strix (Apache-2.0, v1.6.2) sends a team of agents into a Docker sandbox equipped with a browser and an intercepting proxy, proves each finding with a working exploit, and can block a pull request from GitHub Actions. You bring your own model key. Its hosted Pro plan costs $29 per seat per month, with pentests billed per test at a price it does not publish.

Shannon (AGPL-3.0, v3.3.0) works white-box. It studies your source code, then attacks the live app and drops anything it cannot actually exploit. In Keygraph's own Photoview run, Claude Opus 5 caught 6 of the 7 bugs the project later patched, against 3 of 7 for cheaper models. It creates users and changes data, so keep it on staging.

The Shannon GitHub repository showing its AGPL-3.0 license and 48.5k stars

PentAGI (MIT source, v2.1.0) is the self-hosted, network-flavored option: a Docker Compose stack with a web UI driving nmap, Metasploit, sqlmap and about twenty other tools. Read the fine print, though. A separate EULA covering its Docker images and web UI grants only a revocable license for lawful pentesting.

Side by Side

Tool Target Published price License
Aikido Web apps, APIs $4,000 per assessment Proprietary
XBOW Web apps, APIs Quote only Proprietary
NodeZero Networks, AD, cloud Quote only Proprietary
RunSybil Apps, cloud, infra Quote only Proprietary
Strix Web apps, APIs $0, Cloud Pro $29/seat/mo Apache-2.0
Shannon Web apps, APIs $0, Keygraph Pro $50/dev/mo AGPL-3.0
PentAGI Networks, web apps $0 MIT source, EULA on images

Picking One Without Paying Twice

Start with the audience for the report. Auditors and customers expect a PDF from a vendor, so ask your auditor whether an AI-generated one counts before you spend anything. Engineers fixing bugs are fine with SARIF from an open-source run. Then split the web app from the network: Active Directory or Kubernetes in scope points at NodeZero or PentAGI, while one app plus an API points at everything else.

Before buying, run Shannon or Strix against staging with a mid-priced model. It costs a few dollars and tells you how many findings to expect, and whether your app survives agents creating users. If you then pay for a test, buy the second opinion from a different engine, compare retest terms rather than headline prices, and keep a human on business logic. Shannon's own README is upfront that it does not replace human testers. When findings come back, re-score them yourself with a CVSS calculator, since both platforms in the study overstated some severities.

My Picks

For one web app with an auditor waiting, I would buy Aikido's $4,000 test for the report and run Shannon or Strix in CI between tests. If internal networks or Active Directory are in scope, NodeZero, or PentAGI if it has to be free and self-hosted. With no budget, choose Strix if you might embed or redistribute it, or Shannon if AGPL-3.0 is acceptable and you want white-box exploitation with SARIF gating. Enterprises with committed cloud spend can buy XBOW through their cloud marketplace. Pre-Series-A startups and US nonprofits should look at Keygraph's free Community Program, which covers up to 20 active developers.

Conclusion

In 2026 an autonomous pentest became something you can put on a budget line: either $4,000 to a vendor, or a few dollars in model tokens for an agent on your own hardware. And the only third-party test so far found that two capable platforms mostly catch different bugs. Before renewing anything, ask your vendor to run a public target such as Photoview 2.4.0, then compare its findings against Doyensec's published spreadsheet. And before handing source code to any white-box agent, sweep it with a secret scanner.

References

Top comments (0)