Google’s Mantis caught my attention because it points to a problem most AI security demos quietly walk around: finding a vulnerability is not the same as proving one exists.
That distinction matters. Security teams already live with noisy scanners, half-useful alerts, and findings that require someone experienced to separate a real exploit path from a theoretical complaint. Adding an LLM can improve that workflow. It can also make the scanner more fluent while being wrong in more elaborate ways.
That is not progress. That is just a better-written interruption.
The interesting part of Mantis is not the label “agentic vulnerability scanning.” Everyone is attaching agentic to things now. Apparently software is not allowed to have a normal workflow anymore. The useful part is the structure around grounding: repository context, history, threat models, reviewer stages, critic stages, sandboxed reproduction, and patches that are tied back to evidence.
That shape makes sense because vulnerability detection is not one job. It is several jobs that often get collapsed into one vague “scan the code” box.
The reproduction step is the part I would care about most.
If a system can show a working crash, a failing test, an exploit path, or a concrete data-flow issue, the review conversation changes. The security team is no longer reading model confidence. They are reading evidence. That does not remove human judgment, but it gives the human something useful to judge.
Most teams do not suffer because they have too few alerts. They suffer because the alerts arrive without enough context. Someone has to open the repository, understand the service boundary, trace input handling, inspect validation, check the framework behavior, look at previous fixes, decide whether the path is reachable, and then argue with the scanner’s output. That is expensive work.
AI can help with that work, but only if the system is designed as a workflow rather than a magic box.
For example, a scanner that says:
Possible SQL injection in UserController.java
is not enough. A better system should be able to say:
This route accepts user input here.
The value reaches this query builder here.
This sanitizer does not cover this pattern.
This is a minimal reproducer.
This is the failing test.
This is the proposed fix.
This is the residual uncertainty.
That last line matters. Residual uncertainty is not weakness. It is honesty. A security tool that pretends every finding is equally certain creates bad incentives. Engineers start ignoring it, security teams start tuning it down, and eventually the tool becomes background noise.
The useful promise of an agentic scanner is that it can break the work into stages and let each stage challenge the previous one. One agent may identify suspicious flows, whereas another may criticize the finding, and another may try to reproduce it. Another may propose a patch. A final stage may check whether the patch changes behavior outside the intended area.
That is much closer to how a careful human review works.
It also changes how I would think about model selection. Not every stage needs the biggest model available. Classification, deduplication, clustering similar findings, and summarizing repository structure are not the same as tracing a subtle authorization bypass through five layers of application code. A practical system should spend reasoning budget where reasoning is actually needed.
That is a very SRE-ish way to think about AI, and I mean that as a compliment. The model is not magic. It is one component in a workflow with cost, latency, permissions, state, logs, and failure modes.
The cost piece is easy to ignore in a research announcement and painful to ignore in a real engineering organization. If every pull request triggers a deep multi-agent investigation across a large monorepo, the bill will get interesting very quickly. The system needs triage. It needs cheap filters, expensive analysis only when justified, and clear rules for what runs synchronously in CI versus what runs asynchronously in a security pipeline.
There is also a timing question. Some checks belong directly in the developer loop. They should run fast, fail clearly, and produce a result while the developer still remembers what they changed. Other checks are better as background analysis. A deep investigation across service boundaries may be valuable, but it probably should not block every commit unless the organization is prepared for that operational cost.
This is where the design should separate developer feedback from security investigation.
Developer feedback needs speed and clarity. Security investigation needs depth and evidence. Trying to make one workflow satisfy both usually produces a system that is too slow for developers and too shallow for security teams. That is a familiar failure mode. We have seen it with static analysis, dependency scanning, data-quality checks, and policy-as-code.
AI does not remove that trade-off. It just makes it easier to hide for a while.
I would probably separate the workflow like this:
This keeps the expensive reasoning stages closer to the findings that deserve them.
There is still plenty of operational work hiding behind the nice diagram. The scanner needs sandboxing. It needs restricted network access. It needs deterministic handoffs between stages. It needs audit logs. It needs a way to prevent an aggressive false-positive filter from suppressing weak-looking findings that are actually real. It also needs ownership, because once a tool starts filing security bugs or proposing patches, someone has to decide what “good enough” means.
This is where security automation often becomes uncomfortable. A tool that only reports findings is easy to ignore. A tool that opens patches is harder to ignore, but also more dangerous. The patch might fix the immediate issue while changing behavior somewhere else. It might silence a test instead of fixing the cause. It might introduce a different vulnerability. It might be correct technically but wrong for the product’s authorization model.
Authorization bugs are a good example. They rarely live in one obvious line of code. The check may depend on route configuration, middleware behavior, tenant context, cached permissions, database filters, and assumptions in the UI. A scanner that sees only the controller method may miss the real boundary. A model with repository context may do better, but only if the system gives it the right evidence and then forces it to prove the path.
The same is true for deserialization, SSRF, file access, and dependency confusion. The interesting question is often not “does this function look suspicious?” It is “can untrusted input actually reach this dangerous capability under realistic conditions?” That requires reachability analysis, test construction, environment modeling, and sometimes domain knowledge about how the application is deployed.
This is where agents can be useful, but also where they can become overconfident. A fluent explanation of a possible exploit is not the same as exploitability. I would rather have a scanner that says “I cannot prove this yet” than one that files a confident but ungrounded critical finding.
So I would not give an AI scanner unlimited write access to a production codebase. I would start with evidence generation, then move to suggested patches behind review, then allow more automation only for narrow, well-tested classes of changes.
Something like:
- Phase 1: identify and explain
- Phase 2: reproduce with tests
- Phase 3: suggest patches
- Phase 4: auto-fix low-risk patterns
- Phase 5: measure escaped defects and false negatives
The important part is that automation earns trust through measured behavior. Not vibes. Not a launch post. Measured behavior.
The metrics should reflect that. I would track more than “number of vulnerabilities found.” That metric is too easy to game. A useful program would track confirmed true positives, false-positive rate by category, time from finding to reproduction, time from reproduction to patch, developer review burden, escaped vulnerabilities, and whether suggested fixes survived regression testing.
Those numbers tell you whether the system is improving the security process or merely producing activity.
There is also a governance concern. If the scanner learns from internal code, writes summaries to disk, executes reproducers in sandboxes, and uses multiple models, the organization needs to know where sensitive data goes. Security tooling often has broad repository access by design. That makes data handling, retention, model-provider boundaries, and access logs part of the architecture, not an appendix.
I would treat an AI security scanner almost like a privileged internal service:
That is more work than running a command-line scanner. But if the system is going to reason over sensitive code and propose security patches, the extra discipline is not optional.
The bigger issue is that AI security tooling can easily become another alert generator. Teams do not need more findings. They need better evidence, better prioritization, and a shorter path from suspicion to validated fix.
That is where agentic workflows may actually help. Not because an agent can read a lot of files, although that helps. Not because it can write a patch, although that is useful. The real value is in connecting the steps that human reviewers already perform manually:
- understand context
- form a hypothesis
- challenge it
- reproduce it
- and only then act
There is also a cultural side to this. Developers will trust the system faster if it explains itself in the language of the codebase. Security teams will trust it sooner if there is reproducible evidence. Engineering leaders will trust it faster if it reduces mean time to validated fix without flooding teams with noise. Those are different success metrics, and a serious platform needs to satisfy all three.
My current view is simple: AI belongs in security scanning when it is treated as an evidence-generation system.
The agent can suggest. The harness has to prove. That is enough for the first serious version.
Only after that, let us talk about autonomy.
| References:
- Inspired by InfoQ’s coverage of Google Mantis: https://www.infoq.com/news/2026/09/google-mantis-vulnerability-scan/
- Related trend context: InfoQ’s September 2026 AI, ML, and data engineering coverage highlighted agentic testing, security, and production-readiness themes.
Top comments (4)
Phase 2 writes the reproducer and phase 3 writes the patch. When one generator authors both, the reproducer can be fitted to the patch, and the success criterion for phase 3 ends up authored by the thing being graded. The failure has a cheap and concrete shape: a reproducer that asserts on an exception type or on a substring of an error message, then a patch that rewrites the message. Assertion goes green. Input path is still reachable. Closing that gap is a property of the handoff, and no amount of model quality substitutes for it. The phase-2 artifact has to be frozen and content-addressed before the patch stage can read it, and the post-patch run has to execute those same bytes.
The abstention state has a similar problem one layer down. "I cannot prove this yet" has to be written somewhere, and the shape it gets written in decides whether it survives. Stored flatly as "not reproduced", it becomes a deduplication key, so the next run matches against it and skips the re-investigation. The honest answer hardens into a permanent one. What needs to persist alongside it is the reason, and the common reasons are not the same kind of fact. No reachable entry point found is a claim about the code. A sandbox missing a dependency is a claim about the harness. Budget exhausted at stage four is a claim about scheduling. Only the first should be sticky. Let the other two persist and a bad afternoon in the harness reads back later as a clean bill of health for the code. Phase 5 will not surface that, because escaped defects are counted against what shipped, and a finding suppressed by deduplication never reaches the denominator.
The tiering argument has a direction problem. Small models get classification and clustering, the large one gets the five-layer authorization trace. That splits on reasoning difficulty, which is defensible on its own terms. But triage is also the one stage where dropping a finding is irreversible, since nothing downstream can recover what the filter threw away, whereas an over-eager deep stage burns money and reviewer attention and both of those come back. So the cost-of-error asymmetry points the opposite way from the difficulty asymmetry. One way to satisfy both is to let the cheap stage decide ordering and never membership. It ranks, the budget cuts the tail, and a dropped finding becomes a recorded queue position while no verdict ever gets written, which makes "what did the filter suppress" answerable by replaying the tail. Clustering sits awkwardly in the cheap tier for a related reason. Deciding that two findings share a root cause is the reachability question again, in different clothes.
After phase 3 lands a patch, is the phase-2 reproducer re-run as the byte-identical artifact, or regenerated?
Reproduce-before-patch improves confidence, but automated scanners still have limits when attacks require multi-step adversarial intent. Scanning finds known-bad patterns; adversarial testing explores novel attack chains both need to be part of the assessment scope.
"A better-written interruption" is the exact failure mode of adding an LLM to a noisy scanner: fluency scales, correctness doesn't, and the triage burden lands on the same tired reviewer with nicer prose. The reproduction step is the right thing to anchor on - a working crash, a failing test, or a concrete data-flow trace converts the conversation from "is this real" to "how do we fix it," and every stage before it is just building confidence in a claim. Findings tied back to evidence is the difference between a security tool and a security horoscope.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.