DEV Community

Manos Saratsis
Manos Saratsis

Posted on Originally published at dromeas.ai

What Is AI SAST? How Agentic Code Review Differs From Pattern-Matching Static Analysis

Originally published on the Dromeas blog.

In short: classic SAST matches known-bad patterns, fast and repeatably. AI SAST adds reasoning about data flow and intent across files, which catches behavioral bugs rules miss, but it's less deterministic. It's best treated as a layer on top of rules, not a replacement.

Classic SAST AI SAST
How it finds issues Rules and known-bad patterns Reasons about data flow, intent and context
Same code, same result? Yes Not always, so it needs cross-checking
Strong at Injection, secrets, known CWEs Logic, authorization and cross-file bugs
Weak at Business logic, behavior Explaining itself unless evidence is required

"AI SAST" is everywhere this year. Checkmarx has a learn page on it, Augment Code wrote a 2026 guide, and plenty of other AppSec vendors have picked up the label. Some of that describes a real change in how code gets checked. Some of it is a rule engine with a chatbot bolted on.

I build a code review product, so I have a stake in this. Here's how I'd tell the two apart.

What classic SAST is good at

Static application security testing reads your source code without running it and matches it against a library of known bad patterns. A SQL query built by gluing strings together. A hardcoded API key. An outdated crypto function. It doesn't need to know what your code is for. It just needs to recognize the shape of a known problem.

That's useful. It's fast, it's cheap to run across a big codebase, and it's deterministic: same code in, same result out. As a CI gate that blocks obviously risky patterns before a person ever looks at the PR, rule-based SAST is still a good tool.

The limit is built in, though. A rule engine only finds what someone already wrote a rule for. Even the SAST vendors say this: if the rules are incomplete, things get missed.

What it can't reach

A lot of the bugs showing up in AI-generated code aren't known bad patterns at all. We went into this in our post on behavioral bugs. It's code that runs, passes the linter, looks fine in review, and then does the wrong thing for an input nobody tried. An off-by-one on a boundary. A null check on the wrong variable. Two functions that disagree about whether a list has already been deduplicated.

Nothing about that code is syntactically wrong, so there's nothing for a pattern to match. To catch it you have to understand what the code is supposed to do and compare that with what it actually does.

New Relic's 2026 State of AI Coding report shows what this looks like in practice. 94% of the tech leaders they surveyed rated AI-generated code as higher quality than human code at review time. And 82% had at least one major production failure tied to AI code in the previous six months. Code that looks good in review and breaks in production is exactly the kind of bug static analysis was never going to find.

What AI SAST actually adds

Augment Code's guide makes a useful distinction between two kinds of tools:

  • AI-assisted: the rule engine still does the detecting, and a model helps with triage, prioritization and fix suggestions.
  • AI-native: a model is the detector. It reads the code and reasons about what it does, more like a reviewer than a linter.

The second kind is the interesting one. A model can follow data across several functions and files, work out what a change affects downstream, and judge whether something is a real problem in context. Multi-step issues that only show up when you trace the whole path are within reach in a way they aren't for a fixed rule set.

Where it falls short

This is the part the marketing tends to skip.

It isn't deterministic. Run the same review twice and you can get two different answers. That's a problem if you need consistent results for an audit.

It costs more and it's slower. A model call per file or per PR costs more than a rule sweep, and it takes longer.

One model is still one opinion. A single model judging intent will produce false positives. They just look different from a rule engine's.

Developers already feel this. In Sonar's January survey, 96% said they don't fully trust that AI-generated code is functionally correct, and 38% said reviewing AI code takes more effort than reviewing a colleague's. If people don't fully trust a model's code, it's fair to ask why they'd trust a single model's verdict on a vulnerability.

Questions to ask a vendor

  • Is the model doing the detecting, or just triaging what the rules found?
  • What happens if you run it on the same PR twice?
  • Is more than one model checking a finding before it reaches you?
  • Does it look at what a change affects elsewhere in the codebase, or only at the diff?

Where Dromeas fits

This is why we don't send every finding through one model. Each PR, trunk commit and local diff reviewed by Dromeas is checked by a council of models (Anthropic, OpenAI's GPT-5, Google's Gemini 2.5 Pro, DeepSeek V4 Pro) that verify each other's findings before anything reaches you. They reason over a code map of your repo (files, symbols, call graph), so they see what a change actually touches, not just the diff on its own.

Known bad patterns (OWASP Top 10, CWE, supply chain) are covered on our code security side. The behavioral bugs that pattern matching can't reach are what Bug Tracing is for. They're aimed at different mistakes, and you want both.

AI SAST is a real shift. Just make sure whoever's selling it to you has a plan for checking the model too.

— Manos

Sources: Checkmarx, AI SAST · Augment Code, AI SAST 2026 guide · New Relic, 2026 State of AI Coding · Sonar survey, Jan 2026

Top comments (0)