DEV Community

Cover image for A CVE Is Not an Exploit: Why "Reachable" Matters More Than "Vulnerable"
Fernando Gallardo
Fernando Gallardo

Posted on

A CVE Is Not an Exploit: Why "Reachable" Matters More Than "Vulnerable"

Every week, some dependency scanner flags a critical CVE in your stack, and every week the same ritual plays out: bump the version, re-run the pipeline, close the ticket. Nobody asks the harder question, could anyone have actually reached that code path from outside our system?

That question is the difference between patching because a scanner told you to, and knowing whether you were exposed to remote code execution while you weren’t patched.

The finding

We ran this against a testbed repo from an earlier iteration of our own agent stack, one that has since moved off agno entirely, which is exactly why it’s a clean example to walk through publicly. The scanner flagged agno==2.3.4, a Python package used for agent orchestration, with a critical-severity issue: an arbitrary code execution vulnerability in the model-execution component. The root cause is a classic one, a field_type parameter gets passed into eval() without being constrained to a safe set of expected values. If an attacker controls the content of field_type, they don’t get a type name back; they get arbitrary Python code executed with the privileges of your process.

That’s the CVE. Here’s the part a dependency list doesn’t tell you: does anything in your codebase pass untrusted data into that parameter?

What a version bump doesn’t answer

A dependency scanner (Dependabot, npm audit, Snyk’s base SCA, whatever you’re running) does one job well: it tells you a known-vulnerable version is present in your dependency tree. That’s necessary information. It’s also, on its own, close to useless for prioritization, most teams carry dozens of “critical” CVEs in transitive dependencies they never directly exercise. Triaging by severity score alone means you either patch everything immediately (unsustainable) or patch nothing until someone notices (dangerous), because severity tells you how bad the vulnerability is, not how exposed you are.

Exploitability, for a real system, needs two things to line up:

  1. An entry point an attacker controls: A public endpoint, a webhook payload, a file or PR someone uploads, a message from a bot integration. Anything that crosses your trust boundary.

  2. An unbroken data path from that entry point to the vulnerable sink, in this case, to whatever constructs the field_type value that reaches eval().

Neither of those shows up in a CVE database. They live in your call graph.

What we actually traced

When we ran that repo through Ixtli Code Graph, our whole-program static analysis engine, it surfaced the blast radius as it stood at the time: 25 functions across 4 files called into agno directly, including orchestration and agent code (analyze_code, analyze_dependencies, analyze_iac, create_orchestrator, get_file_type, and roughly twenty more).

That list is the starting point, not the conclusion. The next question, the one that actually determines exploitability, is whether any of those callers sit downstream of an untrusted input. For a service like ours, that means: does content from a PR, a repo, or a webhook payload we accept from a third party ever flow, unsanitized, into the arguments we pass to agno‘s model-execution component? If the answer is yes for even one of those 25 functions, the CVE isn’t theoretical, it’s a live path from an external input to code execution on our infrastructure. If the answer is no across all of them, if every call site only ever receives internally-generated, well-typed values then the CVE is present but currently inert, and you can prioritize accordingly instead of treating it as a fire drill.

That’s the distinction reachability analysis gives you that a flat dependency list can’t: Not “this package is vulnerable“, but “this package is vulnerable, and here is the exact set of places you need to check to know if it matters“.

Why this also isn’t a pentest’s job

A pentest simulates an attacker working against your system as deployed: what can be reached from outside, what an authenticated user can escalate into, what business logic can be abused. That’s a real and necessary discipline, and it catches things static analysis never will, session handling flaws, business logic abuse, chained privilege escalation across services.

But a pentest has a scope, and scope is defined by what someone thought to test. An internal reporting function, an import script that concatenates strings without sanitizing them, an admin panel nobody exercised during the testing window, these don’t fail a pentest because they’re secure. They don’t get tested at all, because nobody thought to include them. The vulnerability isn’t disproven; it’s just outside the sample.

Whole-program code analysis and pentesting are answering different questions: “what can an attacker do with what they can see?” versus “what does this code actually do with whatever reaches it?” Confusing one for a substitute of the other is exactly where risk hides, vulnerabilities that ship to production, sail through the pre-release pentest, and sit there for months, not because no one looked, but because no one was looking for that.

The practical takeaway

If your security process stops at “no critical CVEs in the dependency report,” you know less than it feels like you know. The questions worth asking on every critical finding are:

  • Which functions in my code actually call into the vulnerable component?

  • Do any of those functions sit downstream of an input I don’t control?

  • If yes, that’s your real priority queue, not the CVSS score.

A vulnerability list tells you what exists. A reachability map tells you what to fix first, and why.


ixtli.app

Top comments (0)