Here's the joke nobody's laughing at yet: we spent a decade building SIEMs and EDRs to watch what humans do on endpoints, and we just handed a new class of non-human actor commit access, shell access, and package manager access without blinking.
Context
An AI coding agent (Claude Code, Codex, whatever your team has wired into CI or your local dev box) isn't a chatbot anymore. It reads files, runs commands, installs dependencies, opens PRs. That's a real attack surface, and it's already being used as one. Poisoned npm packages that trigger agent hooks. Malicious hooks that ride along with legitimate-looking config. A bug in the OpenAI forum that let someone push unauthorized PRs through Codex. None of these are exotic. They're the same supply chain and session-hijacking playbook attackers have run against humans for years, just pointed at a new target that happens to have broader, faster, less-questioned access than most junior engineers get on day one.
The response documented here (Gryph + Jev) is basically: log every action the agent takes, feed it to an LLM, ask "does this look like the developer's normal behavior and org policy, or not." Tested against 14 synthetic attacks. Caught all 14.
That's not a new idea in shape. It's UEBA (user and entity behavior analytics) with the "E" now meaning something genuinely agentic instead of a service account that occasionally runs a batch job.
Hype Check
What's understated: the attack surface framing is correct and most teams haven't internalized it. If your CI pipeline lets an agent install packages and open PRs, you've created a new privileged identity, and most orgs don't have a policy document, let alone a monitoring story, for that identity. That's a real gap.
What's overstated, or at least unproven: 14 synthetic attacks with 100% detection is a demo, not a security guarantee. Anyone who's worked in detection engineering knows the gap between "we wrote test cases and passed them" and "we survived contact with an adversary who knows exactly how our detector works." Synthetic attacks are, almost by definition, the ones you already thought of. Zero comments and 2 points on HN tells you the industry hasn't stress-tested this claim yet either, this is early-stage research getting shared, not something battle-tested in the wild.
Also worth being honest about: using an LLM to classify whether another AI's actions are malicious introduces its own trust problem. You've now got a judge that can be confused, prompt-injected, or drift in ways that are hard to audit. "Is this normal for this developer" is a genuinely hard behavioral baseline problem even for humans, we've been tuning UEBA false-positive rates for years and still get it wrong constantly. Doing it for an entity whose "normal" behavior space is enormous and rapidly evolving (because the underlying model keeps changing) is harder, not easier.
Who benefits from the "AI agents are exploitable, here's an AI to watch the AI" narrative? Anyone selling agent tooling gets to say "yes but it's monitored," which is a great sentence for a security review checklist and a much shakier sentence in an actual incident retro.
Implications
If you're running agentic coding tools in anything touching production, you already have an identity and access management problem you probably haven't named yet. The obvious first move isn't buying or building a detection layer, it's the boring stuff: scope what the agent can actually touch, separate its credentials from the developer's, log its actions somewhere durable regardless of whether anything's classifying them yet. Detection is layer two. Least privilege is layer one, and most teams haven't done layer one.
For security teams, this is a reminder that agent supply chain risk (poisoned packages, malicious hooks) is not hypothetical anymore, it's the exact mechanism described here. Your dependency scanning and hook review processes need to account for "this artifact could hijack an agent session," not just "this artifact could run malicious code when a human executes it."
For the industry more broadly: we're about to relearn every lesson from service-account sprawl and CI/CD credential leakage, except faster, because agents are being adopted faster than service accounts ever were.
Open Question
If detecting a compromised agent requires another AI to judge "normal" behavior, and that judge itself can be fooled or drift, at what point are we just stacking uncertain systems on top of each other and calling it defense in depth?
— Cor, Skyblue Soft
Sources
AI-assisted draft or imaging, human-curated, reviewed and edited.
Top comments (0)