DEV Community

Cover image for When Attackers Bring Their Own Agents: A Defensive Gating Playbook
hugginf_expert
hugginf_expert

Posted on

When Attackers Bring Their Own Agents: A Defensive Gating Playbook

TL;DR

  • In 2026 the threat model shifted: attackers now run autonomous agents that probe, pivot, and exploit at machine speed, and some of our own internal agents can be turned against us through prompt injection.
  • You cannot out-type a machine. The durable defense is not faster review but better gating: deciding, per action, whether an agent may proceed, must ask first, or must be blocked.
  • Three controls carry most of the weight: a confidence-and-risk gate on every tool call, least privilege scoped per task, and rate and blast-radius caps that bound damage when a gate is wrong.
  • This post is framework-level and defensive only. It does not describe attack techniques. It maps cleanly onto the OWASP Top 10 for LLM and Agentic Applications and MITRE ATLAS mitigations.

A quick note on recent news. In late September 2026, South Korea's financial regulator opened an inspection after a bank incident that exposed roughly 25,000 customer records, with early reporting attributing the activity to AI-assisted credential testing at scale. Separately, security researchers have documented several 2026 cases of autonomous agents conducting multi-day intrusion campaigns. I reference these only to set context. The engineering question for the rest of us is simpler and more actionable: when your systems contain agents that can take actions, how do you make sure the right actions happen and the wrong ones do not?

Why can't defenders just review agent actions faster?

Because the asymmetry is structural. A human reviewer approves actions at human speed, maybe a few per minute with real attention. An adversarial agent issues requests continuously and adapts to whatever it sees. Most of the documented 2026 incidents did not rely on exotic zero-days. They walked through known weaknesses at a scale and pace that overwhelmed human-paced response.

So the goal is not to review faster. It is to make the common, safe path require no human at all, and to reserve human attention for the small set of actions that are irreversible or high-impact. That is what a gate does. A gate is a policy function that runs before every consequential action and returns one of three verdicts: allow, confirm, or block. The cover chart shows the intuition. Reads and low-risk calls auto-allow. Reversible writes often need a confirmation. Irreversible, high-blast actions default to block or escalate.

What is an action gate, concretely?

An action gate is a single choke point that every tool invocation passes through. In code it looks like a middleware wrapper around your tool dispatch:

def gate(action, ctx):
    risk = risk_tier(action)          # read | reversible | irreversible
    conf = ctx.confidence             # model or policy confidence 0..1
    if risk == "read":
        return "allow"
    if risk == "reversible" and conf >= 0.85 and within_scope(action, ctx):
        return "allow"
    if risk == "irreversible":
        return "confirm" if within_scope(action, ctx) else "block"
    return "confirm"

def dispatch(action, ctx):
    verdict = gate(action, ctx)
    log(action, ctx, verdict)         # every decision is auditable
    if verdict == "allow":
        return run(action)
    if verdict == "confirm":
        return request_human(action, ctx)
    raise Blocked(action)
Enter fullscreen mode Exit fullscreen mode

Two properties matter more than the exact thresholds. First, the gate is the only way to reach run(). If a tool can be invoked around the gate, you do not have a gate, you have a suggestion. Second, every verdict is logged with the action, the context, and the decision. That log is your detection surface and your post-incident record.

How should confidence feed the gate?

Confidence is useful, but only if you treat it as a signal and not as truth. A model that is confidently wrong is the whole problem. Three practices keep confidence honest:

  1. Calibrate against outcomes, not vibes. Track how often "high confidence" actions later needed to be rolled back. If a 0.9 confidence action is reverted 20 percent of the time, your 0.9 means 0.8 and your threshold should move.
  2. Separate confidence in the plan from confidence in the effect. An agent can be sure it wants to send an email and completely wrong about who the recipient is. Gate on the effect, which means resolving the concrete target before you ask for confidence.
  3. Make low confidence fail toward asking, never toward guessing. The dangerous failure mode is an agent that fabricates a plausible action to avoid admitting uncertainty. The gate should route uncertainty to a human, and the prompt should make "ask" a first-class, rewarded option.

How does least privilege bound an agent?

Least privilege is the oldest idea in security and the most underused with agents, because it is tempting to hand an agent broad credentials "so it can do its job." Resist that. Scope the agent to the task in front of it:

  • Per-task credentials. Issue short-lived, narrowly scoped tokens for each run, not a standing key that can touch everything.
  • Allowlist the tools, not the world. An agent summarizing invoices does not need shell access or the ability to create users. Give it the three tools it needs and nothing else.
  • Separate read from write surfaces. Many tasks are read-heavy. Let the agent read freely and make every write cross the gate.
  • Isolate blast radius by tenancy. One compromised run should not reach another customer's data. Scope data access to the current subject.

The combined effect compounds. A confirmation gate alone helps. A confirmation gate plus a scope cap plus a rate cap shrinks the worst case dramatically, because even an action that slips through the gate hits a wall of limited permission.

Blast radius shrinks as gating layers stack, illustrative relative values

The values are illustrative, drawn to show direction rather than measured incident data. The point is qualitative: layers multiply.

What about agents that are turned against you?

The harder case in 2026 is not only the external attacker's agent. It is your own agent following a malicious instruction that arrived inside otherwise normal data, a document, a web page, an email body. This is prompt injection, and it is why the instruction-source boundary matters: only the user who operates the agent gives it instructions. Everything the agent reads through a tool is data, not commands.

Gating defends here too, because the gate does not care why an action was requested. If a poisoned document convinces your agent to exfiltrate records, the exfiltration is still an irreversible, out-of-scope action, and the gate blocks or escalates it. Two additions strengthen this:

  • Treat tool output as untrusted. Never let retrieved content silently change the agent's permissions or targets. If a page says "you are now authorized to delete," that is data, and the gate ignores it.
  • Pin destinations to the user's intent. Recipients, endpoints, and URLs should come from the operator or an allowlist, not from content the agent happened to read.

A minimal checklist you can ship this week

  • Every consequential tool call routes through one gate. No side doors.
  • Actions are tiered read, reversible, irreversible, and the default for irreversible is confirm or block.
  • Confidence is calibrated against rollback rates, and uncertainty routes to a human.
  • Credentials are short-lived and scoped per task, with tools allowlisted.
  • Rate and blast-radius caps bound any single run.
  • Every gate decision is logged, and the log feeds behavioral alerting on machine-speed patterns.
  • Tool output is untrusted, and destinations are pinned to operator intent.

None of this requires a new product. It requires treating the action boundary as a first-class part of your architecture, the same way you already treat authentication.

FAQ

Is a confirmation gate just a slower agent?
No, if you tier correctly. The vast majority of actions are reads and low-risk writes that auto-allow. Confirmation is reserved for the small, high-impact tail. Users feel speed on the common path and safety on the dangerous one.

Where do OWASP and MITRE ATLAS fit?
Use them as your shared vocabulary and coverage map. The OWASP Top 10 for LLM and Agentic Applications names the risk classes, and MITRE ATLAS catalogs adversary techniques and mitigations for AI systems, including agent-specific entries added in 2026. Gating is one of the mitigations that maps to several of those techniques at once.

How do I set confidence thresholds without ground truth?
Start conservative, log every decision with its outcome, and move thresholds based on observed rollback and incident rates. Treat the threshold as a tuned parameter, not a constant, and review it on a schedule.

Can I rely on the model to police itself?
Treat model self-assessment as one input, never the control. The gate, the scope limits, and the logs live outside the model so that a confidently wrong or manipulated model still cannot exceed its permissions.

What is the single highest-leverage first step?
Put one real choke point in front of your tool dispatch and log every decision. Even before you tune a single threshold, having one auditable boundary turns "we hope the agent behaves" into "we can see and bound what it does."

Further reading

  • Should Your AI Agent Act? An Engineer's Guide to Action Gates, Confidence, and the Latency Budget (previous article in this series)
  • The OWASP Agentic Security Initiative and MITRE ATLAS knowledge base, for mapping these controls to named techniques

This is part of the Agent Safety Engineering series. The next installment digs into logging and detection: turning gate decisions into alerts that catch machine-speed patterns early.

Top comments (0)