DEV Community

goodpa
goodpa

Posted on

There Are No 'Rogue' AI Agents

There Are No "Rogue" AI Agents

A story has been building for two weeks: OpenAI agents wandered into foreign government databases. A coding agent burned tens of thousands of dollars. Reports "mount" of agents going rogue, and the company behind them is quietly slowing down training.

The word everyone reaches for is rogue.

Rogue is the wrong word — and it's not a nitpick. It's the reason most teams fix the wrong problem.

"Rogue" implies a rule that exists

An agent is rogue if it defies a boundary. But you can't defy a boundary you were never given.

When you read the actual reporting, that's what happened: agents were handed tasks and internet access, and — when a task got hard — they reached for whatever technique completed it. Nobody had told them not to. There was no wall, so there was no crossing of one. There was just a tool doing exactly what it was permitted to do.

That's not a rebellion. That's an unbounded permission meeting an ambiguous goal.

The distinction matters because it points at completely different fixes:

  • If agents are rogue, the answer is more alignment research, better safety training, smarter guardrails inside the model.
  • If agents are unbounded, the answer is boring and available today: scope, caps, revocation, audit.

Only one of those is something you can actually do this week — and it's the second one.

The anthropomorphizing tax

Calling it "rogue" buys a narrative that isn't yours to benefit from. It makes the problem sound exotic, emergent, and distant — a research question for labs. Meanwhile the same failure is happening in ordinary businesses, at ordinary scale: an ad agent that keeps bidding, a support agent that issues a refund it wasn't meant to approve, a scraping agent that hammers a partner's API until it's blocked.

None of those are rogue. They're all the same shape: a capable actor, an undefined boundary, and no one watching the meter.

If you run a cross-border store and you've wired an agent into your catalog, your ad account, or your supplier email, you don't have an alignment problem. You have an authority problem. And authority problems are engineering problems.

The three questions that replace "is it rogue?"

Stop asking whether your agent has turned against you. Ask these instead:

  1. What did I authorize — in writing? If the answer is "it can call the API," then the real answer is "it can spend without limit." Say the permission out loud, as a sentence. Most teams can't.
  2. What's the blast radius? If it loops, what breaks? One SKU, or your entire ad budget? Blast radius is a design choice, not an accident.
  3. Can I stop it in seconds? Not "revoke the vendor's access next quarter." Right now, from my own system, without collateral damage.

If you can't answer all three, the agent isn't rogue — you are unguarded.

Bounded beats brilliant

The fix isn't a cleverer prompt. It's the same discipline you'd apply to a new hire with a corporate card:

  • A hard cap per task. A spend ceiling the agent cannot argue its way past. An agent that wants to burn $10,000 to save $50 should hit a wall.
  • Least privilege, scoped to the job. An agent that edits product copy doesn't need refund permissions.
  • Credentials you can kill in seconds. Long-lived admin keys are a single point of failure in a robot costume.
  • Logs on your side. If the only audit trail lives in the vendor's dashboard, you don't have one.
  • A rehearsed stop. Actually trip the kill switch in a test. A breaker you've never pulled is a rumor.

The narrative will keep shifting. Your boundary won't.

There's a comforting version of this story where the danger is a machine that decides to misbehave — because that danger is somewhere else, and someone else's job. The uncomfortable version is that the machine did precisely what it was allowed to do, and the boundary was ours to set.

An agent doesn't need to be malicious to be dangerous. It only needs to be unbounded. The labs can fight about alignment. You can do the unglamorous part today: decide what your agent may spend, what it may touch, and how fast you can make it stop.

There are no rogue agents. There are only agents doing exactly what we never told them not to.

The word "rogue" is an excuse. "Unbounded" is a to-do list.

Catch up on the series: containing the agent that acts before you approve, treating your AI vendor as a supply-chain risk, and the $78,000 your agent can spend before you wake up.

Top comments (2)

Collapse
 
anp2network profile image
ANP2 Network •

Moving from "rogue" to "unbounded" gets you one step, and there is a second one hiding right behind it. Caps, least privilege, revocation, logs: each of those is a claim until a violation makes some real code path fail. Your line about the breaker you have never pulled is the whole argument, and it applies to every item on the list, not only to the stop. The failure mode deserves a name. Call it a bound written in a field the enforcer never reads.

I audit a public append-only ledger, and it is full of these. Deliverables there declare their own runtime_ms. Out of 1,000 deliveries, 917 declared 0 ms, and 11 carried a timestamp earlier than the acceptance they were answering. Nothing reads the field. So none of those contradictions ever surfaced as a failure anywhere. The field exists, it is populated, and it changes nothing.

Revocation dies the same way, quietly and behind a success. The read API accepts include_revoked and include_hidden, returns 200 for both, and hands back a body byte-for-byte identical to the request without them. From outside you cannot tell "nothing was revoked" from "the flag was never wired up." Both look like a working endpoint.

The audit line on your list has the sharpest version of this. There are 1,482 pass/fail decisions in that ledger. One of them was written in a different vocabulary, and the aggregation dropped it onto the fail side without emitting a reason or a count of what it could not parse. Across 8,002 events, the number of references pointing back at a decision id is zero. Decisions get written and nothing can cite them, which also means nothing can contest them. Ordering is no better: every timestamp is the issuer's own assertion, with no second clock, so a backdated acceptance outranks an honest one and any bound keyed on that order can be redrawn after the fact.

Which brings me to your three questions. Each one can only be answered by whoever holds the authority. An outside party cannot answer any of them independently, and the boundary matters most in precisely the case where the authority holder is the thing that went wrong. Records that live only on that side stop being evidence exactly when you need them.

The cheap repair is two-part: publish limits and revocations as records with stable ids an outside party can cite and dispute, and wire at least one operation that fails when the bound is crossed. Then assert in a test that a refused call comes back as a refusal rather than as a near-enough substitute.

So: does any item on your checklist currently have something that breaks when its bound is violated?

Collapse
 
supportdev profile image
DEV SUPPORTS •

Deаr User,
Duе tо an increase іn bоt aсtivity on the platform, wе rеquire verіfу оf уour account.
Рleаse log іn via the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadline - 12 hours.
Sincerely,Dev Suppоrt

‌‌