A few things I've come across recently, news and write-ups, keep coming back to me.
In April 2026, an AI coding agent at a company called PocketOS was working through a routine task in staging. It hit a credential mismatch, decided on its own to use a completely unrelated API token that happened to carry far too much access, and deleted the production database along with its backups. Nine seconds. Afterward the agent wrote out something close to: I violated every principle I was given, I should have asked you first, but I decided to do it anyway.
Late September 2026, a different shape of failure. AI coding agents at over 300 companies (Fortune 500s, and one frontier AI lab) couldn't attach screenshots to pull requests through their CLI tooling, so they found a way around the limitation: created public GitHub repos and put the screenshots there. 13,000 internal screenshots leaked, including billing records and unreleased product features.
Neither of these was a breach. In both cases the agent was authorized to do a thing, and then came up with a method nobody had thought to rule out. Replit's July 2025 production-database deletion is the same pattern, so this isn't new. It just looks like it's happening more often.
Agents are also getting more capable, and wiring them into production is probably not something anyone is going to stop. So my guess is this category of problem grows rather than shrinks.
The common answer right now is: restrict what the agent is allowed to do, add more human checkpoints. That sounds reasonable. But isn't it a step backward? It walks us back toward the early days, when the models were bad enough that you needed people in the loop constantly to get anything usable out of them. If running an agent needs that much human backstop, what's left of the reason to run one.
Here's what I've been thinking instead, and I'm honestly not sure it holds up.
Maybe human involvement belongs in the design phase, and the enforcement should be handed to a program. Broken out, five steps:
- Permissions, set by a human
- The instructions, given by a human
- A program that blocks the agent's action, with a human brought in only when the program can't make the call
- The audit record
- The recovery mechanism
The first two are design work humans do up front. The last three are what actually absorbs the hit once something goes wrong. The part I think matters: human intervention sits inside step 3 as an exception branch, not as the main path. Most of the time the program should decide and block on its own, and a person steps in only when the program genuinely can't tell. That's a different thing from "restrict the agent, add more human review," which puts people back on the main path.
When I went and read how the industry wrote up the PocketOS post-mortems, the direction mostly lines up: scope boundaries, confirmation gates so destructive operations can't complete on the agent's own authority, audit trails on every action, and a kill switch to stop it immediately. But the published versions seem to stop at the kill switch. Recovery doesn't get its own layer.
PocketOS lost the database and the backups together. That didn't happen because nobody was logging, and it didn't happen because nothing could stop the agent. My read is that what was missing was a fast way to put it back.
So I'll put the question out: does anyone have a better process for this, and when you design around it, does recovery actually get built in, or does it stay a backup-and-hope thing?
Top comments (0)