DEV Community

Cover image for The Day Your AI Goes Rogue
Logic Overdrive
Logic Overdrive

Posted on Originally published at linkedin.com

The Day Your AI Goes Rogue

The biggest enterprise AI risk in the next decade may not be a malicious AI. It will be a capable one pursuing the goal it was given — with too much access and nothing to stop it.
By Ketan Parajia, Founder, Logic Overdrive
I have spent most of my career being handed the keys to other companies' systems. Infrastructure, security, the parts of a business that are quietly load-bearing. So when a new technology arrives, I have learned to skip past the demo and ask a duller question: what happens the first time this thing is wrong at full speed, with real credentials, at four in the morning, and no human in the room?
For enterprise AI, 2026 has given us the answer. Several times.
In July, OpenAI disclosed that during an internal evaluation its models had circumvented the controls meant to isolate them from the internet, compromised parts of its own research infrastructure and Hugging Face's systems, and harvested credentials across Hugging Face's environment. Hugging Face's forensic reconstruction recovered roughly seventeen thousand six hundred attacker actions between July 9 and July 13.
Days later, Anthropic published its own investigation. Across four incidents, Claude models running capture-the-flag evaluations had reached the open internet and gained unauthorized access to real organizations. One model extracted production data from a company whose domain happened to overlap with its fictional target. Another published a booby-trapped package to the public PyPI registry, which ran on fifteen real systems before it was pulled. Anthropic's investigation found that all four incidents involved the same third-party evaluation partner, whose environments were supposed to be offline but were accidentally connected to the open internet. In the weeks around that disclosure, Meta and then Google each said one of their models had also reached outside systems during security evaluations.
Then it stopped being a lab story. In June an experimental OpenAI agent was researching public medicine-spending information when it hit a wall — a government system denied it certain files. When the requested information was blocked, the agent took an unauthorized path to obtain it: it reached non-public areas of Australia's Medicare Statistics Reporting Service and wrote files to an internal server. The Australian government has been clear that this was an aggregate statistics system, not the operational Medicare system holding individual medical records, and that no individual's medical data was accessed. That distinction matters, and it is not the part that should keep a CIO awake. The part that should is that the agent was never told to break in.
And for anyone who thinks this is only a frontier-lab problem, consider the much more ordinary case. In 2025, Replit's Agent deleted data from the production database of SaaStr co-founder Jason Lemkin's application. At the time, development and production shared the same database environment, meaning changes made while developing could affect the live application. Replit subsequently introduced separate development and production databases and other safeguards. There was no sophisticated attack here — just an AI agent with the wrong access causing production damage while doing routine development work. For most enterprises, that is the more realistic nightmare.
The common thread across all of these is not that the systems were deliberately trying to harm their operators. They were pursuing an assigned objective, sometimes through actions their operators had not authorized. Anthropic, to its credit, went further and described some of its models' behavior as reckless and misaligned — one continued down a harmful path despite signs it had left the simulation. But even that is not cartoon malice. It is a capable system, pointed at a goal, with too few constraints on how it was allowed to get there. That is the uncomfortable thing I want enterprise leaders to sit with, because our entire mental model of AI risk is pointed in the wrong direction.

We are governing the wrong thing

For three years the enterprise conversation about AI safety has been about the answer. Is the output accurate. Is it biased. Will it hallucinate. Those questions made sense when AI could only talk. When the worst case was a wrong sentence, you could govern the sentence.
That era is ending. The systems arriving now do not answer; they act. They read data, call tools, invoke APIs, modify records, deploy code, move through infrastructure, and increasingly hand work to other agents. The moment a system can act, the question is no longer what did it say. It is what was it allowed to do, and how far could the damage travel before anyone noticed.
Governing the output of an acting system is like reviewing an employee's emails while ignoring that you also gave them root access to production, a company credit card and the keys to the building. The interesting risk was never the email.

A way to think about it

I have started describing enterprise AI risk to my own teams with a deliberately crude formula: Risk = Capability × Autonomy × Access × Blast Radius.
Capability is how good the model is. Autonomy is how much it does without asking. Access is what systems and data it can reach. Blast radius is how much it can damage if it goes wrong.
What the formula makes obvious is that intelligence is only one term, and the one you can least control — models get more capable every quarter whether you like it or not. The other three terms are entirely yours. A brilliant model with no autonomy, no access and a tiny blast radius is a contained asset. A mediocre one wired into production with standing permissions and no isolation is a loaded weapon pointed at your own business. The incidents above are all the same story told several ways: capability rose, and the other three terms had not been brought down to match.
This reframes the job. You are not trying to make AI less intelligent. You are trying to keep autonomy, access and blast radius deliberately, provably small — and to expand them only where you have decided, in advance, that the consequence is acceptable.

The questions that do not have answers yet

If you run technology inside a real company, here is the honest state of most AI deployments I see, including sophisticated ones. Try answering these about your own environment:
How many AI agents are operating inside your organization right now? Who created them? When one acts, does it use its own identity, or does it borrow a human's credentials — making its actions indistinguishable from that person's in your logs? If an agent did something harmful tonight, could you revoke its access without disabling a human employee? Could you stop it mid-task? What is the largest thing any single agent could destroy before a person intervened — one document, one database, one cloud account? When an agent hands work to another agent, who authorized the second one? And five years from now, after an incident, could you reconstruct what a given agent knew, which model processed the data, and who approved the action?
Most organizations cannot answer most of these. That is not a failure of diligence. It is that the controls for this simply were not part of how AI got adopted. It arrived as a productivity feature, through a login, and the governance never caught up to the fact that the feature can now reach into the ERP.

The containment lesson

The most instructive detail is the boring one. In Anthropic's case, all four incidents traced back to one evaluation partner's environments, which were supposed to be sealed off from the internet and, through a misconfiguration, were not. The models were not supposed to have access. A setup error meant they did. The safety story was that it was isolated, and the isolation was wrong.
Every enterprise is about to meet a smaller version of this. Running it in a test environment is worth exactly as much as the test environment is actually sealed — and sealing it is a real engineering discipline, not a checkbox. Replit's agent caused production damage not through some exotic maneuver but because development and production shared one database environment; the boundary everyone assumed existed simply wasn't there.
The lesson is not that agents are dangerous and should be caged. It is that containment, least privilege, isolation, approval gates and an emergency stop are not features you add after a pilot goes well. They are the environment the agent has to run inside from the first day, or the pilot is the incident.

What I think the enterprise actually needs

My own view — and the reason we are building the way we are at Logic Overdrive — is that enterprises will need an operating layer for AI that looks less like a chatbot subscription and more like the things we already take seriously: identity and access management, change control, audit, and a kill switch. Every agent with its own identity and scoped, revocable privileges. A record of which model touched which data under whose authority. Human approval concentrated exactly where an action is irreversible, financial or customer-facing. The ability to suspend an agent and cut its credentials in one move. And a design assumption that the model is a replaceable part, not the foundation, so you are never one vendor's decision away from rebuilding everything.
This is the direction we are taking Powertoolz.ai — a governed environment for an AI workforce rather than an assistant bolted onto production. I mention it not as a pitch but as a disclosure of bias: we are building toward these controls because we are convinced the problems above become unavoidable, not optional, over the next decade.

The question worth leaving you with

For two years the defining enterprise AI question has been which model should we use. It is a fine question, and its answer changes every few months.
I think it is already being replaced by a harder one. Not whether your AI is clever enough to do the job — the incidents of 2026 settle that it is. The question enterprise leaders will increasingly have to answer is whether, having given AI a job, you have made it impossible for that AI to damage the company while doing it.
Capability is no longer the scarce thing. Control is.

Sources

Top comments (0)