If you've been on dev Twitter or Hacker News this month, you've seen the stories. An agent asked to clean up a folder deletes tens of thousands of files. A coding agent launches hundreds of parallel runs nobody asked for and burns through a five-figure bill. Someone's OpenClaw setup emails their clients without approval. Someone else writes "never touch production" in CLAUDE.md, and the agent reads every secret in the .env anyway.
None of these agents were "hacked" in the movie sense. They were doing roughly what they were built to do: take actions with real tools. The problem is that nothing between the agent and the tool asked a simple question first:
Is this action part of the job?
Permissions aren't the same as intent
The usual advice is to give agents fewer permissions. That helps, but it runs into a wall fast. A coding agent needs to delete files sometimes. An inbox agent needs to send email. A deploy agent needs to touch production. Take those away and the agent is useless; leave them and any one bad step can do damage.
What's missing is a check on intent, not just access. "Delete files in /tmp/build" during a cleanup task is fine. "Delete the home folder" during the same task is not. Both use the same permission.
What we're building: Watchdog
That's the idea behind Watchdog, one of the two products at Alovia AI. You give your agent a one-line mission, like "triage my inbox and draft replies, never send." Before each action runs, Watchdog checks it against that mission. On-mission actions go through. Off-mission ones get stopped, and you see why.
It sits around the agent's traffic rather than inside the model, so it works with any model, open-source or frontier, and with setups like Claude Code, Codex, OpenClaw or Hermes.
We tested it against AgentDojo-style prompt-injection attacks, where a tool result tries to hijack the agent into doing something else. In our full configuration, 0 of 809 attacks got through. In a lighter policy-only mode, it stopped about half with zero false positives on legitimate tasks.
The other side: Shield
The same shift is hitting websites from the other direction. Site owners are posting about AI crawlers multiplying their Vercel bills, residential-proxy scrapers making up 99% of a small forum's traffic, and bots testing stolen cards through free trials.
Shield sits in front of your site, alongside Cloudflare rather than replacing it, and stops abusive bots, AI scrapers and fake signups before they reach your origin, without putting a captcha in front of real people.
Try it
We're a small team in private beta, with 20 free seats for people running agents or sites that are getting hammered. If either of these sounds like your week, grab one at aloviaai.com or drop a comment with what broke. I read every one.
Chan, founder of Alovia AI
Top comments (0)