DEV Community

Cover image for I Gave AI Hands on a Virtual Machine — Here's What Changed
Bijan
Bijan

Posted on AI-assisted

I Gave AI Hands on a Virtual Machine — Here's What Changed

I think AI is the closest thing to what Google search was when it first went public — it didn't replace anyone's job, it made a whole category of work faster, and "knowing how to use it well" became its own skill. Some people reach for the WordPress comparison instead, when talking about vibe coding and the worry it put into web developers a decade ago. Either analogy lands the same way for me: this is a new capability layer, not a replacement for the person using it.

I'm an ethical hacker — I've worked both defense (SOC) and offense (red team, pentesting, bug hunting), which means my day covers everything from planning and scripting to configuring machines and finding and exploiting vulnerabilities, ethically. AI's role in that day has changed a lot over the past while, and the change is worth being specific about, because "I use AI" doesn't actually say much on its own.

How I used to use it: five models, zero shared context

For a while I ran different providers for different jobs, because each one genuinely had different strengths. Gemini for fast, deep search. Claude for coding. ChatGPT for translation. Grok and DeepSeek sitting in reserve for whenever I ran out of tokens on the primary set.

It worked, but it had two real costs. First, models don't share context — handing code from Claude to another model to continue working on it could quietly break things, because the second model has no idea what assumptions the first one made. Second, every switch meant starting from zero: re-explaining the same context, constraints, and goals to a model that had no memory of the conversation you'd just had somewhere else.

If you're running into the same problem, the practical fix is less exciting than it sounds: pick the provider whose pro tier actually covers most of what you need, and stop routing around token limits by hopping providers mid-task.

How I use it now: an agent with hands, on a machine I can delete

The bigger shift wasn't switching tools — it was giving one of them the ability to act directly. I run Claude Code on a virtual machine, which means I'm not copying output back and forth or manually following step-by-step instructions myself. The AI has hands on the machine. It's closer to a teammate than a tool at that point.

In practice: I prompt once, define the skills a task needs, and reuse them across scenarios instead of re-explaining from scratch each time. That parallelism is the actual productivity gain — while I'm building a POC for one vulnerability, the agent can be running recon on a second target at the same time. Less of my own time spent context-switching, more happening concurrently.

The guardrail matters as much as the capability. Full, unsupervised machine access for an AI agent is a genuinely bad idea — not hypothetically, just practically, given how easy it is to get lazy about permissions once something is working well. Running it on a disposable VM means if it does go sideways, the fix is deleting the machine, not cleaning up a production environment.

The part most "AI in security" takes skip: it doesn't have the creative jump

Permission friction is real — approving actions one at a time feels like fielding a constant stream of questions from a very capable but literal-minded assistant, and the alternative (broad access) risks burning through actions, and tokens, fast. Output still needs verifying every time; nothing here removes that step. And running this workflow well is as much a management skill as a technical one — you're now responsible for your own work and for checking someone else's, even if that someone else is an AI agent.

But the sharpest limitation isn't about permissions or verification overhead. It's that AI doesn't have the creative jump a human hacker makes. It'll test a site for XSS, but only within the parameter space it's been given — it won't naturally go looking for the deep, unconventional payload structure a human would try specifically because it's unconventional. That's a real gap in offensive security work, where a meaningful chunk of what makes a finding interesting is exactly the part a rules-following process doesn't generate on its own. Trusting an agent's negative result — "nothing found" — without accounting for that gap is how you end up with false confidence, not just false positives.

Where that leaves me

AI is a tool, not a replacement — it speeds up real chunks of the work, from generating an image for a blog post to helping debug code, but the parts of this job that depend on human creativity and judgment aren't going anywhere. I built in the guardrail (isolated VM, gated permissions) because I don't fully trust unsupervised autonomy yet, and I don't think that trust should be assumed by default. The upside is real. So is the reason to keep a hand on the wheel.

Top comments (0)