I run coding agents with shell access all day, and for a long time there was a low-grade dread underneath it. The permission prompts get annoying enough that everyone eventually turns them off, and then you are one confidently wrong rm away from a bad afternoon.
What fixed it was not better prompting. It was giving the agent its own computer.
The setup is a VM on the same Mac. The agent gets a filesystem with nothing of mine in it, credentials that only work on throwaway repos, and a snapshot taken before every session. If it does something destructive I do not debug it. I roll back. Thirty seconds and the mistake never happened.
That part is easy to describe and, on Apple Silicon, surprisingly annoying to actually build. Here is what I learned doing it.
What Apple's Virtualization framework will actually boot
This is the first thing that trips people up, so it is worth being blunt about it.
Virtualization.framework boots two things: macOS on ARM and Linux on ARM. That is the entire list. There is no Windows on ARM, and no amount of configuration will produce one. It is not a licensing gate you can argue your way past, there is simply nothing to enable. Anything you read that claims otherwise is describing UTM's QEMU backend or VMware Fusion, both of which use their own hypervisor rather than Apple's.
For the agent sandbox case this does not matter, because you want Linux anyway. Linux guests are smaller on disk, boot faster, and you are not burning one of your two permitted macOS guests on a machine that mostly runs npm install.
The two guest types are configured differently in ways the docs do not really foreground:
switch os {
case .macOS:
// Boots the macOS bootloader, and needs a separate
// auxiliary storage file next to the disk image.
bootLoader = .macOS
platform = .mac(auxiliaryStoragePath: auxPath)
case .linux:
// Boots EFI, and needs a variable store that
// persists across reboots. Lose this file and the
// guest forgets where its bootloader is.
bootLoader = .linuxEFI(variableStorePath: efiVars)
platform = .generic
}
The EFI variable store is the one people lose. It is a small file that lives next to the disk image and holds the boot entries. Copy the disk image somewhere without it and you get a guest that boots to an EFI shell and looks broken.
The serial console will change your guest's behavior
This one cost me an evening, and I have never seen it written down.
Attaching a serial console to a Linux guest is the obvious move when the guest fails silently. You get hvc0 piped to a log file and you can finally see what the kernel is doing.
Except Ubuntu's installer detects the serial port and decides that is where you want to be. Subiquity moves the installation UI off the graphical display and onto the serial console, in a reduced text mode, and your VM window sits there showing what looks like a hang. The guest is fine. You just moved its face.
So the serial console has to be opt in, not always on:
// Only attach the console when actively debugging.
consoleLogPath: os == .linux
&& UserDefaults.standard.bool(forKey: "Serial")
? bundle.consoleLogURL.path : nil
The general lesson generalizes past this one bug: on this framework, adding a device is never free. Every device you attach is visible to the guest and the guest may make decisions about it.
Getting code in and out without defeating the point
My first version of this mounted my real project directory into the VM through VirtioFS. It worked immediately and it was completely pointless, because an agent that can write to my actual project folder is just my main machine with extra steps.
The version that works: copy in, work, copy the diff out. The share is read only, or there is no share at all and everything moves over SSH.
If you do want a share, the tag matters and differs by guest:
// Linux guests: pick your own tag, then mount it:
// mount -t virtiofs myshare /mnt/share
let tag = "myshare"
// macOS guests: use Apple's automount tag and the
// share shows up at /Volumes/My Shared Files with
// no mount command at all.
let tag = VZVirtioFileSystemDeviceConfiguration
.macOSGuestAutomountTag
The macOS automount tag is genuinely nice and almost nobody knows it exists.
Rosetta inside a Linux guest
If your toolchain has an x86-64 binary in it somewhere, and it usually does, you can share Rosetta into an ARM Linux guest:
switch VZLinuxRosettaDirectoryShare.availability {
case .installed: // attach the share
case .notInstalled: // offer installRosetta { ... }
case .notSupported: // hide the feature entirely
@unknown default: // treat as unsupported
}
Mount it in the guest with mount -t virtiofs rosetta /mnt/rosetta, register it with binfmt_misc, and x86-64 binaries start running.
Two caveats before you build anything load bearing on this. Apple has signalled that Rosetta 2 is being wound down in future macOS releases, and has not committed to the Linux virtualization path surviving that. Check Apple's current guidance rather than mine. And handle @unknown default as unsupported rather than crashing, because the day this enum grows a case is the day your app stops launching.
Things I got wrong
Snapshot before, not after. Obvious in hindsight. Cost me a session.
Real credentials get used. If the agent can reach it, treat it as in scope. Scope the tokens to throwaway repos and assume anything reachable from that VM is reachable by the agent.
Disk images only grow. The image expands toward whatever you provisioned and does not shrink when you free space inside the guest. Do not oversize "just in case". Keep one clean base image you never boot, clone off it, and delete the clone when you are done. One image that stays reasonable beats five that have all ballooned.
Two macOS guests per host, maximum. That is Apple's license terms, not a technical limit. Nothing stops you, which means it is on you. Irrelevant for Linux guests, which is another reason to use them here.
What actually changed
I run with permissions fully open now. The blast radius is a disk image I can throw away, so the prompt that used to make me think twice does not anymore.
The odd side effect is that I review the final diff more carefully than I used to. I am not spending attention on "is this rm safe", so there is attention left over for whether the code is any good.
Getting started
Any of these will get you there. Apple's framework directly if you want to write the roughly two hundred lines yourself, and honestly it is a good weekend, the API is one of the better ones Apple ships. UTM if you want something free with a GUI and do not mind the setup. lume or Tart if you would rather script it.
Disclosure: I got annoyed enough at the setup friction that I built a Mac app for this, called Kyvenza. Every code sample above is from its engine. It is 49 dollars one time with a 7 day trial, and it will not run Windows either, for the reason in the second section. The approach is the point of the post though, and UTM does it for free if you do not mind the assembly.
Still curious whether anyone has a cleaner pattern for handing credentials to a sandboxed agent. That is the part I like least.
Top comments (1)
The cleanest credential pattern I’ve found is to keep long-lived credentials out of the guest entirely.
Give each VM session its own identity, then have a host-side broker exchange that identity for a short-lived token scoped to one repository and the minimum required operations. Bind the token to the session and expire it aggressively. Destroying or rolling back the VM should also revoke the session, rather than restoring a snapshot containing a still-valid credential.
I’d keep the broker outside the writable guest and avoid injecting credentials into the base image, shell history, environment files, or snapshots. For higher-risk actions, the broker can require an explicit approval tied to the repository, operation, and current commit.
The VM limits filesystem damage, but outbound access is still part of the blast radius. An agent with an unrestricted network and a scoped Git token can still exfiltrate the repository. Destination allowlists and an immutable host-side network log make the credential boundary much more meaningful.