GitHub shipped computer use for Copilot on October 1, in public preview for Copilot CLI and the Copilot app on macOS and Windows. Copilot can read accessible app content and visual context, click controls, enter and edit text, press keys, scroll, drag, and move work between applications. The target is explicit: legacy and GUI-only software with no API, no CLI, and no MCP integration.
I spend my days on agent configuration, so my first question was not "what can it do" but "what stops it". The changelog is short but it names the permission model, and that model matters more than the demo. Here is what it tells us, and where it leaves your rules file.
What actually gates a GUI session
Three layers, in order of hardness:
- OS permissions. On macOS the feature walks you through Accessibility and Screen Recording grants. No grant, no computer use. That is the hardest stop in the whole stack.
- Per-app approval. Copilot asks before controlling an app, and you can review or reset the always-allow list later. This is a real gate, but it is per app, not per action.
- The org kill switch. Organization-managed settings can disable the feature entirely. If you admin a tenant, this is the one setting to find before anyone else does.
Compare that with a terminal agent. A hook can intercept every Bash call, inspect the command, and deny it. A config file can scope which paths are writable. Those controls operate at the granularity of a tool call. A GUI session has no equivalent of a tool call to intercept. The unit of action is "Copilot is driving Safari now", and once that is approved, the individual clicks and keystrokes inside that approval are not things your repo-level rules can veto.
Why this breaks the config-file mental model
Rules files work because tools have names. Bash, Edit, an MCP tool called search_docs, whatever. You write policy against those names: allow this, deny that, ask before the other. I wrote about where this enforcement split belongs in config files vs hooks, and the whole framework assumes the tool surface is enumerable.
Computer use collapses the tool surface into "whatever the accessibility tree and the screen expose". There is no tool name for "the tenth button in the payment dialog". So the vocabulary of allowlists and denylists, the thing most agent configs are built on, has no purchase here. What replaces it, per the changelog itself, is description: computer use "works best when you describe the outcome you want, the applications involved, and any important constraints." That sentence is doing a lot of load-bearing work. Constraints in natural language are the primary control surface for GUI sessions. That is rules-file territory, just without teeth.
What your instructions can still do
Natural language is not nothing. Model behavior follows stated constraints most of the time, and a session that opens with explicit constraints will behave differently from one that opens with "just do the expense report". If you enable this, write a GUI section into whatever instructions file your setup loads:
# Computer use sessions
- Allowed apps: Safari, Numbers, Mail. Nothing else.
- Never type into: password manager, banking sites, anything 2FA.
- Financial actions over $50: stop and show me the screen first.
- If a dialog appears that was not in the plan, stop, screenshot, ask.
- Never install software or grant new permissions to any app.
That will shape behavior. It will not guarantee it, the way a hook that blocks rm -rf guarantees it. Know which one you are writing. Treat instruction-level guardrails as behavior shaping, and treat the OS permissions and the always-allow list as the actual fence.
The part nobody demos: the exfiltration path
Screen Recording plus Accessibility means the agent can read everything on screen and drive everything with a dialog. Now recall that the hot post on DEV this week is literally titled "Prompt Injection Is the New SQL Injection (and We're Not Ready)", and that we have already watched agents leak secrets through tokens meant to beat scanners. A GUI agent that reads a poisoned page can act on it with your hands. The browser is both the tool and the attack surface. I would not enable this on a machine that has a password manager extension active, and I would keep the always-allow list empty until a workflow has earned trust a few times over.
A checklist before you flip it on
- Enable only on a machine you can afford to mis-click. Not your daily driver with production credentials.
- Walk the macOS permission prompts consciously. You are granting Accessibility and Screen Recording, read what that means.
- Run
/computer showafter each session to confirm state,/computer offwhen done. The commands are cheap; make them reflexes. - Keep the always-allow list empty at first. Approve per session until a workflow is boring.
- Write the GUI constraints section before the first real task, not after the first incident.
- If you admin an org, set the managed disable policy now, opt in per team, not the reverse.
The pattern worth watching
First agents got the terminal. Then they got your files, then your tools via MCP, and now the screen itself. Each step moved the enforcement point further from your repo and closer to the OS. Your instructions still travel with the agent, but the fence keeps getting built somewhere you do not version control. That is exactly why I keep the permission model of every new agent surface in the same config kit as the rules themselves, so the two cannot drift apart quietly.
Putting a GUI policy in the same place as the rest
If you maintain agent configs for a team, this release is a good prompt to add a "GUI sessions" section before someone enables the feature with the defaults. The kits in AgentConfig Studio ($29) keep rules, hooks, and permission notes in one versioned place, so a GUI constraint block slots in next to the Bash policy instead of living in someone's memory. If you want to see which files actually load in a repo before extending them, Verify First (€19) walks the tree and tells you. The free Next.js sample kit shows the structure without the paid pieces.
The changelog is one page. Read it before the demo video, because the demo shows what it can do and the changelog shows what stops it. On an agent you are about to hand your keyboard to, the second question is the one that matters.
Top comments (0)