DEV Community

openamer
openamer

Posted on

Driving a real desktop without a VM, a sandbox, or stealing your cursor

Driving a real desktop without a VM, a sandbox, or stealing your cursor

Every "computer use" agent demo works in a browser tab. The hard part is doing it on a real Windows machine — behind the user's live session — without hijacking the cursor or breaking the thing you're automating.

Most agent frameworks stop at the browser. They drive a headless page or a containerised desktop and call it "computer use". That works until the task touches a real application on the user's actual machine: a native window, a file picker, a login that only exists in the user's profile, an OS dialog. Then a headless browser has nothing to say.

OpenAmer is an open-source desktop agent (Apache-2.0, Windows-first) whose whole point is the opposite: it operates the real desktop, in the background, while the person keeps using the machine. This post is about the four design decisions that made that work — and the scars each one left.

1. Control the desktop; don't take it over

The naive way to automate a GUI is to move the real mouse and type on the real keyboard (SendInput, pyautogui, that family). It works — and it makes the machine unusable while it runs: the cursor jumps, your typing lands inside the agent's focus, and you can't do anything in parallel.

The path we use instead is a background control path: capture the target window's state and deliver synthetic input to it without contending for the physical pointer. The user's session keeps its own cursor and focus; the agent works on its own view of reality. The hard parts are unglamorous and specific:

  • Coordinate space. Screen pixels vs. window-relative vs. per-monitor DPI scaling — a click computed in one space and delivered in another lands about 100px off. We hit this repeatedly until DPI became a first-class part of every coordinate, not an afterthought.
  • Focus. Some native controls only accept input when they believe they have focus, so "background" sometimes means borrowing it briefly and transparently rather than never.
  • Timing. A native dialog appears asynchronously, so the agent has to wait on state (does this window exist yet?), never on a fixed sleep.

2. No VM, no container, no cloud

Screenshot-driven agents are usually run in a disposable VM: safe, isolated, and completely disconnected from the user's real files, logins and apps. That defeats the purpose for a personal agent. The whole value is that it can touch your real inbox, your real repo, your real browser profile.

So this runs as a process on the host, in the user's own session, with the user's own credentials — no VM image, no remote inference. That is a security trade and we treat it like one: capabilities are explicit and bounded (a capability with no grant simply does not run), actions that cross a boundary are surfaced rather than silently taken, and every action is written to an outcome ledger with pass/fail so the agent's own claims are checkable afterwards.

3. Drive a real browser profile, not a throwaway one

Half of "desktop" work is web work, and web work needs the user's logins. A fresh headless Chromium has none of them. So the agent controls a real Chrome over the DevTools protocol (CDP) against a persistent profile, inheriting the sessions the user already has — GitHub, a dashboard, an internal tool — instead of trying to defeat a login wall or a captcha.

The trade-off is real: that profile is a live, authenticated identity, so browser tasks run with the same care as anything else that touches a logged-in session. We check login state before acting on it, and verify a result against the real page rather than assuming it.

4. Verify against the world, not the model's self-report

The single most useful discipline in a desktop agent is refusing to trust the model's own "done". A model asked "did that succeed?" will answer confidently and often wrongly. So the checkable unit is not the model's sentence — it's a row in the outcome ledger, written by the deterministic layer, plus, where it matters, a re-read of the actual world (does the file exist? did the profile change? is the text on the page?).

"We posted it" is worth nothing. "Here is the URL, fetched anonymously, HTTP 200, expected byte size" is worth something. That shift — from prose to evidence — is what makes the rest of it safe to run unattended.

Where it lands

Background desktop control is not a demo problem, it's an integration problem: DPI, focus, async dialogs, authenticated profiles, and a verification layer that doesn't take the model's word for it. Get those four right and you get something that does real work on a real machine while you keep using it.

Code and docs: https://github.com/openamer/openamer (Apache-2.0, Windows-first, runs locally).

If you've automated a real desktop: what got you first — DPI, focus, or async dialogs? I'd like to compare scars.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.