Anthropic's Claude computer use API, released in October 2024, takes a screenshot, decides where to click, and returns raw pixel coordinates. That means the hardest part of a screen agent is not the model. It is everything around it: capturing the screen, scaling coordinates correctly, verifying that actions worked, and recovering when they didn't.
This guide covers the architecture that holds up in practice. It focuses on the parts that break first when you move from a demo to a script you can run unattended.
The Core Loop: Observe, Decide, Act, Verify
Every screen agent, whether it is Claude computer use, OpenAI's Operator, or an open-source stack like Microsoft's OmniParser paired with a local model, runs the same loop:
- Observe: capture a screenshot, and optionally the accessibility tree.
-
Decide: send the observation and the goal to a model that returns an action, such as
click(x, y),type("text"), orscroll(dy). - Act: execute the action with an input library.
- Verify: capture again and confirm that the state changed as expected.
Most failed prototypes skip step 4. The model clicks a button, a modal animates in 300ms later, and the next screenshot captures a half-rendered frame. The agent then reasons about a screen that no longer exists.
A minimal executor using pyautogui and mss:
import mss, pyautogui, time
from PIL import Image
def screenshot(max_w=1280):
with mss.mss() as s:
raw = s.grab(s.monitors[1])
img = Image.frombytes("RGB", raw.size, raw.rgb)
scale = img.width / max_w
img = img.resize((max_w, int(img.height / scale)))
return img, scale
def click(x, y, scale):
pyautogui.click(int(x * scale), int(y * scale))
time.sleep(0.5) # let the UI settle before re-observing
The scale factor matters. Models see a downscaled image. On a Retina or 4K display, the coordinates they return must be mapped back to physical pixels, or every click lands off target.
Takeaway: Build the verify step and coordinate scaling first, before you write a single prompt.
Grounding: Getting the Model to Click the Right Pixel
General vision models are good at describing a screen and noticeably worse at pinpointing a 24-pixel icon. You have three grounding strategies.
- Pure pixel prediction. Claude computer use and similar models return coordinates directly. This is the simplest option and works well for large, labeled targets.
- Set-of-Mark prompting. Detect UI elements first, overlay numbered boxes on the screenshot, and ask the model to answer "click element 14." OmniParser outputs these bounding boxes from a fine-tuned YOLO detector plus an icon-captioning model. Choosing a number is far easier for a model than estimating exact coordinates.
-
Accessibility tree. On Windows, use UI Automation through
pywinauto. On macOS, use the AX API. On the web, use Playwright'spage.accessibility.snapshot(). These give you exact element bounds and names with no vision required.
The reliable pattern is hybrid. Use the accessibility tree when it exists, and fall back to Set-of-Mark on canvas apps, games, or remote desktops where the tree is empty.
Takeaway: Before sending a screenshot to a model, check whether an accessibility tree or DOM can give you element coordinates for free.
Sandboxing: Never Run an Agent on Your Real Desktop
An agent with mouse control can delete files, send emails, or approve payments. Anthropic's own reference implementation runs inside a Docker container with a virtual X display (Xvfb), a lightweight window manager, and VNC for observation. Copy that setup.
docker run -p 5900:5900 -p 6080:6080 \
-e ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
-it ghcr.io/anthropics/anthropic-quickstarts:computer-use-demo-latest
This gives you a disposable Linux desktop you can watch at localhost:6080. If the agent misbehaves, kill the container.
Add guardrails beyond isolation:
- Cap the number of steps, for example 30 actions per task.
- Require human confirmation for irreversible actions such as submitting, deleting, or making a payment.
- Keep credentials out of the prompt context.
- Treat on-screen text as untrusted input. A webpage that says "ignore previous instructions" is a real prompt-injection vector.
Takeaway: Run every agent experiment in a container with a step limit and a confirmation gate on destructive actions.
Making It Reliable Enough to Ship
Demo agents succeed sometimes. Production agents need to succeed predictably. Four techniques close most of the gap.
Use deterministic paths where you can. If a step is identical every run, such as logging in or opening a menu, script it with Playwright or pyautogui. Reserve the model for steps that genuinely need judgment. This cuts cost, latency, and failure surface.
Log every step as a trajectory. Save the screenshot, model output, and action for each step as JSONL. When a run fails, you can replay it exactly. You also accumulate an evaluation set.
Benchmark against public suites. OSWorld provides 369 real computer tasks across Ubuntu and Windows apps. Human performance on it is around 72%, while models released in 2024 initially scored far lower. For browser-only agents, WebArena serves the same purpose. Running a subset of either tells you whether a prompt change actually helped.
Wait for state, not time. Replace fixed sleep() calls with polling. Compare consecutive screenshots, and proceed once the pixel difference drops below a threshold or the expected element appears.
Takeaway: Log trajectories from day one, and measure changes against a fixed task set instead of eyeballing demos.
The Skills Employers Are Screening For
Teams hiring AI agent engineers are not mainly looking for prompt writing. The recurring requirements combine several areas:
- Classic automation: Selenium, Playwright, and OS input APIs.
- Computer vision basics: object detection, OCR with Tesseract or PaddleOCR, and image diffing.
- LLM tool-calling schemas.
- Evaluation discipline.
If you can explain why your agent's click accuracy dropped on a 4K monitor and how you fixed it, you stand out from candidates who have only wrapped an API.
Takeaway: Build one portfolio agent that demonstrates grounding, sandboxing, and an evaluation table, not just a screen recording.
Today, pull the Anthropic computer-use demo container, give it one repetitive task you actually do, such as exporting a weekly report from a web dashboard, and log every step to JSONL. By the end of the session, you will have a working sandbox, real failure cases to study, and the first entry in your evaluation set.
Top comments (2)
Good breakdown. One gotcha in the executor snippet on macOS: on a Retina display
mssreturns the image in physical pixels (2x), butpyautogui.clicktakes logical points. Soscale = img.width / max_wmaps back to pixels, and every click lands at roughly double the intended position. Computing the scale froms.monitors[1]["width"](points) instead ofimg.widthfixes it. Multi-monitor setups add one more trap: a display to the left of the main one has negative x.On verify: when the accessibility tree is available, checking the element itself (its value changed, the dialog title appeared, the button is gone) is cheaper and less flaky than screenshot diffing, and it avoids the half-rendered-frame problem you describe.
Your "deterministic paths" point goes further than login scripts. We do this at Auten (computer-use over MCP, I'm on the team): once a task succeeds, the steps are saved and the next run replays them with no model call. When an app changes and a step breaks, it fixes itself. Most repetitive tasks end up never touching the model after the first run, which covers cost, latency and a good part of the reliability gap at once.
The consecutive screenshot diff loop runs into two sharp edge cases in production.
First is persistent micro-animation. A blinking text caret, a spinning progress indicator in the bottom corner, or an OS status bar clock keeps the delta above the settle threshold indefinitely, burning through the step budget. Masking out dynamic regions or requiring element-level bounding box stability avoids getting stuck in an endless loop.
Second is transient blank frames during repaints. Heavy desktop web apps often flash an unpainted white canvas for a frame or two while mounting the DOM. If two consecutive captures sample that intermediate repaint, the diff drops to zero immediately and passes the settle check before the actual interactive elements exist. Requiring two consecutive zero-diff frames spaced 150ms apart stops the runner from acting on empty canvases.