AI agents that drive a desktop by moving the mouse and typing keystrokes operate blind. They send coordinates to the OS and hope the right button was under the cursor. When the UI shifts, a dialog steals focus, or a HiDPI scaling factor changes the math, the agent keeps going -- clicking empty space, typing into the wrong field, or triggering a destructive action it cannot see.
This article shows how to close that loop. You will learn:
- How to capture the screen, draw persistent annotations, and feed the annotated frame back to the agent
- Why coordinate spaces (logical vs. physical pixels) break most naive overlays
- A minimal Python wrapper you can drop into any automation stack
Capture the Frame Before You Act
The first mistake is acting without a recent screenshot. Grab the frame, then decide. This also gives you a baseline image to annotate after the action.
import mss
import numpy as np
from PIL import Image
def grab_screen(monitor_index: int = 1) -> Image.Image:
"""Return the requested monitor as a Pillow RGB image."""
with mss.mss() as sct:
monitor = sct.monitors[monitor_index]
raw = sct.grab(monitor)
return Image.frombytes("RGB", raw.size, raw.rgb)
mss is faster than pyautogui.screenshot and gives you direct access to each monitor's geometry. Keep the monitor index configurable -- multi-monitor setups are where coordinate bugs hide.
Draw Annotations That Survive Scaling
Drawing on the screenshot is easy. Drawing so the overlay matches what the user sees requires respecting the OS scaling factor. On Windows and macOS, the logical coordinate space (what pyautogui returns) differs from the physical pixel space (what the screenshot contains).
from PIL import ImageDraw, ImageFont
from dataclasses import dataclass
from typing import Tuple
@dataclass
class Annotation:
kind: str # "arrow", "box", "text"
coords: Tuple[int, int, int, int] # x1, y1, x2, y2 in logical coords
color: Tuple[int, int, int] = (255, 0, 0)
width: int = 3
label: str = ""
def annotate(
base: Image.Image,
annotations: list[Annotation],
scale_factor: float = 1.0
) -> Image.Image:
"""Draw annotations scaled to the physical screenshot."""
img = base.copy()
draw = ImageDraw.Draw(img)
font = ImageFont.load_default()
for a in annotations:
x1, y1, x2, y2 = [int(c * scale_factor) for c in a.coords]
if a.kind == "box":
draw.rectangle([x1, y1, x2, y2], outline=a.color, width=a.width)
elif a.kind == "arrow":
# Simple arrow: line + triangle head
draw.line([x1, y1, x2, y2], fill=a.color, width=a.width)
head = 10 * scale_factor
draw.polygon([
(x2, y2),
(x2 - head, y2 - head),
(x2 - head, y2 + head),
], fill=a.color)
elif a.kind == "text":
draw.text((x1, y1), a.label, fill=a.color, font=font)
return img
The scale_factor argument is the bridge. On a 200 % scaled display, pass 2.0. On standard density, pass 1.0. Detect it once at startup (see the next section) and reuse it.
Detect the Scale Factor Automatically
Hard-coding the scale factor works until the user drags the window to a different monitor. Query the OS instead.
import sys
import ctypes
def get_windows_scale_factor(hwnd: int = 0) -> float:
"""Return DPI scale factor for the monitor containing hwnd."""
if sys.platform != "win32":
return 1.0
try:
dpi = ctypes.windll.shcore.GetDpiForMonitor
# MDT_EFFECTIVE_DPI = 0
x = ctypes.c_uint()
y = ctypes.c_uint()
monitor = ctypes.windll.user32.MonitorFromWindow(hwnd, 2) # MONITOR_DEFAULTTONEAREST
ctypes.windll.shcore.GetDpiForMonitor(monitor, 0, ctypes.byref(x), ctypes.byref(y))
return x.value / 96.0 # 96 DPI = 100 %
except Exception:
return 1.0
def get_macos_scale_factor() -> float:
"""Return the backing scale factor of the main screen."""
if sys.platform != "darwin":
return 1.0
try:
from AppKit import NSScreen
return NSScreen.mainScreen().backingScaleFactor
except Exception:
return 1.0
def current_scale_factor() -> float:
if sys.platform == "win32":
return get_windows_scale_factor()
if sys.platform == "darwin":
return get_macos_scale_factor()
# Linux: assume 1.0 unless you parse xrandr / wayland output
return 1.0
Call current_scale_factor() once at startup and cache it. If your agent runs long enough for the user to hot-plug a monitor, re-check before each capture.
Feed the Annotated Frame Back to the Agent
The agent needs to see the result of its own action. The simplest loop:
- Grab screen
- Agent proposes action (click at logical
x, y) - Execute action
- Grab screen again
- Draw annotation at the logical coordinates the agent used
- Send the annotated image to the model with the prompt "Did the click land where you intended?"
def verify_action(
agent,
target_x: int,
target_y: int,
scale: float
) -> bool:
"""Execute, annotate, and ask the agent to confirm."""
# 1. baseline (optional, for diff)
before = grab_screen()
# 2. act
import pyautogui
pyautogui.click(target_x, target_y)
# 3. wait for UI to settle
import time
time.sleep(0.3)
# 4. capture result
after = grab_screen()
# 5. annotate where we *thought* we clicked
ann = [Annotation(
kind="arrow",
coords=(target_x - 30, target_y - 30, target_x, target_y),
color=(0, 255, 0),
label="click target"
)]
annotated = annotate(after, ann, scale_factor=scale)
# 6. ask agent (pseudo-code -- adapt to your LLM client)
response = agent.ask(
images=[annotated],
prompt="The green arrow marks where I tried to click. "
"Is that the correct element? Reply YES or NO."
)
return response.strip().upper().startswith("Y")
This pattern turns a blind action into a verified step. The agent can now self-correct: if it answers NO, you can retry with adjusted coordinates or escalate to a human.
Failure Modes You Will Hit
| Situation | Symptom | Mitigation |
|---|---|---|
| User drags window to second monitor | Annotations appear offset | Re-query current_scale_factor() before each capture |
| Remote desktop / VM |
mss captures black frames |
Use the platform's native capture API (Windows Graphics Capture, macOS ScreenCaptureKit) |
| High-frequency actions | Screenshot latency dominates | Batch annotations: capture once, draw multiple, then ask |
| Transient UI (tooltips, menus) | Annotation covers the very thing you're verifying | Draw after the UI settles; use a short poll loop instead of fixed sleep |
| Multi-monitor with different DPI | Single scale factor is wrong | Capture per-monitor and scale per-monitor |
The most insidious bug is the mixed-DPI setup: a 1080p monitor next to a 4K panel. mss.monitors gives you each monitor's physical rectangle. Map the agent's logical coordinates to the correct monitor before scaling.
Key Takeaways
- Capture before you act -- a recent screenshot is the only ground truth.
-
Scale factors are not optional -- logical coordinates from
pyautoguido not match screenshot pixels on HiDPI displays. - Annotate in logical space, draw in physical space -- the conversion belongs in your drawing layer, not the agent's reasoning.
- Close the loop explicitly -- ask the model to confirm the annotated result; treat NO as a first-class signal to retry or abort.
- Test on mixed-DPI setups -- that is where the coordinate math breaks silently.
Source
Show HN: Let your AI agents paint big arrows, boxes and text on your screen -- added a complete capture-annotate-verify loop with automatic DPI detection, multi-monitor handling, and failure-mode analysis the original repo does not cover.
Support this work
These write-ups are researched and published with no paywall, sponsor, or tracking. If one saved you an afternoon, a small tip keeps them coming.
USDT, USDC or USDD · TRC-20 (Tron)
TFTNsfyomKrnUutRjBTGVULp19ByW29KbY
Top comments (2)
I built simple mouse‑clicking agents before and kept seeing silent mis‑clicks. This article helped me realize I forgot about logical‑vs‑physical‑pixel coordinate conversion and that blind actions need explicit visual feedback loop.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.