Some systems will never get an API.
There are internal desktop applications built years ago.
There are vendor portals with nothing but forms.
There are administrative consoles that require five clicks through menus before you reach the setting you need.
There are old applications that work perfectly well for the humans using them but have no modern automation interface.
This is where computer-use agents become interesting.
Instead of integrating with an API, you give the model access to a desktop and let it interact through screenshots, mouse actions and keyboard actions.
Anthropic's current stable computer-use toolset is:
computer_toolset_20260801
It replaces the earlier beta computer-use tool versions for supported models on the Claude API. The stable toolset requires no computer-use beta header.
It also introduces an important change in how the agent loop works: Claude can return multiple computer actions in a single response, known as a batch action. Zoom is also a member of the stable toolset and is enabled by default.
That makes this more than a release-note change.
It changes how I would design a GUI agent.
The task
For a safe demonstration, I don't want an agent clicking around a real production system.
Instead, imagine a deliberately simple internal application:
┌──────────────────────────────────────────┐
│ Legacy Admin Console │
├──────────────────────────────────────────┤
│ │
│ Customer: ACME │
│ Environment: Production │
│ Service: billing-api │
│ │
│ Current state: Disabled │
│ │
│ [ Enable Service ] │
│ │
└──────────────────────────────────────────┘
There is no API.
There is no CLI.
The only supported way to perform the operation is to open the application and click the button.
That's exactly the kind of task where computer use makes sense.
What changed?
The older integration looked conceptually like:
{
"type": "computer_20251124",
"name": "computer",
"display_width_px": 1024,
"display_height_px": 768
}
It also required the corresponding beta header.
The stable toolset looks like:
{
"type": "computer_toolset_20260801"
}
There is no name, no display-size configuration in the tool entry, and no computer-use beta header for this stable version. Anthropic's migration documentation explicitly calls out these changes.
The stable toolset exposes members such as:
screenshot
left_click
type
key
scroll
zoom
cursor_position
and others.
Claude's response identifies the member directly:
{
"type": "tool_use",
"name": "left_click",
"toolset_name": "computer",
"input": {
"coordinate": [640, 420]
}
}
This differs from the older model where the action was encoded inside an input.action field.
The basic agent loop
At its simplest, the application does this:
┌──────────────┐
│ User request │
└──────┬───────┘
│
▼
┌──────────────┐
│ Claude │
└──────┬───────┘
│
tool_use blocks
│
▼
┌──────────────┐
│ Your runtime │
└──────┬───────┘
│
mouse / keyboard
│
▼
┌──────────────┐
│ Desktop GUI │
└──────┬───────┘
│
screenshot
│
└──────────► Claude
Anthropic calls the repeated cycle of Claude producing tool calls and the application returning their results the agent loop.
A simplified implementation looks like this:
import anthropic
client = anthropic.Anthropic()
tools = [
{
"type": "computer_toolset_20260801"
}
]
messages = [
{
"role": "user",
"content": "Open the legacy admin console and enable the billing service."
}
]
while True:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
tools=tools,
messages=messages,
)
messages.append({
"role": "assistant",
"content": response.content,
})
tool_uses = [
block for block in response.content
if block.type == "tool_use"
and getattr(block, "toolset_name", None) == "computer"
]
if not tool_uses:
break
results = []
for block in tool_uses:
result = execute_computer_action(block)
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"toolset_name": "computer",
"content": result,
})
messages.append({
"role": "user",
"content": results,
})
The real implementation needs to handle screenshots as image results and implement every computer action you expose.
But the loop is fundamentally this:
Observe → Decide → Act → Observe
Batch actions are a bigger deal than they look
Suppose Claude sees a browser window.
It might decide:
1. Click address bar
2. Type the URL
3. Press Enter
4. Take screenshot
Previously, an application could be designed around one computer action per round trip.
The stable toolset allows Claude to return several actions in one response.
For example:
[
{
"type": "tool_use",
"name": "left_click",
"toolset_name": "computer",
"input": {
"coordinate": [640, 60]
}
},
{
"type": "tool_use",
"name": "type",
"toolset_name": "computer",
"input": {
"text": "http://legacy-app.local"
}
},
{
"type": "tool_use",
"name": "key",
"toolset_name": "computer",
"input": {
"text": "ENTER"
}
},
{
"type": "tool_use",
"name": "screenshot",
"toolset_name": "computer",
"input": {}
}
]
Anthropic's documentation calls this a batch action and specifies that the application should execute the actions sequentially, not concurrently.
That last point is important.
If Claude says:
click → type → screenshot
you cannot execute the three operations simultaneously.
The second action depends on the first.
Your loop has to process every action
This is one of the easiest migration mistakes.
Don't do this:
block = response.content[0]
execute(block)
A response can contain multiple tool_use blocks.
Instead:
for block in response.content:
if block.type == "tool_use":
execute(block)
Anthropic explicitly warns that every block in a batch needs a corresponding tool_result. If an earlier action fails, later actions should be reported as not executed rather than blindly continuing.
That means your executor needs logic similar to:
failed = False
for block in tool_uses:
if failed:
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"toolset_name": "computer",
"is_error": True,
"content": "Not executed: an earlier computer action in this turn failed."
})
continue
try:
result = execute(block)
except Exception as exc:
failed = True
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"toolset_name": "computer",
"is_error": True,
"content": str(exc)
})
continue
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"toolset_name": "computer",
"content": result
})
This isn't just defensive programming.
A GUI is stateful.
If the first click fails, the next click might land on an entirely different part of the screen.
Zoom changes the game for dense interfaces
One of the most useful additions is zoom.
The model can request a region of the screen:
{
"type": "tool_use",
"name": "zoom",
"toolset_name": "computer",
"input": {
"region": [100, 200, 400, 350]
}
}
This is useful when the full desktop contains something difficult to read:
- a dense table
- a small status label
- a tiny error message
- a dialog with multiple options
- a configuration panel
Zoom is enabled by default in the stable toolset.
One subtle point is that coordinates remain based on the original full screenshot, even after a zoom operation.
So your application should not accidentally reinterpret zoomed coordinates as coordinates relative to the cropped image.
Building a safe sandbox
I would not test this against a production desktop.
A basic Linux sandbox can provide:
Docker
├── Xvfb
├── lightweight window manager
├── test application
├── browser
└── agent runtime
Anthropic's own computer-use documentation describes a virtual display based on Xvfb and a Linux desktop environment as part of the computing environment.
For a reproducible demo, create a small application containing intentionally awkward GUI elements:
Legacy Console
[Server]
server-01
[Status]
STOPPED
[Actions]
[ Start ] [ Restart ] [ Delete ]
Confirmation dialogs:
"Are you sure?"
Now you have something that can be safely broken.
That's exactly what you want for an agent experiment.
The safety layer
Computer use should never mean:
Claude has unrestricted control of my computer.
The runtime should be treated as a security boundary.
I would start with four controls.
1. Application allow-list
Only allow the agent to interact with:
legacy-console
browser
text editor
Don't expose a general employee workstation.
2. Step budget
For example:
MAX_ACTIONS = 40
If the agent reaches 40 actions, stop.
This protects against loops such as:
click
screenshot
click
screenshot
...
3. Destructive-action confirmation
A button such as:
Delete server
should not be treated like:
Open menu
The application can intercept the click and ask a human for confirmation.
4. Audit screenshots
Save:
timestamp
agent action
screenshot
tool result
That gives you a visual audit trail.
Batch actions create a new safety consideration
Batching makes the agent faster.
It also means that several actions can arrive before your application gets another chance to reason about the state.
Anthropic explicitly notes that if a human needs to confirm consequential actions, the confirmation should happen before each block is executed, because a batch can contain multiple actions.
For example:
Claude:
click "Delete"
click "Confirm"
screenshot
Your runtime should not simply execute all three.
Instead:
click Delete
↓
intercept
↓
human approval
↓
click Confirm
↓
screenshot
Batching is an optimization, not a reason to bypass authorization.
A useful experiment: batching versus one action per turn
If I were publishing benchmark results, I would measure at least:
| Metric | Sequential | Batched |
|---|---|---|
| Number of model turns | measure | measure |
| Wall-clock time | measure | measure |
| Input tokens | measure | measure |
| Output tokens | measure | measure |
| Computer actions | measure | measure |
| Failed actions | measure | measure |
Don't publish invented numbers.
Run the same task multiple times under the same conditions.
For example:
Task:
Open the test application.
Navigate to Settings.
Change the polling interval.
Save.
Verify the new value.
Run it:
10 × sequential
10 × batched
Then report:
median
p90
success rate
Median matters because GUI agents can occasionally take a very different path from one run to another.
The failure gallery is more valuable than a perfect demo
A good computer-use article shouldn't only show the successful run.
Show the failures.
For example:
Failure 1 — Wrong click
The model clicks beside a button because two controls are visually close.
Failure 2 — Misread field
The model interprets:
1O0 ms
as:
100 ms
instead of recognizing that the first character is O.
Failure 3 — Unexpected dialog
A confirmation window appears that wasn't present during the previous run.
Failure 4 — Changed layout
The application moves a button after a window resize.
These failures demonstrate why computer use should be treated as an automation system rather than a magical replacement for deterministic APIs.
When computer use is the wrong choice
This is perhaps the most important conclusion.
If an API exists:
Use the API.
If a CLI exists:
Use the CLI.
If an RPA platform provides deterministic selectors:
Consider RPA.
Computer use is most interesting for the long tail:
Best option
│
▼
┌──────────────────┐
│ Stable API exists│
└────────┬─────────┘
│ yes
▼
API
│ no
▼
┌──────────────────┐
│ Reliable CLI? │
└────────┬─────────┘
│ yes
▼
CLI
│ no
▼
┌──────────────────┐
│ GUI-only system │
└────────┬─────────┘
│
▼
Computer use
The strength of computer use is also its weakness.
It can interact with almost anything a human can see.
That means it is much less deterministic than an API.
What I would use it for
The sweet spot is probably not:
"Automate everything."
It is:
"Automate the systems that are too expensive to integrate traditionally."
Examples include:
- legacy enterprise applications
- internal admin consoles
- old desktop software
- visual verification workflows
- cross-application tasks
- temporary automation while an API integration is being built
That makes computer use particularly interesting for the "long tail" of enterprise automation.
Final takeaway
The stable computer-use toolset isn't just "the beta without a beta header."
It changes the integration model.
You now have:
- a stable
computer_toolset_20260801 - multiple actions in one model turn
- zoom as a standard member
- explicit toolset/member configuration
- a cleaner tool-call structure
- an agent loop designed around batches
But the fundamentals haven't changed.
The agent still needs:
Observe
↓
Decide
↓
Act
↓
Verify
↓
Repeat
And the application still needs to decide what the agent is actually allowed to do.
The best computer-use implementation isn't the one that gives Claude the most control.
It's the one that gives Claude exactly enough control to complete the task safely.
Top comments (0)