DEV Community

Vishwa Santhosh
Vishwa Santhosh

Posted on

Claude Computer Use Is Out of Beta: I Gave an Agent a GUI-Only Task

Some systems will never get an API.

There are internal desktop applications built years ago.

There are vendor portals with nothing but forms.

There are administrative consoles that require five clicks through menus before you reach the setting you need.

There are old applications that work perfectly well for the humans using them but have no modern automation interface.

This is where computer-use agents become interesting.

Instead of integrating with an API, you give the model access to a desktop and let it interact through screenshots, mouse actions and keyboard actions.

Anthropic's current stable computer-use toolset is:

computer_toolset_20260801
Enter fullscreen mode Exit fullscreen mode

It replaces the earlier beta computer-use tool versions for supported models on the Claude API. The stable toolset requires no computer-use beta header.

It also introduces an important change in how the agent loop works: Claude can return multiple computer actions in a single response, known as a batch action. Zoom is also a member of the stable toolset and is enabled by default.

That makes this more than a release-note change.

It changes how I would design a GUI agent.


The task

For a safe demonstration, I don't want an agent clicking around a real production system.

Instead, imagine a deliberately simple internal application:

┌──────────────────────────────────────────┐
│ Legacy Admin Console                     │
├──────────────────────────────────────────┤
│                                          │
│ Customer:        ACME                    │
│ Environment:     Production              │
│ Service:         billing-api             │
│                                          │
│ Current state:   Disabled                │
│                                          │
│ [ Enable Service ]                       │
│                                          │
└──────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

There is no API.

There is no CLI.

The only supported way to perform the operation is to open the application and click the button.

That's exactly the kind of task where computer use makes sense.


What changed?

The older integration looked conceptually like:

{
  "type": "computer_20251124",
  "name": "computer",
  "display_width_px": 1024,
  "display_height_px": 768
}
Enter fullscreen mode Exit fullscreen mode

It also required the corresponding beta header.

The stable toolset looks like:

{
  "type": "computer_toolset_20260801"
}
Enter fullscreen mode Exit fullscreen mode

There is no name, no display-size configuration in the tool entry, and no computer-use beta header for this stable version. Anthropic's migration documentation explicitly calls out these changes.

The stable toolset exposes members such as:

screenshot
left_click
type
key
scroll
zoom
cursor_position
Enter fullscreen mode Exit fullscreen mode

and others.

Claude's response identifies the member directly:

{
  "type": "tool_use",
  "name": "left_click",
  "toolset_name": "computer",
  "input": {
    "coordinate": [640, 420]
  }
}
Enter fullscreen mode Exit fullscreen mode

This differs from the older model where the action was encoded inside an input.action field.


The basic agent loop

At its simplest, the application does this:

             ┌──────────────┐
             │ User request │
             └──────┬───────┘
                    │
                    ▼
             ┌──────────────┐
             │    Claude    │
             └──────┬───────┘
                    │
              tool_use blocks
                    │
                    ▼
             ┌──────────────┐
             │ Your runtime │
             └──────┬───────┘
                    │
              mouse / keyboard
                    │
                    ▼
             ┌──────────────┐
             │ Desktop GUI  │
             └──────┬───────┘
                    │
                 screenshot
                    │
                    └──────────► Claude
Enter fullscreen mode Exit fullscreen mode

Anthropic calls the repeated cycle of Claude producing tool calls and the application returning their results the agent loop.

A simplified implementation looks like this:

import anthropic

client = anthropic.Anthropic()

tools = [
    {
        "type": "computer_toolset_20260801"
    }
]

messages = [
    {
        "role": "user",
        "content": "Open the legacy admin console and enable the billing service."
    }
]

while True:
    response = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=4096,
        tools=tools,
        messages=messages,
    )

    messages.append({
        "role": "assistant",
        "content": response.content,
    })

    tool_uses = [
        block for block in response.content
        if block.type == "tool_use"
        and getattr(block, "toolset_name", None) == "computer"
    ]

    if not tool_uses:
        break

    results = []

    for block in tool_uses:
        result = execute_computer_action(block)

        results.append({
            "type": "tool_result",
            "tool_use_id": block.id,
            "toolset_name": "computer",
            "content": result,
        })

    messages.append({
        "role": "user",
        "content": results,
    })
Enter fullscreen mode Exit fullscreen mode

The real implementation needs to handle screenshots as image results and implement every computer action you expose.

But the loop is fundamentally this:

Observe → Decide → Act → Observe
Enter fullscreen mode Exit fullscreen mode

Batch actions are a bigger deal than they look

Suppose Claude sees a browser window.

It might decide:

1. Click address bar
2. Type the URL
3. Press Enter
4. Take screenshot
Enter fullscreen mode Exit fullscreen mode

Previously, an application could be designed around one computer action per round trip.

The stable toolset allows Claude to return several actions in one response.

For example:

[
  {
    "type": "tool_use",
    "name": "left_click",
    "toolset_name": "computer",
    "input": {
      "coordinate": [640, 60]
    }
  },
  {
    "type": "tool_use",
    "name": "type",
    "toolset_name": "computer",
    "input": {
      "text": "http://legacy-app.local"
    }
  },
  {
    "type": "tool_use",
    "name": "key",
    "toolset_name": "computer",
    "input": {
      "text": "ENTER"
    }
  },
  {
    "type": "tool_use",
    "name": "screenshot",
    "toolset_name": "computer",
    "input": {}
  }
]
Enter fullscreen mode Exit fullscreen mode

Anthropic's documentation calls this a batch action and specifies that the application should execute the actions sequentially, not concurrently.

That last point is important.

If Claude says:

click → type → screenshot
Enter fullscreen mode Exit fullscreen mode

you cannot execute the three operations simultaneously.

The second action depends on the first.


Your loop has to process every action

This is one of the easiest migration mistakes.

Don't do this:

block = response.content[0]
execute(block)
Enter fullscreen mode Exit fullscreen mode

A response can contain multiple tool_use blocks.

Instead:

for block in response.content:
    if block.type == "tool_use":
        execute(block)
Enter fullscreen mode Exit fullscreen mode

Anthropic explicitly warns that every block in a batch needs a corresponding tool_result. If an earlier action fails, later actions should be reported as not executed rather than blindly continuing.

That means your executor needs logic similar to:

failed = False

for block in tool_uses:
    if failed:
        results.append({
            "type": "tool_result",
            "tool_use_id": block.id,
            "toolset_name": "computer",
            "is_error": True,
            "content": "Not executed: an earlier computer action in this turn failed."
        })
        continue

    try:
        result = execute(block)
    except Exception as exc:
        failed = True

        results.append({
            "type": "tool_result",
            "tool_use_id": block.id,
            "toolset_name": "computer",
            "is_error": True,
            "content": str(exc)
        })
        continue

    results.append({
        "type": "tool_result",
        "tool_use_id": block.id,
        "toolset_name": "computer",
        "content": result
    })
Enter fullscreen mode Exit fullscreen mode

This isn't just defensive programming.

A GUI is stateful.

If the first click fails, the next click might land on an entirely different part of the screen.


Zoom changes the game for dense interfaces

One of the most useful additions is zoom.

The model can request a region of the screen:

{
  "type": "tool_use",
  "name": "zoom",
  "toolset_name": "computer",
  "input": {
    "region": [100, 200, 400, 350]
  }
}
Enter fullscreen mode Exit fullscreen mode

This is useful when the full desktop contains something difficult to read:

  • a dense table
  • a small status label
  • a tiny error message
  • a dialog with multiple options
  • a configuration panel

Zoom is enabled by default in the stable toolset.

One subtle point is that coordinates remain based on the original full screenshot, even after a zoom operation.

So your application should not accidentally reinterpret zoomed coordinates as coordinates relative to the cropped image.


Building a safe sandbox

I would not test this against a production desktop.

A basic Linux sandbox can provide:

Docker
 ├── Xvfb
 ├── lightweight window manager
 ├── test application
 ├── browser
 └── agent runtime
Enter fullscreen mode Exit fullscreen mode

Anthropic's own computer-use documentation describes a virtual display based on Xvfb and a Linux desktop environment as part of the computing environment.

For a reproducible demo, create a small application containing intentionally awkward GUI elements:

Legacy Console

[Server]
server-01

[Status]
STOPPED

[Actions]
[ Start ] [ Restart ] [ Delete ]

Confirmation dialogs:
"Are you sure?"
Enter fullscreen mode Exit fullscreen mode

Now you have something that can be safely broken.

That's exactly what you want for an agent experiment.


The safety layer

Computer use should never mean:

Claude has unrestricted control of my computer.

The runtime should be treated as a security boundary.

I would start with four controls.

1. Application allow-list

Only allow the agent to interact with:

legacy-console
browser
text editor
Enter fullscreen mode Exit fullscreen mode

Don't expose a general employee workstation.

2. Step budget

For example:

MAX_ACTIONS = 40
Enter fullscreen mode Exit fullscreen mode

If the agent reaches 40 actions, stop.

This protects against loops such as:

click
screenshot
click
screenshot
...
Enter fullscreen mode Exit fullscreen mode

3. Destructive-action confirmation

A button such as:

Delete server
Enter fullscreen mode Exit fullscreen mode

should not be treated like:

Open menu
Enter fullscreen mode Exit fullscreen mode

The application can intercept the click and ask a human for confirmation.

4. Audit screenshots

Save:

timestamp
agent action
screenshot
tool result
Enter fullscreen mode Exit fullscreen mode

That gives you a visual audit trail.


Batch actions create a new safety consideration

Batching makes the agent faster.

It also means that several actions can arrive before your application gets another chance to reason about the state.

Anthropic explicitly notes that if a human needs to confirm consequential actions, the confirmation should happen before each block is executed, because a batch can contain multiple actions.

For example:

Claude:
    click "Delete"
    click "Confirm"
    screenshot
Enter fullscreen mode Exit fullscreen mode

Your runtime should not simply execute all three.

Instead:

click Delete
      ↓
intercept
      ↓
human approval
      ↓
click Confirm
      ↓
screenshot
Enter fullscreen mode Exit fullscreen mode

Batching is an optimization, not a reason to bypass authorization.


A useful experiment: batching versus one action per turn

If I were publishing benchmark results, I would measure at least:

Metric Sequential Batched
Number of model turns measure measure
Wall-clock time measure measure
Input tokens measure measure
Output tokens measure measure
Computer actions measure measure
Failed actions measure measure

Don't publish invented numbers.

Run the same task multiple times under the same conditions.

For example:

Task:
Open the test application.
Navigate to Settings.
Change the polling interval.
Save.
Verify the new value.
Enter fullscreen mode Exit fullscreen mode

Run it:

10 × sequential
10 × batched
Enter fullscreen mode Exit fullscreen mode

Then report:

median
p90
success rate
Enter fullscreen mode Exit fullscreen mode

Median matters because GUI agents can occasionally take a very different path from one run to another.


The failure gallery is more valuable than a perfect demo

A good computer-use article shouldn't only show the successful run.

Show the failures.

For example:

Failure 1 — Wrong click

The model clicks beside a button because two controls are visually close.

Failure 2 — Misread field

The model interprets:

1O0 ms
Enter fullscreen mode Exit fullscreen mode

as:

100 ms
Enter fullscreen mode Exit fullscreen mode

instead of recognizing that the first character is O.

Failure 3 — Unexpected dialog

A confirmation window appears that wasn't present during the previous run.

Failure 4 — Changed layout

The application moves a button after a window resize.

These failures demonstrate why computer use should be treated as an automation system rather than a magical replacement for deterministic APIs.


When computer use is the wrong choice

This is perhaps the most important conclusion.

If an API exists:

Use the API.

If a CLI exists:

Use the CLI.

If an RPA platform provides deterministic selectors:

Consider RPA.

Computer use is most interesting for the long tail:

                    Best option
                        │
                        ▼
              ┌──────────────────┐
              │ Stable API exists│
              └────────┬─────────┘
                       │ yes
                       ▼
                      API

                       │ no
                       ▼
              ┌──────────────────┐
              │ Reliable CLI?    │
              └────────┬─────────┘
                       │ yes
                       ▼
                      CLI

                       │ no
                       ▼
              ┌──────────────────┐
              │ GUI-only system  │
              └────────┬─────────┘
                       │
                       ▼
                 Computer use
Enter fullscreen mode Exit fullscreen mode

The strength of computer use is also its weakness.

It can interact with almost anything a human can see.

That means it is much less deterministic than an API.


What I would use it for

The sweet spot is probably not:

"Automate everything."

It is:

"Automate the systems that are too expensive to integrate traditionally."

Examples include:

  • legacy enterprise applications
  • internal admin consoles
  • old desktop software
  • visual verification workflows
  • cross-application tasks
  • temporary automation while an API integration is being built

That makes computer use particularly interesting for the "long tail" of enterprise automation.


Final takeaway

The stable computer-use toolset isn't just "the beta without a beta header."

It changes the integration model.

You now have:

  • a stable computer_toolset_20260801
  • multiple actions in one model turn
  • zoom as a standard member
  • explicit toolset/member configuration
  • a cleaner tool-call structure
  • an agent loop designed around batches

But the fundamentals haven't changed.

The agent still needs:

Observe
   ↓
Decide
   ↓
Act
   ↓
Verify
   ↓
Repeat
Enter fullscreen mode Exit fullscreen mode

And the application still needs to decide what the agent is actually allowed to do.

The best computer-use implementation isn't the one that gives Claude the most control.

It's the one that gives Claude exactly enough control to complete the task safely.

Top comments (0)