Mobilerun just achieved a 100% official score on AndroidWorld, Google Research's benchmark for evaluating AI agents on Android.
For context, Google's open-source ARTEMIS reports a 99%+ task completion rate on AndroidWorld. Mobilerun achieved a full 100%, completing every one of the benchmark's 116 tasks.
This is a meaningful milestone for mobile AI agents. But the interesting part isn't just the number, it's what the benchmark actually tests.
What is AndroidWorld?
AndroidWorld is a benchmark designed to measure how well AI agents can perform tasks inside Android applications.
It contains 116 tasks spanning 20 Android apps, with workflows that require agents to understand instructions, interact with mobile interfaces, perform multiple actions, and ultimately reach the correct outcome.
These aren't simply isolated UI actions. An agent might need to:
- Navigate through an application
- Locate the right UI element
- Enter or modify information
- Perform several actions in sequence
- Handle the current state of the application
- Verify that the requested change actually happened A tool successfully executing a tap or typing text doesn't necessarily mean the task was completed. AndroidWorld's official score evaluates the result of the task, rather than simply whether individual actions were executed.
116/116: A 100% official score
For our latest evaluation, we ran AndroidWorld using mobile-harness, our set of mobile tools and instructions for running mobile tasks through coding agents. The evaluation used:
- Tasks:116
- Full scores:116
- Official score:100%
- Recorded interactions: 1,419
- Device: Pixel 6 emulator
- OS: Android 13 / API 33
- Model:gpt-6-astra / low
Every one of the 116 tasks received an official score of 1.0. That gives us a perfect benchmark score: 116 / 116 = 100%
We also recorded the interactions, screenshots, tool operations, answers, judgments, and timing for the evaluation. The individual trajectories are available on our benchmark explorer, so you can inspect what happened task by task rather than relying on a single aggregate number.
How does this compare with ARTEMIS?
Google's ARTEMIS is an open-source system for natural-language Android automation. It combines UI understanding, accessibility information, OCR, visual targeting, execution profiles, diagnostics, and MCP integrations for AI coding assistants.
System- AndroidWorld result:
- mobilerun / mobile-harness: 100%
- Google ARTEMIS: 99%+
The difference is small numerically, but reaching a perfect score means completing every task in the 116-task suite.
One important caveat: these are not identical evaluation configurations. Our run used mobile-harness with gpt-6-astra / low, app cards and shared memory, while ARTEMIS has its own architecture and evaluation configuration. So the comparison should be understood as a comparison of reported AndroidWorld results, not a controlled head-to-head experiment.
The role of mobile-harness
The result wasn't achieved by building another completely isolated mobile agent. Instead, we wanted to explore a different question:
What happens when you give a general-purpose coding agent the right tools, mobile-specific knowledge, and memory to operate mobile apps?
That's the idea behind mobile-harness. It provides mobile tools and instructions that allow agents such as Claude Code, Codex, and OpenClaw to operate Android and iOS environments. The harness combines three important components:
1. Mobile tools
The agent gets the ability to interact with the device rather than merely describe what should happen. It can observe the mobile environment, interact with UI elements, and execute workflows.
2. App cards
App cards provide structured knowledge about applications and their workflows. Instead of starting every task completely from scratch, the agent can use information learned about how particular apps work.
3. Shared memory
Mobile tasks often involve recurring patterns. An agent that discovers how an application behaves shouldn't necessarily have to rediscover the same information every time.
Shared memory allows useful lessons from execution to be carried forward. Together, these components give a general-purpose coding agent a more structured way to operate mobile environments.
Execution is not the same as completion.
One of the things we focused on during the benchmark was outcome verification. Imagine an agent is asked to add an item to a list. It could:
- Find the input field.
- Type the requested text.
- Press a button.
- Receive a successful tool response.
But that doesn't necessarily prove the item was actually saved. The application could have rejected the input. The wrong button could have been pressed. The screen could have changed without the intended state being persisted.
That's why our evaluation uses both accessibility information and screenshots to understand the state of the application and calls for reading back changes where appropriate. The benchmark's official scoring then determines whether the intended outcome was actually achieved.
From 63% to 91.4% to 100%
This result is also part of a longer progression for us. Earlier in our AndroidWorld work, Droidrun achieved 63.0% across the 116-task benchmark. We later reached 91.4%, after changing our agent architecture to use a dynamic manager-executor feedback loop instead of relying on a rigid sequence of planned actions.
Now, with mobile-harness, we've reached 100%. That progression reflects something we've learned repeatedly while working on mobile agents, and that is reliable mobile automation isn't just about choosing a better model, the surrounding system also matters.
Tools, state, application knowledge, memory, recovery, verification, and the way an agent interacts with its environment can all affect whether an instruction actually becomes a successful outcome.
Explore the full benchmark: https://mobilerun.ai/benchmark/
Explore the github repo: https://github.com/droidrun/mobile-harness
Top comments (0)