Most "AI agents" stop at text. They plan, they explain, they hand you a snippet to run yourself.
Auten is an MCP server that closes that last gap: it gives your agent a real screen to work on —
your computer (Mac, Windows, Linux) and your Android phone — so it clicks, types, and reads the
screen like a person sitting at the machine. Even in apps with no API.
I wired it into Claude Code. Here is what that actually looks like.
One line to install
Auten ships as an MCP server on npm, so any MCP client can use it:
claude mcp add auten -- npx -y @autenai/mcp
That is it for Claude Code. Cursor, Codex, and other MCP clients use the same package with their
own config. Nothing else to run on the machine — the piece on the device is thin (hands and eyes),
the brain is your agent.
Giving it a goal, not a script
The thing that surprised me: you do not script coordinates or write selectors. You give the agent a
goal, and it reads the accessibility tree to find what it needs.
A real task I ran: "open System Settings and turn on Night Shift from sunset to sunrise." The agent
listed the on-screen elements, clicked into Displays, found the Night Shift button by its label,
opened the schedule dropdown, and picked the option. When I asked it to do the same on my phone —
"turn on battery saver" — it drove the Android Settings app the same way.
No API exists for either of those. It worked because Auten operates the real UI, not an integration.
The parts that make it usable, not just a demo
A computer-use agent is only useful if it is repeatable, cheap, and safe. Three things do that:
- Learns a task once, replays with no AI call. Record a task once and Auten saves it as a skill with parameters. Next time Auten replays it deterministically — seconds, no model tokens.
- Self-heals when the UI shifts. When a button moves or a layout changes and a saved skill breaks, it repairs the step from where it failed instead of dying.
- Secrets stay local. Passwords and keys are typed on your machine through a local login path; the agent never sees the value.
Where it is honest about limits
This is early. It is genuinely good at UI navigation, form filling, and multi-step tasks across apps,
and it is best when you let it confirm the result after each important step rather than firing blind.
It is not magic — give it a clear goal, check the outcome, and it compounds: the second run of any
task you teach it is deterministic.
If you are building agents that should do real work instead of just chatting, this is the missing
hand. Try it: auten.ai/mcp
Top comments (0)