Claude Code Android automation does not need adb, a USB cable or a local emulator if the phone lives in the cloud: you install a skill, add an access token, and Claude Code sends tasks in plain English to a real Android device, then polls them until they finish or need you. Setup is four commands and one token.
The device in this tutorial comes from Airtap, which runs a real Android phone for you in an isolated cloud container, and the bridge is its open-source skill at github.com/airtap-ai/airtap-skill. Every command below comes from that repo's README and skills/airtap/SKILL.md, checked against the main branch in September 2026. This is for developers who already use Claude Code or Codex and want it to operate apps on a remote phone. Building or testing your own APK is out of scope.
TL;DR
claude plugin marketplace add airtap-ai/airtap-skill
claude plugin install airtap@airtap
export AIRTAP_PERSONAL_ACCESS_TOKEN="<token from Settings in the Airtap app>"
python3 scripts/airtap.py task create --message "<what to do>" --receiver-id cloud
How does Claude Code Android automation work without ADB?
Most Android tools for Claude Code are MCP servers that drive a phone plugged into your machine through adb, so Claude Code decides every tap. This skill splits the work differently. Claude Code writes the task in plain language, and Airtap's own agent, working on the device, does the tapping.
The path looks like this:
- Claude Code picks up the
airtapskill and runs its bundled CLI,scripts/airtap.py. - The CLI calls the Airtap API over HTTPS with your token as a Bearer header.
- The task goes to a receiver:
cloud(the hosted phone, and the default) or a physical Android phone you linked through the Autopilot app. - The on-device agent works through the app screen by screen and writes progress messages back to the task.
- Claude Code polls those messages and decides what to tell you.
The trade-off is control. With adb you script every coordinate. Here you hand off one task per request and manage its lifecycle, which is most of what follows.
What you need before you start
- Claude Code or Codex. Cursor, Windsurf, Cline and GitHub Copilot also work through
npx skills. -
python3withrequestsandpython-dotenv(the skill's wholerequirements.txt). - Node.js if you use the
npx skillsroute. - An Airtap account. Sign in at airtap.ai/app with Google. No paid plan is needed; a per-account daily limit applies and resets at 00:00 UTC.
Step 1: install the skill
Claude Code
claude plugin marketplace add airtap-ai/airtap-skill
claude plugin install airtap@airtap
Per the Claude Code plugin docs, a plugin you install from the shell loads the next time you start Claude Code, or right away if you run /reload-plugins in an open session. claude plugin list confirms it's installed.
Codex
npx skills add airtap-ai/airtap-skill -a codex
Cursor, Windsurf, Cline or GitHub Copilot
The -a flag of npx skills picks the agent:
npx skills add airtap-ai/airtap-skill -a cursor
npx skills add airtap-ai/airtap-skill -a windsurf
npx skills add airtap-ai/airtap-skill -a cline
npx skills add airtap-ai/airtap-skill -a github-copilot
The public repo; the install commands above come from this README.
If the dependencies are missing, run this from the skill directory:
pip3 install -r requirements.txt
Step 2: add your personal access token
In the Airtap web app, open Settings, create a personal access token and copy it. Then give it to the skill in one of two ways.
Set an environment variable before you launch your agent:
export AIRTAP_PERSONAL_ACCESS_TOKEN="your-personal-access-token"
Or write it into scripts/.env from the skill directory:
python3 scripts/airtap.py --add-token "your-personal-access-token"
Two repo details matter. Values in scripts/.env override a shell variable of the same name, so a stale .env beats a fresh export. And SKILL.md tells the agent not to ask you to paste the token into chat, so run this yourself in a terminal.
Skip this step and the first API call fails fast. This is what receiver get-list prints with no token set:
Error: Missing auth token. Run --add-token or add AIRTAP_PERSONAL_ACCESS_TOKEN to .env before running this command.
Step 3: run a first task on the cloud phone
You don't type CLI commands yourself. Ask Claude Code something like "Use the Airtap skill to open the Play Store on the cloud phone and list the apps on the home screen," and the skill turns it into:
python3 scripts/airtap.py task create \
--message "Open the Play Store and list the apps on the home screen" \
--receiver-id cloud
The response includes a taskId, which every later command needs. Also note:
- The cloud phone takes about 60 seconds to set up the first time.
- Some apps come pre-installed. Sign in to Google Play on the phone for the rest.
- You log in to your own accounts yourself, once. The web dashboard shows the cloud phone's live screen.
To target your own linked phone instead, run receiver get-list and pass its ID. SKILL.md tells agents to stay on cloud unless you ask for a specific device.
Step 4: watch the task, answer it, or stop it
A task can take a few seconds or several minutes, so the repo treats it as a long-running job. Monitor it with task poll:
python3 scripts/airtap.py task poll --task-id "task_abc123"
It calls task get-details every 10 seconds (change that with --interval-secs). It stops by itself once the task completes, fails, is cancelled or needs input, and it returns the full task payload plus a _poll summary. In that payload, messages is the conversation in order, and most progress updates come from the newest type: "agent" message.
Send more input, or resume a paused task:
python3 scripts/airtap.py task add-user-message --task-id "task_abc123" --message "Continue"
Stop it:
python3 scripts/airtap.py task cancel --task-id "task_abc123"
If you leave out --model-id, tasks run on airtap-1.0-flash, which is quicker and fine for short jobs. For long multi-step flows, pass --model-id airtap-1.0. It's slower, but more reliable on complex work. Both task create and add-user-message also accept --image-file if a screenshot or reference picture helps.
What do the task states mean for your agent?
task poll exits on six states. Three are final, and three mean a human is needed. This table maps each one to what SKILL.md tells the agent to do:
taskState |
What happened | What your agent should do | Command |
|---|---|---|---|
COMPLETED |
Task finished | Summarize the latest agent messages | none |
FAILED |
Task could not finish | Report the state and the last agent message | none |
CANCELLED |
Task was cancelled | Report it | none |
WAITING_FOR_USER_INPUT |
The agent needs a clarification | Ask you, forward the answer | task add-user-message |
WAITING_FOR_USER_INTERVENTION |
Someone has to act on the device (a login, for example) | Tell you exactly what to do, wait for your OK, resume | task add-user-message |
WAITING_FOR_USER_CONTINUE |
Task hit its step limit | Summarize progress, ask whether to continue | task add-user-message |
Claude Code already reads these rules from SKILL.md. To pin the same behavior for every agent in a repo, add this to CLAUDE.md or AGENTS.md:
## Android tasks (Airtap skill)
- Use receiver `cloud` unless I name a device.
- After `task create`, run `task poll` and wait. Do not create a second task for the same request.
- On any WAITING_* state, stop and ask me. Never answer on my behalf.
- Never ask me to paste the access token into chat.
How should you word Android tasks for an agent?
The on-device agent reads the screen at each step, so a vague task still runs, but it wanders. Habits that help:
- Name the app and the finish line. "In the Weather app, read tomorrow's high for Chicago and stop" is better than "check the weather".
- Say what must not happen: "do not place the order", "do not send the message".
- Put the output format in the task so the summary is easy to parse.
Three task messages written that way:
Open the Play Store, search for "Google Keep", and report the rating and last-updated date of the official app. Do not install anything.
In Google Maps, find the three closest pharmacies to Union Square, San Francisco that are open now. Return name and closing time.
Open the Gmail app, count unread emails from the last 24 hours, and list the five senders. Do not open or archive any email.
If a task depends on where you are, update your location first:
python3 scripts/airtap.py user update-location --payload '{"latitude":37.7880,"longitude":-122.4075}'
Troubleshooting the Airtap skill
| Symptom | Likely cause | Fix |
|---|---|---|
Missing auth token... |
No token in the environment or scripts/.env
|
Step 2 |
| The new token is ignored | An old scripts/.env overrides the shell variable |
Edit or delete scripts/.env
|
Airtap API request failed with status 404 |
SKILL.md: on any 404 or request failure, get an updated skill before retrying |
claude plugin marketplace update airtap, or npx skills update airtap
|
| Claude Code doesn't list the skill | Plugin not loaded yet |
/reload-plugins, then check claude plugin list
|
| Skill missing on claude.ai/code | Cloud sessions don't load plugins you installed locally | Run Claude Code locally |
ModuleNotFoundError: No module named 'requests' |
Dependencies missing | pip3 install -r requirements.txt |
python3 not found on Windows |
The interpreter is called python or py there |
Swap the command name |
| Task stuck in a WAITING state | It's waiting on you | See the state table |
When a request fails without a clear message, check the CLI's debug log at airtap/airtap.log in your system temp directory.
Limits and security notes
- This is not a test harness. You get task results and progress messages, not per-tap assertions. For testing your own app, use adb-based tools.
-
Logins stay with you. According to Airtap's homepage, passwords are never stored or shown to the agent, and the agent is blocked from reading secure fields. That also means a task can pause in
WAITING_FOR_USER_INTERVENTIONuntil you type something yourself. - Live screen view is only for the cloud phone. Physical devices linked through Autopilot have no live view yet, so poll output is all you see there. Autopilot is Android only.
- Treat the token like any API key. Keep
scripts/.envout of git and out of anything you screenshot. - A loop that keeps re-creating failed tasks eats the daily soft cap. Limit retries.
Quick answers
Can Claude Code drive an iPhone this way? No. The receivers are Android: the cloud phone or your own Android phone through Autopilot.
Can it use my real phone instead of the cloud one? Yes. Link it with the Autopilot Android app, find its ID with receiver get-list, and ask for that device by name.
Once your first task poll comes back COMPLETED, the setup is done. What's left is wording, and the three task messages above are a reasonable template for the next one.
Top comments (0)