Emulators in CI are fine until your app cares about being on a real device. Once a sensor check or hardware-specific behaviour enters the picture, a green pipeline stops meaning much.
I run a rack of real Android phones in Singapore that people drive over adb, and this came out of running it. Here is the workflow I'd start from, including the parts that fail quietly the first time you try it on a hosted runner.
name: android device tests
on: [push, pull_request]
jobs:
device-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: android-actions/setup-android@v3
- name: build debug apk
run: ./gradlew assembleDebug
- name: connect
env:
ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
run: |
for i in 1 2 3 4 5; do
adb connect "$ADB_ENDPOINT" && sleep 2
if adb devices | grep -q "${ADB_ENDPOINT}.*device"; then
echo "connected"; break
fi
echo "retry $i"; adb disconnect "$ADB_ENDPOINT"; sleep 5
done
adb devices
That's the first half. The rest of the job installs, tests and releases, and I'll get to it below.
The shape of it
A GitHub Actions runner is a throwaway Linux box with no phone attached. To test on a real device, the runner connects out to a phone over adb, installs your build, runs the tests, and drops the connection. The phone lives elsewhere and stays up between runs. The runner is the disposable part.
You need three things:
- platform-tools on the runner, so you have a current
adb - the phone's
host:portstored as a secret - a test runner, such as Espresso or instrumented tests through Gradle, or Appium
Getting the phone reachable is a setup step, not part of the test. Keep it separate in your head and in your job, because it fails differently.
Keep the endpoint in a secret
Put the adb endpoint in a repository secret called ADB_ENDPOINT and read it from the environment in each step. That keeps it out of the workflow file and out of logs. If your provider also hands you a token for authorising the adb connection, store it the same way.
One caution: GitHub masks secret values in logs, but a masked value can still leak if you transform it before printing. Don't pipe it through sed or base64 into something that echoes.
The rest of the job
After the connect step, add install, test, and disconnect:
- name: install build
env:
ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
run: adb -s "$ADB_ENDPOINT" install -r app/build/outputs/apk/debug/app-debug.apk
- name: run tests
env:
ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
run: ANDROID_SERIAL="$ADB_ENDPOINT" ./gradlew connectedDebugAndroidTest
- name: disconnect
if: always()
env:
ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
run: adb disconnect "$ADB_ENDPOINT"
The if: always() on the disconnect step is not optional. If a test fails and the job stops, you still want to release the device cleanly so the next run doesn't trip over a stale connection.
The connect retry is the whole game
The most common cause of flaky device CI is adb connect racing the network. A hosted runner cold-starts every time, so the first adb connect often returns before the tunnel is really up, and adb devices lists the phone as offline.
The loop in the first block handles that: connect, wait, check that the state is device, and only then move on. If you want the job to fail loudly when all five attempts are used up, add a final check after the loop that exits non-zero when the endpoint still isn't in the device state.
Treat a missing device as a setup failure, not a test failure. That way "the phone wasn't reachable" and "the app is broken" show up as different red marks, and you stop wasting an afternoon debugging the app when the network was the problem.
Gotchas specific to CI
-
adb server version doesn't match: the runner's adb is stale. Usesetup-androidso you get current platform-tools. - Install fails with a signature error: a previous build was signed differently. Run
adb uninstall <package>before installing. - Tests hang, then time out: default adb timeouts are tuned for a cable, not a network hop. Raise them. For Appium, bump
adbExecTimeout. - Device
offlineright after connect: that's the race described above. Retry with a state check.
Two PRs, one phone
This one hits teams as they grow. Two pull requests trigger at the same moment, both jobs connect to the same phone, and they corrupt each other's state. Installs get replaced mid-test, and tests fail in ways that look like app bugs and aren't.
A concurrency group makes device jobs queue instead of colliding:
concurrency:
group: android-device
cancel-in-progress: false
cancel-in-progress: false matters here. It means a running test finishes rather than being killed by a newer push, which is what you want when a real device is mid-test. A cancelled job on a real handset can leave the app half-installed or a test mid-write, and the next run inherits that.
If you'd rather lock the phone itself than rely on the workflow-level queue, there's a worked version of the full flow (tunnel, lock, connect, test, release) in this examples repo.
Scaling to a matrix of phones
Once one phone works, a matrix across several devices is a small step. Give each phone its own secret and fan out:
strategy:
matrix:
phone: [ADB_ENDPOINT_A, ADB_ENDPOINT_B, ADB_ENDPOINT_C]
Then pull the right endpoint per job with secrets[matrix.phone]:
env:
ADB_ENDPOINT: ${{ secrets[matrix.phone] }}
Each job now drives its own handset in parallel. Note that the concurrency group above would serialise the whole matrix, so give each matrix entry its own group, for example android-device-${{ matrix.phone }}, so different phones run side by side while jobs for the same phone still queue.
Why bother with real hardware
A passing run on an emulator tells you the app works on an emulator. A passing run on a real phone on a real mobile network tells you it works for an actual user. For anything that reads device signals, that is the difference between a meaningful test and a comforting one.
The cost is that your pipeline now depends on a network path to a physical device, which is why the retry, the always-disconnect step, and the concurrency group are not extras. Without them the first week of real-device CI is mostly chasing flakes that have nothing to do with your code.
Top comments (0)