DEV Community

Xavier Fok for cloudf.one

Posted on Originally published at cloudf.one AI-assisted

Run Android Tests on Real Phones in GitHub Actions

Emulators in CI are fine until your app cares about being on a real device. Once a sensor check or hardware-specific behaviour enters the picture, a green pipeline stops meaning much.

I run a rack of real Android phones in Singapore that people drive over adb, and this came out of running it. Here is the workflow I'd start from, including the parts that fail quietly the first time you try it on a hosted runner.

name: android device tests
on: [push, pull_request]

jobs:
  device-tests:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: android-actions/setup-android@v3

      - name: build debug apk
        run: ./gradlew assembleDebug

      - name: connect
        env:
          ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
        run: |
          for i in 1 2 3 4 5; do
            adb connect "$ADB_ENDPOINT" && sleep 2
            if adb devices | grep -q "${ADB_ENDPOINT}.*device"; then
              echo "connected"; break
            fi
            echo "retry $i"; adb disconnect "$ADB_ENDPOINT"; sleep 5
          done
          adb devices
Enter fullscreen mode Exit fullscreen mode

That's the first half. The rest of the job installs, tests and releases, and I'll get to it below.

The shape of it

A GitHub Actions runner is a throwaway Linux box with no phone attached. To test on a real device, the runner connects out to a phone over adb, installs your build, runs the tests, and drops the connection. The phone lives elsewhere and stays up between runs. The runner is the disposable part.

You need three things:

  • platform-tools on the runner, so you have a current adb
  • the phone's host:port stored as a secret
  • a test runner, such as Espresso or instrumented tests through Gradle, or Appium

Getting the phone reachable is a setup step, not part of the test. Keep it separate in your head and in your job, because it fails differently.

Keep the endpoint in a secret

Put the adb endpoint in a repository secret called ADB_ENDPOINT and read it from the environment in each step. That keeps it out of the workflow file and out of logs. If your provider also hands you a token for authorising the adb connection, store it the same way.

One caution: GitHub masks secret values in logs, but a masked value can still leak if you transform it before printing. Don't pipe it through sed or base64 into something that echoes.

The rest of the job

After the connect step, add install, test, and disconnect:

      - name: install build
        env:
          ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
        run: adb -s "$ADB_ENDPOINT" install -r app/build/outputs/apk/debug/app-debug.apk

      - name: run tests
        env:
          ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
        run: ANDROID_SERIAL="$ADB_ENDPOINT" ./gradlew connectedDebugAndroidTest

      - name: disconnect
        if: always()
        env:
          ADB_ENDPOINT: ${{ secrets.ADB_ENDPOINT }}
        run: adb disconnect "$ADB_ENDPOINT"
Enter fullscreen mode Exit fullscreen mode

The if: always() on the disconnect step is not optional. If a test fails and the job stops, you still want to release the device cleanly so the next run doesn't trip over a stale connection.

The connect retry is the whole game

The most common cause of flaky device CI is adb connect racing the network. A hosted runner cold-starts every time, so the first adb connect often returns before the tunnel is really up, and adb devices lists the phone as offline.

The loop in the first block handles that: connect, wait, check that the state is device, and only then move on. If you want the job to fail loudly when all five attempts are used up, add a final check after the loop that exits non-zero when the endpoint still isn't in the device state.

Treat a missing device as a setup failure, not a test failure. That way "the phone wasn't reachable" and "the app is broken" show up as different red marks, and you stop wasting an afternoon debugging the app when the network was the problem.

Gotchas specific to CI

  • adb server version doesn't match: the runner's adb is stale. Use setup-android so you get current platform-tools.
  • Install fails with a signature error: a previous build was signed differently. Run adb uninstall <package> before installing.
  • Tests hang, then time out: default adb timeouts are tuned for a cable, not a network hop. Raise them. For Appium, bump adbExecTimeout.
  • Device offline right after connect: that's the race described above. Retry with a state check.

Two PRs, one phone

This one hits teams as they grow. Two pull requests trigger at the same moment, both jobs connect to the same phone, and they corrupt each other's state. Installs get replaced mid-test, and tests fail in ways that look like app bugs and aren't.

A concurrency group makes device jobs queue instead of colliding:

concurrency:
  group: android-device
  cancel-in-progress: false
Enter fullscreen mode Exit fullscreen mode

cancel-in-progress: false matters here. It means a running test finishes rather than being killed by a newer push, which is what you want when a real device is mid-test. A cancelled job on a real handset can leave the app half-installed or a test mid-write, and the next run inherits that.

If you'd rather lock the phone itself than rely on the workflow-level queue, there's a worked version of the full flow (tunnel, lock, connect, test, release) in this examples repo.

Scaling to a matrix of phones

Once one phone works, a matrix across several devices is a small step. Give each phone its own secret and fan out:

strategy:
  matrix:
    phone: [ADB_ENDPOINT_A, ADB_ENDPOINT_B, ADB_ENDPOINT_C]
Enter fullscreen mode Exit fullscreen mode

Then pull the right endpoint per job with secrets[matrix.phone]:

env:
  ADB_ENDPOINT: ${{ secrets[matrix.phone] }}
Enter fullscreen mode Exit fullscreen mode

Each job now drives its own handset in parallel. Note that the concurrency group above would serialise the whole matrix, so give each matrix entry its own group, for example android-device-${{ matrix.phone }}, so different phones run side by side while jobs for the same phone still queue.

Why bother with real hardware

A passing run on an emulator tells you the app works on an emulator. A passing run on a real phone on a real mobile network tells you it works for an actual user. For anything that reads device signals, that is the difference between a meaningful test and a comforting one.

The cost is that your pipeline now depends on a network path to a physical device, which is why the retry, the always-disconnect step, and the concurrency group are not extras. Without them the first week of real-device CI is mostly chasing flakes that have nothing to do with your code.

Top comments (0)