This is a guest post from Oskar Kwaśniewski, CTO and co-founder of TesterArmy, where agents test mobile apps, web apps and websites like real users.
A pull request sits in the queue for two days. The diff is small, the description is clear, but nobody has five minutes to open the simulator and tap through the screen that changed. Eventually it merges anyway, and the first person to actually exercise that flow is a user in production.
Last week Expo showed what happens when you put an agent inside EAS Workflows: a GitHub issue labeled "repro" becomes a pull request with a fix and recordings that prove it works. That's the repair half of the story. This post covers the other half: the pull request that hasn't merged yet, and what I've started calling agentic CI. Instead of running the same fixed list of jobs on every change, the pipeline works out what to build and what to test based on what actually changed.
When I talk to devs about their workflow, the bottleneck isn't writing code anymore. Agents can produce plenty of pull requests. What's missing is the confidence to click merge, because reviewing a UI change still means a human opening the app and tapping through it, and that human rarely has the time.
The fix: an agent that reads the pull request description and diff on every PR, understands what changed, opens the app on a simulator, taps through it like a real user would, and leaves you a recording of what it did and what happened.
What agentic CI means here
Regular CI runs a fixed list. Every pull request triggers the same jobs in the same order, whether it renamed a variable or rewrote onboarding. Agentic CI means that list adapts per pull request. EAS Workflows builds the app and reuses the last native build when only JavaScript changed. The TesterArmy agent reads the pull request and decides what to actually click through on the simulator.
Right now TesterArmy runs on cloud iOS Simulators and Android emulators, not physical devices (real device support is on the way). You can run it alongside whatever e2e framework you already have.
Part one: an agent writes the workflow for you
The TesterArmy docs page for Expo EAS has a section called "Instructions for AI agents." It's a prompt. Paste it into whatever coding agent you already use in the repo (Claude Code, Codex, Cursor), and it handles the setup: inspects eas.json and .eas/workflows, adds two build profiles, writes a new workflow file without touching your existing ones, and never commits a secret.
Those two profiles are the only requirement on the Expo side. TesterArmy needs an iOS Simulator .app, not an .ipa, and a single Android .apk, not an .aab.
// eas.json
{
"build": {
"testerarmy-ios-simulator": {
"ios": {
"simulator": true
}
},
"testerarmy-android-apk": {
"android": {
"buildType": "apk"
}
}
}
}
The workflow reads three values from your EAS environment: TESTERARMY_API_KEY, TESTERARMY_PROJECT_ID and TESTERARMY_GROUP_ID. Store the API key as an EAS secret, never as an EXPO_PUBLIC variable. A fourth, TESTERARMY_DYNAMIC_AGENT_ENABLED, is optional and defaults to true.
Before running any of this, you need a mobile project in TesterArmy with at least one saved test. Run that test once from the dashboard against an uploaded build first, so you know the upload and the flow work before CI starts depending on them. More detail on this in the docs.
Part two: EAS builds the app and skips what it can
The basic workflow triggers on push to main, on pull requests to main, and on workflow_dispatch. It builds the iOS Simulator app and Android APK, downloads each artifact with eas/download_build, uploads to TesterArmy, and runs your saved test group once per platform. Start here. A full native build on every run is slower, but it removes variables while you're wiring things up.
Once that's stable, let EAS skip the native build when possible. Expo's fingerprint job hashes the native runtime. The get-build job checks for an existing build with that hash. If one exists, the repack job repackages the app's metadata and JavaScript bundle on top of it, no full native rebuild needed. If no matching build exists, it falls back to a normal build.
jobs:
fingerprint:
name: Calculate app fingerprints
type: fingerprint
environment: preview
get_ios_build:
name: Find matching iOS Simulator build
needs: [fingerprint]
type: get-build
params:
platform: ios
profile: testerarmy-ios-simulator
simulator: true
fingerprint_hash: ${{ needs.fingerprint.outputs.ios_fingerprint_hash }}
wait_for_in_progress: true
repack_ios:
name: Repack iOS app
needs: [get_ios_build]
if: ${{ needs.get_ios_build.outputs.build_id }}
type: repack
params:
build_id: ${{ needs.get_ios_build.outputs.build_id }}
build_ios:
name: Build iOS Simulator app
needs: [get_ios_build]
if: ${{ !needs.get_ios_build.outputs.build_id }}
type: build
environment: preview
params:
platform: ios
profile: testerarmy-ios-simulator
The Android half mirrors this with the APK profile. The upload job runs after whichever of repack or build produced a build_id. One caveat straight from Expo's own docs: the fingerprint job is built for CNG projects. If you commit your android or ios directories directly, this won't work, so stick with the basic workflow.
This is the part that actually keeps the loop short. The native build is the slow step in any EAS workflow, and most pull requests on a mature Expo app touch JavaScript only. With repack in place, the native build happens once per runtime change, and every other pull request goes straight from repack to upload to test.
Part three: an agent works out what to test
Regression tests on every PR confirm you didn't break something that used to work. What's new here is a job that runs only on pull requests, against the same uploaded build, and figures out its own test plan.
It reads the files and diff to understand the intent behind the change, writes an ordered list of natural-language steps specific to that change, runs them on the simulator, and posts a GitHub check plus a pull request comment. The comment first shows the planned steps as a table, then updates in place with per-step results. Every run comes with a video.
run_ios_dynamic_agent:
name: Run iOS TesterArmy dynamic agent
needs: [upload_ios_app]
if: ${{ github.event_name == 'pull_request' }}
environment: preview
env:
APP_ID: ${{ needs.upload_ios_app.outputs.app_id }}
COMMIT_SHA: ${{ github.sha }}
PR_NUMBER: ${{ github.event.pull_request.number || '' }}
PR_TITLE: ${{ github.event.pull_request.title || '' }}
PR_DESCRIPTION: ${{ github.event.pull_request.body || '' }}
steps:
- uses: eas/checkout
- name: Run dynamic PR agent
run: |
npx --yes testerarmy@latest pr run-dynamic \
--project "$TESTERARMY_PROJECT_ID" \
--platform ios \
--app-id "$APP_ID" \
--pr-number "$PR_NUMBER" \
--pr-title "$PR_TITLE" \
--pr-description "$PR_DESCRIPTION" \
--commit-sha "$COMMIT_SHA" \
--output .testerarmy/dynamic-result.json
Two behaviors worth knowing before you turn this on.
First, a skip judge reads the diff before anything runs. If the pull request changes nothing a user can see (a docs edit, a config bump), the run gets reported as skipped, the check shows "Tests skipped" with the reason, and the job passes. On mobile, the judge decides per platform, so an iOS-only change still gets tested on iOS while the Android run gets skipped. Once the judge decides to test, the run always executes.
Second, a real failure fails the EAS job, but the GitHub check is advisory by default. Failed runs conclude neutral, and GitHub treats neutral as passing. If you want a failed test to actually block a merge, turn on "Block merges on failed tests" in the project's PR Testing tab and mark the "TesterArmy / Exploration Test" check as required in your branch protection settings.
Close the loop with your coding agent
If Claude Code, Cursor, or Codex writes your pull requests, you can make every one of them arrive planner-ready. Add the excerpt from the docs to AGENTS.md or CLAUDE.md once, and from then on every pull request description ends with a testing-instructions section the testing agent can execute directly.
## TesterArmy testing instructions
Every pull request description must end with a `## TesterArmy testing
instructions` section. A QA agent reads the full PR description and
writes a test plan from it, so write instructions the agent can execute:
- State the entry point: the exact route, path, or deep link where the
change is visible.
- Give numbered steps to reach and exercise the feature, in the order a
user would take them. Name buttons, tabs, and fields by their visible
labels.
- State the expected behavior after each meaningful action.
- If the flow needs an account, reference the test account by its label.
Never paste credentials or secrets.
- List required setup the planner cannot infer: feature flags, mock or
sandbox mode, seed data, or a configuration link that must be opened
first.
- Add an "Out of scope" line for adjacent surfaces the plan should not
test.
That's the whole loop. A coding agent writes the change and the instructions. EAS builds or repacks. A testing agent plans from the instructions and the diff, taps through the app, and leaves a video on the pull request. Nobody writes a test script.
What this looks like at fifty pull requests a day
Juno builds a health assistant for people living with chronic illness, on iOS and Android, with React Native and Expo. Seven people ship one to two updates a week, which on their side means fifty to a hundred pull requests a day.
Every pull request gets about six TesterArmy runs at roughly five minutes each, across both platforms, covering onboarding, symptom logging, medication, notifications, paywall, and account flows. The flows are written as plain-language user journeys, not selectors. In July 2026 the team merged 501 pull requests. In August, 791. Marshall Gould, Juno's co-founder and CEO, put it more bluntly than I would: "We are shipping PRs ten times faster."
Where to start
- Clone the example repository. It has the workflow file, both profiles, and the dynamic agent job wired up and ready to run.
- Get a saved test passing from the dashboard first, then turn on the dynamic agent, then layer in fingerprint and repack.
- Check the Expo EAS docs for the agent prompt and troubleshooting section.
If it snags on your app, open an issue on the example repository and tell us where.
This post is based on content from the Expo blog. Follow @expo for more React Native content.
Top comments (0)