In my previous post I tested a single skill from my React Native harness: a refactoring agent. This time I ran the whole thing end to end for the first time. One confirmation from me, then the agent planned a feature, built it, tested it, tapped through it on the iOS simulator, and ran a review pass on its own code.
Here's how it's wired and what actually happened.
One checkpoint, then hands-off
The skill is called feature-pipeline. Before touching code, it sends me a single message with the proposed scope, the design direction, and how data, state and errors will be handled. I approve once, and from then on it doesn't stop to ask:
- Build the UI and logic.
- Test it.
- Verify on device: drive the running app on the simulator.
- Review: call the harness's other agents as real subagents (refactoring, test coverage, architecture review, plus instrumentation agents if the feature needs them).
- Report: one consolidated list of what was done and what's still open.
It never merges to main and never touches deployment or store submission. Those are hard boundaries, not settings.
The part I'm most excited about: how the agent sees the app
Step 3 uses another skill, rn-app-driver. Most tools that let an agent use a mobile app work from screenshots: take a picture, guess where the button is, tap coordinates. That's slow and expensive in context.
The driver reads the app's React component tree straight from Hermes instead. It takes around 10 ms and 1–3k characters, and looks like this (from my headache tracker):
screen: Home
@Home/scroll ScrollView [0,62,402,723]
"Saturday" "September 26"
HomeActions
@HomeActions/log-attack Pressable "Log attack" icon:Plus [24,561,354,84]
TabBar
@TabBar/history Pressable "History" icon:TabHistory [101,796,100,44]
Every @id is something the agent can act on, with stable ids across runs and a real frame in points. Taps are real touch events on the simulator (via AXe), aimed at those frames, so the agent never guesses coordinates. Screenshots are still there for visual checks, just not for navigation.
A few rules make it usable rather than scary:
- One action per command. Each action waits for the UI to settle and prints the new screen, and the agent checks it before deciding the next step. No planning five taps ahead on screens it hasn't seen.
- A safety policy. Anything whose label looks like it changes data (save, delete, log, edit, and so on) stops with an error until I explicitly confirm that specific action.
- Everything is logged. Each session writes an action log, and a walked flow can be saved as a test case and replayed later. The replay exits with 0 or 1, so it can eventually become a regression gate in CI.
The run
The target was a new "Exercises list" screen in my gym rep counter app, another side project.
- Build: about 5–7 minutes from approval to working code.
- On-device check: a few more minutes, and it took several runs. Verifying a screen like this needs the app in exactly the right state with the right data, and getting there reliably is the hard part.
- Review agents: found only code smells, nothing structural.
- Result: a report with recommendations prioritized by importance. No PR.
That last one was my decision, not a failure. With this much access, a run can end in many ways, and for now I want the irreversible steps (committing, pushing, opening a PR) to stay with a human. A full cycle with builds and pushes is on the list to test later, deliberately and separately.
What surprised me
The speed and autonomy, honestly. Scoping was sensible, it picked the right agents, and it really did use the app on the simulator instead of just reading its own code back.
But it made mistakes, and the interesting part is that it pointed them out itself in the final report as things worth fixing. Which raises the obvious question: if the agent already knows what's wrong, why stop after one pass?
What I learned
- Read the structure, not the pixels. For navigating an app, the component tree beats screenshots on speed, cost and reliability. Screenshots are for checking how things look.
- State is the real problem of on-device testing. Tapping is easy. Reliably getting the app into the exact state a check needs is not.
- An agent that lists its own mistakes deserves a second pass. Running the implement-and-review cycle several times, not once, is my next priority.
- Keep the irreversible steps human, for now. Autonomy up to a tested, reviewed result is great. What happens with that result is still my call.
What's next
- Let the pipeline iterate on its own recommendations for a few cycles and compare the result with a single pass.
- Start saving walked flows as replayable test cases and keeping the logs, so the next posts can show real traces and numbers instead of summaries.
- Test the full cycle, including commits and pushes, on a project where that's safe.
If you're letting agents drive your app on a simulator: how do you deal with test data and state? That's the part I'd most like to compare notes on.
Top comments (0)