DEV Community

Cover image for My harness taught the agent how to build. Now it needs to see.
Valentine Tikhomirov
Valentine Tikhomirov

Posted on

My harness taught the agent how to build. Now it needs to see.

At the end of my last post I said the next step was letting my pipeline iterate on its own review findings. I didn't get there. I ran into something more basic first: the part of the harness that lets the agent use the app.

That part, rn-app-driver, reads the React component tree and taps the simulator so the agent can check its own work. It did its job. But the more I used it, the more it felt like the wrong shape for the problem. So I pulled it out of the harness and started building a separate tool around it.

This post is about why, and about a few things an agent simply can't learn from reading your code.

Where the driver stopped fitting into a skill

Only the agent could see it. The driver ran inside the agent's session. While it was tapping through my app, I was looking at a terminal. I only found out what it actually did by reading the log afterwards.

It only worked in one agent. It was built as a Claude Code skill. I also use Codex, and I'd like the same capability to work with whatever agent comes next.

Safety by word list doesn't hold. After the last post, two readers pointed out the same weak spot: the driver decided whether an action was risky by its label. "Delete" and "Save" stopped for confirmation, but a button labelled "Done" could change data and sail right through. They were right. The real fix isn't a better word list. It's putting the decision in front of a human who can see the screen.

A component name isn't a location. The driver told the agent "this is HomeActions/log-attack". What the agent actually needs next is "this is HomeActions.tsx:42", so it can open the right file instead of searching for it.

Things an agent can't see from your code

While building this, I kept measuring what the running app tells you that the source doesn't. Three findings stood out. All measured on one Mac, iOS simulator, RN 0.87 with React 19.

1. JS errors leave no trace in the iOS system log. I recorded 90 seconds of normal app use: 3,498 native log events. Every single one came from Apple's frameworks. Zero from the app, React Native or Hermes, including during deliberate JS errors. If your agent "checks the device logs", it's checking the wrong place for anything that happens in JavaScript.

2. In React 19, component locations can be off by 20 lines. React 19 removed _debugSource, the old way to map an element back to its file and line. The replacement points elements rendered inside .map(...) at the start of the callback, not at the tag. In a real app that was 8 out of 22 points I tested, with the actual tag 14–24 lines further down. An agent trusting that line edits the wrong place.

3. "The screen changed after the tap" proves nothing. The simulator's video stream keeps sending around 53 frames per second even when nothing on screen moves. So "a new frame arrived" is not evidence that your tap landed. What did work: comparing frame size against the idle background. A real UI reaction produced frames of 4–29 KB, idle ones were well under 2 KB.

None of this is visible in a diff. All of it matters when an agent claims "I checked, it works".

What I'm building

A desktop app for React Native development where the agent and I look at the same running app:

  • For the agent: the component tree with real file:line locations, logs and errors from the app, the simulator screen, taps, gestures and typing. It connects over MCP, so it works with Claude Code and Codex today, and with any MCP-capable agent in principle.
  • For me: the live simulator, what the agent is doing on it, and an approval card for risky actions. The agent asks, I see the screen and the action, I decide.
  • One rule I won't bend: the tool never types into my terminal for me. It can paste text, but Enter is always mine. I learned this the hard way when a test typed a message plus Enter into a terminal running an agent in auto mode, and kicked off a model turn nobody asked for.

The first stage, the part the agent talks to, already works and has been tested with both Claude Code and Codex. The desktop app is in progress: live iOS simulator at around 48 fps, and clicking anywhere on the screen finds the component under that point in about 5 ms.

The harness isn't going anywhere. It stays the "how to build" part: skills, review agents, pipelines. The new tool becomes its eyes and hands.

An honest result

To check that the tool actually helps, I planted a bug in my gym app and asked both agents to find it. They found it in 6 out of 6 runs. Then I realised they found it by reading the code, without needing the tool at all.

That's the lesson I keep coming back to. If you want to show that runtime tooling helps an agent, you have to test it on bugs that aren't visible in the source: wrong state, timing, a layout that only breaks on device. Otherwise you're measuring how well the model reads code.

What's next

I'll write up the most surprising findings separately, with numbers:

  • streaming the iOS simulator into an Electron app at ~50 fps (screenshots topped out at 2.3 fps);
  • why an environment variable from your terminal never reaches your MCP server in Codex;
  • a trap for anyone writing MCP servers: what Claude Code actually shows the model when a tool returns both text and structured content.

And a question for you: how do you check what your agent did in the app right now? Do you open the simulator yourself after every run, or trust the summary?

Top comments (0)