I have never signed off on a mobile release feeling confident about the device matrix. Not once. Nobody I've worked with has either.
The reason is boring. Walking the same screen across fifteen screen sizes, three OS versions and two accessibility scales is a full day of someone's life, and it's a day nobody budgets for. So you test on three devices, you ship, and you let the reports tell you about the other twelve.
That's the constraint four separate releases quietly removed this week.
Callstack took Apex generally available — a coding model specialised in React, strongest on RN and Next.js. Codex shipped Mobile Dev: Simulator, Instruments, logcat and Metro in one surface. Stim Desktop 0.1 watches and steers an agent's React Native work with live devices, replay and logs. Medula shipped headless devtools built for agents.
Nobody coordinated that. Four teams hit the same wall in the same quarter.
And the wall was never code generation. Agents write RN components fine. The wall was that the agent couldn't see the device — and neither could we, at any width that mattered.
So the shift isn't "agents write my tests". It's that device breadth stopped being expensive. Twenty simulator configurations running the same journey is now a cheap thing to kick off, not a sprint you have to argue for.
What I'd stay honest about: breadth is not confidence. An agent exploring twenty devices and reporting back is twenty chances to be told something plausible and wrong. It only becomes real when the run leaves an artefact behind — a flow file your CI replays on every PR. Then the second run is deterministic and the twentieth is free.
That's the whole play. Pick the matrix you've been quietly ignoring. Have the agent explore it once. Keep what it finds as tests you own.
We spent a decade calling the device matrix a QA problem. It was a cost problem, and the cost just moved.
Top comments (0)