<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mohammad Haidar</title>
    <description>The latest articles on DEV Community by Mohammad Haidar (@moehydr).</description>
    <link>https://dev.to/moehydr</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4114147%2Fa4109668-257e-41e5-b418-c525d6386b1d.jpeg</url>
      <title>DEV Community: Mohammad Haidar</title>
      <link>https://dev.to/moehydr</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/moehydr"/>
    <language>en</language>
    <item>
      <title>Our Test Runner Couldn't See Inside a WebView. Then We Gave It AI.</title>
      <dc:creator>Mohammad Haidar</dc:creator>
      <pubDate>Thu, 10 Sep 2026 12:52:21 +0000</pubDate>
      <link>https://dev.to/moehydr/our-test-runner-couldnt-see-inside-a-webview-then-we-gave-it-ai-4mh4</link>
      <guid>https://dev.to/moehydr/our-test-runner-couldnt-see-inside-a-webview-then-we-gave-it-ai-4mh4</guid>
      <description>&lt;p&gt;&lt;em&gt;Test IDs are a nightmare in a hybrid app, and in some containers there is nothing to attach one to. We added AI selectors to our test runner so it could find elements from pixels instead.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every end-to-end test begins with a negotiation. You want to tap a button, so the button needs a stable identity, so someone opens the app code and adds a test ID. Multiply that by every element in every flow and the real cost becomes clear. It was never writing the tests. It was instrumenting the app so the tests could find anything at all.&lt;/p&gt;

&lt;p&gt;In a hybrid app, that bill comes three times over.&lt;/p&gt;

&lt;h2&gt;
  
  
  The nightmare
&lt;/h2&gt;

&lt;p&gt;Our app ships as one binary assembled from several kinds of container. Some screens are native. Some are cross-platform. Some are embedded web. A single user journey walks through all of them without noticing.&lt;/p&gt;

&lt;p&gt;The test runner notices. A test ID has to be added in whichever layer actually draws the element, in that layer's language, with that layer's build and release cycle. One flow crossing three container types meant three separate instrumentation changes before anyone could write its first line.&lt;/p&gt;

&lt;p&gt;Some elements have nothing to attach an ID to. A system dialog ships without test IDs and always will. And then the webviews, where it stopped being expensive and became impossible. Inside an embedded web container the native view tree often collapses into a single opaque node. No button, no label, no hierarchy to walk. The runner would face a screen full of content a human reads perfectly well and report that there was nothing to select.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we did before
&lt;/h2&gt;

&lt;p&gt;Nothing good. We tapped fixed percentages of the screen and hoped no layout would shift. We wrote flows that stopped politely at the boundary of the thing they were meant to test. We left comments marking journeys as blocked, waiting on instrumentation that never arrived.&lt;/p&gt;

&lt;p&gt;Whole regions of the app went untested, not because nobody cared but because the vocabulary for describing them did not exist. Every workaround was a promise to come back later, and none were kept.&lt;/p&gt;

&lt;h2&gt;
  
  
  Until we stopped needing selectors
&lt;/h2&gt;

&lt;p&gt;There is exactly one thing every container in our app has in common. Native, cross-platform, embedded web, and the system's own dialogs all end up as pixels. A screenshot does not care which layer drew it.&lt;/p&gt;

&lt;p&gt;So we stopped asking the view tree where things were and started asking an AI to look at the screen.&lt;/p&gt;

&lt;p&gt;We maintain a fork of our test runner, which already shipped a couple of AI-backed assertions, so the extension point was there. We added three commands of our own, all built on a vision model that receives the screenshot and answers questions about it.&lt;/p&gt;

&lt;p&gt;The first takes a natural-language query and hands back a point you can tap. You describe the element the way you would describe it to a colleague looking over your shoulder, and colour, order, and position are all fair game, because the AI is looking at the same picture you are. No identity on the element, no change to the app, and no dependency on a hierarchy existing.&lt;/p&gt;

&lt;p&gt;The second command takes a reference image instead of words, matching a component snapshot against a live screen. The third evaluates a condition in plain language, so a flow can branch on what the AI sees rather than on the presence of a selector.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making an AI selector survive CI
&lt;/h2&gt;

&lt;p&gt;This is where most AI testing demos quietly end. Asking a model where the button is and tapping its answer looks wonderful once and falls apart across a suite that runs nightly. Screens are dense, elements are small, and a coordinate off by a few percent lands on the neighbouring row.&lt;/p&gt;

&lt;p&gt;So the AI gets up to three passes at the same screen, each one narrowing the answer from the last.&lt;/p&gt;

&lt;p&gt;Take a screen with three radio buttons stacked vertically. None of them carries a test ID, and the only thing separating the one you want from the other two is that it is red. The test says exactly that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;extractPointWithAI&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;query&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;the third radio button, the red one&lt;/span&gt;
    &lt;span class="na"&gt;outputVariable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RED_OPTION&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;tapOn&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;point&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${RED_OPTION}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two lines, and this is what runs underneath them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;          THE SCREEN                      WHAT THE VIEW TREE SAYS
   ┌──────────────────────────┐
   │                          │             RadioButton
   │    ◯    Option A         │             RadioButton
   │                          │             RadioButton
   │    ◯    Option B         │
   │                          │             three identical nodes.
   │    ◯    Option C         │             no ids. no colour.
   │         ^^^^^^^^ red     │             nothing to select on.
   │                          │
   └──────────────────────────┘
                 │
                 │   screenshot  +  serialized view hierarchy
                 ▼
   ┌────────────────────────────────────────────────────────┐
   │  PASS 1    "the third radio button, the red one"       │
   │                                                        │
   │            reasons: three stacked controls, the        │
   │            lowest one is tinted red                    │
   │            answer:  22%,68%                            │
   └────────────────────────────────────────────────────────┘
                 │
                 │   crop 30% of the screen around 22%,68%
                 ▼
   ┌────────────────────────────────────────────────────────┐
   │  PASS 2    ┌────────────────────┐                      │
   │            │                    │   same question,     │
   │            │   ◯   Option C     │   ten times the      │
   │            │       ^^^^^^^ red  │   pixels per element │
   │            └────────────────────┘                      │
   │                                                        │
   │            answer:  44%,52%  of the crop               │
   │            mapped:  21%,66%  of the full screen        │
   └────────────────────────────────────────────────────────┘
                 │
                 │   the point  +  the full screenshot
                 ▼
   ┌────────────────────────────────────────────────────────┐
   │  PASS 3    "is 21%,66% the element I asked for?"       │
   │                                                        │
   │            answer:  yes  (or a corrected point)        │
   └────────────────────────────────────────────────────────┘
                 │
                 ▼
          tapOn:  point: 21%,66%

   any pass fails  →  keep the previous pass's answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pass one gets the view hierarchy alongside the screenshot, which matters more than it sounds. Even when the tree cannot produce a unique selector, as with three identical radio nodes, it usually still carries the real strings and the real bounds, so the model reads labels instead of transcribing them off an image. Colour and fill it cannot tell you at all. That part only exists in the pixels. When the tree is opaque, the screenshot alone still works, just with less help.&lt;/p&gt;

&lt;p&gt;Pass two exists because a full screen gives the model very few pixels per element. Cropping to a region around the first answer asks the identical question with an order of magnitude more detail, then maps the refined answer back to full-screen coordinates.&lt;/p&gt;

&lt;p&gt;Pass three shows the model the point it picked and asks whether it got it right. An AI grading its own answer sounds circular. It catches real mistakes.&lt;/p&gt;

&lt;p&gt;Every pass fails soft, falling back to the previous answer rather than failing the test, and each appends to a reasoning log attached to the command. When something breaks at three in the morning, the report says what the AI thought it was looking at, not merely that a tap missed. Passes are configurable and the default is one, so precision is there when a flow needs it and nobody pays for extra AI calls otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it unlocked
&lt;/h2&gt;

&lt;p&gt;The radio buttons are the easy case to picture, and there are harder ones where a test ID would not have helped even if someone had added it.&lt;/p&gt;

&lt;p&gt;The handle of a price range slider is a position on a track. No identity expresses "drag from here to about a third of the way along", because what the test needs is a coordinate, not a name. A seat map renders differently on every run, so no fixed identity will ever point at an available seat, and the flow has to branch on what is actually on screen.&lt;/p&gt;

&lt;p&gt;Those were never naming problems. There was nothing to name. Plenty of elements sit in a perfectly readable hierarchy and still have no useful identity, because the difficulty is not naming. It is meaning.&lt;/p&gt;

&lt;h2&gt;
  
  
  The discipline that keeps AI honest
&lt;/h2&gt;

&lt;p&gt;An AI selector is the last resort, not the default. Our guidance ranks AI below text and below test IDs, and that ordering does real work.&lt;/p&gt;

&lt;p&gt;The clearest evidence is that the migration runs both ways. When product copy churns and an exact-string assertion breaks for the third time, a team rewrites it as a description of what the screen should show, and it stops breaking on wording. But when a component finally gets a stable identity, teams go back and replace the AI call with the plain selector.&lt;/p&gt;

&lt;p&gt;That is the right instinct. A plain selector is fast, free, and deterministic. An AI call is none of those things. It is for the cases where no selector can exist, or where creating one costs more than the test is worth. Reaching for AI first trades a maintenance problem for a slower, more expensive one.&lt;/p&gt;

&lt;p&gt;The same discipline shows up as something we built that nobody uses. The command that matches a reference image works fine and has never appeared in a single test. When a sentence does the job, nobody goes hunting for a screenshot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The AI command we are building now
&lt;/h2&gt;

&lt;p&gt;It asserts that everything on screen is written in the expected language, with an ignore list for what is meant to stay untranslated, like brand and station names. The AI reports back every string that looks wrong, the language it thinks that string is in, and why.&lt;/p&gt;

&lt;p&gt;We ship in a lot of languages, and localization gaps are the classic bug no assertion catches, because there is nothing to select. A missing translation is not a missing element. It is a present element with the wrong words inside it, and the only way to find one used to be a human who reads that language looking at the screen.&lt;/p&gt;

&lt;p&gt;That is the real shape of the win. Not that AI made the tests smarter, but that it let them ask questions our old vocabulary could not express.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;The hidden cost of end-to-end testing is instrumentation, and hybrid apps pay it several times over. In the containers where no selector exists, that cost is not high. It is infinite, and the tests simply do not get written.&lt;/p&gt;

&lt;p&gt;Put an AI behind the selector and the container type stops mattering, because pixels are the one interface every layer shares. Keep it as the fallback, since the best test ID is still a real test ID. The point is that you no longer need one to write the test.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>e2e</category>
      <category>testing</category>
      <category>reactnative</category>
    </item>
    <item>
      <title>My Most Useful Repo Is Empty, and My AI Agent Lives in It</title>
      <dc:creator>Mohammad Haidar</dc:creator>
      <pubDate>Mon, 07 Sep 2026 15:36:34 +0000</pubDate>
      <link>https://dev.to/moehydr/my-most-useful-repo-is-empty-and-my-ai-agent-lives-in-it-2fc3</link>
      <guid>https://dev.to/moehydr/my-most-useful-repo-is-empty-and-my-ai-agent-lives-in-it-2fc3</guid>
      <description>&lt;p&gt;&lt;em&gt;How an empty host repo binds two codebases for an AI coding agent: instructions, skills, memory, and verification that ends on a device.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A native module moved from the Android shell into the React Native repo, and one method got a new name on the way. The JS kept calling the old one. The build passed, the app started, and a whole screen quietly stopped working. Nothing in either repo could have flagged it, because the name lived in both and neither side knew about the other.&lt;/p&gt;

&lt;p&gt;That is the failure I built this setup to catch. It starts from a simpler observation.&lt;/p&gt;

&lt;p&gt;AI-assisted coding is great inside one repo. The agent reads the repo's instructions, sees the whole tree, and its mistakes stay local.&lt;/p&gt;

&lt;p&gt;But what if your app ships from two repos, and every real change has to land in both, in the right order? Open the agent one level up and it sees every project on your machine, with nothing telling it how these two relate.&lt;/p&gt;

&lt;p&gt;My answer is a third repo. Empty of application code. It holds the binding between the other two, the instructions that only make sense across them, the skills, and the memory. The agent runs from there. I use Claude Code. The pattern ports to any agent that reads instruction files.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup, and the gap
&lt;/h2&gt;

&lt;p&gt;One Android APK, built from a React Native monorepo and a native Kotlin shell. The React Native side publishes its bundle and its bridge modules' native code as an npm package. The Android side pins that package by exact version and points Gradle into &lt;code&gt;node_modules&lt;/code&gt;. No over-the-air updates. The JS a user runs is whatever version the pin says.&lt;/p&gt;

&lt;p&gt;Between them sits a hand-maintained bridge. Module names, method names, event names, and storage keys exist on both sides and must match exactly. No codegen checks them. When one side drifts, nothing fails at build time. The call resolves to &lt;code&gt;undefined&lt;/code&gt; at runtime. That is the story at the top.&lt;/p&gt;

&lt;p&gt;Each repo's docs cover its own half. Nobody owns the gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  The host repo
&lt;/h2&gt;

&lt;p&gt;The whole thing fits in one screen:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;host/
├── CLAUDE.md            the contract between the two repos
├── notes/               cross-session state, written for the agent
└── .claude/
    ├── settings.json    points at the two repos
    ├── skills/          bridge wiring, local builds
    ├── agents/          the review agent
    └── agent-memory/    what that agent has already learned
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The settings file at &lt;code&gt;.claude/settings.json&lt;/code&gt; is the whole trick:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"additionalDirectories"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"../rn-app"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"../android-shell"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A session started here sees exactly those two trees and nothing else. The host's &lt;code&gt;CLAUDE.md&lt;/code&gt; imports each repo's own &lt;code&gt;CLAUDE.md&lt;/code&gt;, so repo-internal rules stack underneath. On top, the host adds what only makes sense across the boundary: how the repos bind, the four identifier kinds that must match, where native modules live during an ongoing migration, areas owned by other teams, and a list titled "things this file does not yet cover" so the agent says "not specified" instead of inventing.&lt;/p&gt;

&lt;p&gt;Plus the one rule that exists only because there are two repos: plan before editing, because order matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills that compose
&lt;/h2&gt;

&lt;p&gt;Procedures live in skills, markdown files the agent pulls in when a task matches their description. The interesting part is how they chain.&lt;/p&gt;

&lt;p&gt;The bridge-wiring skill knows the two paths for a native module: the short one entirely inside the React Native repo, and the legacy one that touches the shell. It defaults to the short path and flags the legacy one. It also carries the migration checklist, with one rule: never rename a module while moving it.&lt;/p&gt;

&lt;p&gt;The local-build skill knows how to run a debug build against Metro, and how to build a release-style APK against an unpublished bundle by pointing the shell's &lt;code&gt;package.json&lt;/code&gt; at a local path instead of a published version.&lt;/p&gt;

&lt;p&gt;A review agent, with its own memory of past pitfalls, reads a pull request through one team's ownership boundaries and comments only on changed lines that team owns.&lt;/p&gt;

&lt;p&gt;On their own, each is a checklist. Together they close a loop the repos cannot close by themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verification that reaches the device
&lt;/h2&gt;

&lt;p&gt;Ask for a new native module. The wiring skill writes the TS wrapper and the Kotlin class and registers it. Then, instead of stopping at "done", the local-build skill takes over: build the bundle locally, link it into the shell, assemble the APK, install it, and exercise the feature.&lt;/p&gt;

&lt;p&gt;That last step is where the silent &lt;code&gt;undefined&lt;/code&gt; would otherwise surface in QA days later. Now it surfaces in the same session, on the actual APK a user would get. The agent reports what it saw, not what it expects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory
&lt;/h2&gt;

&lt;p&gt;Because the agent runs from the host, its memory is scoped to cross-repo work rather than to either repo. Notes on long-running efforts that span both sides sit next to the instructions, each with a status and a last-updated date. The review agent keeps its own store of pitfalls it has already found. The next session starts from what was verified, not from zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does not fix
&lt;/h2&gt;

&lt;p&gt;The bridge is still hand-maintained. The loop catches mismatches. It does not prevent them. Codegen would, and that is a bigger change than a markdown file.&lt;/p&gt;

&lt;p&gt;The rules only work if they are read. When the agent slips back into one-repo thinking, the fix is usually a sharper sentence, not a longer file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;If one artifact is built from more than one repo, the contract between them lives nowhere. Give it a home your tools read, scope the agent to exactly the repos involved, and let the skills compose into a loop that ends on a device.&lt;/p&gt;

&lt;p&gt;The empty repo is just that, with a working directory attached.&lt;/p&gt;

</description>
      <category>reactnative</category>
      <category>android</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
