You own a web app and two mobile apps. Someone above you says "unify the automation." There are two obvious answers and both are wrong: build one framework with a shared abstraction over web, iOS and Android, or let three engineers own three unrelated suites and call the union of them a strategy.
The shape of a working answer looks more like the Allego engagement, which we've been on since 2013: 2.4k+ web automation tests, 1.4k+ iOS tests, 1.1k+ Android tests, 5,000+ automated scenarios in total, run by 7 automation engineers under one test management system and 30+ CI jobs. Three separate suites. One strategy. Nobody wrote a page object that drives both a browser and a UIKit view.
This piece is about which decisions belong in the shared layer and which ones you have to make three times.
One strategy is not one suite
Split the thing you are unifying into a decision layer and an execution layer.
The decision layer is where "one strategy" is real, and it holds: what gets automated and at which level, what state a test is allowed to create and how it creates it, what "covered" means and who audits it, where results land, when suites run and what a red run blocks. None of that is platform-specific. All of it is expensive to decide twice, and catastrophic to decide three times with three different answers.
The execution layer is how a test drives the application: locators, waits, gestures, permission dialogs, app install and reset, device provisioning. This layer does not transfer, and the attempt to make it transfer is where most "unified framework" projects die.
The Allego numbers above describe an end state, not a design document — I can't tell you the team drew this line on a whiteboard in 2013. But the end state is the shape: one team, one ledger, three drivers. Worth noting what else that engagement ended up with: a second automation framework, deliberately introduced alongside the first, to automate cases previously considered unautomatable — it covered 70% of them. If unification meant "one framework," that second suite would be a failure. It isn't. It's the strategy working.
The layer that transfers: data, contracts, coverage rules
A login flow written once for web is worth nothing on iOS. A routine that provisions a tenant with three users, a seeded catalog and a paid subscription is worth exactly as much on all three, because it never touches a UI.
This is the part of a cross-platform strategy that pays back the most and gets discussed the least. Concretely, four things transfer:
State creation through the API, not the UI. Every test on every platform arrives at its first assertion through the same HTTP calls. The shape is boring on purpose:
// one fixture module, consumed by the web runner and both mobile runners
export async function seedActiveSubscriber(api: ApiClient) {
const org = await api.post('/orgs', { plan: 'pro', seats: 5 });
const user = await api.post(/orgs/${org.id}/users, { role: 'admin' });
await api.post(/orgs/${org.id}/billing, { status: 'active' });
return { org, user, credentials: user.credentials };
}
The web suite calls it before page.goto. The iOS suite calls it before the app launches with a deep link. The Android suite does the same. Nothing about it knows what a driver is.
One identifier scheme. If web calls it checkout_submit and Android calls it btnCheckoutSubmit, your coverage reports across platforms are two documents that cannot be compared. The naming rule is shared even though the locator strategies are not.
One definition of "covered." This one costs nothing to agree on and is nearly impossible to retrofit. Note how differently the number behaves across engagements even when the work is comparable: 95% at Allego, 90% of delivered features at CipherHealth, 85% of the application at Compass, ~40% automation coverage at Sway. Those are four different definitions, not four different levels of rigor. Inside one product you get to pick one.
An API suite as the platform-neutral safety net. The API layer is the only place where a test written once genuinely protects all three clients. Compass ended up with 500+ API auto-tests next to 1.6k+ automation scripts and 200+ mobile scenarios; CipherHealth has 250+ API tests supporting 1.4k+ web scenarios and 100+ mobile ones; Sway reached 100% API endpoint coverage while product-level automation coverage sat at about 40%.
To be straight about the evidence: none of those engagements documented the API suite as being built to serve the mobile suites. That's my argument, not theirs. What the numbers show is a consistent pattern — a substantial API layer sitting under a lopsided web/mobile split. The Sway ratio is the one I'd point at hardest: with one QA engineer, full endpoint coverage and 40% UI coverage is a rational allocation, not a shortfall.
The layer that doesn't: gestures, permissions, and device matrices
The divergence between web and mobile is not cosmetic, and a shared abstraction hides exactly the things that break in production.
There is no browser equivalent of: first-launch install state, OS permission prompts, backgrounding and process death, push notification delivery, biometric auth, offline and flaky-network state, app store update paths, or deep links that only resolve when the app is already installed. You cannot express "deny location permission, background the app for 40 seconds, resume, assert cached state" in an abstraction that also has to make sense to a Chrome driver. When a team tries, what usually survives is the lowest common denominator — click, type, assert text — which is precisely the subset of mobile behavior that was never the risk.
The matrices diverge too, and they scale differently. Browser matrices grow along one axis, roughly by version: QMSsupports 4 main browsers; CipherHealth went from 1 supported browser version to the 3 latest. Mobile matrices grow along two, device times OS version, and the OS axis is not symmetric between the platforms.
The current numbers make the asymmetry concrete. On iOS, Apple measured 79% of all iPhones running iOS 26 on June 7, 2026, and 86% of devices introduced in the last four years. On Android, cumulative Statcounter data from April 2026 puts Android 16 at 22.3% — you have to go back to Android 11 to cover 86.9% of devices. One iOS version is most of your users. One Android version is a fifth of them. Suites that support "the latest two versions" on both platforms are making a much weaker promise on Android than on iOS, and it is worth saying so out loud in the strategy doc rather than discovering it in a crash report.
That is why the OS-version count is a first-class number in these engagements: Allego moved from 1 version of iOS and Android used for autotests to 5 of each; CipherHealth supports the 2 latest versions of both; QMS supports 2 main mobile versions.
Real devices are the other non-negotiable. SimpliField listed "lack of testing on real devices" as an explicit pre-engagement problem and now runs functional testing across 9+ real iOS and Android devices; Grover keeps 5+ devices spanning OS versions, screen resolutions and browsers, alongside 4 localizations. Emulator-only coverage does not catch permission and hardware behavior, and no shared abstraction changes that.
Your three suites will not be the same size
The reflex after "one strategy" is coverage parity: equal automation on all three platforms because that feels fair. The observed ratios say otherwise, and lopsided is the norm rather than the exception.
Compass: 1.6k+ scripts against 200+ mobile scenarios. CipherHealth: 1.4k+ web scenarios against 100+ mobile ones. QMS: more than 300 of its 500+ automation scripts are web. Allego, the most balanced of the group, still runs 2.4k web to 1.4k iOS to 1.1k Android.
Now the counter-evidence, which is the more interesting half. On an end-to-end encrypted messenger we've worked on since 2020, the split is >700 iOS scripts and >700 Android scripts — near-parity, >1.4k automated tests in total, 98% of the 2.5k+ test cases automated. That is a product where mobile is the product; there is no web suite to be lopsided against. SimpliField runs the other way: a mobile-first product with 8,000+ documented test cases and 30+ automation scripts, of which 20 are iOS tests and 20 are Android — a deep manual and documentation layer with a deliberately thin automated one.
None of these sources explain the ratios, so treat them as distributions rather than prescriptions. The usable inference is narrower: parity is a question about where risk and revenue sit, and the answer is almost never "evenly." If your web suite is 10x your mobile suite, that is a finding to defend or fix, not automatically a defect.
Tooling: converge on a language, never on a driver
The practical rule I'd defend: converge on a language and a runner where the humans overlap; do not try to converge on the thing that talks to the app.
Allego runs Cucumber, Watir and WebdriverIO on web and Appium on mobile, across Ruby, JavaScript and Java, plus that second framework. QMS pairs Java/Selenide/Cucumber on web with Appium, Xcode and Android Studio on mobile. Sway pairs Playwright with Appium under a single CodeceptJS runner — the closest thing here to convergence, and note that it converges the runner, not the driver. Grover's unification went a different direction entirely: migrating the automation suite from Ruby to Java and rewriting the logic, which is a language consolidation inside one suite, not a merge across platforms.
Appium shows up as the mobile driver in nearly all of these, and it is worth being precise about why: it is named as the stable, maintainable mobile framework, and nothing in these engagements suggests it delivered code reuse with the web suites. It is a mobile driver, not a bridge.
The ecosystem gap is also just a fact you have to staff for. In the week of 23–29 August 2026, playwright pulled 87.5M npm downloads against 1.39M for appium. Mobile automation is a smaller world with fewer answered questions.
The cost of the split is real: more stacks means more expertise per hire, duplicated helpers, and two or three places to fix the same reporting bug. Small teams pay hardest — QMS runs this with 2 full-stack QA engineers, Grover with 1 full-stack plus 1 manual, Sway with 1. If you have one engineer, "one framework per platform group" may still be one framework more than you can maintain, and thinning the mobile suite deliberately (SimpliField's 30+ scripts) is a defensible answer.
Where the strategy actually runs: CI, reporting, and the test ledger
The cheapest things to unify are the pipeline and the ledger, and they are what make three suites feel like one program.
One test management system, one report format, one place where a coverage number is computed. Allego's engagement included consolidating onto a fully integrated test case management system alongside adding 5 parallel threads and 30+ automation jobs. CipherHealth runs 50+ automated jobs with up to 20 parallel threads, up from 1 thread, and went from 1 local environment to 4 including production. Compass runs CircleCI jobs that fire on pull request creation for specific microservices, schedule regression and smoke runs, and post job results into Slack; Sway runs about 5 GitHub Actions workflows.
Per-platform jobs, independent triggers, independent cadences. Web E2E on every PR is achievable — Compass runs it in 10 minutes per PR, down from 40+. A full Android matrix on every PR is not, and pretending otherwise is how mobile jobs get muted.
One caution on attribution: in every one of these engagements the framework, the thread count, the environment count and the test count all changed together. Compass went from a 7-day regression to 2 days; CipherHealth reports a 10x faster regression run reaching 5 hours with a 10-minute smoke suite. Do not credit any single change with those numbers, including the ones you're planning to make.
What the split costs
Three suites means triple maintenance, three chances for the coverage definition to drift, and cross-platform regressions that are slow because the slowest platform sets the pace. Allego's regression is still 18 hours — down from over 70, with smoke at 4 hours from over 20, and 3.5x faster execution overall — but 18 hours is not a pre-merge gate. The messenger went from ~20 days of regression to 2 days, which is a 10x improvement and still a multi-day mobile regression.
Shared coverage criteria and a shared API suite bound the drift. Nothing in this evidence removes the maintenance cost, and none of these sources reports maintenance effort at all, so anyone telling you a unified strategy lowers total maintenance is guessing.
What you can point at is defect yield: Allego finds 2.5x more critical and blocker defects than before; CipherHealth cut production bugs by 35% across 240+ releases; SimpliField shipped 24 app versions with zero critical or major bugs; QMS saw user complaints drop 90% after five months.
The checklist
If I were starting this on Monday, in this order:
- Decide the coverage definition once, write it down, and apply it identically to all three platforms — even where the resulting numbers embarrass one of them.
- Provision every test's state through one API layer, and build the API suite first; it is the only suite that protects all three clients.
- Keep drivers separate and stop looking for the abstraction. Converge language and runner where the same people work on both.
- Write the device and browser matrices down explicitly, with version counts, and revisit them when the OS distribution moves — a "latest two versions" policy means something very different on iOS than on Android.
- Unify CI reporting and the test ledger before you unify anything else. It is cheap and it is what makes three suites legible as one program.
- Expect suite sizes to differ by a factor of several, and require a risk argument for the ratio rather than a parity target.
The strongest parity in any of this evidence — 700 iOS scripts to 700 Android — came from a product where mobile is the entire product. Parity follows risk distribution. It is a result, not a goal.
Top comments (0)