DEV Community

Cover image for Shopify Left React Native: 4 Questions Before You Follow
Kiell Tampubolon
Kiell Tampubolon

Posted on

Shopify Left React Native: 4 Questions Before You Follow

Yesterday, Sep 10 2026, Shopify announced it is moving every mobile app back to native Swift and Kotlin, away from React Native. When a team that committed that hard reverses that fast, the interesting question is not "is React Native dead". It is: which assumption expired?

I spent this morning reading the decision post, the migration post, and the HN thread (it sat at #2 on the front page). Below: why the one-shot AI port failed, the agent-testing bottleneck most coverage skipped, and a 4-question protocol you can run before you follow anyone.

This is the same lesson I keep hitting this month, outside security: assumptions that silently stopped holding. Read-only database modes that were not read-only. Sandboxes that were not sandboxes.

If a stack decision is an assumption with a date attached, here is the artifact I sketched while reading:

# assumption register: my format, not Shopify's
decision: "Shop app runs on React Native"
made: 2020
founding_assumption: >
  One JavaScript codebase reaches native quality
  faster than two native codebases.
what_would_invalidate_it:
  - "agents make native rebuilds cheap"
  - "parity tax outgrows the savings"
  - "cross-stack hiring stops being the constraint"
status: "expired 2026-09-10, per Shopify"
Enter fullscreen mode Exit fullscreen mode

Why did Shopify pay to rebuild apps that already worked?

In 2020 they went all-in on React Native for three reasons: stop building features twice, let developers cross stacks, chase less parity. What changed: adopting the New Architecture would have forced a revisit of native modules, rendering, and shared/platform boundaries anyway. First they tested whether coding agents could build SwiftUI and Compose directly. One engineer with agents ported most of the RN Shop app to SwiftUI in a week. Not production-ready, but proof that feature-for-feature migration was achievable.

A core team of six built the foundations and main journeys, feature teams joined midway, and it reached stores 12 weeks after the PoC with users signed in, push working, and analytics events intact downstream.

All numbers below are Shopify's self-reported measurements; I have not verified them independently.

Metric React Native Native
iOS cold start to home feed 3200 ms 2466 ms (23% faster)
Android cold start 4433 ms 2233 ms (50% faster)
iOS app size 67 MB 68 MB (+1.5%)
Android app size 293 MB 184 MB (37% smaller)
Crash-free sessions 99.5%+ 99.95%+
Android release build baseline ~75% faster

What did coding agents actually change?

Three costs got cheap:

First, cross-platform translation: the week-long SwiftUI PoC shows one engineer plus agents can carry behavior across a language boundary.

Second, ramp-up outside your primary stack. The post's author (mustafa01ali) answered in the thread: "we invested in ramping up teams on native before going all in". Agents cut that ramp cost.

Third, parity maintenance: Helix checkpoints make "does the native build match" checkable. Their Tardis tool compares event names, counts, and payload fields across the RN and native apps, tolerating differing timestamps and UUIDs.

The honest limit, from Shopify itself: "Native still means building and maintaining software on two platforms, that cost has not disappeared". And: "React Native apps can be fast. Ours are." Native did not get easy. Specific line items got cheaper.

Why did the obvious AI approach fail?

The obvious approach is pointing an LLM at the RN codebase for a native port. Shopify tried it. Verdict: "a huge amount of unmaintainable code that can't be shipped". And Shopify's broader caution on generated code: "Generated code could satisfy feature requirements while still introducing duplication, architectural drift, or performance problems."

What shipped instead was process. Their Helix system migrates screen by screen through ordered checkpoints: each must prove behavior with tests, match the running app in visual review, survive two adversarial AI reviewers, and pass human approval before commit. Feedback is remembered; the loop gets more autonomous. A Pi extension (Pi is their coding agent) orchestrates specialized subagents: inspect RN source, document behavior, prepare platform plans, implement, review parity.

My favorite detail: "Plan acceptance was tied to a hash of its contents: changing a plan invalidated its previous approval." Approval attaches to content, not effort; editing the plan re-triggers review.

What a Helix checkpoint enforces

The posts describe behavior, not schema. My illustrative reconstruction of one checkpoint spec (not Shopify's format):

# illustrative reconstruction, not Shopify's format
checkpoint: 14
screen: "checkout, payment selection"
must_pass:
  tests: "unit + integration green"
  visual_review: "matches running RN app, same state"
adversarial_reviewers: 2
human_gate: "approve before commit"
Enter fullscreen mode Exit fullscreen mode

The fields are guesses; the property is not. "Done" was machine-checkable before a human saw the diff.

What was the real bottleneck?

The least-covered detail I found. Shopify reports agents "can make code changes in seconds, but it takes them several minutes to test the output" through accessibility-tree and screenshot control of simulators. Their conclusion: "It doesn't matter how good the model is if it can't test its work quickly."

The fix was architectural, not a better model. Business logic is fully decoupled from UI and runs headlessly on desktop; agents drive it through a CLI in milliseconds (inspect state, navigate, perform actions). The simulator is reached only via remote-mode commands, no layout or accessibility parsing. Reported payoff: "agents that work autonomously for hours."

What a headless check looks like

Shape, not their tooling:

# illustrative shape of a headless parity check, not Shopify's CLI
agent> navigate: cart, add item, open checkout
agent> inspect state      -> 3 line items, subtotal 41.97
agent> compare events     -> cart.updated x2, checkout.started x1
# the simulator only wakes up when the question is about pixels
Enter fullscreen mode Exit fullscreen mode

If your agents test in minutes, this is the loop to fix before you touch the framework.

How do you re-run a stack decision on your own team?

Four questions, in order, filling in the register above:

  1. Name the assumption your stack decision rests on, in one sentence. Two sentences means two decisions; split them.

  2. Ask what would prove it wrong, and whether that thing has now happened. Shopify's founding assumptions (stop building features twice, cross-stack developers, less parity chasing) each got a concrete 2026 test.

  3. Price the tax the abstraction charges you today, with a number if you can: startup milliseconds, app size, build minutes, review times. Self-reported or not, numbers beat vibes.

  4. Set a review date instead of a verdict. Shopify called RN's future bright in January 2025 and reversed 20 months later; a dated review, reopened when the New Architecture question arrived, would have surfaced this sooner.

Is the HN pushback right?

hermitwriter argues parity is a divergence cost, not an implementation cost: it accrues for years across experiments, analytics, accessibility, edge cases. "You've built prototypes. You're making a claim about a cost that compounds over time based on what it costs at t=0." Counterfactual: if agents make dev cheaper, they make RN dev cheaper too; point the same agents at the RN codebase. The real driver may be the upstream tax: framework upgrades and dependency churn, which predates LLMs.

doc_ick says the main cost of two teams is organizational, keeping two products in sync. SV_BubbleTime: headcount. Shopify has hundreds of engineers; your team of five cannot run this play. joshstrange's App Store reviews took 40, 67, then 98 hours (others report 9-day waits), so over-the-air web updates keep an edge. simonhamp defends abstractions: "It's not the layers that are the problem; they're the point. It's the quality of those layers."

The strongest objection is the simplest: the decision post contains no cost data, so we get outcome metrics, not decision math. My position, held loosely: defensible for Shopify, given a motivated team, agent infrastructure, and native supervision capacity. What persuades me is not "native is back" but that the cost model's inputs changed enough to re-run the decision.

So, for the comments: is "build once" still a moat, or did it become a constraint? Would you rewrite a working app? Is this an AI story or an upstream-tax story in an AI costume?

What happens to React Native now?

Shopify maintained react-native-skia, flash-list, and restyle. Per Simon Willison's writeup, the first two are finding new homes; restyle archives at the end of 2026. If you depend on them, watch who adopts them.

Watch Kotlin Multiplatform too; the thread debates it against Rust-UniFFI, mike_hearn calling KMP "a token optimization". Watch item, not a verdict.

Do not panic-migrate. A working RN app is still a working app. Their conditions were specific: hundreds of engineers, agent infrastructure like Helix and Tardis, and essential native expertise. Without those, the read is "they re-ran their decision", not "you should too".

This theme, assumptions that quietly expire, runs through my last three posts: read-only Postgres MCP servers failing with one SQL keyword, two CVSS 9.8 agent sandbox CVEs landing the same day, and the MCP spec going stateless, turning a planted prompt into a credential.

Takeaways

  • A stack decision is an assumption with a date. Write the date down.
  • One-shot AI ports failed here; ordered, human-gated steps shipped.
  • A working app does not owe anyone a rewrite.

Top comments (0)