DEV Community

Cover image for 72 commits, 42 decisions, 6 hidden bugs: building a React Native partner app with Claude as a team, not a tool
Onkar Deokate
Onkar Deokate

Posted on

72 commits, 42 decisions, 6 hidden bugs: building a React Native partner app with Claude as a team, not a tool

I'm building ParkEase, a peer-to-peer parking marketplace for India: drivers, space owners, valets, car-wash partners and admins, spread across a NestJS API, a worker, an Expo React Native app, an admin panel and a marketing site. Task 14 was the car-wash partner app. That's the screen a washer lives in all day: receive offers, take before and after photos, edit a price menu, check earnings.

I built it with Claude Code, and "built with AI" undersells how it actually worked. The task ended as 72 commits, about 22,800 added lines across 185 files, 1,013 mobile tests, 397 integration tests against real Postgres, 42 recorded decisions, and six defects that had been sitting in already-merged code. It also took far longer than it should have, and some of the mistakes were the AI's own. Here's all of it.

The setup: /flow, gates, and a paper trail

The core of my workflow is a skill I call /flow. It isn't a prompt. It's a routing layer that knows my project: the architecture decision records (ADRs), the coding rules, the task files, and which of about 80 skills and 25 specialist subagents apply at each phase. Every phase has a gate:

Phase Gate
Think brainstorming: no code until I approve a design
Direction (UI only) 2–3 rendered visual directions; I pick
Plan writing-plans: bite-sized TDD tasks
Build test-driven-development: see it fail first
Review a stack of lenses, never one reviewer
Design audit three design lenses, because a code review is not a design review
Verify evidence before any "done"
Ship an 18-line pre-PR gate, pasted into the PR

Two rules in /flow did the most work on this task.

"The governance chain is read from the docs, never inferred." My ADRs override my rules, which override the PRD. When a skill's advice contradicts an ADR, the ADR wins, and Claude has to say so out loud. It came up in the first ten minutes. The UI/UX skill's generator recommended orange with a gaming display font for a car-wash app. The project's palette ADR says cobalt and availability-green, and "brand orange" had already been rejected once, so Claude rejected the suggestion and cited the ADR.

"An agent never satisfies a gate." Subagents do bounded work inside a gate. The gate takes the approval, and the gate checks the result. It's the difference between five green agent reports and a feature that works.

What went well

1. Choosing a direction from pictures, not paragraphs

Before a single screen existed, Claude rendered three directions as tappable mockups on a design canvas. Worklist was dense and literal. Bay made money the headline and turned the photo pair into one object. Docket used a work-ticket metaphor. I picked Bay.

The reasoning behind Bay was genuinely good product thinking. The task file drew the before/after photo gate as an alert card that appears when a transition is already blocked. Bay makes the pair a persistent two-slot frame from the moment a job is accepted, so the gate becomes a standing obligation instead of an interruption. I wouldn't have asked for that. A good designer brings the idea you didn't ask for.

2. Review as a stack found bugs that were already in main

Every task got a spec-plus-quality review and a fix loop. Then the whole branch went through seven specialist lenses, run three at a time: security, silent failures, database, TypeScript, React, test adequacy and spec conformance. Here are six things they, and the live walkthrough, surfaced in code that was already merged:

  • The valet proof-photo upload had never worked. The mobile client posted multipart/form-data; the API parsed { proofPhotoId: string }. Every call returned a 400. The component test injected the broken dependency, so the test couldn't see it.
  • The driver's wash screen failed for every real partner. It read users.name into a non-null schema, and nothing in the codebase ever writes users.name.
  • A business partner could never be verified. Only the ID-document endpoint set pending, and the business form doesn't collect an ID.
  • Valet and washer partners opened the app to "Unmatched Route". The redirect built /(role), and those two route groups had no index.tsx.
  • The idempotency layer's store/release calls were fire-and-forget. A failed write left a key locked for 24 hours.
  • An unregistered partner got a 404 with the error code ERROR, not NOT_FOUND. The client couldn't tell "you haven't signed up yet" apart from "the server broke", so it had no way to route to registration.

None of these was a style nit. Each one would have broken a real user.

3. A ledger of every decision taken on my behalf

The execution skill has a rule I've come to love: rulings, not stalls. When a plan is ambiguous or wrong, the agent decides, and records it:

Ruling: T6-R1 — a pending washer's offers state is "online switch locked with
the reason", not "offers visible, Accept locked". The server makes the task
file's version unreachable (set-availability.command.ts:43 refuses isOnline
unless verified). — Cost if wrong: a pending-preview feed would be a server
change in task 13's dispatch.
Enter fullscreen mode Exit fullscreen mode

By the end there were 42 of these, and every one went into the PR body as a table: decision, reason, cost if wrong. That's the part I actually review. It turns "the AI did stuff" into decisions I can overrule.

4. It opened the app, and that found what tests can't

The UI gate says a screen isn't done until it has been opened at mobile and desktop widths. Claude drove the web preview and measured things:

  • tab labels were 12px text in a 9px box, so they were clipped;
  • at a short viewport, the empty-state illustration covered the online toggle and pushed the only button under the tab bar;
  • the switch's "on" colour was react-native-web's default #009688, not a design token;
  • green, which the design system reserves for "available right now", was being used for "done".

A mechanical design detector, run as a separate isolated pass, independently flagged the same occlusion. The design review scored the screens 23/40. Every issue was fixed and re-checked live afterwards.

What went badly

1. The AI's own plan had bugs, and reviews caught them late

I want to be honest about this, because the "AI writes perfect code" crowd isn't.

  • The plan's upload step sent no idempotency key, so every upload would have failed as "check your connection".
  • The fix for that was then wrong in the other direction. Reusing the key across retries meant the server replayed a stale Cloudinary signature for 24 hours, and Cloudinary rejects signatures older than an hour. That became an ADR: an idempotency key belongs to a side effect, not to an HTTP call.
  • The earnings query's correlated subquery rendered as:
  where "txn_id" = "txn_id"   -- always true
Enter fullscreen mode Exit fullscreen mode

Every earnings line summed every posting in the ledger. It passed two reviews, because every test fixture had exactly one job. A later reviewer caught it with a two-job test.

  • A plan test asserted rupeesToPaise('0.07') → 7, which contradicts its own ₹10 minimum. The justification for it ("319.20 * 100 is a float error") was false; in Node it's exactly 31920. The implementer stopped, flagged the contradiction, and didn't bend the test. That's exactly what you want.

The lesson isn't "AI is sloppy". It's that a plan is a hypothesis, and the review stack is what tests it.

2. Fixes introduced regressions

Two fixes from the final fix wave broke something new, and both were caught by the scoped re-review:

  • memoising the offer card made a recycled FlashList cell carry an "expired" flag onto a live offer;
  • a stale-lock takeover meant a committed-but-unstored request could run twice (think "second payment order").

Both were fixed before merge. It's also a reminder that "the review passed" is only as good as the review of the fix.

3. Wall-clock time, mostly rate limits

The work itself wasn't the slow part. Agents were cut off by session and weekly rate limits more than five times, mid-task. The Sonnet weekly quota, which the implementer and reviewer subagents were running on, ran out on day one. From then on, subagents ran on Opus for anything needing judgment and Haiku for pure transcription. The design of the process saved it: every agent writes its report to disk, every task has a ledger line, and an interrupted agent is resumed with its context intact instead of restarted. Nothing was redone, but the calendar time still stretched across two days.

4. The environment lied, and it nearly cost a false bug hunt

At one point the camera sheet's shutter measured at y=1564 in an 812px viewport, which looked like a serious layout bug. It wasn't. The browser pane was hidden, so requestAnimationFrame was firing zero times a second, and the slide-in animation never ran. Claude measured frames per second before diagnosing, and that check is now a documented learning.

Similarly, the web preview wouldn't start because Metro was crawling 1.6 GB of stale git worktrees. The fix was a blockList in metro.config.js, not deleting my files, which Claude refused to do without asking. When I did ask, it first showed me which worktrees held uncommitted work.

How Opus 5.5 helped

The main session, the "controller" that plans, dispatches and adjudicates, started on Opus 5 and moved to Opus 5.5 early in the task. Everything from the implementation onward ran on it: orchestration, the final seven-lens review, and every ruling. What stood out wasn't raw coding speed. It was judgment in the controller seat:

  • It said when its own plan was wrong. "The defect is in MY plan's Step 4 code, not the implementer's transcription" is a sentence I saw more than once, followed by a ruling and a fix.
  • It verified before it asserted. Before acting on a reviewer's "Critical", it re-read the code. It confirmed set-availability.command.ts:43 before overruling a line in the task file, and it confirmed errorCodeFor returns 'ERROR' for a bare NotFoundException before changing the server.
  • It held the line on destructive actions. It deleted no worktree with uncommitted work until I chose. It didn't edit my CLAUDE.md because a subagent suggested it. It didn't open the PR until the gate was green.
  • It pushed back on scope in both directions. It proposed extending the earnings API when the task was unbuildable without it, and it refused the sparkline because it would have required money arithmetic in the client.

How AI actually helps (and where it doesn't)

It helps most as a team. Seven reviewers reading the same branch through seven different lenses is something I couldn't staff myself. The database lens found a missing index. The React lens found the recycled-cell bug. The test-adequacy lens found that no test anywhere proved an unverified partner gets refused, the single highest-risk gap in the branch.

It helps with the tedious truth-telling. Every PR claim comes with pasted output: lint 16/16, typecheck 15/15, 1,013 tests, 397 integration tests, 72 commits scanned by gitleaks with zero leaks. Where the gate wasn't green (pnpm audit has 33 pre-existing findings), the PR says so.

It does not replace a device. The real camera, the real Cloudinary upload, TalkBack on Android and the Maestro end-to-end flows are all explicitly listed in the PR as not verified here. An honest "I couldn't check this" beats a confident green checkmark.

The token-optimisation playbook

Long agentic sessions die from context bloat before they die from bad code. Here's what we did, all of it visible in the session:

  1. Hand artifacts over as files, not pastes. Each implementer gets a task brief extracted from the plan (never the whole plan). Each reviewer gets a precomputed diff package: commit list, stat and full diff in one file. The controller never pastes a diff into a prompt.
  2. Short return contracts. Every subagent writes its full report to disk and returns under 15 lines: status, commits, one-line test summary, concerns. The detail lives in the file, and the controller reads it only when needed.
  3. Tiered models. Haiku for transcription tasks and small scoped re-reviews, and for running the pre-PR gate commands. Opus for money, security, concurrency and design judgment. The rule of thumb: turn count beats token price, so the cheapest model isn't cheap if it takes three times the turns.
  4. Scoped re-reviews. After a fix, the reviewer sees only the fix diff and the findings list, and gives a verdict of ADDRESSED or NOT ADDRESSED per finding. No re-reviewing the world.
  5. Slice big diffs by surface. The final branch diff was 700 KB. It was split into server source, server tests, mobile source and mobile tests, and each lens read only its slice.
  6. Cap the output. "At most 12 findings, severity-ranked, under 600 words, no praise section." Praise is tokens.
  7. Batch lenses three at a time, so a burst of parallel work doesn't blow the rate limit, the one resource we kept running out of.
  8. Resume, don't restart. A rate-limited agent is resumed with its context rather than re-dispatched from scratch. The largest implementer run used around 520k tokens; redoing that would have been the most expensive mistake available.
  9. A ledger as recovery memory. progress.md records every task's status and every ruling. After an interruption, the ledger and git log beat the controller's recollection, and no completed task was ever re-dispatched.
  10. Measure before you investigate. One requestAnimationFrame count replaced what could have been an hour of chasing a phantom layout bug.

Our /flow also has a local knowledge graph of the codebase (tree-sitter, no embeddings) for "what calls what" questions, which saves ten greps with one query. In this task I'll admit we still grepped more than we queried. That's the next habit to build.

The takeaway

The most valuable output of this task wasn't the 22,800 lines. It was six bugs that were already shipped, a 42-row table of decisions I can overrule, and a PR that tells me exactly what it did not verify.

AI didn't make this fast. Rate limits and my own process made sure of that. What it did was make the work inspectable: every claim has evidence, every decision has a cost-if-wrong, and every gap has a row in the backlog. For a solo developer building something that moves real money, I'll take inspectable over fast.

ParkEase is built with Claude Code, Expo SDK 57, NestJS 11 on Fastify, Drizzle, and PostgreSQL 18 + PostGIS. The task 14 PR: https://github.com/Deonkar/parkease/pull/14

Top comments (0)