Most React apps don't fail because of React. They fail because of decisions made in week one that nobody revisited: a folder structure that made sense for 5 components, a state approach that made sense for 2 screens, a "we'll add tests later" that never arrived.
This is the playbook I'd follow as a tech lead starting a greenfield app. The principle is start minimal, but make the cheap decisions correctly now, because they're expensive to change later.
1. Decide What You're Optimizing For (Before Touching a Terminal)
Write down three things in a one-page ARCHITECTURE.md:
- Who will work on this? (2 devs vs. 4 teams changes everything)
- What's the expected lifespan and growth? (internal tool vs. customer-facing product)
- What are the non-negotiables? (accessibility, performance budget, security posture, browser support)
This isn't bureaucracy. It's the document you point to in code review when someone proposes adding a fourth state library.
Tip: Keep a lightweight ADR (Architecture Decision Record) folder. One short markdown file per decision: context, decision, consequences. In six months, "why did we choose X?" has an answer.
2. Tooling: Pick Boring, Modern Defaults
| Concern | Recommendation | Why |
|---|---|---|
| Build tool | Vite (or Next.js if you need SSR/SEO) | Fast, maintained; Create React App is deprecated |
| Language | TypeScript, strict mode | Catches an entire class of bugs; the best documentation you'll have |
| Package manager |
pnpm (or npm), with a committed lockfile and pinned Node version (.nvmrc / engines) |
Reproducible installs |
| Routing | React Router or TanStack Router | Don't hand-roll it |
| Server state | TanStack Query | Caching, retries, dedupe, loading/error states |
| Client state |
Start with useState/useReducer + Context, add Zustand only when needed |
Resist premature global state |
| Forms | React Hook Form + Zod | Performance + shared validation schemas |
| Testing | Vitest + React Testing Library + MSW + Playwright | Fast unit/component tests, realistic network mocking, a few e2e smoke flows |
Key insight: most "state management" pain is actually server state pain. Using TanStack Query removes a huge amount of Redux-style boilerplate and stale-data bugs before you've written any of it.
Turn on "strict": true in tsconfig.json on day one. Retrofitting strictness later is miserable; starting with it is free.
3. Folder Structure: Organize by Feature, Not by File Type
The classic components/, hooks/, utils/, services/ layout collapses at around 50 components, because everything is "shared" and nothing has an owner.
Prefer feature-based structure:
src/
app/ # app shell: providers, router, global config
features/
auth/
api/
components/
hooks/
types.ts
index.ts # public API of the feature
orders/
...
shared/
ui/ # design-system-level components (Button, Modal)
lib/ # generic helpers (date, formatting)
config/
tests/ # test utilities, setup, mocks
Rules that keep it healthy:
- A feature exposes a public API through
index.ts. Other features import from there, never from its internals. - Features may depend on
shared/, but shared never depends on features. - Features shouldn't import from each other directly. If they must, that's a signal to extract something to
shared/or rethink boundaries.
Enforce these with ESLint (eslint-plugin-boundaries or no-restricted-imports) rather than good intentions. Rules that aren't automated erode.
4. Get the Component Fundamentals Right
- Keep components small and single-purpose. If you need "and" to describe it, split it.
-
Separate logic from presentation. Custom hooks for logic (
useOrders), components for rendering. - Colocate styles, tests, and types with the component.
-
Compose, don't configure. Prefer
childrenand composition over a component with 14 boolean props. -
Derive, don't sync. If a value can be computed from props/state, compute it during render instead of storing it in another
useStateplus auseEffect. Many "weird bug" tickets come from this pattern. -
Treat
useEffectas an escape hatch, not a default. It's for synchronizing with external systems, not for transforming data or reacting to user events. -
Don't memoize everything. Use
useMemo/useCallback/React.memowhere profiling shows a problem. (If you're on the React Compiler, even less manual memoization is needed.)
5. Design for Failure From Day One
Apps rot fastest at their edges: network failures, empty data, unexpected shapes.
- Error boundaries at the app level and around major feature areas, so one failure doesn't white-screen everything.
- Every async view has four states: loading, error, empty, success. Make this a code-review checklist item.
-
Validate data at the boundary. TypeScript types disappear at runtime. Parse API responses with Zod (or similar) so a backend change fails loudly and early, not as a mysterious
undefined is not a functionthree screens deep. - Centralize your API client. One place for base URL, auth headers, error normalization, and retry policy.
6. Make It Debuggable (Future You Will Be Grateful)
Debuggability is a feature you build, not something you discover you lack at 2 a.m.
- Source maps uploaded to an error-tracking tool (Sentry, Datadog RUM, etc.), with release versions tagged.
- Structured logging behind a thin logger wrapper, so you can change the destination and filter noise without touching call sites.
- Correlation IDs: send a request ID header from the client so frontend errors can be matched to backend logs.
- React DevTools + TanStack Query DevTools wired up in development.
- Predictable data flow. One-directional, with minimal global mutable state. Bugs are far easier to reproduce when state transitions are explicit.
- Feature flags for risky changes, so rollback is a toggle and not a redeploy.
- Web Vitals reporting (LCP, INP, CLS) from real users, so performance regressions show up in dashboards, not in complaints.
7. Quality Gates: Automate Everything That Can Be Automated
If a standard depends on a human remembering it, it will fail eventually. Encode it in tooling.
Local (pre-commit):
-
ESLint (flat config) with
typescript-eslint,eslint-plugin-react-hooks,jsx-a11y - Prettier for formatting (no more style debates)
- Husky + lint-staged to run these on staged files only
- Conventional Commits (optional, but enables automated changelogs)
CI (every PR, blocking):
- Install with the frozen lockfile
- Type-check (
tsc --noEmit) - Lint
- Unit/component tests with a coverage threshold you can actually sustain
- Build
- CodeQL scan
- Dependency audit
Pragmatic testing strategy: a testing trophy, not a pyramid. Mostly integration-style component tests that exercise behavior the way a user would (React Testing Library), a thin layer of unit tests for pure logic, and a handful of Playwright tests for critical journeys (login, checkout, the thing that makes money). Test behavior, not implementation details, and your tests will survive refactors.
8. Build Test-First Where It Pays Off (TDD, Pragmatically)
TDD isn't about hitting a coverage number. It's a design tool: writing the test first forces you to define the behavior and the public interface before you write the implementation. The result is smaller units, fewer hidden dependencies, and code that's easy to change because it was built to be tested.
The loop is red, green, refactor:
- Red: write a failing test that describes one behavior.
- Green: write the simplest code that makes it pass.
- Refactor: clean up with the safety net in place.
Where TDD pays off the most in a React app:
- Business logic and pure functions (pricing rules, validators, reducers, data mappers). These are the cheapest and highest-value tests.
-
Custom hooks, using
renderHook. Test the contract (inputs, outputs, state transitions), not the internals. - Bug fixes. Reproduce the bug as a failing test first, then fix it. The bug can never silently return.
- Component behavior, written from the user's point of view: "when I click Submit with an empty email, I see an error and no request is sent."
Where to be flexible: exploratory UI work and visual layout. Spiking a design test-first is slow, so sketch it, then lock in the behavior with tests once the shape is clear.
Habits that make TDD stick:
- Test through the UI the way users do:
getByRole,getByLabelText,userEvent. AvoidgetByTestIdas a first resort, because role queries double as accessibility checks. - Mock at the network boundary with MSW (Mock Service Worker), not by mocking your own hooks or modules. Your tests then exercise the real data flow, and the same mocks work in dev, unit tests, and Playwright.
- Keep tests fast (Vitest in watch mode gives sub-second feedback). Slow tests don't get run, so TDD dies.
- One behavior per test, with names that read like specs:
it("shows an error when the email is invalid"). - Don't test implementation details (internal state, private functions, exact render counts). If a refactor breaks a test but not the behavior, the test was wrong.
- Gate PRs on "new behavior ships with tests" in the definition of done, instead of chasing a vanity coverage percentage.
9. Adhering to CodeQL and Security Standards
CodeQL performs semantic static analysis on your code. For JavaScript/TypeScript it commonly flags issues like DOM-based XSS, unsafe URL handling, prototype pollution, ReDoS-prone regexes, and hard-coded credentials.
Set it up:
- Enable GitHub's code scanning with CodeQL (default setup is a few clicks; use advanced setup via workflow if you need custom queries or build steps).
- Run on every PR and on a weekly schedule (new queries ship over time).
- Use the
security-extendedquery suite if your risk profile justifies the extra noise. - Make high/critical findings block merges. Treat triage as part of the PR, not a quarterly cleanup.
Write code that doesn't trigger the alerts in the first place:
| Risk | Habit |
|---|---|
| XSS | Avoid dangerouslySetInnerHTML. If unavoidable, sanitize with DOMPurify and isolate it in one reviewed component |
Unsafe redirects / javascript: URLs |
Validate and allowlist URLs before putting user input in href, window.location, or src
|
| Secrets in the bundle | Anything prefixed VITE_/NEXT_PUBLIC_ ships to the browser. Never put secrets there |
| Token storage | Prefer HttpOnly, Secure, SameSite cookies over localStorage for auth tokens |
| Dependency risk | Enable Dependabot/Renovate, run npm audit in CI, review new dependencies before adding |
eval / new Function / unvalidated postMessage
|
Don't. Validate event.origin if you must use postMessage
|
| Regex from user input | Avoid, or escape it; review patterns prone to catastrophic backtracking |
| Missing headers | Set a Content-Security-Policy, X-Content-Type-Options, Referrer-Policy at the edge/server |
Pair CodeQL with secret scanning and push protection. They catch different problems.
One honest caveat: CodeQL is excellent at data-flow bugs but isn't a substitute for threat modeling or review. Treat it as a safety net, not a security strategy.
10. Performance Is a Budget, Not a Cleanup Task
-
Route-level code splitting (
React.lazy+Suspense) from the start. It's nearly free early and painful later. - Set a bundle size budget and fail CI when it's exceeded (
size-limitorbundlewatch). - Virtualize long lists (TanStack Virtual) rather than rendering thousands of DOM nodes.
- Optimize images, use modern formats, lazy-load below the fold.
- Be deliberate about dependencies: one
import momentcan quietly cost you 70 KB. Check bundle impact before adding a library.
11. Accessibility and Design System: The Cheapest Time Is Now
- Use semantic HTML first; ARIA is a repair tool, not a starting point.
- Run
eslint-plugin-jsx-a11yand add axe checks to component/e2e tests. - Build a small set of primitives in
shared/ui(Button, Input, Modal, etc.), ideally on accessible headless foundations like Radix UI or React Aria. Features compose these instead of reinventing them, which keeps UI consistent and refactors cheap. - Centralize design tokens (colors, spacing, typography) as CSS variables or Tailwind config.
12. Configuration, Environments, and Delivery
-
Typed, validated env config. Parse
import.meta.envthrough a Zod schema at startup so a missing variable fails at boot, not in production. - Same build artifact, different config across environments wherever possible.
- CI/CD from week one: preview deployments per PR make reviews dramatically better.
- Document the "first hour" experience: a README where a new dev can clone, install, run, test, and deploy. Test it by having a new joiner follow it literally.
13. Make the Codebase AI-Ready
Your team, and increasingly your tools, will use AI assistants to write, review, and refactor code. A codebase that is easy for a new engineer to understand is also easy for an AI to work in. And a codebase with weak guardrails will let AI-generated mistakes slip through at speed. "AI-ready" means clear context in, strong verification out.
1. Give AI the context it can't guess
- Add a root-level instructions file (
CLAUDE.md,AGENTS.md, or equivalent) covering: stack and versions, folder structure, naming conventions, how to run, test, and lint, and the don'ts ("never useany", "no direct imports across features", "don't add dependencies without an ADR"). - Your ADRs and
ARCHITECTURE.mddo double duty here. They explain the why, which is exactly what an AI (or a new hire) can't infer from code. - Keep these files short, current, and tested. Stale instructions are worse than none.
2. Make the code legible
- Strict TypeScript is the best AI guardrail you have. Types give the model, and the compiler, a precise contract.
- Consistent patterns beat clever ones. If every feature follows the same structure, AI output will match it. If every feature is different, AI output will be a fourth variation.
- Descriptive names, small files, and small functions. Models and humans both reason better over focused units.
- Colocate tests, types, and components so the relevant context is in one place.
3. Make verification automatic (the most important part)
AI-generated code looks plausible even when it's wrong, so your pipeline is the real reviewer:
- Type-check, lint, tests, build, and CodeQL run on every PR, with no exceptions for AI-authored code.
-
Fast local feedback (a single
pnpm checkrunning typecheck + lint + tests) so an AI agent can verify its own work in a loop. - Tests as executable specs. TDD and AI pair naturally: write the failing test (or have AI draft it, and review it carefully), then let the AI make it pass. The test defines "correct" independently of the code that was generated.
- Enforce import boundaries and architecture rules with lint, not with documentation alone.
4. Set team policies, not just tools
- Humans own the output. An AI-generated PR is reviewed to the same standard as any other, and the author must be able to explain every line.
- Be careful with dependencies. AI tools can suggest packages that are outdated, unmaintained, or in rare cases non-existent. Verify before installing.
- Protect secrets and proprietary code. Define what may be shared with external AI tools, and keep secrets out of the repo, since secret scanning and push protection matter even more here.
-
Review for security explicitly. AI can reproduce insecure patterns it has seen (unsanitized HTML, token storage in
localStorage). CodeQL and your lint rules are the backstop. - Keep PRs small. AI makes it easy to generate 2,000-line diffs that nobody can review.
5. Consider AI in the product, too (if relevant)
If the app itself will call LLMs, plan for it early: a server-side proxy so API keys never reach the browser, streaming-friendly UI patterns, loading/error/retry states for non-deterministic responses, input/output validation with Zod, rate limiting, and logging for debugging prompts and failures.
14. Scale the Team Before You Scale the Architecture
Resist the urge to reach for micro-frontends, monorepo tooling, or a custom state framework on day one. A well-structured modular monolith carries most teams surprisingly far.
Add complexity only when you feel a specific pain:
| Signal | Possible response |
|---|---|
| Builds/CI getting slow | Monorepo with caching (Turborepo/Nx) |
| Multiple apps sharing UI/logic | Extract packages in a monorepo |
| Independent teams blocked by shared deploys | Consider micro-frontends |
| Heavy SEO/first-load needs | SSR/SSG (Next.js, Remix/React Router framework mode) |
| Complex cross-feature client state | Introduce Zustand/Redux Toolkit |
The pattern: pain first, tool second. Each addition should have an ADR explaining the problem it solves.
15. Habits That Prevent the Six-Month Rot
Technology choices matter less than team habits:
- Definition of done includes tests, types, accessibility, and error states.
- Bug fixes start with a failing test.
- Small PRs. Large ones get skimmed, not reviewed.
- Boy-scout rule, with scheduled refactoring time (e.g., 10-15% of each sprint).
- Monthly dependency upgrades, small and boring, instead of an annual migration crisis.
- Delete dead code and stale flags aggressively.
- Track tech debt visibly, as tickets with owners, not as hallway complaints.
- Write down the "why": ADRs and good PR descriptions.
- Review AI-assisted code like any other code, and keep the instructions file up to date as conventions change.
The Day-One Checklist
- [ ]
ARCHITECTURE.md+ ADR folder - [ ] Vite/Next + TypeScript
strict - [ ] Feature-based structure with enforced import boundaries
- [ ] TanStack Query + centralized API client + Zod validation
- [ ] Error boundaries + loading/error/empty states
- [ ] ESLint, Prettier, Husky, lint-staged
- [ ] Vitest + React Testing Library + a few Playwright smoke tests
- [ ] TDD workflow agreed: tests-first for logic, hooks, and bug fixes
- [ ] MSW set up for network mocking across dev and tests
- [ ] A single
checkscript (typecheck + lint + test) that runs locally and in CI - [ ] CI: typecheck, lint, test, build, CodeQL, audit
- [ ] Dependabot/Renovate + secret scanning
- [ ] Error tracking with source maps + Web Vitals
- [ ] Bundle-size budget + route-level code splitting
- [ ] Typed env config + preview deploys
- [ ]
CLAUDE.md/AGENTS.mddocumenting conventions, commands, and don'ts - [ ] AI usage policy: review standards, dependency vetting, what data can be shared
- [ ] README that a new dev can follow in an hour
Closing Thought
An app that survives six months isn't the one with the cleverest architecture. It's the one where the right things are automated, the boundaries are enforced, and the decisions are written down.
The same things that make a codebase easy for a new teammate to join, such as clear structure, written decisions, strict types, and fast automated checks, are what make it safe to hand to an AI. Invest in the guardrails once, and both humans and machines benefit.
Start small, and make it hard to do the wrong thing.
Top comments (0)