DEV Community

Reid Ji
Reid Ji

Posted on

No method has been shown to predict what software to build. So I'm backtesting mine on three years of store data

Update, Oct 9: I wrote this on Sept 29. The first backtest finished on Oct 1. One signal first looked 13x better than random; compared against needs of the same size, it was about 2x; with finer size bands, an estimated 1.4–1.6x, and that is in-sample, so likely on the high side. Which signals really work will take forward checks. One signal ("Missing Here") was moved to watch-only on Oct 1. Originally published on Medium.

You can build a Chrome extension with AI in an afternoon now. Building isn't the bottleneck anymore. Knowing what to build is.

I'm running a public experiment: mine public store data for unmet needs, let AI research whether each one is worth building, let AI build the good ones, and publish every result — misses included. A week in, I hit the question underneath all of it: why should anyone trust a verdict that says "build this"?

So I read everything I could find on how to pick software opportunities. Short version:

No method has been independently shown to predict what to build. What you can borrow is a vocabulary for describing opportunities and a few robust correlates of success. Whether a specific signal works, you have to test yourself.

The experiment that humbled me

A popular playbook: find abandoned or removed extensions that still have users, and rebuild them. People sell lists of these.

We took the 100 top-scoring candidates and had an AI agent research each one on the web:

  • 0 were worth building. 34 were "maybe, with conditions", 66 were no.
  • 90 already had a better, maintained alternative.
  • 66 served a need that was shrinking or gone.
  • 40% of the top 30 were "maybe", versus 28% of ranks 61–100 — only slightly better, and not one was a clear yes. The score meant much less than it looked.

Abandoned usually means abandoned for a reason. A list tells you what's broken; only validation tells you whether it's worth fixing.

Popular frameworks describe opportunities well — they don't predict them

  • Outcome-Driven Innovation gives you a great definition ("important but underserved outcomes"). Its famous 86% success rate comes from 21 of the vendor's own client projects.
  • Disruptive innovation: a re-examination of the 77 cases cited in the original work found only 9% fit all four elements of the theory.
  • "Timing matters most": the widely quoted 42% comes from a founder scoring 200 companies himself, after the outcomes were known.
  • Indie hacker wisdom (stair-step, audience-first, "4 of my 70+ projects made money") is genuinely useful — and mostly self-reported.

Two things do have real statistical backing:

  1. A meta-analysis of 233 studies: the strongest correlate of new-product success is product advantage (a correlation of .34, the strongest in that meta-analysis). The samples are mostly manufacturing and B2B, so whether it transfers to small indie software is an open question.
  2. A randomized trial with 116 startups: founders trained to treat ideas as falsifiable hypotheses performed better and were more willing to pivot. The real benefit is killing false positives.

And I could not find a single verified case of "picked an idea from store data → made money."

Base rates are brutal

  • Of subscription apps, 17.2% reach $1K/month within a year; 3.5% reach $10K.
  • The median Chrome extension has 17 users; about 70% have fewer than 100.
  • About 60% of Chrome extensions live roughly a year; among potentially infringing extensions, 86% are near-copies of removed ones.

So I assume random picks hit 5–20% of the time. Any method that reports a hit rate without comparing it to a random baseline is telling you nothing.

What we're moving to

Three layers, plus a separate "need" object (signals are live; dimensions and forms are being designed):

  • Signals — why is there a gap here? Seven on trial: Abandoned Users, Leader Complaints, Unfair Pricing, Rising Demand, Underserved Website, Leader Went Bad (acquired, ads added, price hike, ratings sliding), and Missing Here (people use it in another store; Chrome has nothing like it).
  • Dimensions — is it worth it? (being added) Size, willingness to pay, product advantage, build difficulty, acquisition difficulty, timing, evidence strength.
  • Form — what do we ship? (being designed) A single tool, a bundle of small needs, a template, or a funnel into something else.

Needs live on their own because of an early lesson: across 30,000 extensions, the unit of opportunity is a need, not an extension. One need often has dozens of extensions; old ones die and users have already moved.

Every "build this" will also ship with the riskiest assumption, the cheapest test, and a stop-loss line — our practical takeaway from that randomized trial.

Why backtesting

If no theory has been shown to predict, test against history. We have daily full snapshots of the Chrome Web Store going back to September 2023:

  1. Pretend it's a past date and run today's rules to produce that day's opportunity list.
  2. Check whether, within 12 months, a newcomer reached 10,000+ users in those needs.
  3. Compare against randomly picked needs.

(Backtesting is borrowed from quant investing: judge today's rules by how they would have done on past data, not by how convincing they sound.)

That answers which signals beat chance and by how much, how to weight the dimensions, and something I couldn't find any outside study answering: do gaps mined from store data actually get filled by someone who wins?

Pitfalls to fix first: remove duplicate clones before counting competitors; clean the removals first — many are spam, so keep only ones that had real users and were compliant; and deduplicate "hits" too, or one winner's dozens of clones become dozens of hits.

One consequence to design for now

If many people subscribe to the same opportunity list, a shared list speeds up saturation of the very opportunities it finds. How opportunities are distributed (limited, batched delivery, routed by each user's own constraints) may need to be designed together with the backtest, not bolted on after launch.

If you're figuring out what to build

  1. Look at markets as needs, not individual products.
  2. Decide what separately from what form.
  3. Ask "why is there a gap here?" first. No signal, no bet.
  4. Rate product advantage before ease of building, or you'll drift toward easy-but-pointless.
  5. Write the riskiest assumption, cheapest test and stop-loss before you build.
  6. Always compare your hit rate to a random baseline.

The first backtest results are in the update at the top. When the forward checks are in, I'll publish them as they are — including the signals that turn out to be useless.

How do you decide what to build right now? Reply with what you're working on.

I'm Reid. Building is cheap now; knowing what to build isn't. Data picks, AI builds, every result goes public — misses included.

Sources: Wikipedia (Outcome-Driven Innovation); Strategyn; e-Literate, "Cracks in the Foundation of Disruptive Innovation"; Bill Gross talk transcript (Singju Post); Evanschitzky et al. 2012 meta-analysis; Camuffo et al., Management Science; TechCrunch 2024-03-12 on RevenueCat data; DebugBear, "Counting Chrome Extensions"; arXiv 2406.12710 and 2406.00374. Our data: 100-candidate research run and 30,009-extension need clustering, 2026-09-24. Data: Chrome-Stats, processed by us.

Drafted with AI assistance; the decisions and opinions are mine.

If you're weighing a direction right now, we do one free verdict a day. The link is in my profile.

Top comments (0)