DEV Community

Boris Binyaminov
Boris Binyaminov

Posted on Originally published at whittleos.com

We ship a thing that judges startup ideas, and we could not answer whether it works

The problem: we ship a thing that judges startup ideas, and we could not answer whether it works. Not in the marketing sense — in the sense of never having run anything.

The design we settled on. Take companies whose stories are over, reduce each to the one-liner it could have been described by at founding, strip the name, and feed it to the live gate. The known ending and its public source sit in the table next to our verdict.

The scoring rule is the part worth stealing. Only a KILL counts as catching something. A HOLD or a pilot-first is the gate declining to decide, and counting those as saves is precisely the arithmetic that lets everything in this category call itself accurate.

Run on 2026-06-23 against 8 companies — 5 that really failed, 3 that really worked. Solo lens: 4 of the 5 flops killed, and 1 of the 3 successes killed as well. Funded lens: 2 of 5 flops, 0 successes. 3 rows were never killed by either lens.

Then we did not publish a percentage, and that decision is the piece rather than a caveat on it. Four separate defects, any one of which is enough.

The sample is small and hand-picked — chosen because the outcomes are known and famous, which is a selection rule with nothing to do with the population of ideas a user actually brings.

Hindsight leaks. The cases are anonymised, but a model may still recognise a famous shape. We cannot prove it did not, so we cannot treat a hit as clean.

The cells move. Each deal-breaker verdict is a five-sample supermajority vote, which is materially steadier than a single pass, but the weighted total is still one pass and a borderline row can drift a point or two between runs. It is a dated snapshot you can re-run, not a statistic.

And the verdict is founder-bound by design, which makes a single accuracy number incoherent rather than merely imprecise. The success we killed on the solo lens is a payments company: correctly killed for one person, correctly not killed for a team. Same one-liner, two right answers.

What I would tell anyone building an evaluation for a system whose output depends on who is asking. Decide what counts as a catch before you look. Keep every row that goes against you. And notice when your metric is asking a question your system does not answer — that was the real finding here, and it arrived last.

Full table, every row, including the flop that walked through both lenses: https://whittleos.com/guides/do-startup-idea-validators-work

Top comments (0)