We build an AI tool that tells founders whether their startup idea is worth building. That sentence should make you suspicious, and the suspicion is earned by the category.
In March, a study in Science tested 11 large language models on requests for personal advice and found they affirmed the user's actions 49% more often than humans did, including when the action was a bad one. Asking a chatbot whether your idea is good is personal advice with a spreadsheet attached. It tends to say yes.
That is why our verdict is not a model's opinion. The interview feeds six research stages, and the verdict is arithmetic: market size, growth, business model, problem clarity and audience each score against fixed thresholds, and the sum decides go, pivot or no-go. Nobody, including us, gets to nudge it up so the report feels better.
But an honest calculation is still a forecast. The market numbers are estimates, the synthetic focus group is a model of people rather than people, and the founder's own answers are the founder's view. A verdict built carefully on forecasts can still be wrong. So we added a step whose only job is to measure how wrong.
The step: a survey that can prove the verdict wrong
After the verdict, the founder gets a survey generated from their own project and sends the link to people they consider their audience. When answers come back, we put a second score next to the AI score and keep the difference.
That sounds like "ask your friends", and the obvious objection is that friends are polite. The objection is correct, and it is exactly why the questions are built the way they are.
Politeness needs an opinion to attach to
The questions follow The Mom Test: they ask only about the past and the present. When did this last happen to you. What do you use for it today. What do you spend on it now.
The generator is instructed to leave out the questions everyone instinctively writes: would you use this, would you pay for this, do you like the idea. Those invite kindness. "When did this last happen" has no polite answer. Either there was a last time or there wasn't.
Every question is tied to one hypothesis from the validation (a red flag, a green light, the persona, pricing, the market), so every answer moves something specific instead of producing a vibe.
Screening comes first, with the role we are looking for sitting among plausible alternatives. Behaviour questions fix politeness; they do not fix the fact that your friends may not be the audience. Hiding the right screening answer is how you find that out.
Five answers, then a gap
Nothing is scored until five people have answered. Five is not a statistic, and we don't pretend it is. It is enough to see direction: whether the problem you assumed is common, rare or absent among the people you reached.
Each hypothesis then gets a verdict: confirmed, rejected or inconclusive. The headline number is the gap between the score from real answers and the AI score. Within 10 points is green, up to 20 is yellow, beyond that is red, and the direction matters as much as the size: was the model optimistic or pessimistic?
A second round goes out only on what was rejected or unclear. Confirmed hypotheses are not asked again.
The part that made us uncomfortable
The prediction has to exist before the answers do. That is the whole trick, and it applies with or without a tool: if you read survey results without writing your prediction down first, almost any result feels expected afterwards. Hindsight does the validating for you.
In our case the computed score is the written-down prediction, which means the product ends up keeping a record of its own misses. A validator without this step can stay confident forever, however honest its arithmetic. One with it accumulates specific numbers about how far off it was, and has to show them.
We are not going to claim an accuracy figure for that loop here, because a figure like that should come from a lot of real surveys, not from a landing page.
If you want to do it without us
We wrote the method up as a guide with a copyable ten-question template: prediction first, hypotheses, past behaviour only, hidden screening, read after five, keep the gap. It is here: https://gonogo.team/validation-survey
A question for anyone who has run validation surveys: when your results disagreed with what you expected, did you actually write the expectation down beforehand, or did you only notice the disagreement because it was large?
Top comments (0)