DEV Community

Boris Binyaminov
Boris Binyaminov

Posted on Originally published at whittleos.com

A score with no published rubric has no floor to fall to

Ask a chatbot what it thinks of your startup idea and it will help you. That is the whole problem: helpfulness and judgement point in opposite directions, and only one of them was trained for.

Four mechanisms, none of them a defect. You told it the idea is yours, and a model completes the conversation it is in — the same text pasted as a competitor is doing this, what is wrong with it reliably produces a different answer, with nothing about the business having changed.

Agreeable answers were preferred during training. Humans rate the response that engages with their plan above the one that dismisses it. A sensible objective for almost every task, and a disastrous one for a go or no-go.

There is no cost to a false yes. A tool that encourages you loses nothing when you fail eight months later; a tool that kills your idea loses you today. The incentive is not disreputable, it just points away from the thing you wanted.

And scores drift upward when nothing anchors them. A number generated rather than computed from a published rubric has no floor to fall to — nothing says what a low rating looks like, so the model produces something plausible, and plausible-sounding is high.

The concrete instance, using our own idea so nobody has to trust our judgement about somebody else's. Same one-liner, 3 tools, on 2026-06-10.

A competitor returned 72 / 100 and PROMISING — excellent potential, among the top ideas we have seen. Our own adversarial mode returned 15 / 100, and our balanced check returned 0 / 100, short-circuiting on a deal-breaker rather than producing a score at all.

The limit of that evidence, stated before anyone quotes it back at us: one idea, one day, each tool once. It supports exactly one sentence — these were the outputs — and no claim about how any tool behaves in general. We publish the date so it can be re-run, which is the only thing that makes a sample of one worth publishing.

What actually separates a warm answer from a judgement, and none of it is tone — a tool can be blunt and still be flattering. Is the verdict bound to you, or is it a judgement about an idea in the abstract, which is a judgement about nobody. Are the weights published, so the number can be argued with. Can you open the evidence — a link that resolves to somebody saying the thing, not a list of sources considered.

The practical move if you get an encouraging answer: do not discard it, re-ask it in a frame that removes the flattery. Paste the idea as a competitor's product and ask what is wrong with it. Ask what would have to be true for it to fail, then check the cheapest of those yourself. If the verdict survives all of that it might be right. If it evaporates the moment the idea stops being yours, you learned what the first answer was measuring.

The table and the three properties: https://whittleos.com/guides/why-ai-says-your-idea-is-great

Top comments (0)