Search how to build a weighted scorecard and every result stops in the same place: choose criteria, assign importance, multiply, sum. All correct, and all of it ends exactly where the decision starts. Which criteria. What importance.
Why weights stay unpublished
That gap is not an accident of writing. Publishing weights commits you — a reader can say you priced trust twice as high as margin and I think you are wrong, and can notice when your own scores do not follow your stated rules. Unpublished weights cannot be argued with, which is comfortable and useless.
Ours, and the one that counts for nothing
Ours: 7 categories that carry weight, plus one that is rated, displayed, and contributes nothing. The top three are tied at 20 each rather than one dominating, because no ordering between reachability, channel and economics survived contact with real ideas, so none was invented.
Evidence that the problem is real is deliberately weighted below all three. That reads wrong until you see what it prevents: excellent evidence of a problem nobody can be reached about is still a bad business, and weighting evidence at the top would let a well-documented dead end outscore a reachable one.
The decision that matters more than the weights
The implementation decision that matters more than the weights: the model rates each area, and code computes the total. If a language model hands back its own overall number it can talk itself into a pass, and no amount of prompt discipline makes that auditable.
What a point is actually worth
What a point is worth, which is what a published rubric actually buys. Take an idea rated 7 everywhere except distribution, which comes back 3 — a decent product nobody can be reached about, the most common shape there is. The total is 62.
Spend three rating points on the weak area and you get 70: a gain of 8, and the grade moves. Spend the same three on the lightest category and you get 64, a gain of 2, and nothing about the decision changes. The heaviest area carries 4 times the leverage of the lightest.
That ratio is the entire practical value of publishing a rubric. Without it you cannot tell the expensive fix that moves a verdict from the satisfying one that does not, and work goes where it feels productive rather than where it counts.
The honesty clause
The honesty clause, because a published rubric obliges it. The arithmetic is fixed and repeats; the ratings underneath are model judgements sampled at temperature one, so the same idea run twice can score differently. A total is a reading, not a measurement — and anyone claiming otherwise is describing their arithmetic and hoping you hear it as their whole system.
One more thing we had to get right: uncertainty is not a middling score. A category the model cannot judge does not average to the middle. It is marked low-confidence, and enough of those push the verdict to hold rather than letting a confident-looking total emerge from several shrugs.
Every weight, the worked example and the grade bands: https://whittleos.com/guides/how-to-score-a-business-idea
Top comments (0)