These articles come from lessons learned while building Eterna Clarity and the operating system I use to run it.
I recently built a scoring system to help Eterna decide which businesses looked like the strongest opportunities for Custom Clarity. Then I gave it an important rule: if the ranking surprised me, I was not allowed to change the scoring simply because I preferred the old answer.
That became much more important than I expected. By the time I designed the model, I had already researched many of the companies and formed opinions about which ones looked strongest. It would have been very easy to build something that turned those opinions into numbers and then call the result objective.
A number can make an opinion look scientific
Eterna needed a better qualification method because a business can look promising from a distance for all kinds of bad reasons. A weak website does not prove weak operations. Hiring activity can mean growth, turnover or neither. A company can appear digitally unsophisticated while running excellent internal systems that are simply invisible from the outside.
So the new model stopped asking for one vague judgement and started examining several dimensions separately. It also distinguished the apparent quality of the opportunity from the quality of the evidence supporting that conclusion.
That distinction matters. Two companies might both appear to be strong opportunities, but one conclusion could be supported by several independent signals while the other rests mostly on inference. Giving both businesses similar scores without representing that difference would create precision that the research had not earned.
A score tells me what the evidence appears to suggest. The evidence grade tells me how seriously I should take the score.
The strange rankings were the valuable ones
I ran the new system against an existing batch of 25 companies before allowing it to replace the earlier qualification method. For each business, I compared the old ranking, the new ranking and a fresh human review.
The important part was what happened when they disagreed. Instead of adjusting weights until the order looked familiar again, I treated the disagreement as something that needed an explanation.
Sometimes the model might be overvaluing a particular signal. Sometimes the original judgement might have been too generous because the business looked like an easy fit. Sometimes important information could be missing, or a public signal could mean something different from what I first assumed.
All of those possibilities are useful. If I change the model every time it produces an answer I dislike, I eventually get a scoring system that is exceptionally good at agreeing with me.
That is not independent judgement. It is a mirror with arithmetic.
Human judgement still matters
I do not want a scoring model making these decisions by itself. Public information is incomplete, businesses are messy, and context often changes the meaning of a signal.
Human judgement becomes especially valuable when the result looks strange. The mistake is using that judgement as an invisible answer key where every disagreement automatically means the model must be wrong.
That problem exists far beyond sales. A hiring rubric can slowly be adjusted until it ranks the candidates a manager already likes. An investment model can be refined until favourite companies rise back to the top. A product-prioritization framework can become a complicated way of justifying decisions that were already made.
The moment you know the result you want, every change to the scoring method deserves more scrutiny.
A useful question is simple: did I discover a flaw in the model, or am I uncomfortable because the model challenged one of my assumptions?
Test it on cases that did not create it
There was another problem with the first 25 companies. They had already influenced how I thought about qualification.
Their failure modes helped shape the new model. The weird cases I encountered helped determine what the system needed to consider. Even without intentionally fitting the model to those companies, they were part of its education.
So the next test needed businesses I had not used while designing it. I chose a separate set from outside Alberta and planned to run the same method without changing the rules first.
That is a useful test for almost any decision framework. If you develop a hiring rubric by studying your best employees, try it on people who were not part of that analysis. If you create a project-risk framework after three painful failures, see what it says about projects that had nothing to do with those failures.
A framework that explains the examples used to create it may simply be a good description of those examples. The more interesting question is whether the reasoning still works somewhere new.
The score should guide attention, not create certainty
One of the easiest mistakes with scoring systems is treating the ranking as the decision itself.
If a company scores highly, that does not automatically make it a lead. It means the available evidence suggests that company deserves more attention than another one. Further research may strengthen the case, weaken it or reveal that the opportunity was never real.
That has become an important boundary in Eterna's acquisition work. Discovery can be broad and inexpensive. Qualification should be more demanding, and contacting a real business should require stronger evidence again.
The score helps decide where to spend the next unit of research. It does not create entitlement to somebody's attention.
That also makes uncertainty easier to handle. A company does not need to be labelled good or bad before enough is known. Sometimes the right state is simply promising, but poorly evidenced.
A useful model should be able to surprise you
I built the qualification system because I wanted Eterna to make better decisions about where Custom Clarity might genuinely be useful. The most valuable thing it can do is not reproduce my judgement more neatly.
It can force vague impressions into rules that can be inspected. It can apply the same questions more consistently than I might. Most importantly, it can produce a result that makes me stop and look again.
Sometimes that second look will expose a bad assumption in the model. Sometimes it will expose one in mine.
If every result confirms what I already believed, I have learned almost nothing.
Top comments (0)