What ranking 2,163 companies taught me about turning data into decisions
Most companies are not short of information. They are short of an interpretation
layer, and the interpretation layer is where the mistakes hide.
We built one this month and it failed in a way worth describing, because it
failed quietly and it looked correct while it failed.
The job
We had 2,163 companies in one US sector and needed to know which ones to
approach first. Two axes matter: can they pay, and do they have a problem worth
paying to fix. We measured both from public signals. Domain age, sitemap
presence, last publication date, mail provider, registry filings, site condition.
Coverage was uneven and we recorded that rather than smoothing it over. Mail
provider resolved for 95 percent of them. Domain age resolved for 24 percent.
Sitemap data existed for 12 percent.
The first failure was in the measurement
The 24 percent figure bothered me, so I sampled it. Of five companies marked
unmeasurable, four had archives going back more than a decade. One went back
twenty two years. The data existed. We had run eight parallel workers at the
archive and been throttled, then recorded the throttling as absence.
We reran it with two workers and waiting periods. Coverage went from 24 percent
to 88 percent. My time estimate for that rerun was fifteen minutes. It took 179.
The lesson is not about workers. Not measured and does not exist are two
different findings, and a pipeline that collapses them will report confident
nonsense.
The second failure was in the ranking
I scored the companies with a weighted sum. Ability to pay, plus severity of
problem, plus supporting signals.
The top of the list came back with a company scoring 100 on ability to pay and 3
on severity of problem. It ranked in the top twenty. A business with money and no
problem had been placed at the front of a sales list, because a high score on one
axis carried a near zero on the other.
A sum lets one axis substitute for the other. That is exactly wrong here, since
both have to be present or there is nothing to sell. Changing the core to the
square root of the product fixed it. The same company dropped to 40 points and
position 119.
What the data said once it could be trusted
- 1,680 of 1,912 measured companies have no sitemap
- 76 have a blog
- 205 show any publication date, ever
- 1,293 use Google Workspace or Microsoft 365, so these are real operating firms
Almost nobody in this sector publishes anything. That is the finding, and it only
became visible after the measurement layer stopped lying to us twice.
The part that generalises
An intelligence layer is not a dashboard. It is a set of claims about reality, and
each claim can be wrong in two directions: the data can be missing, or the maths
on top of it can be shaped so that a missing value looks like a good one.
Both failures here were mine, both looked like working software, and neither
would have surfaced without going back and sampling the rows the system had
already given up on.
Top comments (0)