Six batches of a pharmaceutical extraction process. Yields sit at σ ≈ 0.02 percentage points, every one of them hugging the specification lower bound. Charging masses are recorded to 0.01 kg.
Nobody on a factory floor weighs to 0.01 kg. The operator said so plainly when I asked: there is always drift, nobody weighs precisely and nobody can, the records are written to match what the process spec requires, and the numbers are computed.
That is the same shape as earnings benchmark-beating, one industry over. What interests me is not the fraud angle. It is a measurement property that fell out of the data and that I think travels past this domain.
How hard a number is to fake backwards scales as 1 / (target × tolerance)
A main component specified at ≥30% with a 5% tolerance gives you a wide band of plausible values. A fabricated number is cheap to produce there, and nearly impossible to tell apart from a measured one. A trace component specified at ≤0.5 ppm has almost no room, so fabricate it and the distribution gives you away. In one incoming-inspection lot, the range ratio across three sub-channels came out around 55×, strictly decreasing with magnitude.
So the numbers easiest to manipulate are the ones hardest to detect. Those read like two separate facts. They are one fact said twice, because both follow from the width of the plausible band.
If that sounds familiar, it should. It is the same reason I have spent this year filing token-accounting bugs in usage trackers. A per-message token count with a loose upper bound drifts quietly for weeks. A hard-capped counter announces its own breakage the first time it is wrong.
The part I want stress-tested
Does 1 / (target × tolerance) hold in your domain?
Clinical trial records, emissions reporting, financial close, safety incident logs, SLA reports, anything with a compliance number and a stated tolerance. If it fails somewhere, that failure is worth more to me than agreement.
I have opened a comment window until September 30 on the three papers behind this: Open comment window on GitHub. They go to v1.0 in October, labelled not peer reviewed, with every comment logged and dispositioned in public. Zero comments will be recorded as zero.
The data is de-identified. Product names, batch numbers, lots, equipment and people are all pseudonyms. The σ, range ratios and multipliers keep their original values, so the results recompute. CC BY 4.0.
Top comments (0)