Munchable scans a food label and says one of four things: Good fit, Caution, Avoid, or Can't assess. The fourth one is a real answer rather than an error state, it has its own grey colour that is deliberately never green, and it has its own icon because a verdict is never carried by colour alone.
For a long time one line of code was responsible for about a seventh of all the Can't assess answers we gave. We deleted it. This is what it was, why it looked right, and what we had to build instead.
The heuristic
The engine reads an ingredients list, maps every word to a canonical id, and some words it cannot map at all. Call that share the unknown ratio. The rule was: above a fifth unknown, a good verdict gets downgraded to Can't assess.
The reasoning is hard to argue with in the abstract. An unread word might be a trigger. If we cannot read a fifth of a label, our green light is a guess, and a health tool has no business guessing in the direction of reassurance.
It shipped, and it was wrong, and the way it was wrong is the interesting part.
What the unknown words actually were
Not ingredients. Here are two real examples of the words that were spending our users' confidence:
voir bouteille
residu sec a 110 c
"See bottle." A dry-residue measurement at 110 degrees. The first is a pointer to the rest of the label. The second is a laboratory figure that happens to be printed inside an ingredients panel on a bottle of mineral water.
Neither is food. Neither could trigger anything. And on a label with eight ingredients and three lines of boilerplate, they were enough to push the ratio over a fifth and convert a perfectly well-read product into "we cannot tell you".
Meanwhile the words that actually decide a verdict, the onion, the cream, the hydrogenated fat, had been read perfectly.
Why the ratio was the wrong measurement
The realisation, once we stopped defending the design:
Our condition rules are positive trigger lists. The engine does not clear a product by understanding every word on it, it raises a flag when it recognises one of the words that matter. An ingredient that triggers nothing needs no row, no clearance, and no reading, because nothing about it can ever reach a user.
So what protects a person is not a ratio over words nobody could read. It is whether the trigger lists are complete, in the languages that labels are actually printed in. The ratio was a proxy for that, and it was a proxy with a fatal property: it was cheap to compute, so it fired on the cheapest thing that resembled the real signal. Boilerplate.
The cap is gone. Here is the entire remaining confidence function:
export function assessConfidence(product: NormalizedProduct): Confidence {
const hasIngredients =
product.ingredientsTags.length > 0 || (product.ingredients?.length ?? 0) > 0;
if (!hasIngredients) return 'unknown';
const unknownN = product.unknownTags.length;
const completed = product.statesTags?.includes('en:ingredients-completed') ?? false;
const realUnmapped = Math.max(0, (product.unknownIngredientsN ?? 0) - product.droppedNoiseN);
if (completed && unknownN === 0 && realUnmapped === 0) return 'verified';
return 'database';
}
Three outcomes. No ingredient data at all is unknown, and there is nothing to assess. A complete list with nothing unread is verified. Everything else is database: usable data, whatever share of it we could not read.
Note the second condition for the top badge. A source can flag ingredients it could not map which never surface as tags at all, so they are invisible to the unknown count. Those are subtracted from the noise the normalizer itself threw away, and what is left blocks verified. The effect is that the confident boilerplate case (the source counted "to preserve freshness", we dropped it) keeps the top badge, while a label with genuinely unread ingredients does not.
The ratio that remains is computed over real ingredients, with the taxonomy ancestors we derive excluded, so expanding "hydrogenated palm oil" into its parents cannot dilute it. A denominator you generate yourself is a denominator you can cheat with by accident.
What we gave up, stated plainly
This is the trade and I will not dress it up. On a label the engine could barely read, a product that contains an unread trigger now reads Good fit where it used to read Can't assess. We made the product more confident in exchange for making it more useful, and that is only defensible if the thing we replaced the heuristic with is real.
What we replaced it with is a measurement, run on the ingredients that are actually on labels rather than on a vocabulary list:
pnpm triggers # the gap report
pnpm triggers --top 400 # how deep to look
pnpm triggers --list gerd # what one condition catches, most common first
Two numbers, in the order they matter. Naming: does the engine have an id for this word at all, because without that nothing downstream can fire. Reach: does any condition have something to say about that id.
Membership is resolved through the taxonomy exactly as the engine does it, so a flour variant counts as caught when the flour root is a key. Which is also the strategic point the report makes visible: deciding the top of the tree is worth far more than deciding its leaves.
That report used to have a third number, and deleting it was part of the same clean-up. It measured the share of named ids somebody had "reviewed", which counted explicit "nothing applies" clearances as progress. It bought model calls to record non-events, and no user ever saw a single one of them. A metric that rewards recording nothing is worse than no metric, because it is indistinguishable from work.
The part that never went through the cap
Allergens. They never passed through the verdict cap, before or after, because they do not work like conditions: they only ever add a warning, they can never clear anything, and an empty allergen result is worded as what it actually is. The result screen says the data Munchable holds did not mention it, which is a different sentence from "this is free from it", and that distinction is the whole reason the wording is in the engine rather than left to a screen.
The general shape of this mistake
If you take one thing from this: a safety heuristic built on a proxy will spend its budget on whatever is cheapest to measure, not on the risk you were worried about. Ours spent a seventh of our answers on mineral water bottles printing laboratory values.
The fix was not a better ratio. It was to go and do the harder work the ratio was standing in for, and to delete the ratio so it could not keep taking credit.
Look at it running
- munchable.app/conditions/gerd-reflux lists what one condition actually checks for. That list is the thing the coverage report measures.
- Does onion cause reflux? is one trigger, with the reason text a real scan would show.
- app.munchable.app scans a barcode and shows you the verdict with its confidence badge underneath, including, when it applies, the grey one.
Related, on the same engine: half our coverage tests assert that the app says nothing, and a preference is not a verdict.
Top comments (0)