Munchable is a gut-health scanner. You point it at a barcode and it tells you whether the product suits the conditions on your profile, using a fixed rule set rather than a language model. A few weeks ago we added a second, optional lens over the same scan: a healthy shopping filter, for the additives and nutrient levels somebody says they would rather not buy.
The first version of that filter was three lines long. Take every additive the product carries, name it, show the list. It was useless, and the reason is worth writing down, because it is the same reason most "flag the bad stuff" features quietly stop being read.
The number that killed the naive version
Citric acid, E330, is on 14,836 products in our catalogue. Lecithin, E322, is on 10,604. Both are additives by every definition, including the regulatory one. A filter that names them fires on most of the aisle, and a warning that fires on most of the aisle is not a warning, it is wallpaper. The user learns to scroll past the row within a day and the feature is dead while still shipping.
So the filter needed to be able to say nothing. That meant a piece of data we did not have.
The taxonomy knew the names and not the jobs
Our ingredient taxonomy knows 629 E-numbers. It can tell you that en:e211 exists and print "E211". What it cannot tell you is that E211 is a preservative: of those 629 ids, essentially none have a parent that is not another E-number, and en:preservative has two children, neither of which is an E-number. The hierarchy is a numbering scheme, not a classification.
That gap is why every condition rule in our engine that touches an additive had already been forced to hardcode individual numbers. Do it once more and it is a pattern. So we built the classification properly, one map, for the whole additive range:
export interface AdditiveEntry {
class: AdditiveClass; // what it does
concern: AdditiveConcern; // how loudly to say so
name?: string; // the substance in plain English
}
335 entries today. The class field is the boring half: preservative, colour, sweetener, emulsifier, and so on, plus two that are not additive classes at all in the regulatory sense (hydrogenated-fat and added-sugar) because hydrogenated palm oil and glucose syrup belong to this filter and to nowhere else in the engine, and one map beats two parallel structures.
The concern field is the load-bearing half.
Three tiers, and only one of them speaks
export const ADDITIVE_CONCERNS = ['flagged', 'notable', 'benign'] as const;
- flagged (168 entries) means the filter names it, if the preference covering its class is switched on. These are the things a person who turns on "no artificial colours" actually has in mind.
- notable (77 entries) never fires. It is listed as one closing line of context under the row, so the additive picture is complete without being alarming. Phosphates, modified starches, anti-caking agents.
- benign (90 entries) is silent. Citric acid, ascorbic acid, the carbonates, the natural gums, the tocopherols, the packaging gases. Naming these is the noise we started out generating.
Within a class the split is not arbitrary, it follows what the shopper means by the word. The synthetic dyes are flagged; the plant and mineral colours are benign, because they are exactly what someone avoiding artificial colour is pleased to find:
'en:e102': f('colour', 'Tartrazine'), // flagged
'en:e129': f('colour', 'Allura red AC'), // flagged
'en:e160c': b('colour', 'Paprika extract'), // benign
'en:e163': b('colour', 'Anthocyanins'), // benign
One preference is off by default even for people who switch the whole filter on: emulsifiers. "Is this an emulsifier" is true of the gums, the lecithins and E471, they are between them on a very large share of the catalogue, and a filter that fires on all of them says nothing by saying everything. It exists for the people who want it and stays out of the way of everyone else.
The map is not allowed to have an opinion
One rule kept this file honest while it was being written: it says what an additive is and how loudly to say so. It makes no claim about harm and carries no health copy.
That is not squeamishness, it is a type-level consequence of what this layer is. A condition rule set answers "does this suit my gut" from clinical guidance. This one answers "is this the kind of food I said I wanted to buy". The user asked not to see artificial colours, the pack has one, and naming it is the entire job. The moment the map starts carrying sentences about what an additive might do to you, it is pretending to be the other thing.
Which is also why the copy templates live in code rather than in the data:
"Allura red AC (E129), an artificial colour"
"Hydrogenated palm fat"
"Maltodextrin, an added sugar"
An E-number always gets its class named, because the number alone tells a reader nothing. A plain ingredient is left to speak for itself, since "Hydrogenated palm fat, a hydrogenated fat" says one thing twice. Added sugars are the exception, because "Dextrose" and "Maltodextrin" are precisely the words a label reaches for when it does not want to say sugar.
What the layer is allowed to do to a verdict
Nothing, in one direction. This filter only ever adds findings. It cannot clear a product and it cannot force an avoid, because nothing in it is a hazard. The strongest thing it can do is keep a product off green.
It also never gets a green tick of its own. When it finds nothing the row is muted grey and reads "Nothing you asked about", not "None found, all clear". The filter read an ingredient list and whatever nutrition figures the row happened to carry. "Nothing you asked about" is a true sentence; anything implying the pack is clean would not be.
That is a deliberate contrast with the allergen layer next to it, which is safety-critical and hedges every line, because "not mentioned" must never be read as "free from". Two layers, two tones, and mixing them up would make both worse.
The nutrient half, and the thresholds we did not invent
Three of the ten preferences read the nutrition panel instead of the ingredient list: high sugar, high salt, high saturated fat. The thresholds are the UK FSA front-of-pack red bands, 22.5 g sugars, 1.5 g salt and 5 g saturated fat per 100 g.
Using a published standard was not laziness. Munchable does not get to invent what "high in sugar" means, and the number a shopper has already seen on the front of a thousand packs is the number our filter should agree with.
Where a figure is missing, the filter says nothing about that nutrient. It does not announce that it could not assess it. Most of the catalogue carries fat and fibre only, and a filter that opened with three lines of what it did not know would bury the part that works.
One label word, one finding
The last bit of polish is a deduplication that comes straight out of having a taxonomy. A normalized product carries every ingredient plus its ancestors, which is what lets hydrogenated palm oil match a rule keyed on en:hydrogenated-oil. It also means glucose-fructose syrup arrives carrying en:fructose, and both are flagged added sugars, so the naive version reported one ingredient twice.
Nearest match wins, the same precedence the condition rules use:
const specific = ids.filter(
(id) => !ids.some((other) => other !== id && ancestorsOf(other).includes(id)),
);
Drop any id that is an ancestor of another id in the same list, and you are left with the specific word the label actually printed.
Go and look at one
The engine behind this is the same one that generates our public answer pages, so you can read its output without installing anything:
- Is carrageenan OK with IBD? and is acesulfame K OK with IBD? are two additives where the condition rules have something to say, separately from this preference layer.
- The full answer index is one page per ingredient and condition we hold a rule for.
- The IBD guide is the condition those two sit under.
The healthy shopping filter itself lives in the app, on the result screen under the verdict, at munchable.app. Turn on the preferences you care about during onboarding and scan something beige from the middle of a supermarket. If the row is quiet, that is the tiering doing its job.
Top comments (0)