DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our taxonomy knew 629 E-numbers and could not tell us one was a preservative

Munchable scans a packaged food and tells you whether it fits your gut condition. Alongside the condition verdict there is a second, much lighter thing it can do: name the additives in a product for someone who has decided they would rather not eat certain classes of them. Sweeteners, artificial colours, that sort of preference.

That feature sat in the backlog for months behind one sentence I could not get past. Our ingredient taxonomy knows 629 E-numbers, and it cannot tell you that any of them is a preservative.

You can see the naming half working at munchable.app/does-e330-cause-reflux, one of 373 ingredient pages at munchable.app/answers. It knows E330 is citric acid, it knows a label might print either, and it knows what the reflux rule thinks. What it did not know, until we built the thing this post is about, is that citric acid is an acidity regulator.

A synonym tree is not a classification

Ingredient taxonomies are built as hierarchies, which makes it easy to assume the hierarchy means what you want it to mean. Ours relates an ingredient to a broader ingredient: cheddar to cheese, cheese to dairy. That relation is what lets a rule about dairy fire on a pack that printed cheddar, and it is the backbone of the whole engine.

Additives are in that same tree, and they are in it as names. I measured what the parent links actually said:

  • 629 E-numbers in the taxonomy.
  • Exactly one of them has a parent that is not another E-number.
  • en:preservative exists as a node. It has two children. Neither is an E-number.

So the graph relates E150a to E150, which is a naming relationship, not a functional one. Nothing anywhere said "E211 is a preservative", and nothing could, because the tree is not that kind of tree. The engine could render the word "sodium benzoate" and had no idea what the molecule is for.

The tell, with hindsight, was in our own condition rules. Every rule that touches an additive listed individual numbers by hand. Nobody wrote "flag preservatives" because there was no way to say it. I had been reading a workaround as a design choice for a year.

Classification alone would have shipped a useless filter

The obvious fix is a map from additive id to function, done once, for the whole range. That is most of the work. It is also where the feature would have failed if we had stopped there, and the reason is in the catalogue statistics rather than in the code.

E330, citric acid, appears on 14,836 products in our catalogue. E322, lecithin, on 10,604. Those are not exotic. They are in bread, in tins, in chocolate, in half the aisle.

A filter that fires on them flags nearly everything a shopper picks up. And a filter that flags everything communicates nothing: the user learns within three products that the badge is noise and stops reading it. That is worse than not shipping the feature, because it also teaches them to distrust the badges that do matter, which in this app includes an allergen warning.

So the map carries two facts per additive, not one. What it is, and how loudly to say so:

/**
 *   'flagged'  the filter names it, when the preference covering its class is
 *              on. These are the things a person switching on "no artificial
 *              colours" has in mind.
 *   'notable'  never fires. Listed as context under the row, so the additive
 *              picture is complete without being alarming: phosphates,
 *              modified starches, anti-caking agents.
 *   'benign'   silent. Citric acid, ascorbic acid, carbonates, natural gums,
 *              tocopherols, the gases. Naming these would be noise.
 */
Enter fullscreen mode Exit fullscreen mode

The tier is the load-bearing part, and it is not derivable from the chemistry. It comes from what the additive is plus how common it is plus what a person switching this preference on actually had in mind. Citric acid is an acidity regulator and it is silent. The classification without the tier is a technically correct feature that no user would keep switched on.

The middle tier is the one I would defend hardest. "Notable" additives never trigger anything; they are listed as context so the picture is complete. Without that tier you get a binary where everything you decline to flag is invisible, and a user who can see phosphates on the label but not in the app concludes the app missed them.

The map makes no claim about harm

One rule in the file, and it is about scope rather than code:

this file says what an additive is and how loudly to say so. It makes NO claim about harm and carries no health copy, because the filter is a preference, not a nutrition verdict. The user asked not to see artificial colours; this file knows which ones those are.

This matters because of where the output lands. Munchable's condition verdicts are backed by published clinical guidance, and they are the thing a user with a diagnosis relies on. The additive filter is a lifestyle preference sitting next to them in the same interface. The moment the preference lane starts implying "this is bad for you", it is making a health claim with none of the evidence behind it, and it contaminates the lane that does have evidence. I wrote about keeping those two lanes separate in a preference is not a verdict, and this file is where that separation is enforced at the data layer rather than in the UI.

Practically: no row in the map contains a sentence. Rows are enums. There is nowhere for a health claim to hide, which is a much more reliable guarantee than remembering not to write one.

How 629 entries get maintained without a review queue

Hand-writing a classification for an entire regulatory range once is fine. Keeping it current as additives are approved, renamed and reformulated is the part that kills these files.

Two inputs, with different privileges:

  1. The hand-written map in the repo. Reviewed, versioned, wins every merge.
  2. A curation overlay, published separately, which may only add ids the hand map does not carry, and whose rows hold enums only.

The constraint on the overlay is what makes it safe to merge automatically rather than through an approval queue. It cannot change an existing classification, it cannot introduce prose, and the deterministic code is still the thing that decides what the user sees. AI proposes, the engine disposes. If a proposal is wrong, the blast radius is one additive that nobody had classified at all, and the fix is a line in the hand map, which outranks it.

A coverage audit then closes the loop in the other direction: the script that reports which ingredients the engine cannot speak about counts a classified additive as covered, so the 334 additives this map settled stop appearing in the curation backlog for ever.

Try it

The naming layer is public and you can poke at it without an account:

The filter itself is in the app. Sign in at app.munchable.app, switch on a preference under Healthy shopping, and scan a few things from the cupboard. The useful test is not whether it flags something. It is whether, after ten products, you still want it on.

Top comments (0)