DEV Community

Daniel Pertu
Daniel Pertu

Posted on

A label that prints "nuts" is evidence. The same word inferred from "almond" is not.

Munchable scans a food barcode and says whether the product fits your gut condition. The conditions are the main job, but the part of the app that has to be right in a different way is the allergy check. A condition verdict that is too cautious costs you a biscuit. An allergy warning that is wrong costs more than that.

This post is about one rule inside that check, and why it needed two extra facts recorded during parsing to be expressible at all: a general allergen word counts as evidence only when the pack itself printed it.

The expansion that makes condition rules possible

A label is a string. Before anything can reason about it, the engine turns it into a set of ingredient ids, and then adds every ancestor of every id from Munchable's ingredient taxonomy. "Almond" arrives, and the normalized product carries the nut family it belongs to as well.

That expansion is the reason the condition rules stay small. A rule can key on one node near the top of a family and cover every specific ingredient underneath it, in every language a label might print, without a row per word. The alternative is a rule table that grows with the vocabulary rather than with the biology.

The same expansion is wrong for allergens

Peanut sources include the specific words you would expect, plus one general one: a pack that prints "nuts" without saying which nuts is evidence that peanuts could be in there. That is how the picker describes the option in the app, word for word:

Peanuts, peanut oil and "nuts" printed without saying which.

Now put the expansion next to it. A pack prints "almond", the normalizer adds the nut family as an ancestor, and the product now carries a tag meaning "nut" that nobody printed anywhere. If the allergen check treats that tag the way it treats a printed word, a bag of almonds warns a peanut-allergic user about peanuts. It is a warning with no evidence under it, on the screen where a user has the least reason to doubt us.

So the check needs to distinguish a word the label printed from a word the taxonomy inferred. Normalization records both:

  • printedTags, the ids that came from the label's own words.
  • derivedFrom, mapping each inferred ancestor to the ingredient that implied it, so a warning can name it.

A general source then has one extra condition on it:

for (const source of ALLERGEN_SOURCES[allergen]) {
  if (!present.has(source.tag)) continue;
  const from = product.derivedFrom[source.tag];
  // "Nuts" on a label is evidence; a nut tag inferred from "almond" is not
  // evidence of peanuts. A general source counts only when the pack printed
  // that word itself.
  if (source.generic && !printed.has(source.tag)) continue;
  hit = {
    allergen,
    level: 'contains',
    via: 'ingredient',
    ingredientTag: from ?? source.tag,
  };
  break;
}
Enter fullscreen mode Exit fullscreen mode

printedTags has to exist as its own set, which is less obvious than it looks. You cannot get the same answer by asking whether a tag is absent from derivedFrom, because derivedFrom records the first ingredient that implied a tag. A pack that lists almonds early and prints "nuts" further down ends up with a nut tag that is both printed and recorded as derived. One of those two facts is about evidence and the other is about provenance, and conflating them is how you lose a real warning rather than gain a false one.

Two evidence paths, in a deliberate order

The ingredient list is the first path. The second is the label's own allergen lines, parsed server side into allergen ids: the "Contains" line and the precautionary "May contain" advice, which stay separate because they are different claims.

A "contains" from either path beats a "may contain" for the same allergen. When both paths produce "contains", the ingredient path wins the reference, because it can name the word on the pack: "contains tree nuts: almond" sends you to a line you can go and read, where "the label says so" sends you nowhere.

There is no third level for "absent"

The engine never reports an allergen as absent. It cannot: the absence of a word in the data Munchable holds is not the absence of an ingredient in a factory. So the screens get three states and only two of them are findings:

/**
 * The three states a declared allergy can be in for one product. "Not
 * mentioned" is deliberately not a verdict tier: the data can never promise
 * absence, so it is quiet grey information, never green and never a claim.
 */
export type AllergyLevel = 'contains' | 'may-contain' | 'not-mentioned';
Enter fullscreen mode Exit fullscreen mode

"Not mentioned" draws in muted text with its own glyph, so the state never depends on colour alone, and it never borrows the green the condition verdicts use. A profile with an allergy on it always produces at least one group, which is what keeps the allergy block on screen in every state rather than only in the alarming one.

Where the layer sits in the verdict

Above every condition, and it only ever pushes down.

A declared allergen the data says is in the product makes the product an avoid, whatever the conditions made of it, and however little else of the label could be read. A result that says "can't assess" while listing peanuts is still a product to put back on the shelf. Precautionary advice keeps the product off green and goes no further, because "may contain" is the user's call and the notice puts the words in front of them to make it.

Allergens are also the one layer that never passes through the confidence cap that governs the condition tiers, which is what lets the previous sentence be true on a badly read label. We deleted a confidence ratio earlier this year that was turning a seventh of all scans into "can't assess"; that change could only be made calmly because the allergen path was never downstream of it.

One file owns every allergy word

The result screen and the menu screen both render allergy findings. How a finding is worded, grouped and coloured lives in one module for exactly that reason: they cannot drift apart on the one subject where they have to agree. The engine answers per allergen, that module turns it into the sentence a person reads ("Contains peanuts and milk"), and both screens draw the same thing.

There is also a standing notice on every result for a profile with an allergy set, in full, every time:

Munchable warns about the allergies you list when the label data names them. It cannot see every label, every recipe change or every factory, so never rely on it alone. Always read the pack. Lactose intolerance is not a milk allergy.

The last sentence is there because the app has a lactose intolerance condition and a milk allergy option, one screen apart in the same flow, and somebody was going to pick the wrong one.

Try it

The whole onboarding runs in the browser with no account: app.munchable.app walks you through region, conditions, allergies and shopping preferences, and your answers stay on the device, which is the other half of why the server never sees a condition. Pick Peanuts on the allergies step and read the description, then read the notice on the summary screen before you sign in to scan anything.

The seven conditions each have a public guide at munchable.app/conditions, with the evidence caveat for each one read out of the same rule set the app runs.

The rule in one line: an inferred tag is good enough to decide whether a biscuit suits your gut, and not good enough to tell somebody there are peanuts in it.

Top comments (0)