Munchable's rules engine cannot read. That is a design decision, and it is the most load bearing one in the product.
The engine takes canonical ids, the en:-prefixed tags of a food taxonomy, and decides what they mean for seven digestive conditions. It has no free-text parser anywhere in it. So when somebody photographs an ingredients panel in a supermarket, the text their phone read has to become canonical ids before anything can be stored or evaluated, and the thing that does that conversion is a lookup table.
Not a model. Not fuzzy matching. A table.
Why determinism beats accuracy here
The obvious objection is that a lookup table will miss words a model would catch. True. We took the table anyway, because of something that has nothing to do with per-word accuracy.
Two people in two cities photograph the same packet of biscuits. Their submissions have to be comparable, because comparing them is how the data becomes trustworthy without us having to trust either of them individually. That comparison only exists if the same label text always yields the same tags. The moment the mapping step is nondeterministic, "these two people read the same label" stops being computable, and the entire consensus mechanism underneath community data collapses into a pile of individually plausible rows.
Determinism is not an aesthetic preference in that architecture. It is the thing that makes the architecture possible.
The second reason is the one I would defend in any health product: a fuzzy matcher turns "I do not know this word" into a confident neighbour. For a trigger list that is unfortunate. For an allergen it is the single failure mode that must not exist. A word no language knows stays unknown.
The table, measured
The hand-written part is small and is only an English override layer. The bulk is generated from the taxonomy's own synonyms:
135,692 entries, keyed `<lang>|<normalized name>` -> canonical id
193 distinct language codes
4.7 MB of JSON
The language codes are a good lesson in what real reference data looks like. There are 193 of them and they are not all ISO 639. There is a cz and a dk (which are country codes wearing a language's coat), and one entry whose language is the capitalised string La. You do not get to assume your reference data is clean. You get to key your lookups on whatever it actually says, and make sure the fallback path means a messy key costs you a miss rather than a wrong answer.
Three tries, in this order
export function resolveGeneratedNameTiered(name: string, lang?: string): string | null
- The label's own language. A Spanish label saying "nata" should be read as Spanish before anything else gets a vote.
- English. Second because the English long tail is the largest, and because a surprising share of labels print the English word regardless.
- Any language, first match wins. Last, because an ingredient name that only exists in Polish is still a real ingredient, and refusing it would be fastidiousness rather than safety.
The order is the whole policy. Each tier is strictly more desperate than the one above it, and nothing below tier one is allowed to override tier one.
Unknown words are recorded three times on purpose
The honesty rule in this module is that an unmapped token is never dropped. It survives in three places, and the redundancy is deliberate because three different consumers need it:
- In the ingredient tag list, as a bare slug with no
en:prefix. That is how the engine knows, structurally, that this one is not in the vocabulary. No special case, no side channel: the shape of the id carries the fact. - In an
unknownTokensarray, because the interface has to be able to show a person which words were not understood. - As a count, because confidence is computed from it.
A pipeline that silently discards what it could not parse will tell you it is working perfectly right up until the day someone asks it a question in Portuguese.
And unknowns are not a dead end. Words the table cannot map land in a backlog table and come back later as aliases through curation, which means the unknown list is a work queue with a measurable length rather than a loss. We wrote about the shape of that curated layer in our taxonomy knew 629 E-numbers and could not tell us one was a preservative.
The ceiling on cleverness
Normalization strips diacritics and strips plurals. That is the complete list of transformations we allow before the lookup.
It is tempting to add more. Stemming, a Levenshtein pass, splitting compound words, dropping parentheticals. Every one of those is a chance to merge two things that are not the same thing, and in a food vocabulary the near-misses are vicious: a sugar and a sugar alcohol, a milk protein and a milk-free substitute named after it, an oil and the hydrogenated version of that oil. The ones that matter most to our users are precisely the pairs that look alike.
So the rule is that cleverness goes in the curated layer where a human signs off on it, not in the hot path where it silently applies to everything forever.
See the vocabulary on a live page
The alias sets are not hidden. Every ingredient answer page has a section called "What it is called on a label", which is the lookup table pointed at the reader:
- Does onion cause reflux? says: "Onion is also printed as onions. Munchable recognises all of these as the same ingredient, which is the point of scanning rather than reading: the wording changes between brands and the rule does not."
- Is apple low FODMAP? is the same structure on a different condition.
- The full index is a few hundred of these, and each one is generated by running the real engine rather than by writing prose, which is a separate story: 373 static pages that answer by running the scanner.
If you want to watch the unknown path instead, photograph something obscure in the app and read the confidence line it gives you. It is telling you the size of the gap, which is the only honest thing a parser can do about a word it has never seen.
Top comments (0)