DEV Community

Daniel Pertu
Daniel Pertu

Posted on

A word no language knows stays unknown: 135,692 lookups and no fuzzy matching

Munchable's rules engine cannot read. That is a design decision, and it is the most load bearing one in the product.

The engine takes canonical ids, the en:-prefixed tags of a food taxonomy, and decides what they mean for seven digestive conditions. It has no free-text parser anywhere in it. So when somebody photographs an ingredients panel in a supermarket, the text their phone read has to become canonical ids before anything can be stored or evaluated, and the thing that does that conversion is a lookup table.

Not a model. Not fuzzy matching. A table.

Why determinism beats accuracy here

The obvious objection is that a lookup table will miss words a model would catch. True. We took the table anyway, because of something that has nothing to do with per-word accuracy.

Two people in two cities photograph the same packet of biscuits. Their submissions have to be comparable, because comparing them is how the data becomes trustworthy without us having to trust either of them individually. That comparison only exists if the same label text always yields the same tags. The moment the mapping step is nondeterministic, "these two people read the same label" stops being computable, and the entire consensus mechanism underneath community data collapses into a pile of individually plausible rows.

Determinism is not an aesthetic preference in that architecture. It is the thing that makes the architecture possible.

The second reason is the one I would defend in any health product: a fuzzy matcher turns "I do not know this word" into a confident neighbour. For a trigger list that is unfortunate. For an allergen it is the single failure mode that must not exist. A word no language knows stays unknown.

The table, measured

The hand-written part is small and is only an English override layer. The bulk is generated from the taxonomy's own synonyms:

135,692  entries, keyed `<lang>|<normalized name>` -> canonical id
193      distinct language codes
4.7 MB   of JSON
Enter fullscreen mode Exit fullscreen mode

The language codes are a good lesson in what real reference data looks like. There are 193 of them and they are not all ISO 639. There is a cz and a dk (which are country codes wearing a language's coat), and one entry whose language is the capitalised string La. You do not get to assume your reference data is clean. You get to key your lookups on whatever it actually says, and make sure the fallback path means a messy key costs you a miss rather than a wrong answer.

Three tries, in this order

export function resolveGeneratedNameTiered(name: string, lang?: string): string | null
Enter fullscreen mode Exit fullscreen mode
  1. The label's own language. A Spanish label saying "nata" should be read as Spanish before anything else gets a vote.
  2. English. Second because the English long tail is the largest, and because a surprising share of labels print the English word regardless.
  3. Any language, first match wins. Last, because an ingredient name that only exists in Polish is still a real ingredient, and refusing it would be fastidiousness rather than safety.

The order is the whole policy. Each tier is strictly more desperate than the one above it, and nothing below tier one is allowed to override tier one.

Unknown words are recorded three times on purpose

The honesty rule in this module is that an unmapped token is never dropped. It survives in three places, and the redundancy is deliberate because three different consumers need it:

  • In the ingredient tag list, as a bare slug with no en: prefix. That is how the engine knows, structurally, that this one is not in the vocabulary. No special case, no side channel: the shape of the id carries the fact.
  • In an unknownTokens array, because the interface has to be able to show a person which words were not understood.
  • As a count, because confidence is computed from it.

A pipeline that silently discards what it could not parse will tell you it is working perfectly right up until the day someone asks it a question in Portuguese.

And unknowns are not a dead end. Words the table cannot map land in a backlog table and come back later as aliases through curation, which means the unknown list is a work queue with a measurable length rather than a loss. We wrote about the shape of that curated layer in our taxonomy knew 629 E-numbers and could not tell us one was a preservative.

The ceiling on cleverness

Normalization strips diacritics and strips plurals. That is the complete list of transformations we allow before the lookup.

It is tempting to add more. Stemming, a Levenshtein pass, splitting compound words, dropping parentheticals. Every one of those is a chance to merge two things that are not the same thing, and in a food vocabulary the near-misses are vicious: a sugar and a sugar alcohol, a milk protein and a milk-free substitute named after it, an oil and the hydrogenated version of that oil. The ones that matter most to our users are precisely the pairs that look alike.

So the rule is that cleverness goes in the curated layer where a human signs off on it, not in the hot path where it silently applies to everything forever.

See the vocabulary on a live page

The alias sets are not hidden. Every ingredient answer page has a section called "What it is called on a label", which is the lookup table pointed at the reader:

If you want to watch the unknown path instead, photograph something obscure in the app and read the confidence line it gives you. It is telling you the size of the gap, which is the only honest thing a parser can do about a word it has never seen.

Top comments (0)