DEV Community

Daniel Pertu
Daniel Pertu

Posted on

We deleted the list of retailers from our label parser, because a list only ever saves the first sighting

Munchable reads an ingredients list and scores every word in it against your gut condition. Which means that before it can score anything, it has to answer a duller question for every token on the label: is this an ingredient at all?

A lot of what is printed on a pack is not. There is the brand, the retailer, a slogan, the allergen prose, a storage instruction, the address of a factory, and whatever the optical character recognition made of a crease in the plastic. None of it is food, and all of it arrives in the same array as the food.

The word that forced a decision was a supermarket's own name, printed inside the ingredients list of its own-brand product. A word like that is not an ingredient, cannot be scored, and if it is left unresolved it drags the confidence of the whole scan down, which means a user gets a hedge instead of an answer because of a logo.

The obvious fix, and why we removed it

The obvious fix is a list of retailers and brands in the parser. We had one. It is gone, and the comment where it used to be says why:

// There is deliberately NO hardcoded list of brands or
// retailers. A brand name the layers cannot place goes to the
// model once, the model files it as noise, and the overlay
// answers it for free from then on: the database learns the
// name permanently, where a list in here only saved the first
// sighting and had to be edited and deployed to grow.
Enter fullscreen mode Exit fullscreen mode

Three things are wrong with the list, in increasing order of how much they cost.

It is unbounded. Every supermarket chain in every country we read labels from, plus every house brand under each of them, plus the spelling each one uses on packaging rather than on its shopfront.

It only saves the first sighting. The second time the same word appears, the list helps exactly as much as it did the first time, which sounds like the point until you notice that the alternative saves every sighting after the first.

And it is code, so it grows by a pull request and a deploy. The words we need are discovered by users scanning real packets, at the rate real shopping happens, and a vocabulary that grows at the speed of releases will always be behind a vocabulary that grows at the speed of scans.

What it does instead

Unknown words go through a stack of layers, first one to answer wins. The engine's own knowledge, then glued-on headers, then curated boilerplate patterns, then E numbers, then a multilingual synonym table, then plurals, then preparation and origin modifiers, then single-letter optical character recognition repair, then English compound splitting. We measured that stack before paying for anything, and most of the backlog never needed a model.

A brand name reaches the bottom of the stack, because no amount of string handling can tell you that a word is a company. So it costs exactly one model call, which files it as "not an ingredient", and that decision is written into the taxonomy overlay.

The test for this behaviour is the shortest useful test in the repository:

test('a brand name is free for ever after the model files it once', () => {
  resetTaxonomyOverlay();
  // Nothing knows the word: it goes to the model, which is the one cost.
  assert.equal(resolveDeterministically('en:tesco', 'en'), null);

  // What the model writing "not an ingredient" puts in the overlay.
  const res = loadTaxonomyOverlay({ version: 'brand-test', aliases: {}, extensions: {}, noise: ['tesco'] });
  assert.equal(res.accepted, 1, JSON.stringify(res.rejected));

  // From now on the free layers answer it, and in any label language.
  const after = resolveDeterministically('en:tesco', 'en');
  assert.equal(after?.layer, 'overlay');
  assert.equal(after?.proposal.action, 'noise');
  assert.equal(resolveDeterministically('pl:tesco', 'pl')?.layer, 'overlay');
  resetTaxonomyOverlay();
});
Enter fullscreen mode Exit fullscreen mode

The last assertion before the reset is the one worth reading twice. The overlay is keyed on the bare word rather than on the language-prefixed tag, so a brand filed once from an English label is also answered on a Polish one. Brand names are the one part of a label that does not translate, and the data structure happens to agree.

The deleting of the list only holds up if that learning actually happens, which is why the test exists at all. Remove the overlay step and every label carrying a retailer's name pays a model call forever, and nothing else in the suite would have noticed.

Where the learned decision goes

Into the overlay, which the phone fetches without an app release, versioned by nothing more clever than the latest update timestamp of the curated data. We wrote about that delivery path separately. The practical effect is that one person scanning an own-brand tin of beans in one country makes that word free for everybody, on a device that has not been updated, within a day.

The cost model inverts. A hardcoded list is a fixed asset that depreciates: it covers what it covered on the day it was written. A learned row is bought once at the price of a single model call and then pays out on every future scan of every product carrying that word, in every language, for every user.

The clever bit that had to be made less clever

Not every brand name sits on its own. Plenty of them sit in front of a word the engine knows: a drink brand in front of "flavouring", an oat brand in front of "rolled oats".

The tempting move is to throw the prefix away and alias the whole phrase to the head noun. It is almost always right, and when it is wrong it is quietly wrong:

// A brand or an unknown qualifier in front of a known head ("vimto
// flavouring", "quaker rolled oats"). Aliasing to the head throws the
// prefix away, so it is the LAST resort: keep it and carry on looking
// for a split whose prefix resolves. Without this, "sweet potato
// flour" takes the first split ("potato flour" with an unknown
// "sweet") and is filed as potato flour, when the split one word
// later reads the prefix as sweet potato.
Enter fullscreen mode Exit fullscreen mode

Sweet potato is not potato, and for somebody eating low FODMAP that difference is the entire answer. So the brand-shaped case is held as a fallback and only used once every other split has failed, which costs a few more string operations and buys the thing a parser is for.

Underneath all of it there is a guard that cannot be waived by any layer: a proposal carries the source phrase, and the scan that looks for hidden allergens and triggers sees every word of it. "Fully refined soya bean oil" can never be filed as an alias to oil, however neatly the compound splitter handles it, because the guard finds the soya and refuses to let it drop out of the rule graph. Convenience never gets to remove a word that a rule might need.

What this looks like from outside

You cannot see a brand decision directly, which is rather the point: the work it does is the absence of a shrug in a scan result.

What you can see is the vocabulary it protects. Every ingredient the engine can place has a public page that answers one question by running the production rules, like is onion low FODMAP or is cream low FODMAP, and a few hundred are indexed at munchable.app/answers. A word the resolver cannot place has no page, no score, and no business being on that list.

The app is at app.munchable.app, where onboarding runs before you create an account, and the seven condition rule sets each publish their basis at munchable.app/conditions.

Top comments (0)