Munchable scans a packaged food and tells you whether it fits your gut condition. The conditions are at munchable.app/conditions, and 373 pages of its ingredient reasoning, each one generated by running the engine, are at munchable.app/answers.
The engine can return "not assessed" for an ingredient. That is the honest answer for the long tail of a catalogue with a million products in it, and it is an unacceptable answer for kefir when the user told us they are lactose intolerant.
Each of these came back "not assessed" once, in a build that passed every test:
- kefir, for lactose.
- soya, for IBS.
- tea, for reflux.
Not obscure. Not edge cases. The exact ingredients those users opened the app for. Every unit test was green, because unit tests check that the code does what the code says, and nothing in the suite was checking that the data said anything about kefir.
So the engine now carries a promise list, and the suite fails the build when a promise is broken.
Words, not ids
The promise list is per condition and per allergen, and it is written in label words rather than in internal ids:
export const CONDITION_COVERAGE = {
lactose: [
'milk', 'cream', 'ice-cream', 'yogurt', 'kefir', 'custard', 'whey', 'ghee',
// ... and the rest of what a lactose-intolerant shopper actually meets
],
// ...
} as const;
That choice is doing most of the work. A test written against ids asserts that the rule table has a key, which is the part that was never broken. The real failures live in the layer that gets a printed word onto an id: canonicalisation, language prefixes, plurals, variant spellings, the parent graph. So each promised word is walked exactly as a label word is walked:
const id = canonicalizeTag(`en:${word}`);
if (!isKnownTag(id)) failures.push(`${word} → ${id} is not a taxonomy id`);
else if (!reaches(map, id)) failures.push(`${word} → ${id} reaches no ${map} rule`);
reaches checks the id and its ancestors, because an answer arrived at through a parent is a perfectly good answer. The point is that the test enters the system where the user does.
One more constraint, which is the one that makes the test worth having at all: the promise list may not depend on a job having run. Our curation pipeline widens coverage continuously and publishes an overlay, so coverage measured against a live database tells you what was true this morning. These assertions run against the data in the repo, which means a green suite is a floor that ships with the binary rather than a snapshot of a database.
The half that asserts silence
The other half of the file asserts that the engine says nothing, and I think it is the more valuable half.
Munchable has a lifestyle filter next to the condition verdicts: switch on "no artificial colours" or "no sweeteners" and the app names them when it sees them. The failure mode for a filter like that is not missing something. It is firing on everything, because a badge that appears on every product in the shop teaches the user to ignore all badges, including the allergen warning two rows above it.
Silence is therefore a feature with tests on it:
test('healthy shopping stays quiet about vitamins and mineral salts at any setting', () => {
assert.deepEqual(quietUnder(HEALTH_SILENT_COVERAGE, HEALTH_PREFERENCE_IDS), []);
});
HEALTH_PREFERENCE_IDS is every preference the product has, switched on at once. The assertion is that a maximally twitchy profile still says nothing about vitamin C or a mineral salt, because those are in half the aisle and naming them is pure noise.
Then this one, which is a product decision with a guard rail bolted to it:
test('the commonest emulsifiers never fire for a default profile', () => {
// Lecithin is on 10,604 products and xanthan on 4,274. This is the test that
// fails if `emulsifiers` is ever moved into the default set.
assert.deepEqual(quietUnder(HEALTH_DEFAULT_SILENT_COVERAGE, HEALTH_PREFERENCE_DEFAULTS), []);
});
Switching the filter on writes a default set of preferences, and emulsifiers are deliberately not in it. That decision lives in one line of data and would be entirely reasonable to change on a Tuesday afternoon, which is exactly why the comment names the number of products it would affect. The test is not protecting an invariant of the code. It is protecting a judgement call from being quietly reversed by someone who did not know lecithin is in ten thousand products, and the diff that breaks it gets the explanation in the failure.
And immediately after it, the mirror:
test('but the emulsifiers preference names them when it is switched on', () => {
// ...
assert.equal(named.length, HEALTH_DEFAULT_SILENT_COVERAGE.length,
'a switch that promises the whole class has to deliver it');
});
Those two tests have to be read as a pair. Off by default, complete when on. Without the second one, the easiest way to make the first one pass forever is to make the preference do nothing at all, and a test suite that rewards that is worse than no suite. Every assertion of silence needs a paired assertion that the thing can still speak, or you have written a test that is happiest when the feature is broken.
The allergen half works the same way, with a stricter bar: every promised label word has to raise a contains level warning, not merely something. A may contain where a contains belongs is not a near miss in a product people use to avoid an allergen.
Deliberately not "full coverage"
The promise list is the common cases only, and it is supposed to stay that way. There is no version of this file that covers a million products' worth of ingredients, and chasing one would mean asserting things about words nobody has ever scanned while the shortfall sits somewhere real.
Which is why there is a companion audit script rather than a bigger test. It walks the same candidate lists, reports for each one the id it lands on, which rule map reaches it, whether the allergen layer fires, and whether the result screen would say "not assessed", and marks the misses GAP. The test is the floor nobody may break; the audit is how you decide where the next curation effort goes. Measuring coverage against real products before spending a model's budget on it is the difference between widening answers and widening a database.
Check the promises yourself
The three ingredients at the top of this post each have a public page, produced by the same engine the app runs:
If one of them ever says nothing useful, that is a promise test that should have existed. The full index is at munchable.app/answers, and the aisle version, where the promise is kept against a real pack with twenty ingredients on it, is at app.munchable.app.
Top comments (0)