DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our 373 answer pages contain no written answers, because each one runs the rules engine twice at build time

People do not search for "gut health app". They search for "is onion low FODMAP", and they want the word no above the fold.

Munchable has a rules engine that answers exactly that question, because that is what the app does when you scan a barcode. So the site carries one static page per ingredient and condition pair the engine has something to say about. There are currently 373 of them, each at the site root, with the question itself as the URL:

The interesting part is not the page count. It is that nobody wrote the answers.

The rule

Every verdict on those pages comes out of the engine. The page does not restate the rules in prose that could drift from them; it builds a one-ingredient product and runs the same fitCheck the app runs, then prints what came back.

function probe(tag: string, condition: ConditionId, asMainIngredient: boolean): Probe {
  const ingredientsTags = asMainIngredient
    ? [tag, ...INERT_FILLER]
    : [...INERT_FILLER, tag];

  const result = fitCheck(
    { barcode: '0000000000000', ingredientsTags, statesTags: ['en:ingredients-completed'] },
    { conditions: [condition] },
  );

  const per = result.perCondition[0];
  return { verdict: per?.verdict ?? 'unknown', reason: per?.reasons[0]?.message };
}
Enter fullscreen mode Exit fullscreen mode

It runs twice per page, and the second argument is why. Position on an ingredients list carries information: a legal ingredients list is written heaviest first, so an ingredient at the top is most of the product and the same ingredient near the end is a trace. The engine reads that, which means a single ingredient does not have one answer, it has two. So the page probes it at the front of a list and at the back, and if the two disagree it says so in a sentence instead of picking a side.

That is also where the line on is-olive-oil-ok-with-bam comes from: fat-dense as a main ingredient is an avoid, further down the list a caution. Nobody typed that distinction on that page. It is two function calls and a comparison.

The padding has to be inert, and once it was not

The probe product needs filler so the ingredient under test can sit in the tail of a list rather than being the whole thing. Filler is load-bearing in a way that is easy to miss: anything the rules have an opinion about will contribute reasons of its own, and then the page is answering a question about the padding.

Ours held sunflower oil for a while. It was fine, right up until the bile acid malabsorption rule set learned to read fat-dense ingredients. At that point every single BAM probe came back about the oil, and the page under test had nothing to do with the verdict printed on it. The fix was one line of data and a shouty comment:

// NO FATS OR OILS HERE. This list held `en:sunflower-oil` until the BAM rule set
// learned to read fat-dense ingredients, at which point the filler started
// answering the question.
const INERT_FILLER = ['en:water', 'en:salt', 'en:rice-flour', 'en:rice', 'en:citric-acid'];
Enter fullscreen mode Exit fullscreen mode

The general version of this bug: if you probe a system with a synthetic input, the inert parts of that input are an assumption, and assumptions about a system that is still being developed expire without telling you.

Half the pages would have printed the opposite of the verdict

The engine returns one of four verdicts. The page has to turn that into a heading that answers an English question, and the polarity of the question is not constant:

Condition The question A flagged verdict means
IBS Is onion low FODMAP? No
Lactose Is cheddar high in lactose? Yes
Reflux Does coffee cause reflux? Yes

Derive the heading from the verdict alone and you print the wrong word on every page where the question is phrased in the direction of the problem rather than the direction of safety. So each condition carries a flagMeansYes boolean next to its phrasing, and one table decides the slug, the h1, the title tag and the FAQ question together, which is the only reason a page's URL and the question it asks cannot drift apart.

ibs: {
  slug: (i) => `is-${i}-low-fodmap`,
  ask: (n) => `Is ${n} low FODMAP?`,
  flagMeansYes: false,
  caution: 'It depends on how much',
},
gerd: {
  slug: (i) => `does-${i}-cause-reflux`,
  ask: (n) => `Does ${n} cause reflux?`,
  flagMeansYes: true,
  caution: 'For a lot of people, yes',
},
Enter fullscreen mode Exit fullscreen mode

The caution string is spelled out per condition rather than shared, because the engine raises that verdict for genuinely different reasons. A moderate lactose load is a question of how much; a reflux trigger is a question of who you are. "It depends on how much" is right for the first and useless for the second, which is why the coffee page opens with "For a lot of people, yes" and then explains that the evidence for trigger foods is low certainty, so the app raises a caution and never an avoid on that axis.

One condition went from zero pages to the most pages

The pages are generated, so their distribution is a fact about the rule sets rather than an editorial choice. Today:

Condition Answer pages
Bile acid malabsorption 67
IBS and low FODMAP 64
SIBO 64
Lactose intolerance 52
IBD 48
GERD and reflux 44
Gastroparesis 34

Bile acid malabsorption is at the top of that table and used to generate nothing at all. It reads fat, the axis keys on the fat figure printed on the label, and a one-ingredient probe has no nutriment panel, so every BAM probe returned unknown and the generator correctly produced no pages. It only started answering once the rule set gained an ingredient-level view of fat density. A page family appearing as a side effect of a rules change, with no content work at all, is the whole return on this design.

The inverse is also enforced: a pair the rules have nothing to say about gets no page. "Onion is not relevant to bile acid malabsorption" is thin content and nobody searched for it.

Three routing hazards that come with a root dynamic segment

These pages live at app/[question]/page.tsx, a dynamic segment at the site root. That is the right URL shape, because the question is the page and a prefix would make the slug worse, but it has sharp edges.

A root segment answers everything. Without dynamicParams = false, that route catches every unmatched top-level path, and every typo becomes a soft 404 that a crawler happily indexes. With it, only the generated slugs exist and the rest is a real 404.

export const dynamic = 'error';
export const dynamicParams = false;
Enter fullscreen mode Exit fullscreen mode

A generated slug could collide with a real route. Next resolves static routes before dynamic ones, so if a slug ever came out as privacy or answers, the real page would win and the answer page would vanish silently. There is a reserved-path set checked at module load, which throws the build rather than losing a page. No slug can collide today, since all of them start is- or does-, and that is exactly the kind of fact that quietly stops being true.

The key is attacker-controlled, so the lookup is a Map. An object literal keyed by URL segment answers lookup['__proto__'] with something inherited from Object.prototype, which sails through an if (!page) return notFound() guard and then throws further down. A Map returns undefined like it should.

160 characters, and two vocabularies for the same verdict

The page heading can afford to say "is flagged as an avoid". A search snippet cannot: at the longest ingredient names, the page's own answer sentence is over the limit before it has said anything about what to do next. So there are two phrasings of each verdict, one for the page and a compressed one for the meta description, plus a fixed closing call to action sized to what is left once the worst case has been paid for.

const ANSWER_LEAD      = { avoid: 'is flagged as an avoid', /* ... */ };
const META_ANSWER_LEAD = { avoid: 'is an avoid',            /* ... */ };
const META_CLOSE = ' Scan a barcode to check a real product.';
Enter fullscreen mode Exit fullscreen mode

A test asserts the longest possible description still fits. The snippet has to deliver the answer before the click, and a snippet cut off mid-verdict delivers the opposite of one.

None of this works if nothing links to it

373 pages that only appear in a sitemap get crawled slowly and ranked badly, so the internal links are generated from the same data:

  • The index groups the questions by condition, because a reader has one condition and that group is their shortlist.
  • Each condition guide lists the ingredients its own rule set flags, which is a query against the engine rather than a hand-kept list.
  • Each answer page links to its siblings: is onion OK for SIBO and does onion cause reflux are both one click from the onion page.

Nothing in that graph is maintained by hand, which matters more than it sounds: hand-maintained hub pages are the first thing to rot when a rule set grows, and a rotted hub is indistinguishable from a page that was never built.

The transferable bit

If your content is downstream of a program, generate it from the program. The test is simple: change a rule, and ask how many pages a human now has to go and edit. If the answer is not zero, the pages and the program will disagree eventually, and your readers will find the disagreement before you do.

Munchable is at munchable.app. All 373 answer pages are static and public, no account needed.

Top comments (0)