CogniPrep ships an llms.txt. It is the file you put at the root of a site so an
assistant that lands there gets a short, plain description of what the thing is
instead of guessing from the marketing page. Ours is about 14 KB and it reads
well. It is also wrong, and it has been wrong for months.
Open it and read the fifth line:
https://cogniprep.app/llms.txt
CogniPrep is a practice platform for candidates facing pre-hire assessments. It
covers 14 assessment providers with 85 browser-playable test simulations.
Now open the catalogue page in another tab:
The heading under the title says "All 56 are covered". There are 56 provider
cards, each one labelled with its own count. Paste this into the console on that
page and you get the total:
const n = document.body.innerText.match(/(\d+)\s+games?\b/g).map(Number);
console.log(n.length, n.reduce((a, b) => a + b, 0));
// 56 372
Fifty six providers, 372 test simulations. The file that exists specifically so a
machine does not have to guess is understating the product by 42 providers and
287 tests.
The file disagrees with itself too
The interesting part is not that it drifted. It is that nothing could have caught
the drift, because the number in the opening line is not derived from anything,
including the rest of the same file.
llms.txt has 18 ### headings. Four of them are pricing sections. The other 14
are providers, and each one carries its own count in its heading, like
### Arctic Shores (14 games). Add those 14 headings up:
const t = document.body.innerText; // on /llms.txt
const secs = [...t.matchAll(/^### (.+)$/gm)].map(m => m[1]);
const counts = secs.map(s => (s.match(/\((\d+)\s+(?:tests?|games?)\)/) || [])[1]).filter(Boolean).map(Number);
console.log(secs.length, counts.length, counts.reduce((a, b) => a + b, 0));
// 18 14 89
Eighty nine. The opening line says 85. The headline and the body of the same
file, both written by hand, were last true at different times.
The coincidence that makes this easy to miss: the headline says 14 providers and
there really are 14 provider sections. The document is internally consistent on
the only number a reader is likely to check by counting the headings, and wrong
on both of the numbers that matter.
How it got stale while being edited
Three games went into one provider in a commit from this week, and that commit touched llms.txt. The
Aon heading went from ### Aon / cut-e (6 tests) to ### Aon / cut-e (10 tests),
and a sentence listing the new material was added under it.
The opening line was not touched, because the opening line is 23 lines above the
heading being edited and nothing links the two. A diff that correctly updates the
section you are looking at is the most convincing kind of stale file: it has a
recent git log, so it looks maintained.
Every other surface derives this
The thing that makes the file an outlier is that it is the only public surface in
the repository that is written rather than generated.
The provider slugs live in one exported array. /games maps over it to render the
cards. The sitemap maps over the same array for /games/<provider> entries, which
is why the sitemap currently has 375 URLs of which exactly 56 are provider pages:
const xml = await (await fetch('/sitemap.xml')).text();
console.log(xml.match(/<loc>/g).length); // 375
console.log(xml.match(/\/games\/[a-z0-9-]+</g).length); // 56
The per-provider test counts come from per-provider catalogue modules, which is
where the 372 on /games comes from. Adding a provider updates the landing page,
the catalogue, the sitemap, the search index and the per-provider hub, with no
list to remember. There is a test that fails if the slug array and the metadata
object ever disagree on membership or order.
llms.txt is a static file in public/. It is reachable by exactly one code
path, which is Vercel serving a byte range off disk.
There used to be a generated companion, llms-full.txt, built from the same
arrays as the pages. That one could not go stale, and it is gone:
GET https://cogniprep.app/llms-full.txt -> 404
So the generated file was removed and the hand written one survived, which is
exactly the wrong way round.
Why I am not just correcting the numbers
Editing 85 to 372 fixes the file until the next provider ships, which on
current pace is days. The numbers in that opening line are the same two numbers
/games already computes at render time from the same source of truth. A route
that emits llms.txt can interpolate both and never be wrong again, and the
### Provider (N tests) headings underneath it are the same map over the same
array.
The lesson I keep relearning on this codebase is narrower than "generate your
content". It is that a number written by hand next to a number that is derived
will eventually contradict it, and the hand written one is always the one facing
the public. If a file exists to be read by a machine, it is strange to be the only
file a human has to remember to update.
Count the cards yourself on https://cogniprep.app/games and then read
https://cogniprep.app/llms.txt. One of them is generated.
Top comments (1)
We need to produce a casual YouTube comment. Must start with specific reaction or question about this video. So comment about mismatch counts. Use lowercase start. No quotes? Actually can use straight quotes. Avoid punctuation. Keep short, one or two sentences. No URLs. No marketing. Ok. Potential comment: "wait, why does the llms.txt still list only 14 providers when the site already shows 56? did anyone update the file?" That starts with lowercase "wait". That's fine. Ensure no em-dash.