A joke file about office cats has become one of the clearest tests of how weak evidence can spread through generative engine optimization.
In an August 7 article for Search Engine Journal, Mark Williams Cook describes creating cats.txt, a fictional web standard for declaring office cats, their roles, breeds, and even a "PurrLevel." He then tested the file against four types of evidence commonly presented in support of llms.txt.
The result was deliberately absurd. AI crawlers fetched cats.txt. Google indexed it. AI systems retrieved information that existed only in the file. ChatGPT could even be prompted into explaining why the invented standard might help visibility. The point was not that llms.txt can never matter. The point was that these observations do not prove the causal claim being made.
That distinction should influence every GEO strategy a marketer buys or recommends.
Crawling is evidence of access, not impact
The first weak inference is simple: an AI crawler requested a file, therefore the file must be used as a ranking or citation signal.
Crawling only proves that a bot accessed the URL. A crawler can fetch material that is ignored later, stored for another purpose, or never used in a live answer. Williams Cook's cats.txt experiment demonstrated this by attracting requests from several crawlers to a file whose standard had been invented as a joke.
This matters because crawler logs are valuable but easy to overinterpret. A marketer can use logs to verify whether AI agents can access a site, which is why an AI crawler access checker is useful as a technical diagnostic. Access is a prerequisite for some forms of retrieval. It is not proof that the content receives extra weight or improves brand visibility.
The same reasoning applies to robots rules, sitemap discovery, and other technical signals. They help answer whether a system can reach content. They do not by themselves answer whether the content will be selected, trusted, cited, or recommended.
Indexing is not proof of importance
The second weak inference is that Google indexing a file demonstrates that the file is meaningful to search or AI systems. Search engines index many types of accessible text. Indexing means the URL and its contents were discovered and stored in some form. It does not establish that a proposed web standard has special status.
cats.txt was indexable too. That is exactly why the experiment works as a counterexample. If a nonsense standard can satisfy the same test, the test cannot distinguish a useful standard from an invented one.
This is a basic evidence principle that marketers should apply well beyond llms.txt. Before treating an observation as proof, ask whether the same observation could happen if the tactic had no special effect. If the answer is yes, the observation is weak evidence.
Retrieval can happen without special treatment
The strongest sounding claim in the debate is that an AI system can return information that exists only inside an llms.txt file. That can look like direct proof that the system understands and honors the standard.
The problem is that ordinary web retrieval can produce the same outcome. If a text file is indexed or otherwise discoverable, a search backed AI system can retrieve it as a normal URL and use the text in an answer. The system does not need to recognize a special standard for that to happen.
Williams Cook shows the same behavior with cats.txt. AI systems could surface fictional cat details because the information was available on the open web. The file worked as retrievable content, not necessarily as a recognized protocol.
For marketers, this should change how tests are designed. A useful GEO experiment needs a comparison group or another way to separate ordinary retrieval from the claimed special mechanism. Simply observing that the information appeared in an answer is not enough.
ChatGPT agreeing with a tactic is not validation
The funniest part of the experiment is also the most important. When enough web content described cats.txt as useful, ChatGPT could produce a confident explanation of why it might help search and AI visibility.
That illustrates a major problem in AI marketing research. Asking a language model whether a tactic works is not equivalent to testing the tactic. The model may summarize the prevailing discussion around the tactic, including unsupported claims that have been repeated often.
Marketers need an evidence standard that sits above model confidence. The AI search visibility question should be measured through observable outcomes such as inclusion across a defined prompt set, citations, competitive share of voice, qualified traffic where available, branded demand, and ultimately business results. Even those measures need careful experimental design because AI answers can vary between runs.
llms.txt can still be a low cost experiment
The cats.txt argument does not require a dramatic conclusion that nobody should ever publish llms.txt. Search Engine Journal's article explicitly says the issue is the evidence used to sell the tactic, not a claim that the file can never have future value.
That is a useful position for small businesses. An llms.txt generator can make experimentation inexpensive. A company may decide to publish the file because the implementation cost is low and future support is possible. The mistake would be presenting that implementation as a proven way to improve ChatGPT, Gemini, or Perplexity visibility without evidence showing that effect.
This is an important commercial distinction. "Low cost preparation for a possible standard" is a defensible recommendation. "Proven AI ranking lever" requires proof.
GEO needs better experiments, not more rituals
Generative search is moving quickly enough that marketers will continue encountering tactics before strong evidence exists. That is normal in an emerging field. The professional response is to label uncertainty clearly and test claims in ways that could prove them wrong.
For any new GEO tactic, ask what mechanism is being claimed, what outcome should change if the mechanism is real, what else could explain the observation, and what comparison would separate those explanations.
The cats.txt experiment is memorable because it turns a methodological problem into a joke. The lesson is serious. If an invented cat standard can pass the same test used to justify a paid marketing tactic, the test needs to improve before the invoice does.
Originally published on the Mustard Seed blog.
Top comments (0)