DEV Community

Cover image for Google said llms.txt does nothing. I checked what 137,210 domains actually got.
Juan Camilo Auriti
Juan Camilo Auriti

Posted on

Google said llms.txt does nothing. I checked what 137,210 domains actually got.

The llms.txt file was supposed to be the bridge between your website and AI search engines. A markdown file at your root, a curated list of your best pages, a clean map for models that struggle with bloated HTML. Add it, and ChatGPT, Perplexity, and Google AI Overviews would finally find you.

I wanted to believe it. I built an open-source GEO audit tool that checks for llms.txt among its signals. Then the data came in, and it doesn't say what the checklist articles say.

What Google said

In May 2026, Google published its official guide on optimizing for generative AI features in Search. The guide has a mythbusting section. This is what it says about llms.txt and similar files:

You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in generative AI search.

Google's John Mueller had been saying it for over a year before it landed in the documentation. In April 2025, answering a question on Reddit, he put it plainly:

"AFAIK none of the AI services have said they're using LLMs.TXT (and you can tell when you look at your server logs that they don't even check for it). To me, it's comparable to the keywords meta tag — this is what a site-owner claims their site is about … (Is the site really like that? well, you can check it. At that point, why not just check the site directly?)"

The keywords meta tag. The canonical example of a signal search engines learned to ignore because every site claimed to be the best one.

What Google did next

Ten days before the Search guide told everyone to stop bothering, Google's Chrome team shipped an llms.txt check inside Lighthouse as part of its experimental Agentic Browsing audits. The documentation, last updated May 5, 2026, says:

Without this file, agents may spend more time crawling the site to understand its high-level structure and primary content.

Two teams inside the same company, pointing in opposite directions. When pressed on the contradiction, Mueller framed it as: llms.txt is "not done for search." He called it a "temporary crutch, perhaps to save some tokens" for AI coding tools parsing developer documentation. Not something non-developer sites need.

So which is it? I went to the data.

Study 1: Ahrefs, 137,210 domains

Ahrefs published the largest access study to date on June 15, 2026. They looked at every domain in Ahrefs Web Analytics that received traffic in May 2026 and checked each root for an llms.txt file returning HTTP 200.

The numbers:

  • 28% of the 137,210 domains publish an llms.txt file (about 38,000 sites).
  • 97% of those files received zero traffic in May 2026. Nothing fetched them at all. Not AI bots, not generic crawlers, not humans.
  • 96% of the requests that did reach llms.txt files came from bots.
  • 19.5% of those bot fetches came from named AI tools. GPTBot was first, Claude-Code second. Both are training crawlers and coding agents, not answer engines.
  • Zero requests came from AI bots for llms.txt files that didn't exist. AI bots never go looking.

That last point is the one that kills the mental model where a model "checks for" your file before answering a question about you. It doesn't. It fetches whatever URLs its retrieval step already surfaced, which is normally your HTML.

Ahrefs also found something unsettling: the largest single research crawler in the dataset identified itself as prompt-injection-survey/1.0. Someone is systematically studying llms.txt as a prompt injection vector, because AI agents are designed to ingest and trust whatever the file says.

Study 2: SE Ranking, 300,000 domains

SE Ranking ran a separate analysis of nearly 300,000 domains. They tested whether having an llms.txt file correlates with how often a domain gets cited by AI systems.

It doesn't.

They used both statistical correlation (Spearman) and machine learning (XGBoost with SHAP analysis). The result: removing the llms.txt variable from the model improved its prediction accuracy. The file wasn't a weak signal. It was noise.

Adoption was 10.13% across their dataset, roughly 1 in 10 sites. No major AI platform has publicly committed to reading it.

Study 3: OtterlyAI, 90-day server log experiment

OtterlyAI ran a controlled experiment over 90 days. They placed a correctly implemented llms.txt file at the root of a test website and monitored AI bot traffic.

  • 62,100+ total AI bot visits to the site over 90 days.
  • 84 of those visits targeted /llms.txt.
  • That's 0.1% of AI bot traffic.
  • The site's average content page received ~265 AI bot visits in the same window. The dedicated AI file performed 3x worse than a random page.

They also noted that llms.txt "hardly performed better than an average .pdf file." AI crawlers don't treat it as privileged content.

OtterlyAI subsequently removed the llms.txt checker from their GEO audit tool, because the data showed its impact on AI crawler behavior is "marginal at best."

Study 4: Limy, 515 million bot events

Limy's analysis is the largest dataset I've seen on this question. They monitored 515,382,577 AI bot traffic events over a 90-day window across the brands they track.

  • 408 requests targeted /llms.txt directly.
  • Out of 515 million.
  • That's 0.00008% of total AI crawler traffic.

Their word for the share was "statistically negligible." The bots that matter for AI search visibility (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended) overwhelmingly skip the file and crawl HTML directly.

What actually works (according to the same studies)

The studies don't say "do nothing." They say "do something different."

Google's own guide is clear about what matters: create valuable, non-commodity content. If an AI can write your article on its own, it will never cite you. It already has the answer. The guide points to first-hand experience, unique data, and original analysis as the signals that make content worth surfacing.

The Ahrefs study points out that the bots fetching llms.txt are coding agents reading developer documentation. If your site is API docs or SDK reference, a clean llms.txt genuinely helps. If your site is a blog, an ecommerce store, or a landing page, no agent is parsing your pages for fun. The file points at an empty room.

What moves the needle, according to the peer-reviewed GEO paper by Aggarwal et al. (KDD 2024), is content-level edits: adding citations to credible sources, adding statistics, adding quotations. Their benchmark showed these methods can boost visibility in generative engine responses by up to 40%. That is an effect size. llms.txt has no measured effect size.

What I did with my own tool

I maintain geo-optimizer-skill, an open-source Answer Engine Optimization audit tool with over 1,000 stars on GitHub. It checks for llms.txt among its signals, and I'm keeping the check.

But I changed what the check means. It doesn't score your site higher for having the file. It flags the file as present and moves on. The signals that actually affect your score are the ones with evidence behind them: crawlability, semantic structure, content density, citation-worthy passages, and whether your pages render for bots that don't execute JavaScript.

llms.txt is a 20-minute, zero-risk bet. Ship it if you want. But don't put it on your roadmap as a ranking play, don't pay anyone to "optimize" it, and don't let it displace the work that actually has evidence behind it.

The bottom line

Four independent studies covering 137,210 domains, 300,000 domains, 62,100 bot visits, and 515 million bot events all say the same thing. Google's official documentation says the same thing. John Mueller said the same thing over a year ago.

llms.txt is not an AI visibility tactic. It's a developer documentation tool that got mis-sold as an SEO tool. The disappointment around it traces back to that single category error.

If someone is selling you llms.txt as an AI-visibility tactic, the evidence isn't on their side.

Top comments (0)