DEV Community

Cover image for Google's mythbusting guide says stop obsessing over Schema. The data agrees.
Juan Camilo Auriti
Juan Camilo Auriti

Posted on

Google's mythbusting guide says stop obsessing over Schema. The data agrees.

When Google published its official guide on optimizing for generative AI features in Search in May 2026, it included a mythbusting section. The kind of section that upsets people who sell the things it busts.

Here is what Google explicitly lists under things you do not need:

You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in generative AI search.

That covers llms.txt. But the guide goes further. It says standard technical SEO is what counts for AI features, and that no special schema, no special markup, and no special files are required beyond what Google has always recommended.

The GEO agency playbook, meanwhile, still leads with three items: add llms.txt, add more Schema.org markup, and implement "content chunking" for AI retrieval. Let me check each one against the data.

llms.txt: the file 97% of sites publish for nothing

I covered this in detail in the first part of this series. The short version: Ahrefs analyzed 137,210 domains and found that 97% of llms.txt files received zero requests in May 2026. SE Ranking analyzed 300,000 domains and found no correlation between having the file and being cited by AI. Limy tracked 515 million AI bot events and counted 408 requests to llms.txt. Removing the variable from their machine learning model improved prediction accuracy.

Google's John Mueller compared llms.txt to the keywords meta tag. The keywords meta tag is the canonical example of a self-declared signal that search engines learned to ignore because every site claimed to be the best one. The same logic applies here: a file where you describe your own site as important is not a signal any serious retrieval system would trust.

Schema.org: necessary, but not the lever GEO agencies sell

Google's guide is careful here. It does not say Schema is useless. It says you do not need additional or special schema for AI features. Standard structured data that helps Google understand your content continues to matter. What does not matter is the layer of AI-specific Schema that some agencies add to their deliverables.

The evidence backs this up. The WashU study presented at the October 2026 ACM Internet Measurement Conference analyzed 55,393 searches and 61,212 citations. They found that 29.8% of cited domains did not appear in the first-page results at all, which means Google's AI source selection uses a pool distinct from its ranking algorithm. If Schema were the primary lever for AI citation, you would expect the most-marked-up pages to dominate. They do not.

What the WashU study did find is that AIO-cited domains received higher average credibility scores than traditional first-page results. Credibility, in this context, is not about Schema. It is about source quality, domain authority, and content substance. Google's own guide reinforces this: the most important factor is "valuable, non-commodity content" with a "unique point of view."

The Ahrefs analysis of 863,000 keywords and 4 million AI Overview URLs found that only 38% of cited pages appeared in the organic top 10 for the same query, down from 76% in their July 2025 study. A separate BrightEdge analysis put the number at 17%. If Schema markup were the deciding factor, the overlap with top-ranking pages (which typically have the most Schema) would be higher, not lower.

Chunking: the tactic with no public evidence

"Content chunking" for AI retrieval is the idea that you should break your content into self-contained semantic units so an LLM can ingest them individually. It sounds plausible. It has no public evidence behind it.

I searched for peer-reviewed studies, large-scale analyses, or even practitioner log data showing that chunked content gets cited more frequently than unchunked content. I found none. What I found instead is the CXL 100-citation study, which showed that 55% of AI Overview citations come from the first 30% of a page. Kevin Indig's analysis of 18,012 verified ChatGPT citations found the same pattern: 44.2% from the first 30%, dropping sharply after.

That is not chunking. That is front-loading. The data says put your answer near the top, not split your content into pieces. Google's guide does not mention chunking at all, which is consistent with the absence of evidence for it.

What the data says actually works

The peer-reviewed GEO paper by Aggarwal et al. (KDD 2024) is the closest thing this field has to a landmark study. They built a benchmark (GEO-Bench) and measured what happens to page visibility when you edit the page itself. The methods that worked:

  • Adding citations to credible sources within your content
  • Adding statistics
  • Adding quotations from authoritative sources
  • Improving content density and specificity

These are content-level edits, not file-level or markup-level changes. Their benchmark showed visibility boosts of up to 40% in generative engine responses. That is an effect size. llms.txt, extra Schema, and chunking have no published effect size.

Google's guide and the independent data converge on the same conclusion. The work that matters for AI search visibility is the work that has always mattered for search: make content that is original, specific, and useful. The tactics that agencies package as "GEO-specific" are either things you should already be doing (basic Schema, crawlability, semantic HTML) or things with no measured effect (llms.txt, chunking, AI-specific markup).

What this means for your audit

I run geo-optimizer-skill, and we deliberately do not score sites on llms.txt presence, Schema count, or chunking. We score on the signals that have evidence: crawlability, content density, semantic structure, and whether your pages render for bots that do not execute JavaScript.

If your GEO audit leads with llms.txt and Schema density, you are optimizing the easy-to-instrument metrics, not the outcome metrics. The easy ones produce diffs and look like work. The outcome metric (does your content get cited by AI) requires an experiment, not a dashboard.

That is the trap. The same trap as tuning alert thresholds instead of asking whether the alert should exist. The metric that is easy to measure is rarely the one that describes the result.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to