DEV Community

Nick Johnson
Nick Johnson

Posted on

What AI crawlers actually read before your content (and the 3 files most sites get wrong)

#ai

I spent five months making a consumer health site visible to AI assistants. Monthly AI citations went from ~6.6K to ~62K across platforms, and organic traffic from 22K to just under 800K. Everyone assumes the work was content. It mostly wasn't. The first three weeks were spent on files that crawlers read before they ever get to a page, and that's where the citation curve started moving.

Here's what those files are, what we found broken on nearly every site we've audited since, and how to check yours in about ten minutes.

  1. robots.txt: you may be blocking the crawlers you want AI assistants don't use Googlebot. Each has its own user agent, and a lot of sites block them without knowing, usually from a bot-blocking rule someone added during a scraping panic, or a CDN "bot fight mode" that challenges anything that isn't a browser.

The ones that matter right now:
User-agent: GPTBot # OpenAI
User-agent: ChatGPT-User # OpenAI, live browsing
User-agent: OAI-SearchBot # OpenAI search index
User-agent: ClaudeBot # Anthropic
User-agent: PerplexityBot # Perplexity
User-agent: Google-Extended # Google AI training (not AI Overviews)
User-agent: Bingbot # feeds Copilot

Check with:

bash
curl -s https://yoursite.com/robots.txt | grep -iE "gptbot|claudebot|perplexity|google-extended|oai-searchbot"
If you see Disallow: / under any of those, that's your whole problem. Nothing else in this post matters until it's fixed.

One nuance worth knowing: Google-Extended controls training use, not whether you appear in AI Overviews. Blocking it doesn't remove you from Google's AI answers; blocking Googlebot does. Decide deliberately.
Also test the actual response, not just the file. Cloudflare and similar can return a 403 to a crawler even when robots.txt allows it:

bash
curl -s -o /dev/null -w "%{http_code}\n" -A "GPTBot/1.0" https://yoursite.com/
curl -s -o /dev/null -w "%{http_code}\n" -A "PerplexityBot/1.0" https://yoursite.com/

A 403 or a challenge page here means the assistant never sees you. I built a small AI crawler checker that runs this across the common agents if you'd rather not script it.

  1. llms.txt: the file that tells assistants what you're actually about llms.txt is a plain Markdown file at your root that describes your site for LLM consumers: what it is, what it's authoritative on, and where the important pages are. It's a proposed convention, not a standard every crawler honors yet, but it costs nothing and the crawlers that do read it get a far better picture of your site than they'd infer from a homepage. Most sites have one of two problems. Either the file doesn't exist (curl -I https://yoursite.com/llms.txt returns 404), or it's a platform default. Shopify, for example, generates one that explains how an agent can check out via Shop, and says nothing about the brand. We found that on a baby-care brand with 350 well-written articles: the file told assistants how to buy and nothing about what the company knew.

A hand-written one for a health site looks like this:
markdown

ExampleHealth

Clinician-reviewed guides to lab tests and common conditions, with
at-home testing in the US. Every article is reviewed by a licensed
clinician; reviewer credentials are in Person schema on each page.

Core guides

Services

Optional

Keep it short, link the pages you'd want quoted, and write the one-line descriptions yourself. If you want a scaffold, the llms.txt generator I built outputs this structure from a sitemap; the useful part is still the descriptions, which only you can write.

  1. Structured data: your author is invisible unless the markup says so This is the one that surprised me most. Assistants weigh who wrote a page and who reviewed it, especially in health, finance, and legal. But they read that from JSON-LD, not from a byline rendered in HTML. A page that says "Medically reviewed by Dr. Smith" in text, with only Organization schema in the source, has no reviewer as far as a crawler is concerned.

Check any article:

bash
curl -s https://yoursite.com/blog/some-post | grep -o '"@type"[^,]*' | sort | uniq -c
If the output is just Organization and maybe WebSite, your Article, Person, and FAQPage markup is missing. Here's the minimum that makes a reviewed article legible:

json
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "TSH Blood Test: What It Measures",
"datePublished": "2026-10-06",
"author": { "@type": "Organization", "name": "ExampleHealth" },
"reviewedBy": {
"@type": "Person",
"name": "Jane Smith, MD",
"jobTitle": "Endocrinologist",
"url": "https://example.com/team/jane-smith"
},
"publisher": { "@type": "Organization", "name": "ExampleHealth" },
"mainEntityOfPage": "https://example.com/blog/tsh-test"
}
Plus a separate FAQPage block for the FAQ section, because assistants lift question-and-answer pairs almost verbatim. Validate with Google's Rich Results Test or Schema.org's validator before shipping.

What happened when we fixed these three
Citations started climbing in week three, before any new content shipped. The content pipeline that came after (about 500 articles a month, each answering one question with the answer in the first 90 words) landed on a site crawlers already understood and already trusted, so new pages earned citations in weeks instead of months.

One honest caveat: most of the 62K citations came from Perplexity and Google's AI Mode. ChatGPT was the hardest surface and citations there actually declined over the period. I don't think anyone can promise ChatGPT citations. What you can do is stop being structurally ineligible, which is where most sites are.

The 10-minute check
bash

1. Are AI crawlers allowed and not challenged?

curl -s https://yoursite.com/robots.txt | grep -iE "gptbot|claudebot|perplexity"
curl -s -o /dev/null -w "%{http_code}\n" -A "GPTBot/1.0" https://yoursite.com/

2. Does llms.txt exist, and did a human write it?

curl -s https://yoursite.com/llms.txt | head -20

3. Is the author/reviewer in the schema?

curl -s https://yoursite.com/blog/top-post | grep -o '"@type"[^,]*' | sort | uniq -c

Top comments (0)