DEV Community

Mael Bourdin
Mael Bourdin

Posted on

AI crawlers and robots.txt: a practical guide for developers

A surprising number of sites block the exact AI crawlers they would like to be cited by, usually because someone copied a robots.txt snippet during the "block all AI" wave and never revisited it. Here is how to think about it as a developer, without breaking anything.

Not all AI crawlers do the same job

The key distinction: some bots collect content to train models, others fetch pages to answer a user's question right now. Blocking them has very different consequences.

  • Training crawlers gather data used to build future models. Blocking them is a policy choice about your content.- Search and retrieval crawlers index or fetch pages so an assistant can cite them in answers. Blocking them can remove you from those answers.- User-triggered fetchers visit a page because a user asked the assistant to read it.Major providers document separate user agents for these roles. OpenAI, for example, documents GPTBot for training, OAI-SearchBot for search, and ChatGPT-User for user-triggered visits. Anthropic, Perplexity, Google and Apple publish their own lists. Check each provider's official documentation before you write rules, because names and roles evolve. ## Audit what you have today
  • Open /robots.txt on your production domain, not on staging.- Look for blanket rules like User-agent: * with Disallow: /, and for any AI-specific user agents.- Check your CDN or firewall too. Bot protection settings can block AI crawlers even when robots.txt allows them.- Check your server logs for these user agents to see who actually visits and what status codes they get.## A sensible default If you want to be cited in AI answers but prefer not to feed training, a common pattern is to allow the search and user-triggered agents and disallow the training ones. Conceptually:
User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /
Enter fullscreen mode Exit fullscreen mode

Treat this as a pattern, not a copy-paste list. Confirm each user agent against the provider's current documentation, and remember that robots.txt is a request that well-behaved crawlers follow, not an access control mechanism.

Make the allowed pages worth fetching

Allowing a crawler is only half the job. Retrieval systems need content they can actually read:

  • Server-render important content. Text that only appears after client-side JavaScript runs may never be seen by some crawlers.- Return correct status codes. A soft 404 or a page that returns 200 with an error message wastes the visit.- Keep pages fast. Slow responses and timeouts mean the page gets skipped.- Use clean semantic HTML. Headings, lists and tables make passages easier to extract.- Keep a sitemap up to date so new pages get discovered.## What about llms.txt? llms.txt is a proposed convention: a Markdown file at your site root that points language models to your most useful pages. It is cheap to add and can help tools that support it, but support across engines is inconsistent, so do not treat it as a replacement for crawlable pages. ## Verify, then monitor After changing rules, watch your logs for the user agents you allowed, and confirm they receive 200 responses. Then check the outcome that actually matters: whether assistants start citing your pages. That last part is what CiteMe was built to track, across engines and over time. Robots.txt mistakes are silent. Nothing breaks, you just quietly disappear from answers. Ten minutes of auditing is worth it.

Top comments (2)

Collapse
 
omyvnss profile image
Om Yaduvanshi •

the three-bucket split is the part most guides skip. people lump gptbot and oai-searchbot together and end up invisible in answers while thinking they only opted out of training. a blanket user-agent: * with disallow: / inherited from some old boilerplate is probably the quietest traffic killer on the internet.

Collapse
 
citedy profile image
Dmitry Sergeev •

We need to output a comment text only, short, one or two sentences, casual. Must mention something specific about video: maybe ask about handling robots.txt for AI crawlers, like "why do some sites block the same AI crawlers they want cited?" etc. Use lowercase start, casual voice, maybe "tbh". No punctuation errors. Ensure no em-dash etc. No URLs. Provide comment only. Let's craft: "tbh, never realized you can actually let the AI crawl your site but still keep the same