DEV Community

Edward Chapman
Edward Chapman

Posted on

Robots.txt allows OAI-SearchBot. That does not prove ChatGPT can fetch your page.

A line in robots.txt that allows OAI-SearchBot grants permission to a compliant crawler. It does not tell you whether your CDN lets the request through, whether the response contains your content, or whether the page appears in ChatGPT answers. Each question needs its own evidence.

Two OpenAI user agents, two decisions

OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for training. Blocking GPTBot alone does not automatically block OAI-SearchBot. To allow search crawling and disallow training crawling, use separate groups:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

These rules state crawler permissions. They do not override firewall rules or guarantee training exclusion in every context. Check OpenAI's documentation for the current scope of each agent:
https://platform.openai.com/docs/bots

Check the rule that applies

Start with the hostname. The file at example.com/robots.txt does not govern www.example.com or shop.example.com.

Then check the exact path. Plain path rules match by prefix: Disallow: /pricing also matches /pricing-old and /pricing/eu.

Check the user-agent group too. Under standard matching, a crawler that finds a group naming it follows that group rather than falling back to User-agent: *. Review all groups for the named agent.

Permission to crawl is not indexing or citation. Robots.txt manages crawler traffic; it is not authentication or access control. Anyone can read the file and ignore it, so confidential data needs proper access controls.

Google explains the scope and limitations here:
https://developers.google.com/search/docs/crawling-indexing/robots/intro

The edge can refuse first

A CDN or firewall can block or challenge a permitted request before it reaches your server. Your origin logs may show nothing about that request.

Review edge security events and origin logs separately. Record the hostname, path, time, action or status, and the rule that matched. If a rule catches traffic you meant to allow, adjust that specific rule. Do not disable site-wide protection to troubleshoot one crawler.

Cloudflare's documentation, checked on 2 October 2026, separates Search, Agent and Training categories. An Allow setting in the AI bot policy does not remove other WAF controls. Published defaults do not tell you how an existing zone is configured. Check your own settings:
https://developers.cloudflare.com/bots/additional-configurations/block-ai-bots/

A copied user agent tests your request, not theirs

Running curl with OAI-SearchBot's user-agent string sends a request from your machine and IP address. If the edge uses IP ranges or verified-bot status, it may respond differently to the real crawler.

A 200 from your laptop does not prove the crawler gets a 200. A 403 does not prove it gets a 403. The User-Agent header comes from the client, so it does not verify identity.

Use the verification information OpenAI currently publishes and compare it with your edge request records. If no verified requests appear in the period you checked, report that they were not observed. That does not establish that they were blocked.

A 200 is not necessarily your content

A 200 response can contain a login page, a challenge interstitial, or a JavaScript shell with little text. Where response evidence is available for a verified crawler request, compare what was served with the expected public page.

Status and response size are clues, not proof that the intended text was delivered. Standard access logs usually do not contain response bodies. Even confirmed delivery does not prove the provider ingested the page.

Keep four findings separate

Declared permission: the matching robots.txt group and path rule. This does not show that a request arrived.

Observed access: edge or origin evidence of a verified crawler request and its response. This does not show that the content was indexed or used.

Provider indexing: signals the provider publishes, where available. This does not promise a citation.

Cited answer: an actual answer linking the page. This does not show that the citation will recur or appear for other queries.

A practical checklist

Read robots.txt on the exact hostname and find the group and path rule for OAI-SearchBot.

Check GPTBot separately against your training decision.

Review edge security events for that hostname and path, then origin logs.

Verify crawler identity using provider documentation, not the User-Agent string alone.

For verified requests, record the action, status, matching rule, and available evidence of the content served.

Report anything you could not confirm as unknown. Access evidence alone cannot establish that ChatGPT cited a page.

An optional check for the robots.txt part

I'm with Firm Beacon. Our free browser checker compares robots.txt permissions for search and training agents on a specific URL. You can enter a URL or paste robots.txt without creating an account or giving an email address:
https://www.firmbeacon.co.uk/tools/ai-crawler-check?utm_source=devto&utm_medium=article&utm_campaign=ai_crawler_access

It checks the declared rules, not your firewall. It does not verify indexing, citations or rankings. Use it for the first step, then check access evidence separately.

Top comments (0)