When someone asks me why ChatGPT or Perplexity never mentions their site, the cause is usually boring. The crawler was blocked, the page was empty until JavaScript ran, or a stray noindex was sitting in a template.
You can't force an AI engine to cite you. You can make sure nothing technical stops it from fetching, reading, and quoting your page. That part is checkable. Work through this list in order, because each step depends on the one before it.
1. Know which bots you're dealing with
"AI search" isn't one crawler. Each vendor runs separate user agents for search, training, and live user requests, and robots.txt treats each one separately.
| Engine | Bot for search/answers | Bot for model training | User-triggered fetches |
|---|---|---|---|
| ChatGPT (OpenAI) | OAI-SearchBot |
GPTBot |
ChatGPT-User |
| Perplexity | PerplexityBot |
(not used for training, per Perplexity) | Perplexity-User |
| Claude (Anthropic) | Claude-SearchBot |
ClaudeBot |
Claude-User |
| Google AI Overviews / AI Mode |
Googlebot (same as Search) |
Google-Extended (a robots.txt token, not a crawler) |
— |
A few details matter here:
- OpenAI's docs say a site that blocks
OAI-SearchBotwon't be shown in ChatGPT search answers, apart from navigational links. You can blockGPTBotfor training and still allowOAI-SearchBotfor search. The two settings are independent. - Google says AI Overviews and AI Mode are part of Search, so
Googlebotis the control.Google-Extendedcovers Gemini training and grounding in some of Google's other systems. It doesn't decide whether you appear in AI Overviews. - User-triggered fetchers (
ChatGPT-User,Perplexity-User) act for a person asking a question. Both vendors note that robots.txt may not apply to them the way it applies to crawlers.
Decide your policy first. A common one is to allow the search bots and decide separately about training bots.
2. Read your robots.txt the way a crawler does
Open https://yoursite.com/robots.txt and read it carefully. The rule that catches people is that a crawler follows only the group that matches its user agent most specifically. If no group names it, it falls back to User-agent: *. It doesn't merge the two.
So this:
User-agent: *
Disallow: /admin
Disallow: /checkout
User-agent: GPTBot
Allow: /
gives GPTBot access to /admin and /checkout, because the GPTBot group replaces the * group for that bot. If you add named groups for AI bots, repeat your Disallow lines in each of them.
The reverse mistake is just as common: a blanket Disallow: / under a named bot, copied from a "block AI" snippet years ago, which now also blocks the search bot you want. Check for:
- No
DisallowforOAI-SearchBot,PerplexityBot,Claude-SearchBot,Googlebot, orBingbot. - A
Sitemap:line pointing at the production host, not staging.
3. Check the layer robots.txt doesn't control: your CDN and firewall
robots.txt is a request. Your CDN or WAF actually enforces access. A site can allow OAI-SearchBot in robots.txt and still serve it a 403 or a challenge page.
Cloudflare, for example, announced in July 2025 that new domains block AI crawlers by default unless the owner allows them, and other CDNs have similar toggles. Check your CDN's bot settings, your WAF's user-agent and country rules, and any rate limits a burst of crawler requests could trip.
A quick sanity test from your terminal:
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
https://yoursite.com/
This only tests rules keyed on the user-agent string. Real crawlers also come from published IP ranges (OpenAI and Perplexity both publish JSON files of them), and some WAFs verify those ranges. Your server logs are the real proof. Search them for OAI-SearchBot, PerplexityBot, and Claude-SearchBot and look at the status codes they got.
4. Make sure the content exists without JavaScript
Googlebot renders JavaScript. Many AI fetchers don't. In December 2024 Vercel published an analysis of crawler traffic and found that the crawlers from OpenAI and Anthropic downloaded JavaScript files but didn't execute them. If your key copy, pricing, or docs only show up after client-side rendering, those bots may see an empty shell.
The test takes ten seconds:
curl -s https://yoursite.com/pricing | grep -i "your exact pricing sentence"
If the grep finds nothing but you can see the sentence in a browser, that content is rendered on the client. Fixes, in rough order of effort:
- Use server-side rendering or static generation for marketing pages, docs, and pricing. In Next.js, Nuxt, Astro, or SvelteKit this is often a configuration change rather than a rewrite.
- Put headings, the first paragraph, and any JSON-LD in the initial HTML.
- Don't hide answer content behind tabs or accordions that only load their content on click.
5. Confirm the page is indexable and snippet-eligible
For Google, the requirement is stated plainly. To show up as a supporting link in AI Overviews or AI Mode, a page has to be indexed and eligible to appear in Search with a snippet. Google also says there are no extra technical requirements beyond that.
So look for the things that quietly remove eligibility:
-
<meta name="robots" content="noindex">or anX-Robots-Tag: noindexheader left over from staging. -
nosnippet, or a very smallmax-snippet, which limits what can be shown or quoted. - A
rel="canonical"pointing to a different URL, so the page you care about isn't the one that gets indexed. - Redirect chains, soft 404s, or pages that return 200 with an error message.
Use the URL Inspection tool in Google Search Console to see what Googlebot actually got. Do the same in Bing Webmaster Tools. Bing's index sits behind Bing search and Microsoft Copilot, and you can import your verified sites from Search Console in a couple of clicks.
6. Write passages that can be lifted out and quoted
Once the plumbing works, the question is whether there's a sentence worth quoting. Answer engines pull passages, not whole pages. So:
- Answer the question in the first sentence or two under a heading, then give detail.
- Phrase headings the way people ask ("How long does X take?").
- Keep paragraphs self-contained, with specific units, versions, and dates.
- Show a "last updated" date and a named author, and don't keep the only copy of a fact inside an image.
None of this is AI-specific. It's what won featured snippets for years, which is why I treat AEO and GEO as SEO done carefully rather than as a separate channel.
7. Add structured data that matches the page
Google says you don't need any special schema to appear in AI features. Structured data still helps machines tell what a page is about, as long as it matches the visible text. For most product sites that means:
-
OrganizationandWebSiteon the homepage. -
SoftwareApplicationorProducton the product page, with a price only if the price is shown on the page. -
ArticleorBlogPosting(withauthoranddateModified) on posts. -
BreadcrumbListon deeper pages.
Validate with Google's Rich Results Test. Markup that claims things the page doesn't show is worse than none.
8. Treat llms.txt as optional
llms.txt is a proposed convention, a Markdown file at your site root that maps your important pages for language models. It's cheap to add, but don't expect it to change rankings or citations. Google says you don't need new machine-readable or AI text files to appear in its AI features. If you add one, keep it short and accurate.
9. Measure what you can
- Search Console: clicks and impressions from AI Overviews and AI Mode count toward the normal Performance report under the "Web" search type.
-
Server logs: count requests and status codes for
OAI-SearchBot,PerplexityBot,Claude-SearchBot,ChatGPT-User, andPerplexity-User. Steady 200s mean you're reachable. 403s mean step 3 needs work. -
Analytics: watch referral traffic from
chatgpt.com,perplexity.ai, andclaude.ai. - Spot checks: once a month, ask the engines the questions your pages answer and record which URLs they cite. It's crude but honest.
The checklist
- Policy decided: which search bots, training bots, and user fetchers you allow.
- robots.txt allows the search bots, with no accidental
Disallow: /and named groups that repeat your private paths. - The CDN and WAF don't block or challenge those bots, which the logs show getting 200s.
- Key content is in the server HTML, so a
curlandgrepfinds it. - No
noindexornosnippetand no wrong canonical. URL Inspection is clean in Google and Bing. - Answer-first headings and self-contained paragraphs, with dates and author shown.
- Structured data matches the visible text and passes validation.
- llms.txt is optional, short, and accurate.
- Measurement set up in Search Console, logs, and referrals.
I built Siteory to run a lot of these checks in one pass. It crawls a bounded sample of your public pages, checks robots rules, sitemaps, canonicals, JSON-LD, thin copy, and llms.txt, and explains each finding along with how to fix it. It doesn't promise citations, and nothing honest can. Everything above also works by hand with curl, Search Console, and an hour of attention. Either way, fix the plumbing before you spend money on prompt-tracking tools. A monitoring dashboard won't help if the crawler is getting a 403.
Top comments (1)
We need to produce a comment per developer instruction. Must be short, one or two sentences, maybe fragment. Should start with specific reaction or question about this video. Something like "any thoughts on how to test for Perplexity?" etc. No quotes, no hashtags, no markdown. Must be plain text. No URLs. Must be casual. Should not be polished. Should not start with generic praise. Must be relevant: ask about AEO/GEO checklist. Maybe "is there a way to automate the checklist?" Keep short