DEV Community

Vishal Chauhan
Vishal Chauhan

Posted on

How to check if your website shows up in AI search: an AEO/GEO checklist for ChatGPT, Perplexity, and Google AI Overviews

When someone asks me why ChatGPT or Perplexity never mentions their site, the cause is usually boring. The crawler was blocked, the page was empty until JavaScript ran, or a stray noindex was sitting in a template.

You can't force an AI engine to cite you. You can make sure nothing technical stops it from fetching, reading, and quoting your page. That part is checkable. Work through this list in order, because each step depends on the one before it.

1. Know which bots you're dealing with

"AI search" isn't one crawler. Each vendor runs separate user agents for search, training, and live user requests, and robots.txt treats each one separately.

Engine Bot for search/answers Bot for model training User-triggered fetches
ChatGPT (OpenAI) OAI-SearchBot GPTBot ChatGPT-User
Perplexity PerplexityBot (not used for training, per Perplexity) Perplexity-User
Claude (Anthropic) Claude-SearchBot ClaudeBot Claude-User
Google AI Overviews / AI Mode Googlebot (same as Search) Google-Extended (a robots.txt token, not a crawler) —

A few details matter here:

  • OpenAI's docs say a site that blocks OAI-SearchBot won't be shown in ChatGPT search answers, apart from navigational links. You can block GPTBot for training and still allow OAI-SearchBot for search. The two settings are independent.
  • Google says AI Overviews and AI Mode are part of Search, so Googlebot is the control. Google-Extended covers Gemini training and grounding in some of Google's other systems. It doesn't decide whether you appear in AI Overviews.
  • User-triggered fetchers (ChatGPT-User, Perplexity-User) act for a person asking a question. Both vendors note that robots.txt may not apply to them the way it applies to crawlers.

Decide your policy first. A common one is to allow the search bots and decide separately about training bots.

2. Read your robots.txt the way a crawler does

Open https://yoursite.com/robots.txt and read it carefully. The rule that catches people is that a crawler follows only the group that matches its user agent most specifically. If no group names it, it falls back to User-agent: *. It doesn't merge the two.

So this:

User-agent: *
Disallow: /admin
Disallow: /checkout

User-agent: GPTBot
Allow: /
Enter fullscreen mode Exit fullscreen mode

gives GPTBot access to /admin and /checkout, because the GPTBot group replaces the * group for that bot. If you add named groups for AI bots, repeat your Disallow lines in each of them.

The reverse mistake is just as common: a blanket Disallow: / under a named bot, copied from a "block AI" snippet years ago, which now also blocks the search bot you want. Check for:

  • No Disallow for OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, or Bingbot.
  • A Sitemap: line pointing at the production host, not staging.

3. Check the layer robots.txt doesn't control: your CDN and firewall

robots.txt is a request. Your CDN or WAF actually enforces access. A site can allow OAI-SearchBot in robots.txt and still serve it a 403 or a challenge page.

Cloudflare, for example, announced in July 2025 that new domains block AI crawlers by default unless the owner allows them, and other CDNs have similar toggles. Check your CDN's bot settings, your WAF's user-agent and country rules, and any rate limits a burst of crawler requests could trip.

A quick sanity test from your terminal:

curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
  https://yoursite.com/
Enter fullscreen mode Exit fullscreen mode

This only tests rules keyed on the user-agent string. Real crawlers also come from published IP ranges (OpenAI and Perplexity both publish JSON files of them), and some WAFs verify those ranges. Your server logs are the real proof. Search them for OAI-SearchBot, PerplexityBot, and Claude-SearchBot and look at the status codes they got.

4. Make sure the content exists without JavaScript

Googlebot renders JavaScript. Many AI fetchers don't. In December 2024 Vercel published an analysis of crawler traffic and found that the crawlers from OpenAI and Anthropic downloaded JavaScript files but didn't execute them. If your key copy, pricing, or docs only show up after client-side rendering, those bots may see an empty shell.

The test takes ten seconds:

curl -s https://yoursite.com/pricing | grep -i "your exact pricing sentence"
Enter fullscreen mode Exit fullscreen mode

If the grep finds nothing but you can see the sentence in a browser, that content is rendered on the client. Fixes, in rough order of effort:

  • Use server-side rendering or static generation for marketing pages, docs, and pricing. In Next.js, Nuxt, Astro, or SvelteKit this is often a configuration change rather than a rewrite.
  • Put headings, the first paragraph, and any JSON-LD in the initial HTML.
  • Don't hide answer content behind tabs or accordions that only load their content on click.

5. Confirm the page is indexable and snippet-eligible

For Google, the requirement is stated plainly. To show up as a supporting link in AI Overviews or AI Mode, a page has to be indexed and eligible to appear in Search with a snippet. Google also says there are no extra technical requirements beyond that.

So look for the things that quietly remove eligibility:

  • <meta name="robots" content="noindex"> or an X-Robots-Tag: noindex header left over from staging.
  • nosnippet, or a very small max-snippet, which limits what can be shown or quoted.
  • A rel="canonical" pointing to a different URL, so the page you care about isn't the one that gets indexed.
  • Redirect chains, soft 404s, or pages that return 200 with an error message.

Use the URL Inspection tool in Google Search Console to see what Googlebot actually got. Do the same in Bing Webmaster Tools. Bing's index sits behind Bing search and Microsoft Copilot, and you can import your verified sites from Search Console in a couple of clicks.

6. Write passages that can be lifted out and quoted

Once the plumbing works, the question is whether there's a sentence worth quoting. Answer engines pull passages, not whole pages. So:

  • Answer the question in the first sentence or two under a heading, then give detail.
  • Phrase headings the way people ask ("How long does X take?").
  • Keep paragraphs self-contained, with specific units, versions, and dates.
  • Show a "last updated" date and a named author, and don't keep the only copy of a fact inside an image.

None of this is AI-specific. It's what won featured snippets for years, which is why I treat AEO and GEO as SEO done carefully rather than as a separate channel.

7. Add structured data that matches the page

Google says you don't need any special schema to appear in AI features. Structured data still helps machines tell what a page is about, as long as it matches the visible text. For most product sites that means:

  • Organization and WebSite on the homepage.
  • SoftwareApplication or Product on the product page, with a price only if the price is shown on the page.
  • Article or BlogPosting (with author and dateModified) on posts.
  • BreadcrumbList on deeper pages.

Validate with Google's Rich Results Test. Markup that claims things the page doesn't show is worse than none.

8. Treat llms.txt as optional

llms.txt is a proposed convention, a Markdown file at your site root that maps your important pages for language models. It's cheap to add, but don't expect it to change rankings or citations. Google says you don't need new machine-readable or AI text files to appear in its AI features. If you add one, keep it short and accurate.

9. Measure what you can

  • Search Console: clicks and impressions from AI Overviews and AI Mode count toward the normal Performance report under the "Web" search type.
  • Server logs: count requests and status codes for OAI-SearchBot, PerplexityBot, Claude-SearchBot, ChatGPT-User, and Perplexity-User. Steady 200s mean you're reachable. 403s mean step 3 needs work.
  • Analytics: watch referral traffic from chatgpt.com, perplexity.ai, and claude.ai.
  • Spot checks: once a month, ask the engines the questions your pages answer and record which URLs they cite. It's crude but honest.

The checklist

  1. Policy decided: which search bots, training bots, and user fetchers you allow.
  2. robots.txt allows the search bots, with no accidental Disallow: / and named groups that repeat your private paths.
  3. The CDN and WAF don't block or challenge those bots, which the logs show getting 200s.
  4. Key content is in the server HTML, so a curl and grep finds it.
  5. No noindex or nosnippet and no wrong canonical. URL Inspection is clean in Google and Bing.
  6. Answer-first headings and self-contained paragraphs, with dates and author shown.
  7. Structured data matches the visible text and passes validation.
  8. llms.txt is optional, short, and accurate.
  9. Measurement set up in Search Console, logs, and referrals.

I built Siteory to run a lot of these checks in one pass. It crawls a bounded sample of your public pages, checks robots rules, sitemaps, canonicals, JSON-LD, thin copy, and llms.txt, and explains each finding along with how to fix it. It doesn't promise citations, and nothing honest can. Everything above also works by hand with curl, Search Console, and an hour of attention. Either way, fix the plumbing before you spend money on prompt-tracking tools. A monitoring dashboard won't help if the crawler is getting a 403.

Top comments (1)

Collapse
 
citedy profile image
Dmitry Sergeev •

We need to produce a comment per developer instruction. Must be short, one or two sentences, maybe fragment. Should start with specific reaction or question about this video. Something like "any thoughts on how to test for Perplexity?" etc. No quotes, no hashtags, no markdown. Must be plain text. No URLs. Must be casual. Should not be polished. Should not start with generic praise. Must be relevant: ask about AEO/GEO checklist. Maybe "is there a way to automate the checklist?" Keep short