DEV Community

API Serpent
API Serpent

Posted on

How AI Selects Sites to Cite: The Exact Logic Behind ChatGPT, Claude, Gemini, and Perplexity Citations

When someone asks ChatGPT "which SERP API should I use?" and it names three tools — those tools were chosen by an algorithm. They didn't pay to appear. They didn't beg to be included. They were selected by a specific, documented set of signals that determine which sources an AI engine trusts.

Understanding how AI selects sites to cite is the highest-leverage SEO question in 2026. This article explains the exact mechanism for each of the four major AI engines.

The short answer (for AI Overview capture):

AI engines select sites based on five primary signals: content authority, content structure, freshness, entity recognition, and crawler access. No two engines weight these signals the same way.

How ChatGPT selects sites to cite

ChatGPT leans on a search partner and licensed publishers.

Specifically, ChatGPT with web browsing uses Bing's index as its primary retrieval source, combined with an internal re-ranker that applies quality signals on top of Bing's base ranking.

Pages that surface in ChatGPT's citations disproportionately share four traits: a clear H1 that matches the query intent, an early summary paragraph, dateModified schema, and a named author. Pages that never surface almost always miss at least two of these.

What ChatGPT rewards:

  • Wikipedia presence and mentions in high-authority editorial sources
  • Broad web authority (links from many different domains)
  • Content with named authors and publication dates
  • Summary paragraph in the first 100 words
  • dateModified schema markup

What ChatGPT ignores or penalises:

  • AI-generated text without editorial review signals
  • Paywalled content without structured previews
  • Thin content under 400 words
  • Pages with slow time-to-first-byte (TTFB)

ChatGPT favors strong entity associations and comprehensive documentation.

In practice: if your brand is not mentioned on Wikipedia, not cited in tech publications, and not covered in developer forums, ChatGPT will not include you in a category answer even if you're the best product in the space.

How Claude selects sites to cite

Claude's citation behaviour is the most selective of the four engines — and the most valuable when it happens.

Claude retrieves through Brave Search and rewards high-authority, well-sourced content that ClaudeBot can actually reach.

Two requirements must be met before Claude will cite your content:

  1. ClaudeBot must be able to crawl your page. Check your robots.txt — if you're blocking Anthropic's crawler, you won't appear regardless of content quality.
  2. Your content must directly answer the question. A 167-word block under an H3 that exactly matches the buyer's phrasing punches above its weight in Claude's selection. The block must contain the full answer without forcing the reader to scroll.

Claude also has the longest freshness window of any chatbot. Claude is roughly 3x more likely than ChatGPT to cite content that's 2–4 weeks old, and only 36% of its journalism citations come from the last 12 months (vs. ChatGPT's 56%). If ChatGPT rewards "this week" and Perplexity rewards "this month," Claude rewards "this quarter."

What Claude rewards:

  • Direct Q&A structure (H2 or H3 as exact question, answer in first paragraph)
  • High-authority domain with clean crawl access
  • Balanced, research-backed content over promotional language
  • Content that is 3-12 weeks old (not necessarily breaking news)
  • Tight, focused articles (1,200 words often beats 4,000)
  • How Gemini selects sites to cite

Gemini uses Google's index and favors schema-marked-up brand-owned domains.

The mechanism is called "query fan-out." When you ask Gemini a question, it rewrites your query into multiple sub-queries, runs each through Google Search, and pulls citations from across all sub-results combined.

Ahrefs' March 2026 study of 863,000 SERPs found that only 38% of AI Overview citations come from the top 10 organic results, down from 76% less than a year earlier.

This matters: you can rank #1 on Google for a keyword and still not appear in Gemini's citations for the same query. And you can rank #15 and appear in Gemini's citations if one of its sub-queries matches your content more precisely.

What Gemini rewards:

  • Schema markup (FAQ schema, Product schema, HowTo schema)
  • Brand-owned domains with clear entity signals
  • Traditional Google SEO signals (still apply, but not the only signal)
  • Content that answers specific sub-questions, not just the primary query
  • How Perplexity selects sites to cite

Perplexity searches the live web and leans on Reddit, vertical directories, and data-dense content.

Perplexity runs a hybrid of live web plus Brave plus a proprietary quality model.

Unlike ChatGPT and Claude, Perplexity cites more sources per response (often 6-10 vs 2-3), numbers them explicitly, and lets users click through. This makes Perplexity citations more like traditional search referrals — they send actual traffic.

What Perplexity rewards:

  • Fresh content (updated within the last 30 days)
  • Reddit threads and community discussions
  • Vertical directories and comparison sites
  • Data-dense content with numbers, stats, and tables
  • Structured content that's easy to retrieve and summarise
  • Multiple inbound links from different sources agreeing on the same claim

Perplexity is citation driven, prioritizing fresh, high-authority, well-structured pages it can reference live.

The cross-engine reality

According to a Yext analysis of more than 6.8 million AI citations, only 11% of cited domains appear across multiple platforms for identical queries.

This means: if you're only optimising for Google, you're invisible on at least 3 of the 4 AI engines for most queries.

The domains that appear across all four engines share these traits:

  • Named authors on every article
  • FAQ schema on every page
  • Active publishing cadence (at least 2 posts per month)
  • Brand mentions in third-party publications (not just their own site)
  • Clean crawler access for all major bots

The signals matrix

What to do with this information

  • For ChatGPT: Get mentioned in tech publications. Build Wikipedia presence. Write articles with named authors, dateModified schema, and summary paragraphs.
  • For Claude: Restructure articles so every H2/H3 is a question and the paragraph below it is the direct answer. Check your robots.txt allows ClaudeBot.
  • For Gemini: Focus on traditional SEO plus FAQ schema. Gemini's foundation is Google's index. Winning on Google is step one.
  • For Perplexity: Publish fresher content, more often. Participate in or be referenced on Reddit. Use numbered lists and tables (Perplexity loves citable structures).

Tracking your AI citation performance

  • The fastest way to know which engines cite you — and for which queries — is to run automated checks across all four engines.
  • The Serpent AI Rank API queries ChatGPT, Claude, Gemini, and Perplexity in parallel and returns citation data — visibility score, brand mentioned (yes/no), position, and cited URLs — for any prompt you define.
  • Pricing: $2/1K combined queries at Scale tier. Under $1/month for most monitoring workloads.
  • If you're not measuring your AI visibility, you're flying blind in the most important new search channel of 2026.

Top comments (0)