<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Costin Gheorghe</title>
    <description>The latest articles on DEV Community by Costin Gheorghe (@costin_gheorghe_40d2a06d8).</description>
    <link>https://dev.to/costin_gheorghe_40d2a06d8</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4108021%2Fa37839d7-b824-4feb-b265-1f1456aa1bed.png</url>
      <title>DEV Community: Costin Gheorghe</title>
      <link>https://dev.to/costin_gheorghe_40d2a06d8</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/costin_gheorghe_40d2a06d8"/>
    <language>en</language>
    <item>
      <title>AI Crawlers Are Now a Line Item: What They Actually Cost Wikimedia and Small Sites</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Wed, 16 Sep 2026 17:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/ai-crawlers-are-now-a-line-item-what-they-actually-cost-wikimedia-and-small-sites-3ge1</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/ai-crawlers-are-now-a-line-item-what-they-actually-cost-wikimedia-and-small-sites-3ge1</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/ai-crawler-bandwidth-cost" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Most of what gets written about AI crawlers is about visibility — can ChatGPT&lt;br&gt;
see your page, will it cite you, should you welcome the traffic. None of that&lt;br&gt;
is the problem for a growing list of sites. Their problem is that the traffic&lt;br&gt;
showed up, kept showing up, and the bill came due.&lt;/p&gt;

&lt;p&gt;The clearest documented case is Wikimedia. The Wikimedia Foundation's own&lt;br&gt;
engineering blog reported that bandwidth consumption for multimedia downloads&lt;br&gt;
from Wikimedia Commons has surged 50% since January 2024, and it did not&lt;br&gt;
hedge about the cause: the surge was not coming from human readers, but from&lt;br&gt;
"automated, data-hungry scrapers looking to train AI models." TechCrunch's&lt;br&gt;
reporting on the post picked up the same figures. This is not a monitoring&lt;br&gt;
question — we &lt;a href="https://letslaunch.today/blog/monitor-ai-crawler-traffic" rel="noopener noreferrer"&gt;covered that one already&lt;/a&gt; —&lt;br&gt;
it's a cost and reliability question, and it deserves its own answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bots cost more per request, not just more in total
&lt;/h2&gt;

&lt;p&gt;The intuitive model is that bot traffic is a scaled-up version of human&lt;br&gt;
traffic — more requests, same shape, so the cost scales linearly with volume.&lt;br&gt;
Wikimedia's explanation is that this isn't what happened, and the mismatch is&lt;br&gt;
the actual story.&lt;/p&gt;

&lt;p&gt;Human traffic clusters around popular pages, the ones a CDN or edge cache&lt;br&gt;
already has warm. Crawler bots, per Wikimedia's own account, tend to "bulk&lt;br&gt;
read" — they work through large numbers of pages indiscriminately, including&lt;br&gt;
the long tail of obscure pages that were rarely requested before and were&lt;br&gt;
never worth caching. Serving those requests means going back to origin&lt;br&gt;
infrastructure instead of an edge cache, every time. A bot doesn't just add&lt;br&gt;
volume; it adds volume in the expensive shape.&lt;/p&gt;

&lt;p&gt;The numbers back this up directly. Bots accounted for roughly 65% of the&lt;br&gt;
&lt;em&gt;most resource-intensive&lt;/em&gt; traffic Wikimedia serves, despite representing only&lt;br&gt;
around 35% of overall pageviews. Read that gap carefully: the group causing&lt;br&gt;
two-thirds of the expensive work is a third of the requests. If you only&lt;br&gt;
looked at request counts, you would badly underestimate what's driving your&lt;br&gt;
bill.&lt;/p&gt;

&lt;p&gt;Wikimedia was blunt about what this does to infrastructure built for a&lt;br&gt;
different kind of spike: its systems, in the Foundation's own words, were&lt;br&gt;
"built to sustain sudden traffic spikes from humans during high-interest&lt;br&gt;
events, but the amount of traffic generated by scraper bots is unprecedented&lt;br&gt;
and presents growing risks and costs." That's a nonprofit that has spent two&lt;br&gt;
decades engineering for viral human traffic, saying the thing it wasn't built&lt;br&gt;
for is a bot.&lt;/p&gt;

&lt;h2&gt;
  
  
  It isn't only Wikimedia
&lt;/h2&gt;

&lt;p&gt;Wikimedia has the size and the engineering culture to publish a number. Most&lt;br&gt;
sites carrying the same problem don't, which is a reason to take the&lt;br&gt;
anecdotal reports seriously rather than dismiss them for lacking a published&lt;br&gt;
percentage.&lt;/p&gt;

&lt;p&gt;The same body of reporting that surfaced Wikimedia's figures also names&lt;br&gt;
several smaller, independent infrastructure projects describing the identical&lt;br&gt;
pattern in their own words: the git-hosting service SourceHut, Diaspora&lt;br&gt;
developer Dennis Schubert, the repair-guide site iFixit, and the&lt;br&gt;
documentation host Read the Docs have all separately and publicly reported&lt;br&gt;
bandwidth spikes and inflated infrastructure costs they attribute to&lt;br&gt;
AI-crawler traffic — in some cases severe enough to affect site reliability&lt;br&gt;
for real users. Unlike Wikimedia's figures, none of these came with a formal&lt;br&gt;
published percentage, so treat them as real, corroborating, but less precise&lt;br&gt;
than Wikimedia's own numbers. The pattern showing up independently across a&lt;br&gt;
git host, a documentation site, a repair-guide site, and an individual&lt;br&gt;
developer's project is still worth something even without a shared&lt;br&gt;
methodology behind it.&lt;/p&gt;

&lt;p&gt;Wikimedia's response is measurable too: its 2025/2026 annual plan sets an&lt;br&gt;
explicit goal of cutting crawler-generated request rate by 20% and bandwidth&lt;br&gt;
usage by 30%. That's a nonprofit budgeting engineering time specifically&lt;br&gt;
against bot load, not against human growth — which tells you how large the&lt;br&gt;
line item got.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means if you're not running Wikimedia-scale infrastructure
&lt;/h2&gt;

&lt;p&gt;Almost nobody reading this runs a site at Wikimedia's scale, and that's the&lt;br&gt;
point worth making rather than skipping past: the mechanism doesn't require&lt;br&gt;
that scale. A "bulk read" crawler that hits your long tail of rarely-visited&lt;br&gt;
pages, bypassing your cache the same way, produces the same disproportionate&lt;br&gt;
cost on infrastructure a hundredth the size. The absolute dollar figure is&lt;br&gt;
smaller. The ratio — request share versus cost share — is the same shape.&lt;/p&gt;

&lt;p&gt;If you're behind a CDN, you already have more leverage here than you might&lt;br&gt;
realize, and most of it costs nothing to turn on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Start with rate-limiting, not blocking.&lt;/strong&gt; LetsLaunch runs behind&lt;br&gt;
Cloudflare, the same setup &lt;a href="https://letslaunch.today/blog/monitor-ai-crawler-traffic" rel="noopener noreferrer"&gt;we've described before&lt;/a&gt;&lt;br&gt;
for watching crawler traffic — the same dashboard that shows you who's&lt;br&gt;
visiting can also throttle them. Cloudflare's rate-limiting and bot-management&lt;br&gt;
rules can target a specific crawler's request rate directly, which is a much&lt;br&gt;
narrower tool than an outright block. A crawler capped at, say, one request&lt;br&gt;
per second still finishes crawling your site — just slower, and without&lt;br&gt;
hammering your origin in a burst. That's usually the actual goal: less load,&lt;br&gt;
not zero bots.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Know what blocking actually trades away.&lt;/strong&gt; This is the part worth being&lt;br&gt;
honest about, because it points in the opposite direction from most of what&lt;br&gt;
this blog has argued elsewhere. We've spent &lt;a href="https://letslaunch.today/blog/should-you-block-gptbot" rel="noopener noreferrer"&gt;several other posts&lt;/a&gt;&lt;br&gt;
on the case for staying crawlable, because citation eligibility depends on a&lt;br&gt;
crawler being able to reach your page at all. Blocking a retrieval crawler&lt;br&gt;
outright — &lt;code&gt;OAI-SearchBot&lt;/code&gt;, &lt;code&gt;PerplexityBot&lt;/code&gt;, &lt;code&gt;Claude-SearchBot&lt;/code&gt; — solves a&lt;br&gt;
bandwidth problem by creating a visibility problem. If your actual goal is&lt;br&gt;
citation, rate-limiting is almost always the better trade: a slower crawl&lt;br&gt;
still completes, and you stay in whatever index that crawler feeds. Reserve&lt;br&gt;
an outright block for a crawler that's demonstrably causing reliability&lt;br&gt;
problems for real users right now, not for one you're merely annoyed at&lt;br&gt;
seeing in the logs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Proof-of-work challenges are a real, separate category.&lt;/strong&gt; Some smaller&lt;br&gt;
open-source infrastructure projects have adopted tools that put a lightweight&lt;br&gt;
computational challenge in front of a request before serving the page — cheap&lt;br&gt;
enough that a normal browser clears it invisibly, expensive enough that&lt;br&gt;
sustaining a bulk-scraping pattern across thousands of pages gets&lt;br&gt;
computationally costly for whoever is running the scraper. This is a real and&lt;br&gt;
growing defense worth knowing about by mechanism, whether or not you adopt&lt;br&gt;
it — it's a different lever from a robots.txt rule or a firewall block,&lt;br&gt;
because it changes the cost calculus for the crawler operator rather than&lt;br&gt;
just refusing the request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Match the response to the actual problem.&lt;/strong&gt; The uncomfortable truth is that&lt;br&gt;
"should I block AI crawlers" doesn't have one answer, because it's two&lt;br&gt;
different questions wearing the same sentence. A site chasing citation&lt;br&gt;
eligibility should not blanket-apply advice written for a bandwidth crisis —&lt;br&gt;
"just block everything AI-labeled" throws away exactly what you're trying to&lt;br&gt;
get. A site with a genuine bandwidth or reliability problem should not feel&lt;br&gt;
obligated to stay maximally open just because citation matters to sites in&lt;br&gt;
general — if the traffic is actively degrading service for real users, that's&lt;br&gt;
a legitimate reason to throttle or block specific offenders, full stop. Check&lt;br&gt;
&lt;a href="https://letslaunch.today/blog/ai-crawler-cheat-sheet" rel="noopener noreferrer"&gt;our crawler cheat sheet&lt;/a&gt; for what each bot&lt;br&gt;
actually is before deciding which category you're in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is AI crawler traffic actually costing sites money?
&lt;/h2&gt;

&lt;p&gt;Yes, documented directly. The Wikimedia Foundation's own engineering blog&lt;br&gt;
reported a 50% surge in bandwidth for multimedia downloads from Wikimedia&lt;br&gt;
Commons since January 2024, attributed to AI-training scrapers rather than&lt;br&gt;
human readers, with bots responsible for roughly 65% of its most&lt;br&gt;
resource-intensive traffic despite being about 35% of pageviews. Smaller&lt;br&gt;
sites — SourceHut, iFixit, Read the Docs, and independent developer Dennis&lt;br&gt;
Schubert among them — have separately reported the same pattern, without&lt;br&gt;
Wikimedia's level of published detail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should I block AI crawlers to save bandwidth?
&lt;/h2&gt;

&lt;p&gt;Only if you have a genuine, demonstrated bandwidth or reliability problem,&lt;br&gt;
and rate-limiting is usually the better first move over an outright block.&lt;br&gt;
Throttling a crawler's request rate reduces load while letting the crawl&lt;br&gt;
still complete, which keeps you eligible for citation in that crawler's&lt;br&gt;
index. A hard block trades your bandwidth problem for a visibility problem —&lt;br&gt;
worth it if the traffic is actually degrading your site for real users, not&lt;br&gt;
worth it if you're just trying to reduce a number in a dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does bot traffic cost more than the same amount of human traffic?
&lt;/h2&gt;

&lt;p&gt;Because of what gets requested, not just how much. Wikimedia's own&lt;br&gt;
explanation is that crawler bots "bulk read" through large numbers of pages,&lt;br&gt;
including rarely-visited ones that were never popular enough to be cached —&lt;br&gt;
so each of those requests goes back to origin infrastructure instead of&lt;br&gt;
being served from a CDN edge cache. Human traffic clusters around popular,&lt;br&gt;
already-cached pages; bot traffic doesn't, which is why a smaller share of&lt;br&gt;
requests can still account for the majority of the most expensive ones to&lt;br&gt;
serve.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>Cloudflare Says Perplexity Ignored Robots.txt. We Checked the Record.</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Wed, 16 Sep 2026 08:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/cloudflare-says-perplexity-ignored-robotstxt-we-checked-the-record-287p</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/cloudflare-says-perplexity-ignored-robotstxt-we-checked-the-record-287p</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/perplexity-stealth-crawling-robots-txt" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Our own &lt;a href="https://letslaunch.today/blog/ai-crawler-cheat-sheet" rel="noopener noreferrer"&gt;AI crawler cheat sheet&lt;/a&gt; says the named&lt;br&gt;
crawlers "have an established track record of compliance" with &lt;code&gt;robots.txt&lt;/code&gt;.&lt;br&gt;
That line needs a footnote, and the footnote is Perplexity.&lt;/p&gt;

&lt;p&gt;In a blog post published in August 2025, Cloudflare accused Perplexity of&lt;br&gt;
"stealth crawling" — disguising its crawler's identity to keep pulling content&lt;br&gt;
from sites that had explicitly blocked Perplexity's declared crawler, either&lt;br&gt;
through &lt;code&gt;robots.txt&lt;/code&gt; or a network-level block. This is not a vague industry&lt;br&gt;
rumor. Cloudflare published a specific experiment, specific traffic numbers,&lt;br&gt;
and a specific response. Perplexity disputes the characterization. Both of&lt;br&gt;
those things are true at once, and this post is about being precise on which&lt;br&gt;
parts are documented fact and which part is a live dispute.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Cloudflare says it found
&lt;/h2&gt;

&lt;p&gt;Cloudflare's test was simple enough to describe in two sentences. It stood up&lt;br&gt;
brand-new domains that were never linked from anywhere and never submitted to&lt;br&gt;
any search engine or index, so the only way to know what was on them was to&lt;br&gt;
actually fetch and read them. Each domain's &lt;code&gt;robots.txt&lt;/code&gt; disallowed every&lt;br&gt;
crawler, with no exceptions.&lt;/p&gt;

&lt;p&gt;Then, according to Cloudflare's blog post, they asked Perplexity's own AI&lt;br&gt;
assistant about the content on those exact test domains — and got back&lt;br&gt;
detailed, accurate answers about pages the assistant should have had no&lt;br&gt;
legitimate way to see. Nothing was indexed. Nothing was linked. The only path&lt;br&gt;
from "page exists" to "assistant can describe it" was a crawl that ignored the&lt;br&gt;
disallow rule.&lt;/p&gt;

&lt;p&gt;Cloudflare says the traffic behind this broke down into two distinct patterns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Perplexity-User&lt;/code&gt;&lt;/strong&gt;, Perplexity's openly declared, user-triggered agent,
running at roughly 20-25 million requests a day. This is the crawler named
on our own &lt;a href="https://letslaunch.today/blog/ai-crawler-cheat-sheet" rel="noopener noreferrer"&gt;cheat sheet&lt;/a&gt; and it isn't the
accusation — a declared agent fetching a page a user asked about is normal
and expected.&lt;/li&gt;
&lt;li&gt;A second, undeclared crawler, at roughly 3-6 million requests a day,
identifying itself as an ordinary Chrome browser on macOS. According to
Cloudflare, this traffic used IP ranges outside Perplexity's official
published ranges and rotated across different network providers — different
ASNs — in a pattern consistent with avoiding detection as a single
identifiable source.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That second pattern is the entire allegation. Not that Perplexity crawls the&lt;br&gt;
web — every retrieval-based assistant does that, openly, and we cover the&lt;br&gt;
mechanics of it elsewhere. The allegation is that when the declared crawler&lt;br&gt;
was blocked, a second, disguised one kept going anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Cloudflare did about it
&lt;/h2&gt;

&lt;p&gt;Cloudflare didn't just publish a blog post and leave it there. It removed&lt;br&gt;
Perplexity from its "verified bots" program, added detection rules for the&lt;br&gt;
disguised traffic pattern to its managed rule sets, and extended the ability&lt;br&gt;
to block that traffic to every customer — including free-tier sites that&lt;br&gt;
previously had no way to act on this kind of pattern even if they noticed it.&lt;/p&gt;

&lt;p&gt;That's the detail worth sitting with. A &lt;code&gt;robots.txt&lt;/code&gt; disallow line is&lt;br&gt;
something any site can write, for free, unilaterally. Detecting a crawler that&lt;br&gt;
is actively trying not to look like a crawler — rotating IP ranges and network&lt;br&gt;
providers to defeat pattern-matching — is not something most site owners can&lt;br&gt;
do themselves. It needs traffic volume and infrastructure at Cloudflare's&lt;br&gt;
scale to even notice the pattern, let alone block it. According to Cloudflare,&lt;br&gt;
over 2.5 million websites added AI-crawler blocking rules within a month of&lt;br&gt;
the related tooling becoming available — a number worth reading as&lt;br&gt;
approximate, since it's Cloudflare's own reporting on adoption of its own&lt;br&gt;
feature, not an independently audited count.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Perplexity says
&lt;/h2&gt;

&lt;p&gt;Perplexity disputes the characterization. Multiple outlets reported a company&lt;br&gt;
spokesperson pushing back, disputing that the undeclared crawler belonged to&lt;br&gt;
Perplexity at all, or disputing that any blocked content was actually&lt;br&gt;
accessed. We're stating that as what it is: a public rebuttal, not a&lt;br&gt;
concession, and not something we can independently verify from here.&lt;/p&gt;

&lt;p&gt;We're not going to adjudicate that dispute for you. What we can say plainly is&lt;br&gt;
that Cloudflare's evidence is unusually concrete for this kind of claim — a&lt;br&gt;
controlled experiment on domains nobody could have found any other way,&lt;br&gt;
compared to the more common pattern of "our traffic logs look suspicious,"&lt;br&gt;
which is much easier to wave away. That doesn't make the allegation settled.&lt;br&gt;
It makes it a specific, checkable claim rather than a vibe, which is more than&lt;br&gt;
most accusations like this get.&lt;/p&gt;

&lt;h2&gt;
  
  
  This isn't how OpenAI's crawlers have been reported to behave
&lt;/h2&gt;

&lt;p&gt;Worth one paragraph of context: multiple reports have described OpenAI's&lt;br&gt;
crawlers — the same &lt;code&gt;GPTBot&lt;/code&gt; and &lt;code&gt;OAI-SearchBot&lt;/code&gt; covered on our&lt;br&gt;
&lt;a href="https://letslaunch.today/blog/ai-crawler-cheat-sheet" rel="noopener noreferrer"&gt;cheat sheet&lt;/a&gt; — actually stopping when a site&lt;br&gt;
disallows them via &lt;code&gt;robots.txt&lt;/code&gt;. That's a reported contrast, not a claim that&lt;br&gt;
OpenAI's crawlers are flawless everywhere at all times; we've said elsewhere&lt;br&gt;
that &lt;code&gt;robots.txt&lt;/code&gt; compliance across the industry is a norm the major crawlers&lt;br&gt;
have generally followed, not a legal guarantee that binds any of them. This&lt;br&gt;
incident is the sharpest documented exception to that norm we're aware of. It&lt;br&gt;
is not proof the norm doesn't exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this actually changes for a site owner
&lt;/h2&gt;

&lt;p&gt;Nothing about this changes the mechanics we've written about elsewhere — see&lt;br&gt;
&lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;why ChatGPT can't see your site&lt;/a&gt; for&lt;br&gt;
the separate, much more common problem of crawlers simply not rendering your&lt;br&gt;
JavaScript. What it changes is how much you can trust a &lt;code&gt;robots.txt&lt;/code&gt; disallow&lt;br&gt;
line to actually be the whole story.&lt;/p&gt;

&lt;p&gt;The honest takeaway is narrow. A &lt;code&gt;robots.txt&lt;/code&gt; rule is a request you cannot&lt;br&gt;
verify is being honored by every crawler that might read it, and this incident&lt;br&gt;
is the concrete case for why — not a reason to assume the file is theater,&lt;br&gt;
since the record shows most named crawlers do follow it. If you want to check&lt;br&gt;
whether your own site is actually being blocked the way your &lt;code&gt;robots.txt&lt;/code&gt;&lt;br&gt;
claims, our &lt;a href="https://letslaunch.today/free/ai-crawler-check" rel="noopener noreferrer"&gt;AI crawler check&lt;/a&gt; fetches your page as&lt;br&gt;
each named crawler and shows you what came back, which tells you what's&lt;br&gt;
happening to declared traffic. It cannot tell you whether an undeclared&lt;br&gt;
crawler is also getting through — that's the part this whole story shows a&lt;br&gt;
text file was never going to catch, and why Cloudflare's fix lived at the&lt;br&gt;
network layer instead of in anyone's &lt;code&gt;robots.txt&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;And a separate reminder, because it's easy to conflate the two: none of this&lt;br&gt;
tells you anything about whether being crawled — declared or not — makes an&lt;br&gt;
assistant more likely to cite you. That's a different, unresolved question,&lt;br&gt;
and we've written about what actually correlates with citation in&lt;br&gt;
&lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;what makes AI engines cite a source&lt;/a&gt;.&lt;br&gt;
Getting crawled and getting cited are not the same problem, and this story is&lt;br&gt;
only about the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did Perplexity admit to stealth crawling?
&lt;/h2&gt;

&lt;p&gt;No. Cloudflare published detailed findings — including a controlled experiment&lt;br&gt;
on unlinked test domains — alleging that an undeclared, disguised crawler kept&lt;br&gt;
accessing sites that had blocked Perplexity's declared crawler. Perplexity&lt;br&gt;
publicly disputed the characterization through a company spokesperson. This is&lt;br&gt;
a documented, disputed claim, not a settled admission.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does this mean robots.txt doesn't work against AI crawlers?
&lt;/h2&gt;

&lt;p&gt;No. The major named crawlers, including OpenAI's, have a general track record&lt;br&gt;
of respecting &lt;code&gt;robots.txt&lt;/code&gt;, which is what our &lt;a href="https://letslaunch.today/blog/ai-crawler-cheat-sheet" rel="noopener noreferrer"&gt;cheat sheet&lt;/a&gt;&lt;br&gt;
is built around. This incident is the clearest publicly documented exception&lt;br&gt;
to that track record, not evidence that compliance across the industry is&lt;br&gt;
fake. Treat it as a reason to verify rather than a reason to give up on the&lt;br&gt;
file entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  How did Cloudflare detect the alleged stealth crawling?
&lt;/h2&gt;

&lt;p&gt;By creating test domains with a &lt;code&gt;robots.txt&lt;/code&gt; disallowing every crawler, never&lt;br&gt;
linking or indexing them anywhere, and then asking Perplexity's AI assistant&lt;br&gt;
about content on those exact domains. According to Cloudflare, the assistant&lt;br&gt;
answered with accurate detail about pages it had no legitimate way to know&lt;br&gt;
about, which pointed to a crawler ignoring the disallow rule rather than any&lt;br&gt;
public index.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>Is AI Referral Traffic Actually Worth Having? What the 2026 Numbers Say</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Tue, 15 Sep 2026 17:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/is-ai-referral-traffic-actually-worth-having-what-the-2026-numbers-say-58a2</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/is-ai-referral-traffic-actually-worth-having-what-the-2026-numbers-say-58a2</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/should-you-want-ai-referral-traffic" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Everything we've written about AI citations on this blog so far has been&lt;br&gt;
about the mechanism: can a crawler see your page, what correlates with being&lt;br&gt;
cited, which platform is worth your time. All of it stops short of the one&lt;br&gt;
question a founder actually cares about once traffic starts showing up in&lt;br&gt;
the logs as &lt;code&gt;chatgpt.com&lt;/code&gt; or &lt;code&gt;perplexity.ai&lt;/code&gt; in the referrer field: is it any&lt;br&gt;
good?&lt;/p&gt;

&lt;p&gt;That's a different kind of question, and it has a different kind of answer.&lt;br&gt;
Nobody has shown a reliable way to cause more AI citations — see &lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;what&lt;br&gt;
actually gets you cited by ChatGPT and&lt;br&gt;
Perplexity&lt;/a&gt; for why we won't&lt;br&gt;
pretend otherwise. But whether the traffic that does show up is worth having&lt;br&gt;
is a separate claim, and it's one several named analytics firms have actually&lt;br&gt;
measured, on real production traffic, not a lab benchmark. We're not going to&lt;br&gt;
blend their numbers into one tidy statistic — they don't agree with each&lt;br&gt;
other closely enough for that — but they agree on the direction, and that&lt;br&gt;
agreement is worth taking seriously.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, by source
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Seer Interactive&lt;/strong&gt;, a digital-marketing analytics firm, measured&lt;br&gt;
conversion rates by referral source and found ChatGPT-referred traffic&lt;br&gt;
converting at roughly 15.9%, against roughly 1.76% for Google organic search&lt;br&gt;
traffic in the same measurement — call it roughly a 9x difference. That's&lt;br&gt;
specifically ChatGPT compared against organic search, not "AI traffic" as a&lt;br&gt;
category. The same analysis measured Perplexity at around 10.5% and Claude at&lt;br&gt;
around 5.0%. All three beat the organic baseline by a wide margin, but they&lt;br&gt;
don't match each other — ChatGPT, Perplexity and Claude converted at three&lt;br&gt;
meaningfully different rates in the same study, which is consistent with&lt;br&gt;
something we've said elsewhere on this blog about AI platforms generally:&lt;br&gt;
they are not interchangeable, and treating "AI referral traffic" as one&lt;br&gt;
undifferentiated bucket erases a real spread in the underlying data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adobe Analytics&lt;/strong&gt;, in an April 2026 report covering March 2026 e-commerce&lt;br&gt;
data, found AI-referred shoppers converting 42% better than non-AI traffic,&lt;br&gt;
spending 48% more time on product pages, and generating 37% higher revenue&lt;br&gt;
per visit. This is worth flagging clearly as a retail-specific measurement —&lt;br&gt;
Adobe's data comes from shopping behavior, not from SaaS sign-ups, content&lt;br&gt;
engagement or lead forms. A founder running a B2B tool should read this as&lt;br&gt;
"directionally consistent with the pattern," not as a number that transfers&lt;br&gt;
to their own funnel.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Conductor&lt;/strong&gt;, an SEO and content platform, reported in its own 2026&lt;br&gt;
benchmark that ChatGPT accounts for roughly 87.4% of all AI referral traffic&lt;br&gt;
across the industries it measured, with every other assistant splitting the&lt;br&gt;
remainder. We already made a version of this point in &lt;a href="https://letslaunch.today/blog/chatgpt-vs-google-ai-overviews-vs-perplexity" rel="noopener noreferrer"&gt;our comparison of&lt;br&gt;
ChatGPT, Google AI Overviews and&lt;br&gt;
Perplexity&lt;/a&gt; using reach&lt;br&gt;
and user counts — Conductor's number reinforces it from a different angle,&lt;br&gt;
using actual referred visits rather than platform-level user counts. Either&lt;br&gt;
way you measure it, "AI traffic" is, for most sites in most categories, in&lt;br&gt;
practice overwhelmingly a ChatGPT question. Building a strategy evenly split&lt;br&gt;
across four assistants is optimizing for a distribution that doesn't match&lt;br&gt;
what shows up in the logs.&lt;/p&gt;

&lt;p&gt;Here's all three side by side, with what each one actually measured:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What it measured&lt;/th&gt;
&lt;th&gt;Headline figure&lt;/th&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Seer Interactive&lt;/td&gt;
&lt;td&gt;Conversion rate by referral source&lt;/td&gt;
&lt;td&gt;ChatGPT ~15.9% vs. organic ~1.76% (~9x); Perplexity ~10.5%; Claude ~5.0%&lt;/td&gt;
&lt;td&gt;Cross-industry conversion comparison&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seer Interactive&lt;/td&gt;
&lt;td&gt;Pages per session by referral source&lt;/td&gt;
&lt;td&gt;ChatGPT ~2.3 vs. organic ~1.2; Perplexity/Gemini trend lower and more decisive&lt;/td&gt;
&lt;td&gt;Engagement depth, not conversion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Adobe Analytics&lt;/td&gt;
&lt;td&gt;E-commerce shopper behavior, March 2026&lt;/td&gt;
&lt;td&gt;+42% conversion, +48% time on product pages, +37% revenue per visit&lt;/td&gt;
&lt;td&gt;Retail/shopping specifically&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conductor&lt;/td&gt;
&lt;td&gt;Share of AI referral traffic by platform&lt;/td&gt;
&lt;td&gt;ChatGPT ~87.4% of AI referrals&lt;/td&gt;
&lt;td&gt;Referred-visit volume, not user counts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  They don't agree, and that's worth saying plainly
&lt;/h2&gt;

&lt;p&gt;A 9x conversion lift and a 42% conversion lift are not the same claim, and we&lt;br&gt;
are not going to pretend they triangulate to some cleaner "real" number in&lt;br&gt;
the middle. They come from different firms, measuring different traffic in&lt;br&gt;
different time windows, using different methodologies and almost certainly&lt;br&gt;
different definitions of "conversion" — a purchase in Adobe's retail data&lt;br&gt;
means something different from a completed signup in Seer's cross-industry&lt;br&gt;
sample. Multiplying a nine-fold conversion difference against a&lt;br&gt;
forty-two-percent one and reporting an average would manufacture precision&lt;br&gt;
neither source claims to have.&lt;/p&gt;

&lt;p&gt;What's consistent, and what we're comfortable stating plainly, is the&lt;br&gt;
direction: every named source that measured this in 2026 found AI-referred&lt;br&gt;
traffic converting meaningfully better than organic search or non-AI&lt;br&gt;
traffic. The size of that gap varies a lot — by platform, by industry, by&lt;br&gt;
what "conversion" was defined to mean. Treat the specific multiples as&lt;br&gt;
illustrative of a real and repeated pattern, not as a number you can plug&lt;br&gt;
into your own forecast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Engagement isn't the same story on every platform
&lt;/h2&gt;

&lt;p&gt;Seer's data on pages-per-session is worth reading carefully rather than as&lt;br&gt;
one more "AI traffic is better" data point, because it complicates a simple&lt;br&gt;
story. ChatGPT-referred visitors browsed roughly 2.3 pages per session&lt;br&gt;
against roughly 1.2 for organic — more exploration, more time in the site.&lt;br&gt;
But Perplexity- and Gemini-referred visitors trended the opposite way:&lt;br&gt;
fewer pages, more focused, more decisive visits. Both patterns are&lt;br&gt;
consistent with higher-value traffic; they just look different. A founder&lt;br&gt;
looking at a low pages-per-session number from Perplexity referrals&lt;br&gt;
shouldn't read that as underperformance against the ChatGPT number — it may&lt;br&gt;
be the opposite, a visitor who already knew exactly what they wanted and&lt;br&gt;
went and did it. "More engaged" doesn't mean the same shape of behavior&lt;br&gt;
across every platform, any more than the conversion numbers mean the same&lt;br&gt;
percentage.&lt;/p&gt;

&lt;h2&gt;
  
  
  A plausible reason, stated as a guess
&lt;/h2&gt;

&lt;p&gt;None of the sources above measured &lt;em&gt;why&lt;/em&gt; this traffic converts better — they&lt;br&gt;
measured that it does. But there's a reasonable explanation that's&lt;br&gt;
consistent with all of it, and it's worth naming as long as it's labeled for&lt;br&gt;
what it is: an inference, not a finding.&lt;/p&gt;

&lt;p&gt;A visitor who clicks through from a citation inside a generated answer has&lt;br&gt;
typically already had part of their question answered, and their intent&lt;br&gt;
sharpened, by the assistant before they ever land on your page. They didn't&lt;br&gt;
arrive from a broad, exploratory query the way a lot of organic search&lt;br&gt;
traffic does — a meaningful share of organic clicks are early-stage research&lt;br&gt;
that was never going to convert regardless of what the landing page says.&lt;br&gt;
AI referral traffic, by construction, has already passed through one round&lt;br&gt;
of qualification a search click hasn't. That's a plausible mechanism&lt;br&gt;
consistent with a 9x lift and a 42% lift both pointing the same direction.&lt;br&gt;
It is not something Seer, Adobe or Conductor directly measured, and we're&lt;br&gt;
not going to present it as more than the reasonable guess it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this actually changes
&lt;/h2&gt;

&lt;p&gt;It doesn't change the honest answer to "how do I get more AI citations,"&lt;br&gt;
which remains: nobody outside these companies can currently tell you a&lt;br&gt;
reliable way to make that happen on purpose. What it changes is the answer&lt;br&gt;
to a different, more modest question — is it worth doing the low-cost things&lt;br&gt;
anyway, on the chance some AI traffic shows up. Before this data, the honest&lt;br&gt;
case for being crawlable, describing your product consistently across&lt;br&gt;
independent pages, and picking up a free listing somewhere like a &lt;a href="https://letslaunch.today/submit" rel="noopener noreferrer"&gt;LetsLaunch&lt;br&gt;
submission&lt;/a&gt; rested entirely on cost: these things are free or&lt;br&gt;
close to it, so there's no real argument against doing them even with an&lt;br&gt;
unproven mechanism.&lt;/p&gt;

&lt;p&gt;Now there's a second argument, and it's a real one even though it's&lt;br&gt;
narrower than it sounds: if any of that effort does eventually produce a&lt;br&gt;
citation and a click, multiple independent, named sources measuring real&lt;br&gt;
production traffic in 2026 found that click converting meaningfully better&lt;br&gt;
than an average organic visit. That's not a promise you'll get more AI&lt;br&gt;
traffic. It's a reason to think the traffic is worth having if it comes.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much better does AI referral traffic convert than organic search?
&lt;/h2&gt;

&lt;p&gt;It depends heavily on the source and the platform, and the studies don't&lt;br&gt;
agree on a single multiple. Seer Interactive measured ChatGPT referrals&lt;br&gt;
converting at roughly 15.9% against roughly 1.76% for organic — about 9x —&lt;br&gt;
with Perplexity around 10.5% and Claude around 5.0% in the same comparison.&lt;br&gt;
Adobe Analytics, measuring e-commerce specifically, found a 42% conversion&lt;br&gt;
lift for AI-referred shoppers. Both point the same direction; neither&lt;br&gt;
number should be treated as the other's substitute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is AI referral traffic conversion data reliable, or is this the same kind of unproven claim as AI citation tactics?
&lt;/h2&gt;

&lt;p&gt;These are different kinds of claims. Whether a specific action causes more&lt;br&gt;
AI citations is unproven — see our review of the citation research. Whether&lt;br&gt;
AI-referred traffic converts better once it arrives is an observed outcome,&lt;br&gt;
independently measured by multiple named firms (Seer Interactive, Adobe&lt;br&gt;
Analytics, Conductor) on real production traffic rather than a lab&lt;br&gt;
benchmark. The exact multiple varies by study, but the direction is&lt;br&gt;
consistent across all of them, which is a meaningfully stronger evidentiary&lt;br&gt;
position than the citation-mechanism claims we're skeptical of elsewhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is most "AI traffic" really just ChatGPT?
&lt;/h2&gt;

&lt;p&gt;For most sites in most categories, yes. Conductor's 2026 benchmark found&lt;br&gt;
ChatGPT accounted for roughly 87.4% of all AI referral traffic across the&lt;br&gt;
industries it measured, with the remainder split among other assistants.&lt;br&gt;
That matches what reach and user-count data already suggest — see our&lt;br&gt;
comparison of ChatGPT, Google AI Overviews and Perplexity — but Conductor's&lt;br&gt;
number is based on actual referred visits rather than platform-level user&lt;br&gt;
counts, which makes it a stronger version of the same point.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>Every AI Crawler That Hits Your Site, and What to Do About Each One</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/every-ai-crawler-that-hits-your-site-and-what-to-do-about-each-one-489b</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/every-ai-crawler-that-hits-your-site-and-what-to-do-about-each-one-489b</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/ai-crawler-cheat-sheet" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every AI crawler identifies itself with a distinct user-agent string, and the&lt;br&gt;
strings do not map cleanly onto "AI" as one category. Some train models. Some&lt;br&gt;
power a live answer with a citation. Some only fetch a page because a person&lt;br&gt;
pasted its URL into a chat window a minute earlier. A &lt;code&gt;robots.txt&lt;/code&gt; rule written&lt;br&gt;
against "AI bots" as if they were one thing usually does something other than&lt;br&gt;
what its author intended.&lt;/p&gt;

&lt;p&gt;This is the reference list: every crawler we could verify, who operates it,&lt;br&gt;
what it actually does, and the specific &lt;code&gt;robots.txt&lt;/code&gt; line for it. Use it to&lt;br&gt;
write the rule you actually mean, not the closest wildcard you found in&lt;br&gt;
someone else's blog post.&lt;/p&gt;
&lt;h2&gt;
  
  
  The three jobs a crawler can be doing
&lt;/h2&gt;

&lt;p&gt;Before the table, the distinction that the whole list hangs on. A crawler&lt;br&gt;
identifying as AI-related is doing exactly one of three jobs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Training.&lt;/strong&gt; It fetches your pages to become part of a dataset a model is
trained on, sometime later, in a batch you have no visibility into. Blocking
it has no effect on anything that happens today.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval and citation.&lt;/strong&gt; It fetches your pages to build or refresh an
index that a live answer engine — ChatGPT Search, Perplexity, Copilot — reads
from at answer time. Blocking it removes you from those answers, full stop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User-triggered fetch.&lt;/strong&gt; Nobody crawls anything. A person pasted your URL,
or asked about it by name, and the assistant fetched that one page, once,
live, to answer that one question. This isn't a crawl in the ordinary sense
— it's the assistant acting as the user's browser for a single request.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We laid out the training-versus-retrieval split in more depth in&lt;br&gt;
&lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;why ChatGPT can't see your site&lt;/a&gt;; this&lt;br&gt;
post exists to give every crawler its own row instead of two examples.&lt;/p&gt;
&lt;h2&gt;
  
  
  OpenAI
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GPTBot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Trains OpenAI's models. Respects &lt;code&gt;robots.txt&lt;/code&gt;, per OpenAI's own published crawler documentation.&lt;/td&gt;
&lt;td&gt;Block if you want out of training; costs nothing in ChatGPT's answers. Allow if you're fine being training data.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;OAI-SearchBot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Powers ChatGPT Search's retrieval and citation index. A separate crawler from &lt;code&gt;GPTBot&lt;/code&gt;, launched alongside ChatGPT Search in October 2024.&lt;/td&gt;
&lt;td&gt;Allow, unless you want to be invisible to ChatGPT Search entirely.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ChatGPT-User&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;User-triggered: someone inside ChatGPT pasted or asked about your specific URL and the assistant fetched it live. Not a routine crawl.&lt;/td&gt;
&lt;td&gt;Allow. Blocking it means a real person who linked to you gets nothing back, for no gain — there's no dataset being built to opt out of.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  Anthropic
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ClaudeBot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Anthropic's primary training crawler, functionally parallel to &lt;code&gt;GPTBot&lt;/code&gt;.&lt;/td&gt;
&lt;td&gt;Block for the same reason as &lt;code&gt;GPTBot&lt;/code&gt;: opts out of training, changes nothing about Claude's live answers.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Claude-SearchBot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Retrieval and citation crawler, parallel to &lt;code&gt;OAI-SearchBot&lt;/code&gt;.&lt;/td&gt;
&lt;td&gt;Allow if you want to be citable in Claude's answers.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Claude-User&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;User-triggered fetch, parallel to &lt;code&gt;ChatGPT-User&lt;/code&gt;.&lt;/td&gt;
&lt;td&gt;Allow, same reasoning as &lt;code&gt;ChatGPT-User&lt;/code&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;OpenAI and Anthropic both explicitly support the training/retrieval split —&lt;br&gt;
that's the entire reason two separate crawlers exist per company instead of&lt;br&gt;
one. If you want out of training but still want to be citable, that's not a&lt;br&gt;
compromise you're forcing on the system; it's the exact use case the split&lt;br&gt;
was built for.&lt;/p&gt;
&lt;h2&gt;
  
  
  Perplexity
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;PerplexityBot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Perplexity's primary crawler, builds its retrieval and citation index.&lt;/td&gt;
&lt;td&gt;Allow if you want to appear in Perplexity's answers. Block only if you want out entirely — there's no separate training-only variant to opt out of instead.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Perplexity-User&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;User-triggered fetch, parallel to &lt;code&gt;ChatGPT-User&lt;/code&gt;.&lt;/td&gt;
&lt;td&gt;Allow, same reasoning as the other user-triggered fetchers.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  Google
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Googlebot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Ordinary search crawler. Required to rank in Google Search at all.&lt;/td&gt;
&lt;td&gt;Allow. This is not an AI-crawler decision — blocking it removes you from Google Search, not from an AI feature.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Google-Extended&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;An opt-out signal specifically for Google's AI training uses (Gemini, AI features). Separate from &lt;code&gt;Googlebot&lt;/code&gt;.&lt;/td&gt;
&lt;td&gt;Block if you want out of Google's AI training. Blocking it does not affect Google Search ranking — that's the entire point of it existing as its own token.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is the pair people mix up most often, because both names start with&lt;br&gt;
"Google" and only one of them touches your search rankings. &lt;code&gt;Googlebot&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;Google-Extended&lt;/code&gt; are read by Google's systems for entirely different&lt;br&gt;
purposes, and disallowing the wrong one either does nothing or costs you your&lt;br&gt;
search visibility, depending on which way you get it backwards.&lt;/p&gt;
&lt;h2&gt;
  
  
  Apple
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Applebot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Ordinary crawler behind Siri and Spotlight search features.&lt;/td&gt;
&lt;td&gt;Allow, unless you specifically want out of Apple's search surfaces.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Applebot-Extended&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Apple's AI-training opt-out signal, parallel in purpose to &lt;code&gt;Google-Extended&lt;/code&gt;. Separate from ordinary &lt;code&gt;Applebot&lt;/code&gt;.&lt;/td&gt;
&lt;td&gt;Block if you want out of Apple's AI training without touching Siri/Spotlight visibility.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  Microsoft
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Bingbot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Microsoft's ordinary search crawler — and also the retrieval backbone behind Copilot's web-grounded answers. Plays double duty the way &lt;code&gt;Googlebot&lt;/code&gt; does for Google, except Microsoft hasn't split it into two tokens.&lt;/td&gt;
&lt;td&gt;Allow. Blocking it costs you both Bing Search ranking and Copilot citations at once, because there's no separate opt-out for one without the other.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  Meta
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Meta-ExternalAgent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Meta's AI-training crawler, for Llama and Meta AI features.&lt;/td&gt;
&lt;td&gt;Block if you want out of training for Meta's models.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Meta-ExternalFetcher&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;User-triggered fetch, parallel to &lt;code&gt;ChatGPT-User&lt;/code&gt;, for links a user shares with Meta AI.&lt;/td&gt;
&lt;td&gt;Allow, same reasoning as the other user-triggered fetchers.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  Everyone else on the list
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler&lt;/th&gt;
&lt;th&gt;Operator&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Amazonbot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Amazon&lt;/td&gt;
&lt;td&gt;Crawls for Amazon's AI/Alexa-related features.&lt;/td&gt;
&lt;td&gt;Allow unless you specifically want out of Amazon's AI features — no training/retrieval split is documented for it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DuckAssistBot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;DuckDuckGo&lt;/td&gt;
&lt;td&gt;Powers DuckDuckGo's AI-assisted answer feature.&lt;/td&gt;
&lt;td&gt;Allow if you want to be citable in DuckDuckGo's AI answers; this is a retrieval crawler, so blocking it removes you from those answers.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Bytespider&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ByteDance (owner of TikTok)&lt;/td&gt;
&lt;td&gt;Crawls to power TikTok search, content recommendations, and ByteDance's own AI features.&lt;/td&gt;
&lt;td&gt;Your call — bundles search, recommendations and AI training into one crawler with no separate opt-out for any one piece.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;CCBot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Common Crawl&lt;/td&gt;
&lt;td&gt;A nonprofit that crawls the open web and republishes the dataset for anyone to use — researchers and AI labs included. Not itself an AI company, but its dataset is a common ingredient in many models' training data.&lt;/td&gt;
&lt;td&gt;Block if your goal is keeping content out of AI training generally, since the dataset feeds more than one lab at once. Allow if you're fine contributing to open research datasets.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Diffbot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Diffbot (commercial web-data extraction)&lt;/td&gt;
&lt;td&gt;Crawls on behalf of its own customers to supply structured data for AI training, search, and retrieval-augmented generation products.&lt;/td&gt;
&lt;td&gt;Treat like &lt;code&gt;CCBot&lt;/code&gt;: it's a pass-through to unknown downstream customers rather than one named consumer product, so block it if you want a hard boundary on who resells your content.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Timpibot&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Timpi&lt;/td&gt;
&lt;td&gt;Builds a decentralized search engine's index and collects training data for LLMs. Smaller and less well-known than the crawlers above, but real and documented.&lt;/td&gt;
&lt;td&gt;Allow or block on the same logic as any other combined search/training crawler — there's no split to take advantage of.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cohere-ai&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Cohere&lt;/td&gt;
&lt;td&gt;Training/retrieval crawler for Cohere's models.&lt;/td&gt;
&lt;td&gt;Allow if you want to be usable by Cohere's products; block to opt out. LetsLaunch's own &lt;code&gt;robots.txt&lt;/code&gt; names this one explicitly rather than relying on a wildcard.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;MistralAI-User&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Mistral AI&lt;/td&gt;
&lt;td&gt;Training/retrieval crawler for Mistral's models.&lt;/td&gt;
&lt;td&gt;Same logic as &lt;code&gt;cohere-ai&lt;/code&gt;. LetsLaunch's &lt;code&gt;robots.txt&lt;/code&gt; names this one too.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;h2&gt;
  
  
  What LetsLaunch actually does
&lt;/h2&gt;

&lt;p&gt;Our own &lt;code&gt;robots.txt&lt;/code&gt; (&lt;code&gt;src/app/robots.ts&lt;/code&gt; in this codebase) names sixteen&lt;br&gt;
agents explicitly in an &lt;code&gt;AI_SEARCH_AGENTS&lt;/code&gt; group and allows all of them,&lt;br&gt;
including &lt;code&gt;cohere-ai&lt;/code&gt; and &lt;code&gt;MistralAI-User&lt;/code&gt; alongside the bigger names — a&lt;br&gt;
listing cited inside any of these answer engines is worth as much to a maker&lt;br&gt;
as an ordinary search result, so we chose visibility over opting out. The code&lt;br&gt;
comment on that group is direct about the user-triggered fetchers too:&lt;br&gt;
&lt;code&gt;ChatGPT-User&lt;/code&gt;, &lt;code&gt;Claude-User&lt;/code&gt; and &lt;code&gt;Perplexity-User&lt;/code&gt; ignore &lt;code&gt;robots.txt&lt;/code&gt; by&lt;br&gt;
design, so naming them there is a statement of intent, not a control that&lt;br&gt;
actually stops anything.&lt;/p&gt;

&lt;p&gt;That's our call for a directory that exists to get products seen. It is not&lt;br&gt;
a universal recommendation. A publisher with a subscription paywall, or&lt;br&gt;
content it doesn't want folded into a model's training data regardless of&lt;br&gt;
citation upside, has a legitimate reason to block the training crawlers on&lt;br&gt;
this list while leaving the retrieval and user-triggered ones alone — that&lt;br&gt;
split is exactly what separate tokens like &lt;code&gt;GPTBot&lt;/code&gt;/&lt;code&gt;OAI-SearchBot&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;ClaudeBot&lt;/code&gt;/&lt;code&gt;Claude-SearchBot&lt;/code&gt; are for.&lt;/p&gt;
&lt;h2&gt;
  
  
  Writing the actual rule
&lt;/h2&gt;

&lt;p&gt;Robots.txt groups by user-agent are independent — a crawler that finds a group&lt;br&gt;
naming it by name ignores the &lt;code&gt;*&lt;/code&gt; group entirely, so a specific agent needs&lt;br&gt;
its own disallow lines if you want to block just that one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight robot_framework"&gt;&lt;code&gt;User-agent: GPTBot&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Disallow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;User-agent:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;ClaudeBot&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Disallow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;User-agent:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;OAI-SearchBot&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;User-agent:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;*&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Allow:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That example opts out of OpenAI and Anthropic training specifically while&lt;br&gt;
staying allowed everywhere else, including for &lt;code&gt;OAI-SearchBot&lt;/code&gt;'s retrieval&lt;br&gt;
index. Copy the shape, not the specific agents — decide per crawler, from the&lt;br&gt;
tables above, rather than pasting a blanket AI-blocking snippet that treats&lt;br&gt;
every name on this page as the same decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking whether it's actually working
&lt;/h2&gt;

&lt;p&gt;A &lt;code&gt;robots.txt&lt;/code&gt; rule is a request, not an enforcement mechanism — the polite&lt;br&gt;
crawlers on this list respect it, but the file itself can't verify anything.&lt;br&gt;
Once you've written the rule you mean, two different checks matter and they&lt;br&gt;
answer different questions. Our &lt;a href="https://letslaunch.today/free/ai-crawler-check" rel="noopener noreferrer"&gt;AI crawler check&lt;/a&gt;&lt;br&gt;
fetches your page as each named crawler once and shows you the status code and&lt;br&gt;
the content each one received right now — useful for confirming a rule change&lt;br&gt;
took effect. For whether a crawler actually shows up over time, on its own&lt;br&gt;
schedule, we wrote a &lt;a href="https://letslaunch.today/blog/monitor-ai-crawler-traffic" rel="noopener noreferrer"&gt;separate guide&lt;/a&gt; to&lt;br&gt;
grepping logs and reading Cloudflare's AI Crawlers dashboard — that's an&lt;br&gt;
ongoing question a one-time fetch can't answer.&lt;/p&gt;

&lt;p&gt;If you're launching a product and want it reachable by name across these&lt;br&gt;
answer engines in the first place, a &lt;a href="https://letslaunch.today/submit" rel="noopener noreferrer"&gt;LetsLaunch listing&lt;/a&gt; is one more&lt;br&gt;
independent page describing it the same way your own site does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does blocking GPTBot remove me from ChatGPT's answers?
&lt;/h2&gt;

&lt;p&gt;No. &lt;code&gt;GPTBot&lt;/code&gt; only collects training data; &lt;code&gt;OAI-SearchBot&lt;/code&gt; is the separate&lt;br&gt;
crawler that builds the index behind ChatGPT Search's citations. Blocking&lt;br&gt;
&lt;code&gt;GPTBot&lt;/code&gt; opts you out of training runs and has no effect on whether ChatGPT&lt;br&gt;
can cite you — only blocking &lt;code&gt;OAI-SearchBot&lt;/code&gt; does that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does blocking Google-Extended hurt my Google ranking?
&lt;/h2&gt;

&lt;p&gt;No. &lt;code&gt;Google-Extended&lt;/code&gt; is a distinct opt-out token for Google's AI training&lt;br&gt;
uses, separate from &lt;code&gt;Googlebot&lt;/code&gt;, which is the ordinary search crawler required&lt;br&gt;
for ranking in Google Search at all. Blocking &lt;code&gt;Google-Extended&lt;/code&gt; leaves&lt;br&gt;
&lt;code&gt;Googlebot&lt;/code&gt; and your search ranking untouched.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should I block the user-triggered fetchers like ChatGPT-User?
&lt;/h2&gt;

&lt;p&gt;Usually not. &lt;code&gt;ChatGPT-User&lt;/code&gt;, &lt;code&gt;Claude-User&lt;/code&gt;, &lt;code&gt;Perplexity-User&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;Meta-ExternalFetcher&lt;/code&gt; don't crawl your site on a schedule — they fetch one&lt;br&gt;
page, once, because a real person pasted your link or asked about it by name&lt;br&gt;
inside that assistant. Blocking them doesn't opt you out of any training&lt;br&gt;
dataset; it means that specific person gets nothing back.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>ChatGPT vs Google AI Overviews vs Perplexity: Where to Put Limited Effort</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Mon, 14 Sep 2026 17:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/chatgpt-vs-google-ai-overviews-vs-perplexity-where-to-put-limited-effort-3ai</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/chatgpt-vs-google-ai-overviews-vs-perplexity-where-to-put-limited-effort-3ai</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/chatgpt-vs-google-ai-overviews-vs-perplexity" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every "how to get cited by AI" article treats "AI" as one thing with one&lt;br&gt;
mechanism. It isn't. ChatGPT, Google AI Overviews and Perplexity are three&lt;br&gt;
different products, built by three different companies, with wildly different&lt;br&gt;
user counts and — more importantly for a founder deciding where to spend&lt;br&gt;
time — genuinely different ways of deciding what to put in front of a reader.&lt;/p&gt;

&lt;p&gt;This isn't another post about the mechanism of getting cited. We've written&lt;br&gt;
those; see &lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;what actually gets you cited by ChatGPT and&lt;br&gt;
Perplexity&lt;/a&gt; and &lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;why ChatGPT can't&lt;br&gt;
see your site&lt;/a&gt; if that's what you're&lt;br&gt;
after. This one is about the products themselves — who uses each, how big&lt;br&gt;
each actually is, and what each one draws from when it answers — so a small&lt;br&gt;
team can decide where limited effort has any chance of paying off before&lt;br&gt;
spending it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three, side by side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;ChatGPT (search/browsing)&lt;/th&gt;
&lt;th&gt;Google AI Overviews&lt;/th&gt;
&lt;th&gt;Perplexity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reach&lt;/td&gt;
&lt;td&gt;1B+ weekly active users (OpenAI, reported mid-2026)&lt;/td&gt;
&lt;td&gt;Appears in an estimated ~25-45% of Google searches, depending on measurement firm and query type&lt;/td&gt;
&lt;td&gt;~230M monthly active users, 1B+ queries/month (Perplexity's own reported figures, early 2026)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What it's built on&lt;/td&gt;
&lt;td&gt;Its own retrieval + browsing layer on top of a general model&lt;/td&gt;
&lt;td&gt;Google's existing web index, enriched with authority/freshness signals&lt;/td&gt;
&lt;td&gt;Its own retrieval layer, tuned toward recent and community content&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Leans on&lt;/td&gt;
&lt;td&gt;Wikipedia heavily among top cited sources&lt;/td&gt;
&lt;td&gt;The same web index that ranks ordinary Google Search&lt;/td&gt;
&lt;td&gt;Reddit and recent content heavily&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Separate destination?&lt;/td&gt;
&lt;td&gt;Yes — its own product, its own habit&lt;/td&gt;
&lt;td&gt;No — a feature inside Google Search&lt;/td&gt;
&lt;td&gt;Yes — its own product, smaller but growing fast&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The numbers are worth taking one at a time, because none of them are as clean&lt;br&gt;
as a single headline figure suggests.&lt;/p&gt;

&lt;h2&gt;
  
  
  How big is each one, honestly
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT&lt;/strong&gt; is the outlier by scale. OpenAI reported ChatGPT crossing 1&lt;br&gt;
billion weekly active users by July 2026, a figure that's been widely&lt;br&gt;
reported since. That's a number for the whole product, not just its&lt;br&gt;
search/browsing mode specifically — but it establishes the ceiling nothing&lt;br&gt;
else here gets close to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perplexity&lt;/strong&gt; is a different order of magnitude. Per Perplexity's own&lt;br&gt;
reported figures, it passed roughly 230 million monthly active users in&lt;br&gt;
early 2026 and processes over a billion queries a month. That's real and&lt;br&gt;
growing fast — the same reporting put the company's valuation at over $20&lt;br&gt;
billion — but it's a fraction of ChatGPT's reach. Treat Perplexity as a&lt;br&gt;
smaller, more specific audience, not a peer-sized alternative.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Google AI Overviews&lt;/strong&gt; is the hardest of the three to pin to one number, and&lt;br&gt;
we're not going to pretend otherwise. Different measurement firms report&lt;br&gt;
different shares of Google searches that trigger an AI Overview in 2026 —&lt;br&gt;
estimates run from roughly a quarter to nearly half of searches, depending on&lt;br&gt;
methodology and query type. That's a wide range, and anyone quoting one exact&lt;br&gt;
percentage as settled fact is smoothing over a real disagreement between&lt;br&gt;
measurement methods. What isn't in dispute: Google Search's total query&lt;br&gt;
volume dwarfs ChatGPT's or Perplexity's own native search by a wide margin.&lt;br&gt;
That's the actual reason AI Overviews matters as much as it does — not&lt;br&gt;
because it's a large percentage of a small number, but because it's a&lt;br&gt;
meaningful percentage of the largest number in this comparison, applied&lt;br&gt;
inside a feature most people never consciously decide to use. Nobody opts&lt;br&gt;
into AI Overviews the way they open the ChatGPT app; it just shows up above&lt;br&gt;
the results they were already going to scroll past.&lt;/p&gt;

&lt;h2&gt;
  
  
  How each one actually decides what to show you
&lt;/h2&gt;

&lt;p&gt;This is the part that should drive where you spend effort, because the three&lt;br&gt;
don't source answers the same way at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT&lt;/strong&gt; leans heavily on Wikipedia among its top cited sources. That&lt;br&gt;
doesn't mean Wikipedia is the only thing it cites, but it's a consistent&lt;br&gt;
pattern across citation studies, and it means an encyclopedic, neutral,&lt;br&gt;
well-structured presence — the kind of thing Wikipedia itself rewards —&lt;br&gt;
correlates with showing up in ChatGPT's answers more than flashier marketing&lt;br&gt;
copy does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perplexity&lt;/strong&gt; leans the opposite direction: heavily toward Reddit and recent&lt;br&gt;
content. Perplexity is tuned to surface what real people are currently&lt;br&gt;
saying, not what a canonical reference source says. A product with an active,&lt;br&gt;
recent presence in relevant subreddits and forums has a genuinely different&lt;br&gt;
shot at a Perplexity citation than the same product with only a polished&lt;br&gt;
homepage — see &lt;a href="https://letslaunch.today/blog/wikipedia-reddit-ai-citations" rel="noopener noreferrer"&gt;why Wikipedia and Reddit are different, mostly-unreachable&lt;br&gt;
bets&lt;/a&gt; for what that actually takes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Google AI Overviews&lt;/strong&gt; doesn't have an independent retrieval system at all&lt;br&gt;
in the way the other two do — it's built on top of Google's own existing web&lt;br&gt;
index, enriched with authority and freshness signals layered on top of&lt;br&gt;
ordinary ranking. Practically, that means the things that get a page ranked&lt;br&gt;
in ordinary Google Search — indexation, authority, freshness, structure — are&lt;br&gt;
the same things that get a page pulled into an AI Overview. It's not a&lt;br&gt;
separate contest with separate rules; it's the same contest with an extra&lt;br&gt;
layer on top.&lt;/p&gt;

&lt;p&gt;The consequence, worth stating plainly: only about 11-14% of domains cited by&lt;br&gt;
ChatGPT also get cited by Perplexity, per independent 2026 citation-tracking&lt;br&gt;
analyses. Whatever gets you into one of these products' answers is not&lt;br&gt;
reliably the same thing that gets you into the other's. Treat them as&lt;br&gt;
separate bets, because the data says they behave like separate bets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to actually spend effort
&lt;/h2&gt;

&lt;p&gt;Given all of that, here's a decisive ordering for a small team that can't&lt;br&gt;
chase all three equally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Start with being crawlable and readable — this is non-negotiable for all&lt;br&gt;
three.&lt;/strong&gt; If the AI crawlers behind ChatGPT and Perplexity can't get your&lt;br&gt;
actual content out of your page (client-side rendering with nothing in the&lt;br&gt;
initial HTML is the classic failure — see &lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;why ChatGPT can't see your&lt;br&gt;
site&lt;/a&gt;), none of the sourcing-behavior&lt;br&gt;
differences above matter, because you were never a candidate to begin with.&lt;br&gt;
This is the one universal prerequisite. Check yours with &lt;a href="https://letslaunch.today/free/ai-readable-site" rel="noopener noreferrer"&gt;the AI-readable&lt;br&gt;
site tool&lt;/a&gt; before anything else on this list.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you can only do one more thing, do ordinary SEO — for AI Overviews&lt;br&gt;
specifically.&lt;/strong&gt; Because AI Overviews is built on Google's existing index&lt;br&gt;
rather than a separate retrieval system, classic SEO fundamentals — the&lt;br&gt;
things that get you ranked in ordinary Google Search — carry more weight here&lt;br&gt;
than they do for ChatGPT's or Perplexity's citation behavior. This is the one&lt;br&gt;
case on this list where "just do normal SEO" is a legitimate, specific&lt;br&gt;
AI-visibility strategy rather than a vague deflection. Given that AI&lt;br&gt;
Overviews sits on top of the search engine with by far the largest total&lt;br&gt;
query volume of the three, this is usually the highest-leverage single move&lt;br&gt;
available to a small team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat ChatGPT and Perplexity as separate, smaller bets — not extensions of&lt;br&gt;
your SEO work or of each other.&lt;/strong&gt; If you have the bandwidth to chase citation&lt;br&gt;
specifically inside these products, know what you're chasing: a Wikipedia-style,&lt;br&gt;
structured, neutral presence tends to correlate with ChatGPT citation more&lt;br&gt;
than pure marketing copy does; an active, recent presence in relevant&lt;br&gt;
community discussion (Reddit foremost) tends to correlate more with&lt;br&gt;
Perplexity. But given the 11-14% overlap figure above, doing well in one is&lt;br&gt;
not a reasonable basis for assuming you'll do well in the other, and neither&lt;br&gt;
is a reasonable basis for assuming your Google AI Overview presence will&lt;br&gt;
follow. A team with genuinely limited time should pick based on where its&lt;br&gt;
actual audience already is — Reddit-heavy technical audience, lean toward&lt;br&gt;
Perplexity's sourcing habits; broad consumer audience, ChatGPT's reach alone&lt;br&gt;
probably makes it worth the baseline effort regardless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't mistake any of this for a guarantee.&lt;/strong&gt; Nothing here is a mechanism&lt;br&gt;
proven to make any of these three name your product. It's a way to prioritize&lt;br&gt;
limited effort given what's actually observable about how each product finds&lt;br&gt;
its answers — reach, sourcing pattern, and how much they overlap with each&lt;br&gt;
other. That's a real basis for a decision. It is not a promise.&lt;/p&gt;

&lt;p&gt;One thing that helps regardless of which of these three you prioritize: an&lt;br&gt;
independent listing that describes your product consistently with how you&lt;br&gt;
describe it everywhere else. It won't cause a citation on its own — nobody&lt;br&gt;
can currently show that any single action does — but a &lt;a href="https://letslaunch.today/submit" rel="noopener noreferrer"&gt;free LetsLaunch&lt;br&gt;
listing&lt;/a&gt; is a low-cost way to add one more consistent, crawlable&lt;br&gt;
description of your product to the web, which is the one thing that&lt;br&gt;
correlates (not causes) with being named, across all three products discussed&lt;br&gt;
here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Should a founder pick just one of ChatGPT, Google AI Overviews or Perplexity to focus on?
&lt;/h3&gt;

&lt;p&gt;Not entirely, but the priority order should reflect the numbers above. Get&lt;br&gt;
crawlable first, since it's a prerequisite for all three. Then lean toward&lt;br&gt;
ordinary SEO, because it's the one strategy that specifically helps with&lt;br&gt;
Google AI Overviews, which sits on top of by far the largest search volume of&lt;br&gt;
the three. Only after that does it make sense to chase ChatGPT- or&lt;br&gt;
Perplexity-specific citation habits, and those two should be treated as&lt;br&gt;
separate efforts from each other, not one combined push.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Google AI Overviews the same kind of thing as ChatGPT or Perplexity?
&lt;/h3&gt;

&lt;p&gt;Not really, and that distinction matters for where you spend effort. ChatGPT&lt;br&gt;
and Perplexity are both independent products with their own retrieval&lt;br&gt;
systems and their own habit-forming destinations. Google AI Overviews is a&lt;br&gt;
feature inside ordinary Google Search, built on Google's existing web index&lt;br&gt;
rather than a separate retrieval system — which is why classic SEO carries&lt;br&gt;
more weight there than it does for the other two.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does doing well in ChatGPT citations mean I'll also do well in Perplexity?
&lt;/h3&gt;

&lt;p&gt;No, and the data says the opposite is closer to true. Independent 2026&lt;br&gt;
citation-tracking analyses put the overlap between domains cited by ChatGPT&lt;br&gt;
and domains cited by Perplexity at only about 11-14%. ChatGPT leans on&lt;br&gt;
Wikipedia-style sourcing; Perplexity leans on Reddit and recent content.&lt;br&gt;
Treat them as separate bets rather than assuming one strategy transfers to&lt;br&gt;
the other.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>llms.txt in 2026: We Publish One. Here's What It Actually Does.</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Mon, 14 Sep 2026 08:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/llmstxt-in-2026-we-publish-one-heres-what-it-actually-does-4i7a</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/llmstxt-in-2026-we-publish-one-heres-what-it-actually-does-4i7a</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/llms-txt-does-it-work-2026" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We publish an &lt;code&gt;llms.txt&lt;/code&gt; and an &lt;code&gt;llms-full.txt&lt;/code&gt; on this site. If you read the&lt;br&gt;
rest of this post, you might reasonably ask why we bothered.&lt;/p&gt;

&lt;p&gt;Short version: it costs nothing, it does no harm, and it is not the thing&lt;br&gt;
doing any real work. If you are deciding whether to spend an afternoon on it,&lt;br&gt;
the honest answer is spend five minutes instead — and then go read our guide&lt;br&gt;
on &lt;a href="https://letslaunch.today/free/ai-readable-site" rel="noopener noreferrer"&gt;what actually makes a site readable to AI crawlers&lt;/a&gt;,&lt;br&gt;
because that part is not optional the way this one is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What llms.txt was supposed to do
&lt;/h2&gt;

&lt;p&gt;The idea, proposed in 2024, is a plain-text file at your domain root that&lt;br&gt;
hands a language model a curated, markdown-formatted summary of your site —&lt;br&gt;
the parts worth reading, without the navigation chrome and cookie banners a&lt;br&gt;
normal crawl picks up. It reads like &lt;code&gt;robots.txt&lt;/code&gt; for meaning instead of&lt;br&gt;
access: not "you may crawl this," but "here is what matters if you do."&lt;/p&gt;

&lt;p&gt;It's a reasonable idea. The problem is not the idea. The problem is that&lt;br&gt;
essentially nobody who could act on it decided to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Google has actually said
&lt;/h2&gt;

&lt;p&gt;At Google Search Central Live, Google's Gary Illyes confirmed Google does not&lt;br&gt;
support &lt;code&gt;llms.txt&lt;/code&gt; and has no plans to. He compared it directly to the&lt;br&gt;
&lt;code&gt;keywords&lt;/code&gt; meta tag — a field search engines stopped reading over a decade&lt;br&gt;
ago because it's written by the site operator and therefore tells you nothing&lt;br&gt;
they didn't already choose to claim about themselves.&lt;/p&gt;

&lt;p&gt;John Mueller went further in a public reply: &lt;em&gt;"Google doesn't use llms.txt or&lt;br&gt;
llms-author.txt. I don't know of any other crawler / llm confirming they're&lt;br&gt;
using these."&lt;/em&gt; Asked about the related &lt;code&gt;content-signal&lt;/code&gt; robots.txt directive,&lt;br&gt;
he added that no crawler or LLM uses it either — &lt;em&gt;"it was made up by a CDN,&lt;br&gt;
afaik it has no effects whatsoever"&lt;/em&gt; — and that shipping unsupported&lt;br&gt;
directives is maintenance burden with no upside.&lt;/p&gt;

&lt;p&gt;That is about as direct as a platform statement gets. Not "we're evaluating&lt;br&gt;
it." Not "support is coming." Nobody is using it, and nobody we know of is&lt;br&gt;
planning to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually happens to a published llms.txt file
&lt;/h2&gt;

&lt;p&gt;Google's disinterest would matter less if some other assistant had adopted the&lt;br&gt;
file instead. So the more useful question is behavioural: when a domain&lt;br&gt;
publishes one, does anything actually request it?&lt;/p&gt;

&lt;p&gt;Ahrefs answered this at scale in May 2026, analysing server logs and bot&lt;br&gt;
traffic across 137,210 domains. The numbers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;28%&lt;/strong&gt; of domains in the sample publish a valid &lt;code&gt;llms.txt&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;97%&lt;/strong&gt; of those files received zero requests for the entire month.&lt;/li&gt;
&lt;li&gt;Of the 3% that got any traffic at all, &lt;strong&gt;96% of requests came from bots&lt;/strong&gt; —
but a mix of crawlers, agents and scrapers, not specifically the assistants
the file is meant for. Named AI retrieval bots like &lt;code&gt;PerplexityBot&lt;/code&gt; and
&lt;code&gt;OAI-SearchBot&lt;/code&gt; accounted for only about 1% of the requests these files did
receive.&lt;/li&gt;
&lt;li&gt;On domains with &lt;strong&gt;no&lt;/strong&gt; &lt;code&gt;llms.txt&lt;/code&gt; at all, AI bots made zero speculative
requests for one. They don't go looking. A crawler that respects
&lt;code&gt;llms.txt&lt;/code&gt; on the sites that have it would probe for it everywhere, the way
every crawler probes for &lt;code&gt;robots.txt&lt;/code&gt; on every domain it touches. That
probing does not happen, which is a stronger signal than the low fetch rate
on its own: the file isn't being ignored occasionally, it isn't part of any
crawler's routine at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Put together: publishing the file gets you into a small minority (28%) who&lt;br&gt;
did the extra work, and of that minority, 97% got nothing for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we still have one
&lt;/h2&gt;

&lt;p&gt;None of this makes &lt;code&gt;llms.txt&lt;/code&gt; harmful, and the cost of publishing one is close&lt;br&gt;
to zero — it's a static file, not an integration. What we actually rely on to&lt;br&gt;
be readable by ChatGPT, Claude and Perplexity is entirely different, and it's&lt;br&gt;
the same lesson we keep landing on when we &lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;check whether a page is visible to&lt;br&gt;
AI crawlers at all&lt;/a&gt;: the thing that&lt;br&gt;
matters is whether the &lt;em&gt;page itself&lt;/em&gt; — the actual HTML a crawler receives at&lt;br&gt;
the actual URL — contains real content, unblocked, with no JavaScript&lt;br&gt;
required to render it.&lt;/p&gt;

&lt;p&gt;That's a &lt;code&gt;robots.txt&lt;/code&gt; that names AI crawlers explicitly instead of leaving&lt;br&gt;
them to guess, and a page that returns full content on the first request, no&lt;br&gt;
script execution needed. &lt;code&gt;llms.txt&lt;/code&gt; is neither of those things — it's a&lt;br&gt;
side-channel a wildly small handful of crawlers might, unverified, glance at&lt;br&gt;
once. We keep ours published because it costs one file and documents intent.&lt;br&gt;
We do not tell anyone it is doing the job that server-rendered HTML is&lt;br&gt;
actually doing.&lt;/p&gt;

&lt;p&gt;If you are prioritising a Friday afternoon between the two, do the HTML.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does llms.txt help with SEO or AI search rankings?
&lt;/h2&gt;

&lt;p&gt;No measurable evidence that it does either. Google has stated directly it&lt;br&gt;
doesn't use the file for search, and no AI assistant provider has confirmed&lt;br&gt;
using it for anything. It has no known effect on classic search rankings or&lt;br&gt;
on whether an assistant cites you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do ChatGPT, Claude or Perplexity actually read llms.txt files?
&lt;/h2&gt;

&lt;p&gt;Essentially never, based on the available data. A May 2026 Ahrefs study of&lt;br&gt;
137,210 domains found 97% of published &lt;code&gt;llms.txt&lt;/code&gt; files received zero&lt;br&gt;
requests in a month, and named AI retrieval bots made up roughly 1% of the&lt;br&gt;
traffic the remaining 3% did receive. Domains without the file saw zero&lt;br&gt;
speculative requests for one — crawlers aren't checking for it the way they&lt;br&gt;
check for &lt;code&gt;robots.txt&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should I still publish an llms.txt file?
&lt;/h2&gt;

&lt;p&gt;If it takes you five minutes, there's no harm in it — it costs nothing to&lt;br&gt;
serve a static file, and it's a reasonable record of intent. Just don't treat&lt;br&gt;
it as a fix for AI visibility, and don't spend real engineering time on it&lt;br&gt;
ahead of the thing that's actually load-bearing: making sure the HTML your&lt;br&gt;
server sends, with no JavaScript required, actually contains your content for&lt;br&gt;
the crawlers that matter.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>ai</category>
    </item>
    <item>
      <title>GPTBot Says It Visits Your Site. Here's How to Check, Every Week.</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Sun, 13 Sep 2026 17:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/gptbot-says-it-visits-your-site-heres-how-to-check-every-week-39pm</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/gptbot-says-it-visits-your-site-heres-how-to-check-every-week-39pm</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/monitor-ai-crawler-traffic" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We &lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;wrote before&lt;/a&gt; about whether an AI&lt;br&gt;
crawler can read your page at all — rendering, robots.txt mistakes, a 403 from&lt;br&gt;
your own firewall. That is a one-time question. You fetch the page as&lt;br&gt;
&lt;code&gt;OAI-SearchBot&lt;/code&gt;, you look at what comes back, you fix it or you don't, and the&lt;br&gt;
answer stays true until you change something.&lt;/p&gt;

&lt;p&gt;This post is about a different question: whether these crawlers actually show&lt;br&gt;
up, on an ongoing basis, once the one-time check passes. A page that was&lt;br&gt;
readable to &lt;code&gt;GPTBot&lt;/code&gt; last month tells you nothing about whether &lt;code&gt;GPTBot&lt;/code&gt; has&lt;br&gt;
fetched it since. Readability is a fact about your server. Traffic is a fact&lt;br&gt;
about their behavior, and it changes without you touching anything.&lt;/p&gt;
&lt;h2&gt;
  
  
  What you're actually looking for
&lt;/h2&gt;

&lt;p&gt;Every major AI crawler identifies itself with a distinct user-agent string.&lt;br&gt;
The ones worth watching for:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;GPTBot&lt;/code&gt;, &lt;code&gt;OAI-SearchBot&lt;/code&gt;, &lt;code&gt;ChatGPT-User&lt;/code&gt;, &lt;code&gt;ClaudeBot&lt;/code&gt;, &lt;code&gt;Claude-User&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;Claude-SearchBot&lt;/code&gt;, &lt;code&gt;PerplexityBot&lt;/code&gt;, &lt;code&gt;Perplexity-User&lt;/code&gt;, &lt;code&gt;Google-Extended&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;Applebot-Extended&lt;/code&gt;, &lt;code&gt;Bingbot&lt;/code&gt;, &lt;code&gt;Amazonbot&lt;/code&gt;, &lt;code&gt;DuckAssistBot&lt;/code&gt;, &lt;code&gt;Bytespider&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;CCBot&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;LetsLaunch's own &lt;code&gt;robots.txt&lt;/code&gt; (&lt;code&gt;src/app/robots.ts&lt;/code&gt; in this codebase) names&lt;br&gt;
several of these explicitly rather than relying on a wildcard rule, for the&lt;br&gt;
same reason you'd want to grep for them by name: a rule or a log line that&lt;br&gt;
just says "bot" tells you nothing about which one.&lt;/p&gt;

&lt;p&gt;Two of these strings are worth separating in your head before you start&lt;br&gt;
counting anything, because they answer different questions.&lt;/p&gt;
&lt;h2&gt;
  
  
  Training crawlers and retrieval crawlers are not the same signal
&lt;/h2&gt;

&lt;p&gt;We made this distinction in the readability post and it matters just as much&lt;br&gt;
for monitoring: &lt;code&gt;GPTBot&lt;/code&gt; collects training data. &lt;code&gt;OAI-SearchBot&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;PerplexityBot&lt;/code&gt; fetch pages to serve as citations in an actual answer. They&lt;br&gt;
run independently, on their own schedules, for their own purposes.&lt;/p&gt;

&lt;p&gt;If you only track &lt;code&gt;GPTBot&lt;/code&gt;, a busy week of hits tells you your content might&lt;br&gt;
end up in a future training run — and tells you nothing about whether you can&lt;br&gt;
be cited in an answer today. If you want to know about citation potential, you&lt;br&gt;
need to watch &lt;code&gt;OAI-SearchBot&lt;/code&gt;, &lt;code&gt;Claude-SearchBot&lt;/code&gt; and &lt;code&gt;PerplexityBot&lt;/code&gt;&lt;br&gt;
separately, not lump every AI-labeled agent into one counter. A dashboard that&lt;br&gt;
reports "247 AI bot hits this week" without splitting training from retrieval&lt;br&gt;
is reporting a number nobody can act on.&lt;/p&gt;
&lt;h2&gt;
  
  
  Grepping your own logs
&lt;/h2&gt;

&lt;p&gt;If you run your own server or have access to raw access logs, this is a grep&lt;br&gt;
away. Standard combined log format puts the request path, status code, and&lt;br&gt;
user-agent on one line, so a single pipeline gets you counts per crawler:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s2"&gt;"GPTBot|OAI-SearchBot|ClaudeBot|PerplexityBot|Claude-SearchBot"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  access.log &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt;&lt;span class="s1"&gt;'"'&lt;/span&gt; &lt;span class="s1"&gt;'{print $6}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;uniq&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;awk -F'"'&lt;/code&gt; splits the line on quote characters, and in combined log format&lt;br&gt;
the user-agent is the sixth quoted field — the request line and the referer&lt;br&gt;
are the fourth and there's no fifth field before it. &lt;code&gt;sort | uniq -c&lt;/code&gt; turns&lt;br&gt;
that into a count per exact user-agent string, which is enough to tell you&lt;br&gt;
who's showing up and how often.&lt;/p&gt;

&lt;p&gt;To see what a specific crawler is actually requesting, rather than just how&lt;br&gt;
often it shows up:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"PerplexityBot"&lt;/span&gt; access.log | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $7, $9}'&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; | &lt;span class="nb"&gt;uniq&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Field 7 is the request path and field 9 is the status code in the default&lt;br&gt;
combined format (exact field numbers shift if your log adds extra fields, so&lt;br&gt;
check one line by hand before trusting the count). This tells you two useful&lt;br&gt;
things at once: which pages a crawler actually cares about, and whether it's&lt;br&gt;
getting 200s or something else on the way in.&lt;/p&gt;

&lt;p&gt;Run this weekly, or wire it into whatever log-shipping you already have, and&lt;br&gt;
you get a trend instead of a single data point. A one-time curl check answers&lt;br&gt;
"can it read this page." A weekly grep answers "does it come back."&lt;/p&gt;

&lt;h2&gt;
  
  
  The version with no log access required
&lt;/h2&gt;

&lt;p&gt;Most sites don't sit in front of raw logs — they sit behind Cloudflare, which&lt;br&gt;
is also LetsLaunch's own setup. Cloudflare's dashboard has an AI Crawlers&lt;br&gt;
section, under its AI Bots reporting, that shows this without any grep at all:&lt;br&gt;
total requests from AI crawlers, a breakdown of allowed versus blocked&lt;br&gt;
requests, a per-crawler split naming the major players — Google, OpenAI,&lt;br&gt;
Anthropic, Perplexity, ByteDance, Common Crawl, DuckDuckGo among them — and&lt;br&gt;
the paths those crawlers hit most.&lt;/p&gt;

&lt;p&gt;If your site is already on Cloudflare, this is the first place to look before&lt;br&gt;
building anything yourself. It's the same information the grep gives you,&lt;br&gt;
already split by crawler, already trending over time, with nothing to&lt;br&gt;
instrument. The tradeoff is that it only sees what passes through Cloudflare's&lt;br&gt;
edge — if you serve some paths from elsewhere, those don't show up here.&lt;/p&gt;

&lt;p&gt;Cloudflare's own traffic data shows AI bot traffic rising sharply year over&lt;br&gt;
year. We won't put a number on that, because the precise figure varies by&lt;br&gt;
which report and which time window you're reading, and a specific percentage&lt;br&gt;
attributed loosely is worse than no percentage at all. The direction is the&lt;br&gt;
useful part: this is traffic worth watching, not a rounding error, and it is&lt;br&gt;
worth setting up a recurring check rather than a one-time one.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you just want a fast answer today
&lt;/h2&gt;

&lt;p&gt;Setting up log monitoring or checking a Cloudflare dashboard is worth it if&lt;br&gt;
you want a trend. If you just want to know right now whether a specific&lt;br&gt;
crawler can reach a specific page, that's a narrower and faster question, and&lt;br&gt;
it's what our &lt;a href="https://letslaunch.today/free/ai-crawler-check" rel="noopener noreferrer"&gt;AI crawler check&lt;/a&gt; does: it fetches your&lt;br&gt;
page as each crawler, once, and shows you the status code and the text each&lt;br&gt;
one received. It won't tell you whether &lt;code&gt;ClaudeBot&lt;/code&gt; visited last Tuesday. It&lt;br&gt;
will tell you, in the next thirty seconds, whether it could visit at all. Pair&lt;br&gt;
that with our &lt;a href="https://letslaunch.today/free/ai-readable-site" rel="noopener noreferrer"&gt;guide to making a page AI-readable&lt;/a&gt; if&lt;br&gt;
the answer comes back wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a crawler hit does and doesn't prove
&lt;/h2&gt;

&lt;p&gt;Here's the part worth being honest about, because it's the same caution we&lt;br&gt;
applied to citations in the readability post. Seeing &lt;code&gt;PerplexityBot&lt;/code&gt; in your&lt;br&gt;
logs is necessary evidence. It is not sufficient evidence.&lt;/p&gt;

&lt;p&gt;A hit tells you the crawler reached your page and (assuming a 200 and&lt;br&gt;
readable HTML) received your content. It does not tell you that content was&lt;br&gt;
used, indexed, or ever surfaced to a single user. You cannot infer citation&lt;br&gt;
frequency from crawl frequency — a crawler can fetch a page a hundred times&lt;br&gt;
and cite it never, and there's no dashboard, ours included, that closes that&lt;br&gt;
gap.&lt;/p&gt;

&lt;p&gt;What monitoring does give you is the other direction, which is just as&lt;br&gt;
useful: absence of hits is real information. If &lt;code&gt;OAI-SearchBot&lt;/code&gt; has never&lt;br&gt;
touched your domain, you cannot be cited from that index, full stop — no&lt;br&gt;
amount of good content changes that if the crawler was never there to read&lt;br&gt;
it. Monitoring won't tell you you're winning. It will tell you, reliably,&lt;br&gt;
when you've been shut out before you've done anything else.&lt;/p&gt;

&lt;p&gt;Treat crawler-hit counts as a diagnostic, not a growth metric. They rule&lt;br&gt;
things out. They don't rule things in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do crawler visits mean my product will be cited by ChatGPT or Perplexity?
&lt;/h2&gt;

&lt;p&gt;No. A crawler hit only confirms the page was fetched and, if it returned a&lt;br&gt;
readable 200, that the content was received. Whether it's later cited in an&lt;br&gt;
answer is a separate step nobody can currently observe from the outside, and&lt;br&gt;
we know of no reliable way to predict it from crawl frequency alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do I need Cloudflare to monitor AI crawler traffic?
&lt;/h2&gt;

&lt;p&gt;No, but it removes the setup work. Cloudflare's AI Crawlers analytics shows&lt;br&gt;
per-crawler request counts, allowed-versus-blocked splits, and top crawled&lt;br&gt;
paths for any site behind its network, with no log access needed. Without&lt;br&gt;
Cloudflare, the same information is available by grepping your raw access&lt;br&gt;
logs for the relevant user-agent strings.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's the difference between monitoring GPTBot and monitoring OAI-SearchBot?
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;GPTBot&lt;/code&gt; hits reflect training-data collection. &lt;code&gt;OAI-SearchBot&lt;/code&gt; hits reflect&lt;br&gt;
retrieval for citations in ChatGPT's search answers. They're separate&lt;br&gt;
crawlers with separate purposes, so tracking only &lt;code&gt;GPTBot&lt;/code&gt; tells you nothing&lt;br&gt;
about whether you can currently be cited — you have to watch the&lt;br&gt;
search/retrieval agents separately to answer that question.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>Does Schema Markup Help You Get Cited by AI? What Google Actually Says</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Sun, 13 Sep 2026 08:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/does-schema-markup-help-you-get-cited-by-ai-what-google-actually-says-280h</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/does-schema-markup-help-you-get-cited-by-ai-what-google-actually-says-280h</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/schema-markup-ai-overviews" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Somewhere between "add schema markup" and "get cited by ChatGPT" a lot of&lt;br&gt;
advice quietly inserts a step that doesn't exist. The pitch is usually some&lt;br&gt;
version of: implement the right structured data and AI Overviews will start&lt;br&gt;
citing you. It's a tidy claim, it sounds technical enough to be true, and it&lt;br&gt;
isn't what Google's own documentation says.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Google actually requires for AI Overviews
&lt;/h2&gt;

&lt;p&gt;Google's current public documentation on AI Overviews and AI Mode is direct&lt;br&gt;
about eligibility: there are no special extra technical requirements beyond&lt;br&gt;
being indexed and eligible to appear in ordinary Google Search with a snippet.&lt;br&gt;
Not a schema type. Not a markup vocabulary. Not a checklist item labeled "AI&lt;br&gt;
Overviews" anywhere in Google's structured data guidelines.&lt;/p&gt;

&lt;p&gt;That's worth sitting with, because it forecloses a specific, common claim.&lt;br&gt;
There is no "AI schema" to add. If a vendor is selling you a structured-data&lt;br&gt;
package specifically for AI Overviews eligibility, they are selling you&lt;br&gt;
something Google has never said exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  So why does structured data get talked about in the same breath as AI answers at all
&lt;/h2&gt;

&lt;p&gt;Because it's genuinely useful for something adjacent, and the two things get&lt;br&gt;
merged in a lot of marketing copy until they sound like one thing.&lt;/p&gt;

&lt;p&gt;Schema markup's documented job is helping a search engine parse a page's&lt;br&gt;
meaning more precisely — this is a product, here is its price, this&lt;br&gt;
organization is who it claims to be, this block of text answers this&lt;br&gt;
question, this article has this structure. That precision is what makes a&lt;br&gt;
page eligible for ordinary rich results in classic Google Search: star&lt;br&gt;
ratings, price ranges, FAQ dropdowns, breadcrumbs. None of that is new or&lt;br&gt;
AI-specific — it's the same mechanism structured data has always served.&lt;/p&gt;

&lt;p&gt;Google's actual stated requirement isn't that the markup exists. It's that&lt;br&gt;
the markup is accurate and consistent with what's visibly on the page. A&lt;br&gt;
&lt;code&gt;Product&lt;/code&gt; schema block claiming a 4.8-star rating on a page that shows no&lt;br&gt;
reviews at all isn't a growth hack, it's a policy violation, and it's treated&lt;br&gt;
as one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The inference worth making carefully
&lt;/h2&gt;

&lt;p&gt;Here's the part that actually justifies writing a post about this, because&lt;br&gt;
if the answer were simply "schema does nothing for AI," there'd be nothing&lt;br&gt;
else to say.&lt;/p&gt;

&lt;p&gt;We've already established, in our &lt;a href="https://letslaunch.today/blog/chatgpt-vs-google-ai-overviews-vs-perplexity" rel="noopener noreferrer"&gt;comparison of ChatGPT, Google AI Overviews&lt;br&gt;
and Perplexity&lt;/a&gt;, that AI&lt;br&gt;
Overviews doesn't run its own independent retrieval system. It sits on top of&lt;br&gt;
Google's existing web index — the same one that ranks ordinary search&lt;br&gt;
results — with authority and freshness signals layered on top. It isn't a&lt;br&gt;
separate product reading a separate corpus; it's a feature inside Google&lt;br&gt;
Search, drawing on the index Google Search already built.&lt;/p&gt;

&lt;p&gt;Follow that through and a reasonable inference falls out: anything that&lt;br&gt;
helps Google's systems parse and trust a page more precisely — including&lt;br&gt;
accurate, consistent schema markup — plausibly supports the same underlying&lt;br&gt;
signals AI Overviews draws from, even though no AI-specific markup exists to&lt;br&gt;
add. Not because schema has an AI mode, but because AI Overviews is drawing&lt;br&gt;
on the same index that schema was always helping Google understand.&lt;/p&gt;

&lt;p&gt;Notice the hedging in that paragraph, because it's load-bearing. "Plausibly&lt;br&gt;
supports" is not "causes." We have no controlled study showing that adding&lt;br&gt;
schema markup increases how often AI Overviews cites a given page. What we&lt;br&gt;
have is a documented mechanism (schema improves parsing) and a documented&lt;br&gt;
architecture fact (AI Overviews sits on the index that parsing feeds), and a&lt;br&gt;
chain of reasoning between them that we have not seen tested. That's a&lt;br&gt;
different, weaker claim than "schema gets you cited by AI," and marketing&lt;br&gt;
copy that skips the hedge is overstating what's known.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this breaks down: FAQPage schema specifically
&lt;/h2&gt;

&lt;p&gt;This is the case where the temptation to conflate the two things is&lt;br&gt;
strongest, so it's worth being explicit rather than implicit about it.&lt;/p&gt;

&lt;p&gt;FAQPage schema is real and it does something documented: implemented&lt;br&gt;
correctly, it can earn an expandable FAQ dropdown directly in Google's&lt;br&gt;
search results, next to your listing. That's a legitimate, specific,&lt;br&gt;
well-documented rich-result feature. Nobody is disputing it.&lt;/p&gt;

&lt;p&gt;What it is not is a citation tactic for AI assistants. We've said this&lt;br&gt;
plainly elsewhere on this site, in our &lt;a href="https://letslaunch.today/free/ai-readable-site" rel="noopener noreferrer"&gt;free AI-readable-site&lt;br&gt;
tool&lt;/a&gt;: we have seen no evidence that adding FAQ&lt;br&gt;
markup produces a lift in how often an assistant cites you, and we're not&lt;br&gt;
going to imply otherwise here either. The FAQ dropdown in classic search&lt;br&gt;
results and being quoted by an AI assistant are two different outcomes, and&lt;br&gt;
a lot of "schema for AI visibility" content treats getting one as evidence&lt;br&gt;
you're working toward the other. It isn't. Implement FAQPage schema if you&lt;br&gt;
want the search-results dropdown — that's a real, achievable thing. Don't&lt;br&gt;
implement it believing it's an AI-citation lever, because nothing we've seen&lt;br&gt;
supports that being true.&lt;/p&gt;

&lt;h2&gt;
  
  
  A concrete example of taking schema seriously without the AI angle
&lt;/h2&gt;

&lt;p&gt;We run structured data on every product page on this site, and the reason&lt;br&gt;
has nothing to do with AI citation. Product names, taglines and descriptions&lt;br&gt;
on LetsLaunch are supplied by makers — user-generated content we don't&lt;br&gt;
control the contents of. If you serialize that straight into a&lt;br&gt;
&lt;code&gt;&amp;lt;script type="application/ld+json"&amp;gt;&lt;/code&gt; tag with a plain &lt;code&gt;JSON.stringify&lt;/code&gt;, a&lt;br&gt;
maker whose product name happens to contain a literal &lt;code&gt;&amp;lt;/script&lt;/code&gt; sequence&lt;br&gt;
breaks out of the script block, because an HTML parser closes a &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt;&lt;br&gt;
tag at the first literal &lt;code&gt;&amp;lt;/script&lt;/code&gt; it sees regardless of what the&lt;br&gt;
surrounding JSON says. That's not a hypothetical — it's the kind of thing&lt;br&gt;
that shows up the first time somebody pastes a tagline with the wrong&lt;br&gt;
characters in it.&lt;/p&gt;

&lt;p&gt;That's why every piece of JSON-LD on this site goes through a function&lt;br&gt;
called &lt;code&gt;serializeJsonLd&lt;/code&gt;, which escapes &lt;code&gt;&amp;lt;&lt;/code&gt;, &lt;code&gt;&amp;gt;&lt;/code&gt; and &lt;code&gt;&amp;amp;&lt;/code&gt; (and a couple of&lt;br&gt;
line-terminator characters that are legal JSON but not legal&lt;br&gt;
mid-statement in older JavaScript) before anything reaches the page. It's&lt;br&gt;
defensive engineering against user-supplied content breaking a page, not a&lt;br&gt;
citation-boosting tactic. We mention it here because it's the honest&lt;br&gt;
version of "we take structured data seriously" — the reason is correctness&lt;br&gt;
and security, and any downstream SEO benefit is a side effect of doing the&lt;br&gt;
markup right, not the reason we did it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual sequence, in order
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Structured data should be accurate and match what's visibly on the page.
That's Google's stated requirement, and it's also just not lying to a
parser that other systems (possibly including AI ones) may lean on later.&lt;/li&gt;
&lt;li&gt;Implement schema for what it's actually documented to do — rich results
in classic search — not for an AI-specific eligibility gate that doesn't
exist.&lt;/li&gt;
&lt;li&gt;Treat any AI-citation benefit as a plausible side effect of a page being
better understood by the same index AI Overviews sits on top of, not as a
measured, provable outcome you can point to.&lt;/li&gt;
&lt;li&gt;If you want an assistant to describe your product accurately, the more
defensible lever remains describing it the same way everywhere it
appears — your own site, and any independent page that lists you,
including a &lt;a href="https://letslaunch.today/submit" rel="noopener noreferrer"&gt;LetsLaunch listing&lt;/a&gt; if you want one more. We've
written about why an assistant needs to actually receive that text in the
&lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;first place&lt;/a&gt; before any of this
matters at all.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Does adding schema markup get you cited by ChatGPT or AI Overviews?
&lt;/h2&gt;

&lt;p&gt;There's no evidence of a direct, measured effect, and Google's own&lt;br&gt;
documentation names no AI-Overviews-specific markup to add in the first&lt;br&gt;
place. What's defensible is a chain of inference: schema helps Google parse&lt;br&gt;
a page more precisely, and AI Overviews sits on top of the same index that&lt;br&gt;
parsing feeds — so accurate schema plausibly supports the same underlying&lt;br&gt;
signals, without anyone having measured a specific AI-citation lift from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is FAQPage schema worth implementing?
&lt;/h2&gt;

&lt;p&gt;Yes, for what it actually does: it's a real, documented way to earn an&lt;br&gt;
expandable FAQ dropdown in ordinary Google search results. It is not a shown&lt;br&gt;
way to get an AI assistant to cite you more often — we've stated that&lt;br&gt;
directly elsewhere on this site and we're repeating it here rather than&lt;br&gt;
softening it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do I need special schema markup to be eligible for Google's AI Overviews?
&lt;/h2&gt;

&lt;p&gt;No. Google's current documentation states there are no special extra&lt;br&gt;
technical requirements for AI Overviews or AI Mode eligibility beyond being&lt;br&gt;
indexed and eligible to appear in ordinary Google Search with a snippet.&lt;br&gt;
There is no AI-Overviews-specific schema type or vocabulary to add.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>Should You Block GPTBot? The Decision, Not the Panic</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Sat, 12 Sep 2026 17:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/should-you-block-gptbot-the-decision-not-the-panic-oc4</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/should-you-block-gptbot-the-decision-not-the-panic-oc4</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/should-you-block-gptbot" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Somebody on a founder Slack posts a robots.txt snippet that blocks GPTBot,&lt;br&gt;
someone else says "wait, doesn't that get you removed from ChatGPT," and the&lt;br&gt;
thread ends in fifteen replies of confident, contradictory advice. It shouldn't&lt;br&gt;
take fifteen replies. This is a narrow, mechanical question with a narrow,&lt;br&gt;
mechanical answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one fact that resolves most of the confusion
&lt;/h2&gt;

&lt;p&gt;Blocking &lt;code&gt;GPTBot&lt;/code&gt; has no effect on whether ChatGPT can cite you in an answer.&lt;/p&gt;

&lt;p&gt;That's it. That's the fact everyone in that thread needed. ChatGPT's search&lt;br&gt;
and citation feature is powered by a different crawler, &lt;code&gt;OAI-SearchBot&lt;/code&gt;, and&lt;br&gt;
we've covered the split in detail in&lt;br&gt;
&lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;why ChatGPT can't see your site&lt;/a&gt;:&lt;br&gt;
&lt;code&gt;GPTBot&lt;/code&gt; collects training data, &lt;code&gt;OAI-SearchBot&lt;/code&gt; fetches pages for the index&lt;br&gt;
behind ChatGPT's citations, and OpenAI operates them as independent crawlers.&lt;br&gt;
Disallowing one in &lt;code&gt;robots.txt&lt;/code&gt; says nothing to the other. If your actual goal&lt;br&gt;
is "I want ChatGPT to be able to find and quote my page," this entire post is&lt;br&gt;
irrelevant to you — go read that one instead, because the crawler you care&lt;br&gt;
about isn't the one this post is about.&lt;/p&gt;

&lt;p&gt;If your actual goal is "I don't want OpenAI training a model on my content,"&lt;br&gt;
keep reading, because that's a real and separate question, and it deserves an&lt;br&gt;
honest answer about what blocking &lt;code&gt;GPTBot&lt;/code&gt; actually buys you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What blocking GPTBot actually does
&lt;/h2&gt;

&lt;p&gt;Per OpenAI's own published crawler documentation, and consistent with&lt;br&gt;
independent server-log analyses of how major AI crawlers behave, &lt;code&gt;GPTBot&lt;/code&gt;&lt;br&gt;
generally respects a &lt;code&gt;Disallow&lt;/code&gt; rule in &lt;code&gt;robots.txt&lt;/code&gt; that names it specifically.&lt;br&gt;
The same is true of the other big named crawlers — Anthropic's &lt;code&gt;ClaudeBot&lt;/code&gt;,&lt;br&gt;
Google's &lt;code&gt;Google-Extended&lt;/code&gt;. This is worth being precise about, because it is&lt;br&gt;
an industry norm, not a legal requirement. Nothing compels a crawler operator&lt;br&gt;
to honor &lt;code&gt;robots.txt&lt;/code&gt;; a rule aimed at a named user-agent can in principle be&lt;br&gt;
ignored, or the crawler could lie about its own identity. But the major, named&lt;br&gt;
crawlers have an established track record of compliance, and it's checkable —&lt;br&gt;
you don't have to take anyone's word for it. Pull your own server logs and&lt;br&gt;
look for the user-agent string after you make the change; if requests from it&lt;br&gt;
stop, the rule is being honored. If you want the full list of every named&lt;br&gt;
crawler and the rule that matches each one, see our&lt;br&gt;
&lt;a href="https://letslaunch.today/blog/ai-crawler-cheat-sheet" rel="noopener noreferrer"&gt;AI crawler cheat sheet&lt;/a&gt; — this post is the&lt;br&gt;
decision framework for one crawler on that list.&lt;/p&gt;

&lt;p&gt;Given that it works, what you actually get is narrower than most people&lt;br&gt;
assume going in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You opt your pages &lt;strong&gt;out of being used as training data going forward&lt;/strong&gt;,
starting whenever OpenAI's crawler next respects the rule.&lt;/li&gt;
&lt;li&gt;You get &lt;strong&gt;nothing retroactive&lt;/strong&gt;. If your content was already fetched in a
past crawl and used in a training run that already happened, blocking the
crawler today does not undo that. There is no unlearning button on the other
end of a &lt;code&gt;robots.txt&lt;/code&gt; edit.&lt;/li&gt;
&lt;li&gt;You get &lt;strong&gt;no change to ChatGPT search or citations&lt;/strong&gt; — covered above, worth
repeating because it's the thing people actually came here worried about.&lt;/li&gt;
&lt;li&gt;You get &lt;strong&gt;no change to Google Search&lt;/strong&gt; — &lt;code&gt;GPTBot&lt;/code&gt; has nothing to do with
Google, and blocking it doesn't touch your rankings there.&lt;/li&gt;
&lt;li&gt;You get &lt;strong&gt;no change to any other AI product's retrieval&lt;/strong&gt; — Perplexity,
Claude's web search, anything else that fetches your page live to answer a
question is a separate crawler making a separate request, governed by
whatever rule you have (or don't have) aimed at it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the actual decision on the table is small and specific: do you want this&lt;br&gt;
one company not training on this content from this point forward. Everything&lt;br&gt;
else people worry about when they think about this — search visibility, being&lt;br&gt;
"in ChatGPT," getting delisted somewhere — is not actually affected by this&lt;br&gt;
switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case for blocking it
&lt;/h2&gt;

&lt;p&gt;The underlying reason opt-out mechanisms like this exist at all is a real,&lt;br&gt;
ongoing disagreement in the industry: some publishers and rights-holders have&lt;br&gt;
raised copyright and compensation concerns about AI companies training models&lt;br&gt;
on their content without consent or payment. That's been a live, publicly&lt;br&gt;
reported point of contention between publishers and AI companies, and it's&lt;br&gt;
the actual origin of both the training/retrieval split and the&lt;br&gt;
&lt;code&gt;robots.txt&lt;/code&gt;-based opt-outs that OpenAI, Anthropic and Google all support.&lt;br&gt;
Whether training on public web content should require consent or payment is&lt;br&gt;
not a question this post is going to answer for you — it's genuinely&lt;br&gt;
unsettled, reasonable people land on different sides, and your view on it is&lt;br&gt;
a fine reason to block the crawler on its own, independent of anything else&lt;br&gt;
in this post.&lt;/p&gt;

&lt;p&gt;If that's your reason, block it. You don't need a further business case. It's&lt;br&gt;
a values call about your own content, and it's yours to make.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case for not blocking it
&lt;/h2&gt;

&lt;p&gt;The case against is more speculative, and worth stating as speculative rather&lt;br&gt;
than dressing it up as a strategy.&lt;/p&gt;

&lt;p&gt;Some founders and publishers want their product to be part of what a model&lt;br&gt;
"knows" about when it's later asked general questions — the idea being that&lt;br&gt;
training data shapes a model's baseline familiarity with a topic or brand,&lt;br&gt;
which is a different thing from live retrieval and citation. That's a&lt;br&gt;
plausible-sounding motivation. It is not a proven one. We've written before,&lt;br&gt;
in &lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;what makes AI engines cite a source&lt;/a&gt;,&lt;br&gt;
about how thin the actual evidence is for anything that reliably causes a&lt;br&gt;
model to recommend or cite a specific product — and being present in a&lt;br&gt;
training run is a different, further-upstream bet than the citation question&lt;br&gt;
that post covers, with even less established evidence behind it. Nobody has&lt;br&gt;
shown that a product being in GPTBot's training data causes ChatGPT to&lt;br&gt;
mention that product later. It might contribute something to how the model&lt;br&gt;
talks about a category in general terms. It might do nothing measurable at&lt;br&gt;
all. There's no study to point at either way.&lt;/p&gt;

&lt;p&gt;So "leave it open in case it helps the model know about us" is a bet, not a&lt;br&gt;
fact, and it should be weighed as one — against a concern (uncompensated&lt;br&gt;
training use of your content) that, for some site owners, is not speculative&lt;br&gt;
at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The middle path both companies explicitly support
&lt;/h2&gt;

&lt;p&gt;Most people asking this question don't actually want an all-or-nothing&lt;br&gt;
choice, and they don't have to make one. OpenAI and Anthropic both explicitly&lt;br&gt;
support splitting training from retrieval at the crawler level, because&lt;br&gt;
they operate them as separate user-agents in the first place:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI&lt;/strong&gt;: disallow &lt;code&gt;GPTBot&lt;/code&gt; (opt out of training), explicitly allow
&lt;code&gt;OAI-SearchBot&lt;/code&gt; (stay eligible for citation in ChatGPT's answers) and
&lt;code&gt;ChatGPT-User&lt;/code&gt; (stay reachable when a person pastes your link into a chat).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anthropic&lt;/strong&gt;: the same shape — disallow &lt;code&gt;ClaudeBot&lt;/code&gt;, allow
&lt;code&gt;Claude-SearchBot&lt;/code&gt; and &lt;code&gt;Claude-User&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google&lt;/strong&gt;: disallow &lt;code&gt;Google-Extended&lt;/code&gt; if you want to opt out of Gemini/AI
feature training, but never touch ordinary &lt;code&gt;Googlebot&lt;/code&gt; — that one indexes
you for Google Search itself, and blocking it removes you from Google
entirely, which is a completely different and much larger consequence than
anything else in this post.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the setting most sites end up at once they understand the split:&lt;br&gt;
opted out of training, still eligible to be cited, still reachable by a&lt;br&gt;
person with a link. It isn't a compromise so much as just correctly stating&lt;br&gt;
what you want, at the level of granularity the crawlers actually offer.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to decide, in practice
&lt;/h2&gt;

&lt;p&gt;Skip the philosophy question for a second and ask two narrower ones:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Do I have an actual objection to my content training a model, independent
of any effect on my visibility?&lt;/strong&gt; If yes, disallow &lt;code&gt;GPTBot&lt;/code&gt; — that's the
whole ask, and blocking it delivers exactly that with no side effects on
search or citation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Am I blocking it because I think it'll somehow hurt me if I don't, or
help me if I do, in terms of getting found?&lt;/strong&gt; If that's the actual reason,
go re-read the mechanics above — there isn't one. Training and retrieval
don't talk to each other. Decide on the merits of the first question
instead, because the second one isn't actually load-bearing here.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you're not sure which of those you are, that uncertainty is itself useful&lt;br&gt;
information: it usually means the real answer is "block &lt;code&gt;GPTBot&lt;/code&gt;, allow the&lt;br&gt;
search and user-agent crawlers," and move on. It costs you nothing you were&lt;br&gt;
actually using, and it settles the one part of this that was a genuine,&lt;br&gt;
values-based decision rather than a guess about SEO mechanics.&lt;/p&gt;

&lt;p&gt;Getting found in the first place is a separate problem from any of this — one&lt;br&gt;
more listing that names your product consistently is a small, low-cost input&lt;br&gt;
to that, and &lt;a href="https://letslaunch.today/submit" rel="noopener noreferrer"&gt;submitting to LetsLaunch&lt;/a&gt; takes a few minutes if you&lt;br&gt;
want one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does blocking GPTBot remove my site from ChatGPT?
&lt;/h2&gt;

&lt;p&gt;No. &lt;code&gt;GPTBot&lt;/code&gt; collects training data; ChatGPT's search and citation feature is&lt;br&gt;
powered by a separate crawler, &lt;code&gt;OAI-SearchBot&lt;/code&gt;. Blocking &lt;code&gt;GPTBot&lt;/code&gt; opts you out&lt;br&gt;
of training runs and has no effect on whether you can be cited in ChatGPT's&lt;br&gt;
answers, because that depends entirely on &lt;code&gt;OAI-SearchBot&lt;/code&gt; staying allowed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does blocking GPTBot undo past training?
&lt;/h2&gt;

&lt;p&gt;No. A &lt;code&gt;robots.txt&lt;/code&gt; rule only affects future crawls. If your content was&lt;br&gt;
already fetched and used in a training run before you added the disallow&lt;br&gt;
rule, blocking the crawler today does not remove it from a model that has&lt;br&gt;
already been trained.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is GPTBot legally required to respect robots.txt?
&lt;/h2&gt;

&lt;p&gt;No. Respecting &lt;code&gt;robots.txt&lt;/code&gt; is an industry norm, not a legal obligation —&lt;br&gt;
nothing compels a crawler to honor a disallow rule aimed at it, and a rule&lt;br&gt;
naming a specific user-agent could in principle be ignored or spoofed. In&lt;br&gt;
practice, the major named crawlers, including &lt;code&gt;GPTBot&lt;/code&gt;, &lt;code&gt;ClaudeBot&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;Google-Extended&lt;/code&gt;, have a documented track record of compliance, and you can&lt;br&gt;
verify it yourself by checking your own server logs for the user-agent after&lt;br&gt;
adding the rule.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>What Actually Gets You Cited by ChatGPT and Perplexity? Reading the Research</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Sat, 12 Sep 2026 08:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/what-actually-gets-you-cited-by-chatgpt-and-perplexity-reading-the-research-2oh8</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/what-actually-gets-you-cited-by-chatgpt-and-perplexity-reading-the-research-2oh8</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Search "how to get cited by ChatGPT" and you'll find no shortage of people&lt;br&gt;
selling a method. Almost none of them link to a study. The ones that do&lt;br&gt;
usually link to the same one, and usually overstate what it found.&lt;/p&gt;

&lt;p&gt;We went and read it, plus what's been published since. This is the honest&lt;br&gt;
version: one real result worth knowing, a real and important gap in what it&lt;br&gt;
proves, and a short list of things that correlate with being named by an&lt;br&gt;
assistant — stated as correlations, because that's all any of this currently&lt;br&gt;
is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one study doing real work here
&lt;/h2&gt;

&lt;p&gt;In November 2023, researchers from Princeton, the Allen Institute for AI,&lt;br&gt;
Georgia Tech and IIT Delhi published &lt;em&gt;GEO: Generative Engine Optimization&lt;/em&gt; —&lt;br&gt;
the paper that put a name to this whole field. It later appeared at KDD 2024.&lt;/p&gt;

&lt;p&gt;They built GEO-BENCH, a set of roughly 10,000 queries across nine domains, and&lt;br&gt;
tested nine content-editing strategies inside a generative-engine prototype to&lt;br&gt;
see which ones changed how often a source got cited in the answer. Two&lt;br&gt;
strategies stood out clearly above the rest: &lt;strong&gt;adding quotations from&lt;br&gt;
authoritative sources&lt;/strong&gt;, and &lt;strong&gt;adding relevant statistics&lt;/strong&gt;. Both produced&lt;br&gt;
visibility gains in the range of 22-41% depending on domain. Citing sources&lt;br&gt;
within the content helped too, by a smaller margin. Keyword stuffing — the&lt;br&gt;
one tactic everyone already knew how to do from a decade of classic SEO —&lt;br&gt;
measurably &lt;em&gt;hurt&lt;/em&gt;, performing worse than making no changes at all.&lt;/p&gt;

&lt;p&gt;That's a genuinely useful finding, and it's the most rigorous data point in&lt;br&gt;
this entire field as of writing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap almost nobody mentions
&lt;/h2&gt;

&lt;p&gt;Read that paragraph again and notice what it does not say. It does not say&lt;br&gt;
"we changed a real webpage and ChatGPT started citing it." It says a&lt;br&gt;
benchmark, built by the researchers, running against a generative-engine&lt;br&gt;
&lt;em&gt;prototype&lt;/em&gt;, showed a visibility lift when content included quotes and&lt;br&gt;
statistics.&lt;/p&gt;

&lt;p&gt;That's a controlled measurement of a mechanism, not a field result. It tells&lt;br&gt;
you that language models, when generating an answer, respond to the &lt;em&gt;presence&lt;br&gt;
of quotable, specific claims&lt;/em&gt; in candidate source text — which is a&lt;br&gt;
believable and useful thing to know about how these systems work. It does not&lt;br&gt;
tell you that adding a statistic to your homepage will get you cited by the&lt;br&gt;
ChatGPT and Perplexity people actually use, which run different retrieval&lt;br&gt;
pipelines, different ranking layers, and different citation logic than a 2023&lt;br&gt;
research prototype, and which nobody outside those companies can fully&lt;br&gt;
observe.&lt;/p&gt;

&lt;p&gt;Most of the "GEO checklist" content published since 2024 quietly drops that&lt;br&gt;
distinction and presents the benchmark's percentages as if they were measured&lt;br&gt;
on production ChatGPT. We're not going to do that. The mechanism is real and&lt;br&gt;
worth designing content around. The specific percentage is not a promise&lt;br&gt;
about your specific page.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's observable about real platforms, separate from the study
&lt;/h2&gt;

&lt;p&gt;A few things about how ChatGPT, Perplexity and Google AI Overviews actually&lt;br&gt;
pick sources are visible from the outside, without needing lab access:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;They don't all pull from the same pool.&lt;/strong&gt; Independent analyses of what
gets cited show ChatGPT leaning heavily on general reference sources like
Wikipedia, while Perplexity leans toward forum and discussion content like
Reddit, with a strong preference for recent material. These are different
retrieval habits, not one "AI search" behaviour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ranking well in Google is not the same contest.&lt;/strong&gt; Multiple 2026
cross-platform studies put the overlap between an assistant's cited sources
and Google's own top-10 results well under half, with estimates ranging
roughly from one in eight to just over one in three depending on
methodology and query type. The studies disagree on the exact number
because they measure different platforms and query sets — but they agree on
the direction: doing well in classic search does not reliably predict
getting cited by an assistant. It's a related contest, not the same one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Being readable is a precondition, not a strategy.&lt;/strong&gt; None of the above
matters if the crawler behind the assistant can't retrieve your content in
the first place — see &lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;our breakdown of what actually makes a page invisible
to AI crawlers&lt;/a&gt; if you haven't
confirmed yours can be read.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What we're not going to tell you
&lt;/h2&gt;

&lt;p&gt;Some things routinely get claimed in this space with no study behind them at&lt;br&gt;
all, and we'd rather say so than repeat them:&lt;/p&gt;

&lt;p&gt;That FAQ schema increases citation rate. We've seen no evidence for it —&lt;br&gt;
structured data has real, separately-documented uses in classic search, but&lt;br&gt;
nothing published shows it moving AI citation.&lt;/p&gt;

&lt;p&gt;That a directory listing, ours or anyone else's, causes an assistant to cite&lt;br&gt;
you. It's an obvious thing for a directory to want to be true, which is&lt;br&gt;
exactly why it deserves suspicion instead of a marketing page. G2 and&lt;br&gt;
Capterra carry more domain authority than almost any product page that could&lt;br&gt;
ever link to them, and receive close to no AI citations — if raw authority&lt;br&gt;
were the mechanism, they'd dominate every category. They don't.&lt;/p&gt;

&lt;p&gt;That an "AI visibility score" from some tool measures anything stable. Asking&lt;br&gt;
a model whether it knows a product and recording the answer isn't&lt;br&gt;
measurement — between 9% and 28% of such answers flip on a repeated identical&lt;br&gt;
prompt, even at temperature zero. A single check dressed up as a score out of&lt;br&gt;
100 is closer to a coin flip than a metric.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually worth doing with limited certainty
&lt;/h2&gt;

&lt;p&gt;Given all of the above, here's what survives: make sure your content is&lt;br&gt;
retrievable at all — no JavaScript-gated text, correct robots rules, real&lt;br&gt;
HTML, &lt;a href="https://letslaunch.today/free/ai-crawler-check" rel="noopener noreferrer"&gt;checkable in about a minute&lt;/a&gt;. Where you do&lt;br&gt;
write about your own product, prefer specific, quotable claims and real&lt;br&gt;
numbers over adjectives — the one mechanism with an actual controlled&lt;br&gt;
measurement behind it. Describe the product the same way across every page&lt;br&gt;
that mentions it, your own site included, since consistent description&lt;br&gt;
correlates with being named, even though nobody has shown which particular&lt;br&gt;
page causes it. And be skeptical of anyone selling you a guaranteed&lt;br&gt;
mechanism, including us, when we don't have receipts for the claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  What makes ChatGPT or Perplexity cite a source?
&lt;/h2&gt;

&lt;p&gt;The most rigorous evidence is a 2023 Princeton study (GEO, published at KDD&lt;br&gt;
2024) that measured a 22-41% visibility lift inside a research benchmark from&lt;br&gt;
adding quotations and statistics to content, with keyword stuffing performing&lt;br&gt;
worse than no change at all. That's a real, controlled finding about how&lt;br&gt;
generative engines weigh content — but it was measured on a lab benchmark, not&lt;br&gt;
on production ChatGPT or Perplexity, so treat the mechanism as real and the&lt;br&gt;
exact percentage as not transferable to your page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does ranking well in Google get you cited by AI assistants?
&lt;/h2&gt;

&lt;p&gt;Not reliably. Multiple 2026 studies comparing AI-cited sources against&lt;br&gt;
Google's top-10 results found overlap well under half, though the exact&lt;br&gt;
figure varies by study and platform. Classic SEO and AI citation appear to be&lt;br&gt;
related but distinct contests — being crawlable and readable matters to both,&lt;br&gt;
but ranking highly in one doesn't predict success in the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is generative engine optimization (GEO) a proven discipline?
&lt;/h2&gt;

&lt;p&gt;Partially. One mechanism — content with specific quotes and statistics&lt;br&gt;
getting cited more than vague or keyword-stuffed content — has real,&lt;br&gt;
peer-reviewed measurement behind it, from a controlled benchmark. Most other&lt;br&gt;
claims circulating under the "GEO" label, including FAQ schema and directory&lt;br&gt;
submissions as citation drivers, currently have no published evidence behind&lt;br&gt;
them at all.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>Wikipedia and Reddit Aren't One Strategy — They're ChatGPT's and Perplexity's Different Bets</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Fri, 11 Sep 2026 17:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/wikipedia-and-reddit-arent-one-strategy-theyre-chatgpts-and-perplexitys-different-bets-h2i</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/wikipedia-and-reddit-arent-one-strategy-theyre-chatgpts-and-perplexitys-different-bets-h2i</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/wikipedia-reddit-ai-citations" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ask ChatGPT about almost anything and there's a decent chance the answer leans&lt;br&gt;
on a page you didn't write, on a site you don't control, that has never heard&lt;br&gt;
of your product. Wikipedia and Reddit show up so often in these systems'&lt;br&gt;
citations that the pattern isn't really a curiosity anymore — it's closer to&lt;br&gt;
structural.&lt;/p&gt;

&lt;p&gt;We've written before about &lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;what the actual research says about getting&lt;br&gt;
cited by AI engines&lt;/a&gt;, and the honest&lt;br&gt;
answer there was mostly "less than the industry claims, and mostly&lt;br&gt;
correlational." This post is narrower. It's not about how to write content&lt;br&gt;
that gets cited. It's about &lt;em&gt;where these systems are pulling from in the first&lt;br&gt;
place&lt;/em&gt; — and what that structural fact does and doesn't leave open to an&lt;br&gt;
independent founder with no Wikipedia page and no Reddit history.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, and who's reporting them
&lt;/h2&gt;

&lt;p&gt;5W, an AI communications firm, published a synthesis called "The State of AI&lt;br&gt;
Citations 2026," covered by PR Newswire as analyzing more than 680 million&lt;br&gt;
tracked AI citations across ChatGPT, Claude, Perplexity, Gemini and Google AI&lt;br&gt;
Overviews. Two findings from it are worth sitting with.&lt;/p&gt;

&lt;p&gt;First: Wikipedia and Reddit together account for more than 25% of all&lt;br&gt;
ChatGPT citations in the US. Not a niche source. More than a quarter of&lt;br&gt;
everything the most widely used assistant cites, concentrated in two&lt;br&gt;
platforms neither of which is a news outlet, a review site, or anything a&lt;br&gt;
product team would normally target.&lt;/p&gt;

&lt;p&gt;Second, and more pointed: none of the Wall Street Journal, the New York&lt;br&gt;
Times, or Bloomberg appear in ChatGPT's top 20 cited domains. Three of the&lt;br&gt;
most authoritative news organizations in the English language, by any&lt;br&gt;
conventional measure of editorial rigor or domain authority, don't crack the&lt;br&gt;
top 20. If you've been operating on the assumption that getting into a major&lt;br&gt;
outlet is the top of the AI-citation food chain, this is worth updating on.&lt;/p&gt;

&lt;p&gt;Other 2026 citation-tracking analyses — methodologies vary enough between&lt;br&gt;
them that we're not going to force a single number onto a single source — have&lt;br&gt;
separately put Wikipedia's share of ChatGPT's own top-10 cited sources at&lt;br&gt;
roughly 48%. Call it directionally consistent with 5W's finding rather than&lt;br&gt;
the same measurement: however you slice it, Wikipedia is not one source among&lt;br&gt;
many for ChatGPT. It's the biggest one, by a wide margin, in more than one&lt;br&gt;
independent look at the data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wikipedia and Reddit are not the same bet
&lt;/h2&gt;

&lt;p&gt;Here's the part that a lazy read of the above skips past, and it matters more&lt;br&gt;
than the headline numbers.&lt;/p&gt;

&lt;p&gt;Wikipedia's dominance is not uniform across assistants. One comparative&lt;br&gt;
analysis found ChatGPT cites Wikipedia for roughly 12% of its own citations,&lt;br&gt;
Claude cites it for roughly 0.1%, and Perplexity doesn't cite Wikipedia in any&lt;br&gt;
meaningful volume at all. Perplexity's citation behavior looks almost&lt;br&gt;
inverted: the same analysis put Reddit at roughly 47% of Perplexity's own&lt;br&gt;
top-10 source share — Perplexity's Wikipedia-equivalent, structurally, is&lt;br&gt;
Reddit.&lt;/p&gt;

&lt;p&gt;So "get on Wikipedia" and "get into Reddit" are not the same advice wearing&lt;br&gt;
two names. They're bets on two different platforms that happen to both&lt;br&gt;
concentrate around user-generated, non-brand-controlled content, and a tactic&lt;br&gt;
aimed at one does close to nothing for the other.&lt;/p&gt;

&lt;p&gt;This lines up with something we said plainly in &lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;the GEO research&lt;br&gt;
post&lt;/a&gt;: these are different&lt;br&gt;
retrieval systems, not one "AI search." Independent analyses found only about&lt;br&gt;
11% of domains get cited by both ChatGPT and Perplexity. If you optimize for&lt;br&gt;
one assistant's citation habits, you have done close to nothing for the&lt;br&gt;
other's, and Wikipedia versus Reddit is the clearest illustration of that&lt;br&gt;
split we've seen yet.&lt;/p&gt;

&lt;p&gt;It's also worth separating this from a different problem entirely: none of&lt;br&gt;
this matters if the crawler behind the assistant can't retrieve your page's&lt;br&gt;
content to begin with. That's a rendering and access problem, not a source-mix&lt;br&gt;
problem, and we've covered it separately in &lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;why ChatGPT can't see your&lt;br&gt;
site&lt;/a&gt;. Fix that first regardless of&lt;br&gt;
anything below — it's the one part of this that's fully mechanical and fully&lt;br&gt;
in your control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why you can't just "get on Wikipedia"
&lt;/h2&gt;

&lt;p&gt;You cannot get your own product a Wikipedia page on demand, and this isn't a&lt;br&gt;
process gap that a clever workaround fixes — it's the intended design.&lt;br&gt;
Wikipedia's notability and neutral-point-of-view rules exist specifically to&lt;br&gt;
resist self-promotion. A page created, or heavily edited, by someone connected&lt;br&gt;
to the subject is routinely flagged and deleted, and the volunteer editors who&lt;br&gt;
patrol new pages are reasonably good at spotting the pattern: a founder,&lt;br&gt;
employee, or paid editor writing about their own product in the neutral third&lt;br&gt;
person is one of the most common deletion categories on the site.&lt;/p&gt;

&lt;p&gt;The honest path to a legitimate Wikipedia mention runs the other direction&lt;br&gt;
entirely. You get covered by independent, reliable secondary sources — real&lt;br&gt;
press, not a paid placement, not a sponsored post — and &lt;em&gt;then&lt;/em&gt;, sometimes, a&lt;br&gt;
Wikipedia editor with no connection to you decides that coverage is enough to&lt;br&gt;
justify an entry, and writes it. The Wikipedia mention is a lagging effect of&lt;br&gt;
real independent coverage existing, not something you pursue directly. If&lt;br&gt;
someone offers to get your product onto Wikipedia for a fee, what they're&lt;br&gt;
actually offering is a page that gets deleted once a volunteer editor notices&lt;br&gt;
the conflict of interest — which is most of them, most of the time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reddit is reachable, but only slowly and only if you mean it
&lt;/h2&gt;

&lt;p&gt;Reddit is the more directly reachable one of the two, but the qualifier&lt;br&gt;
matters: reachable as participation, not as marketing.&lt;/p&gt;

&lt;p&gt;What actually produces the kind of discussion thread that later becomes&lt;br&gt;
citation fodder is answering real questions in relevant subreddits, over&lt;br&gt;
time, from an account with a real history of doing that — not one created&lt;br&gt;
the week your product launched. A first post that reads as a product pitch&lt;br&gt;
gets read as spam, both by the subreddit's human moderators and, per Reddit's&lt;br&gt;
own well-documented moderation norms, often by automated filters before a&lt;br&gt;
human even sees it. It gets removed before anything resembling a citable&lt;br&gt;
thread ever exists, let alone before a crawler could find it.&lt;/p&gt;

&lt;p&gt;There's no shortcut here that survives contact with how Reddit actually&lt;br&gt;
moderates itself. Genuine participation compounds slowly, over months, and&lt;br&gt;
looks like actually being present in a community rather than treating it as a&lt;br&gt;
distribution channel with a login page.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't change
&lt;/h2&gt;

&lt;p&gt;None of the above is a mechanism you can point at and expect a result from.&lt;br&gt;
It's the same caveat we made in the companion post on citation research: this&lt;br&gt;
is the honest, sourced shape of the problem as it currently stands, not a&lt;br&gt;
guarantee about what happens to your specific product if you do these things.&lt;br&gt;
Anyone offering to get you cited on Wikipedia or Reddit as a paid, one-off&lt;br&gt;
deliverable is not offering a real mechanism — they're offering something that&lt;br&gt;
either gets deleted (Wikipedia) or removed by moderators (Reddit) once anyone&lt;br&gt;
looks at it closely.&lt;/p&gt;

&lt;p&gt;What's actually within reach, and costs you nothing beyond the time to do it,&lt;br&gt;
is the same low-effort thing we've pointed to before: describing your product&lt;br&gt;
consistently across whatever independent pages do exist about it, including a&lt;br&gt;
&lt;a href="https://letslaunch.today/submit" rel="noopener noreferrer"&gt;LetsLaunch listing&lt;/a&gt; if you want one more page that names your&lt;br&gt;
product the same way your own site does. We're not claiming that causes a&lt;br&gt;
citation — nobody has shown that, and we said so plainly in the last post too.&lt;br&gt;
It's unproven, but it's also free, which is a different bar than the one&lt;br&gt;
Wikipedia and Reddit ask you to clear.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does Wikipedia show up so often in ChatGPT's answers?
&lt;/h2&gt;

&lt;p&gt;Because ChatGPT's own citation mix leans heavily on it. A 5W synthesis of over&lt;br&gt;
680 million tracked citations found Wikipedia and Reddit together make up more&lt;br&gt;
than 25% of ChatGPT's US citations, and separate analyses have put Wikipedia&lt;br&gt;
alone at around 48% of ChatGPT's top-10 cited sources — making it the single&lt;br&gt;
largest source by a wide margin. That's a property of how ChatGPT's retrieval&lt;br&gt;
currently weighs general reference content, not a signal that Wikipedia&lt;br&gt;
coverage is easy to obtain or that it transfers to other assistants.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does the same Wikipedia-and-Reddit pattern apply to Perplexity?
&lt;/h2&gt;

&lt;p&gt;Not evenly. One comparative analysis found ChatGPT cites Wikipedia for&lt;br&gt;
roughly 12% of its citations while Claude cites it for roughly 0.1%, and&lt;br&gt;
Perplexity doesn't cite Wikipedia in meaningful volume at all — Perplexity&lt;br&gt;
instead leans on Reddit, reported at roughly 47% of its own top-10 source&lt;br&gt;
share. Wikipedia and Reddit are two separate, platform-specific patterns, not&lt;br&gt;
one rule that applies across every assistant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can I get my product added to Wikipedia to improve AI citations?
&lt;/h2&gt;

&lt;p&gt;Not directly, and not on request. Wikipedia's notability and neutral-point-of-&lt;br&gt;
view rules are built to resist self-promotion, and pages created or heavily&lt;br&gt;
edited by someone connected to the subject are routinely deleted once a&lt;br&gt;
volunteer editor notices. The realistic path is being covered by independent&lt;br&gt;
press first — a Wikipedia mention, where it happens at all, follows real&lt;br&gt;
outside coverage rather than preceding it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
    <item>
      <title>How to Write a Paragraph an AI Engine Can Actually Quote</title>
      <dc:creator>Costin Gheorghe</dc:creator>
      <pubDate>Fri, 11 Sep 2026 08:20:00 +0000</pubDate>
      <link>https://dev.to/costin_gheorghe_40d2a06d8/how-to-write-a-paragraph-an-ai-engine-can-actually-quote-2o81</link>
      <guid>https://dev.to/costin_gheorghe_40d2a06d8/how-to-write-a-paragraph-an-ai-engine-can-actually-quote-2o81</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://letslaunch.today/blog/write-content-ai-can-cite" rel="noopener noreferrer"&gt;LetsLaunch&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We &lt;a href="https://letslaunch.today/blog/what-makes-ai-engines-cite-a-source" rel="noopener noreferrer"&gt;wrote up the actual research&lt;/a&gt;&lt;br&gt;
on what correlates with an AI engine citing a source, and the honest version&lt;br&gt;
is shorter than most "GEO checklist" posts want it to be: one controlled&lt;br&gt;
benchmark found that quotations and statistics got cited more than vague&lt;br&gt;
content, and keyword stuffing measurably hurt. That's a real mechanism,&lt;br&gt;
observed in a lab, and it's the most solid data point that exists.&lt;/p&gt;

&lt;p&gt;This post is not about that research. It's about what to actually do with it&lt;br&gt;
— the writing-craft side, for someone sitting down to write a paragraph today.&lt;br&gt;
Not a proven-percentage technique, because we don't have one to sell you. Two&lt;br&gt;
mechanical, checkable habits that follow honestly from what's known, plus one&lt;br&gt;
you should adopt for a much older reason: it makes your writing better for&lt;br&gt;
the person reading it, independent of whatever ChatGPT does with it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why structure matters even without a proven number
&lt;/h2&gt;

&lt;p&gt;An extraction system — whether it's Google building an AI Overview,&lt;br&gt;
Perplexity assembling an answer, or ChatGPT summarizing a page it retrieved —&lt;br&gt;
is doing something specific: pulling a passage out of your page and presenting&lt;br&gt;
it with reduced or no surrounding context. It might show one paragraph. It&lt;br&gt;
might show one sentence. Whatever it shows has to stand on its own, because&lt;br&gt;
the rest of your page usually doesn't travel with it.&lt;/p&gt;

&lt;p&gt;That's true regardless of what the exact citation-rate lift is, which nobody&lt;br&gt;
outside a handful of labs has measured on production systems. What you &lt;em&gt;can&lt;/em&gt;&lt;br&gt;
reason about, without needing a study, is whether a given passage survives&lt;br&gt;
being extracted. Either it still makes a complete, correct claim once it's&lt;br&gt;
alone on the screen, or it doesn't. That's a property of the text, checkable&lt;br&gt;
by reading it, not a promise about traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test: copy it into a blank document
&lt;/h2&gt;

&lt;p&gt;Here's the mechanical version, and you can run it on anything you've already&lt;br&gt;
published.&lt;/p&gt;

&lt;p&gt;Take one paragraph. Copy it alone into a blank document — nothing before it,&lt;br&gt;
nothing after it. Read it as if you'd never seen the rest of the page. Ask:&lt;br&gt;
does this still make a complete, correct, unambiguous claim?&lt;/p&gt;

&lt;p&gt;If yes, the paragraph is structurally extractable. If it depends on "as&lt;br&gt;
mentioned above," a pronoun with no visible referent, or a claim the previous&lt;br&gt;
paragraph set up but this one doesn't restate, it isn't — and an extraction&lt;br&gt;
system either garbles it, skips it, or quotes something that reads as&lt;br&gt;
confused or wrong when separated from its context. None of that requires&lt;br&gt;
knowing anything about how any particular model works. It's the same test a&lt;br&gt;
human editor would run on a pull-quote.&lt;/p&gt;

&lt;h2&gt;
  
  
  A before-and-after
&lt;/h2&gt;

&lt;p&gt;Here's a paragraph that fails the test, from a page describing a fictional&lt;br&gt;
product:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;It handles this automatically, which saves teams a lot of time compared to&lt;br&gt;
the old way of doing it. That's part of why customers tend to stick around&lt;br&gt;
longer once they switch.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Read alone, this says almost nothing. What does "it" do? What's "the old&lt;br&gt;
way"? What's "a lot of time"? "Tend to stick around longer" than what, by how&lt;br&gt;
much? Every noun refers outward to a sentence that isn't there anymore. A&lt;br&gt;
person skimming a summary would come away with no actual information, and an&lt;br&gt;
extraction system quoting it verbatim would be quoting a sentence that reads&lt;br&gt;
as vague at best and misleading at worst — it implies a specific comparison&lt;br&gt;
without stating one.&lt;/p&gt;

&lt;p&gt;Here's the same claim, rewritten to stand alone:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Acme's scheduling tool re-assigns a missed shift automatically within two&lt;br&gt;
minutes, instead of a manager finding a replacement by phone. Teams that&lt;br&gt;
switched from manual scheduling report needing about 30% less manager time&lt;br&gt;
per week on shift coverage.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now it names the product, states the mechanism, gives a concrete number, and&lt;br&gt;
makes a specific comparison. Someone could read only this paragraph, with&lt;br&gt;
nothing before or after it, and walk away with a correct and complete&lt;br&gt;
understanding of the claim. That's the whole test — not length, not keyword&lt;br&gt;
density, just whether the sentence still works with everything around it&lt;br&gt;
removed.&lt;/p&gt;

&lt;p&gt;Notice this rewrite also happens to be the kind of claim GEO-BENCH found&lt;br&gt;
correlated with more citations in its benchmark: specific, quotable, backed&lt;br&gt;
by a number rather than an adjective. That's not a coincidence — a claim that&lt;br&gt;
stands alone and a claim that's specific enough to quote tend to be the same&lt;br&gt;
sentence. But we're stating that as a plausible connection between two&lt;br&gt;
documented properties, not as a guarantee that this exact rewrite will get&lt;br&gt;
you cited anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put the answer before the build-up
&lt;/h2&gt;

&lt;p&gt;The second habit: when a section implicitly answers a question — "how much&lt;br&gt;
does this cost," "does it support SSO," "how long does setup take" — state&lt;br&gt;
the answer in the first sentence of the section, then explain or qualify it&lt;br&gt;
afterward. Not the reverse, where the section opens with context and works&lt;br&gt;
toward a reveal at the end.&lt;/p&gt;

&lt;p&gt;This isn't from a measured percentage. It's a documented, widely-recommended&lt;br&gt;
practice in web writing generally, for a plain reason: both a human skimming&lt;br&gt;
your page and a system scanning for a relevant passage are doing the same&lt;br&gt;
thing — looking for the sentence that answers the question, not reading in&lt;br&gt;
order for a payoff. A reader who wants to know if you support SSO and finds&lt;br&gt;
that answer buried after three paragraphs of company background has a worse&lt;br&gt;
experience whether or not any AI ever touches the page. Front-loading the&lt;br&gt;
answer serves the actual reader first. That it also gives an extraction&lt;br&gt;
system a clean, early sentence to lift is a second, plausible benefit — not&lt;br&gt;
the reason to do it.&lt;/p&gt;

&lt;p&gt;Compare:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Our platform was built from the ground up with enterprise needs in mind,&lt;br&gt;
and over the years we've worked closely with security teams at companies of&lt;br&gt;
every size to refine an authentication experience that meets a wide range&lt;br&gt;
of compliance requirements. SSO is supported.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;against:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Yes, this supports SSO, via SAML and OIDC. It's configured in Settings →&lt;br&gt;
Security, and works with Okta, Azure AD and Google Workspace out of the&lt;br&gt;
box.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The second version answers the implied question in its first four words. Everything&lt;br&gt;
after that is detail a reader can keep going for or stop reading, having&lt;br&gt;
already gotten what they came for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't promise
&lt;/h2&gt;

&lt;p&gt;Neither habit above comes with a measured lift on real AI engines, and we're&lt;br&gt;
not going to invent one. You may come across posts claiming a specific&lt;br&gt;
percentage improvement from "answer-first" writing, sometimes down to one&lt;br&gt;
decimal place, or a claim that answers need to be a precise word count to&lt;br&gt;
count as citable. We looked. We couldn't trace either kind of number back to&lt;br&gt;
a named study, a named researcher, or any data anyone could check — just&lt;br&gt;
marketing content citing other marketing content. A blog built on "verified,&lt;br&gt;
not estimated" doesn't get to make an exception for a number just because&lt;br&gt;
it's specific-sounding. So we're leaving both out, and we'd suggest treating&lt;br&gt;
any post that states one with unusual precision and no source the same way.&lt;/p&gt;

&lt;p&gt;What we will say: self-contained paragraphs and answer-first sections are&lt;br&gt;
good writing by any standard that predates language models entirely. They&lt;br&gt;
make content easier to skim, easier to excerpt in an email, easier to quote&lt;br&gt;
in a meeting, easier for a tired reader at 11pm to get what they need from&lt;br&gt;
without reading the whole page. If a generative engine also finds that shape&lt;br&gt;
easier to extract cleanly — which the one real study we trust suggests is at&lt;br&gt;
least plausible, since specific and self-contained claims are close cousins —&lt;br&gt;
that's a reasonable bonus on top of writing that was already worth doing. It&lt;br&gt;
is not, and we're not going to pretend it is, a guaranteed technique with a&lt;br&gt;
number attached.&lt;/p&gt;

&lt;p&gt;None of this works if the system pulling the passage can't reach your page in&lt;br&gt;
the first place — see &lt;a href="https://letslaunch.today/blog/why-chatgpt-cant-see-your-site" rel="noopener noreferrer"&gt;why ChatGPT can't see your&lt;br&gt;
site&lt;/a&gt; if you haven't confirmed yours is&lt;br&gt;
readable by a crawler before worrying about how the sentences are shaped.&lt;/p&gt;

&lt;p&gt;And if you're publishing a product description anywhere public, including&lt;br&gt;
&lt;a href="https://letslaunch.today/submit" rel="noopener noreferrer"&gt;a LetsLaunch listing&lt;/a&gt;, the same test applies to that paragraph too&lt;br&gt;
— read it alone, with no page around it, and see if it still holds up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does writing self-contained paragraphs actually increase AI citations?
&lt;/h2&gt;

&lt;p&gt;Nobody has published a study measuring that specific effect on production AI&lt;br&gt;
engines, so we can't claim it does. What's documented is a related, narrower&lt;br&gt;
finding: content with specific quotes and statistics got cited more than&lt;br&gt;
vague content inside a Princeton benchmark. Self-contained, specific writing&lt;br&gt;
is a reasonable, checkable habit that's consistent with that finding and&lt;br&gt;
also makes content easier for a human to skim — not a proven technique with&lt;br&gt;
a measured percentage behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's the quickest way to check if a paragraph is AI-citable?
&lt;/h2&gt;

&lt;p&gt;Copy it alone into a blank document, with nothing before or after it, and&lt;br&gt;
read it as a stranger would. If it still states a complete, correct,&lt;br&gt;
unambiguous claim with no missing context, it's structurally extractable. If&lt;br&gt;
it depends on a pronoun, "as mentioned above," or a fact the previous&lt;br&gt;
sentence supplied, it isn't — rewrite it to carry its own context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should I put the answer first or build up to it?
&lt;/h2&gt;

&lt;p&gt;Put it first, then explain. This isn't based on a measured citation&lt;br&gt;
percentage — it's a documented practice grounded in how both readers and&lt;br&gt;
extraction systems scan text: looking for the sentence that answers the&lt;br&gt;
question, not reading in narrative order. A reader who wants to know if you&lt;br&gt;
support SSO benefits from the answer in the first sentence whether or not any&lt;br&gt;
AI system ever touches the page.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
    </item>
  </channel>
</rss>
