<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mikhail Shadrin</title>
    <description>The latest articles on DEV Community by Mikhail Shadrin (@mikhail_shadrin_dev).</description>
    <link>https://dev.to/mikhail_shadrin_dev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4023230%2F9e82abc3-59a0-4c9b-bf6c-96cd4ede2b89.jpg</url>
      <title>DEV Community: Mikhail Shadrin</title>
      <link>https://dev.to/mikhail_shadrin_dev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mikhail_shadrin_dev"/>
    <language>en</language>
    <item>
      <title>Your site might be blocking AI crawlers by accident. Here's how to check.</title>
      <dc:creator>Mikhail Shadrin</dc:creator>
      <pubDate>Thu, 30 Jul 2026 19:06:56 +0000</pubDate>
      <link>https://dev.to/mikhail_shadrin_dev/your-site-might-be-blocking-ai-crawlers-by-accident-heres-how-to-check-4l</link>
      <guid>https://dev.to/mikhail_shadrin_dev/your-site-might-be-blocking-ai-crawlers-by-accident-heres-how-to-check-4l</guid>
      <description>&lt;p&gt;I build a small tool that checks whether AI models can actually reach a website. A few weeks ago I ran it against our own domain and found we were blocking one of the crawlers we care about, in our own robots.txt. Good reminder that this stuff breaks quietly.&lt;/p&gt;

&lt;p&gt;Here's the part worth knowing even if you never touch a tool for it.&lt;/p&gt;

&lt;p&gt;If your robots.txt blocks an AI crawler, that model won't cite you. Not "less often." Zero. It's a binary gate, and no amount of schema markup, FAQ blocks, or content tweaks changes it. The bot never fetched the page, so there's nothing to quote.&lt;/p&gt;

&lt;p&gt;The trap is that you can block one without meaning to. A lot of the "block the AI bots" advice from 2023-2024 got pasted into robots.txt files and then forgotten. Some CDNs and WordPress security plugins add rules of their own. So the owner assumes the site is open, and one specific model has been locked out for a year.&lt;/p&gt;

&lt;p&gt;How to check it yourself, no tool needed:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open yoursite.com/robots.txt in a browser.&lt;/li&gt;
&lt;li&gt;Look for User-agent lines matching the AI crawlers. The ones that matter right now:

&lt;ul&gt;
&lt;li&gt;GPTBot and OAI-SearchBot (OpenAI / ChatGPT)&lt;/li&gt;
&lt;li&gt;ClaudeBot (Anthropic / Claude)&lt;/li&gt;
&lt;li&gt;Google-Extended (Gemini)&lt;/li&gt;
&lt;li&gt;PerplexityBot&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;If there's a "Disallow: /" under any of those, that model is locked out. If they aren't mentioned at all, you're open to them by default, which is fine.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's the whole check for the part that's genuinely black and white.&lt;/p&gt;

&lt;p&gt;Now the honest caveat, because I'd rather you trust the rest of this. Getting crawled is necessary, not sufficient. It lets you be cited. It doesn't make you cited.&lt;/p&gt;

&lt;p&gt;We ran our on-site "readiness" score against real citations across 44 domains. The correlation was basically nothing (Pearson around -0.08). What predicted citations wasn't markup or content depth, it was brand prominence. Well-known domains got cited, everyone else mostly got zero, no matter how clean their SEO was. Small sample, one model, one snapshot, so don't treat it as a study. But it made me skeptical of every tool that promises "optimize on-page, get cited" (including the pull to oversell my own). The crawl gate is real. The rest is mostly being someone the model already trusts.&lt;/p&gt;

&lt;p&gt;So the practical takeaway is small and boring: make sure you aren't accidentally locked out. That's the one move with a guaranteed effect. Everything after it is a slower game.&lt;/p&gt;

&lt;p&gt;If you'd rather not read robots.txt by hand across a pile of sites, the CLI is open source (causabi-geo on PyPI) and there's a hosted version at causabi.com. Either way, happy to answer questions in the comments.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>ai</category>
      <category>webdev</category>
      <category>opensource</category>
    </item>
    <item>
      <title>We Checked Whether On-Site SEO Predicts AI Citations. The Data Says Mostly No.</title>
      <dc:creator>Mikhail Shadrin</dc:creator>
      <pubDate>Wed, 15 Jul 2026 21:07:09 +0000</pubDate>
      <link>https://dev.to/mikhail_shadrin_dev/we-checked-whether-on-site-seo-predicts-ai-citations-the-data-says-mostly-no-1j8f</link>
      <guid>https://dev.to/mikhail_shadrin_dev/we-checked-whether-on-site-seo-predicts-ai-citations-the-data-says-mostly-no-1j8f</guid>
      <description>&lt;p&gt;Every GEO ("generative engine optimization") tool, including ours until&lt;br&gt;
recently, sells some version of the same pitch: fix your robots.txt, add&lt;br&gt;
Schema.org markup, write FAQ schema, and AI engines will cite you more.&lt;/p&gt;

&lt;p&gt;We build one of these tools — Causabi scans sites for AI-crawler readiness&lt;br&gt;
and generates fix files (robots.txt, llms.txt, JSON-LD, FAQ blocks). As part&lt;br&gt;
of validating our own scoring weights, we ran the numbers on whether the&lt;br&gt;
score actually predicts getting cited. Short version: it mostly doesn't,&lt;br&gt;
once brand prominence is in the picture.&lt;/p&gt;
&lt;h2&gt;
  
  
  What we measured
&lt;/h2&gt;

&lt;p&gt;We scored 44 domains on a 6-category on-site readiness algorithm:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;robots.txt (AI bots allowed or blocked)&lt;/li&gt;
&lt;li&gt;Schema.org (Organization/LocalBusiness JSON-LD completeness)&lt;/li&gt;
&lt;li&gt;FAQ schema (FAQPage markup, 3+ entries)&lt;/li&gt;
&lt;li&gt;content depth/structure&lt;/li&gt;
&lt;li&gt;brand/NAP signals&lt;/li&gt;
&lt;li&gt;freshness (dateModified, recency)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then we checked how often each domain actually got cited by an AI engine&lt;br&gt;
(Claude, via its web-search tool, one measurement window, a fixed prompt set&lt;br&gt;
per domain).&lt;/p&gt;
&lt;h2&gt;
  
  
  What we found
&lt;/h2&gt;

&lt;p&gt;We ran the check twice, and I'll give you both runs because the difference&lt;br&gt;
between them is itself informative:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Run 1 (July 2): Claude only, n=44 usable domains.&lt;/strong&gt; Score vs. citation
rate: Pearson r = -0.078, Spearman ρ = -0.028. 86% of domains got zero
citations regardless of score.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run 2 (July 12): Claude + Gemini, n=41 usable.&lt;/strong&gt; Spearman ρ = +0.084
(p = 0.60), Pearson r = +0.148 (p = 0.35). Zero-citation share dropped to
58.5% — Gemini names domains and brands much more freely than Claude.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Note the sign flipped between runs (-0.03 → +0.08). At these sample sizes&lt;br&gt;
and p-values that's noise, and that's exactly the point: there is no&lt;br&gt;
statistically significant relationship between on-site readiness score and&lt;br&gt;
citation rate in either direction. If the correlation were real and strong,&lt;br&gt;
two runs ten days apart wouldn't disagree on the sign.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The domains that &lt;em&gt;did&lt;/em&gt; get cited clustered almost entirely by brand
prominence — well-known domains got cited at a noticeably higher rate
(~0.16 of prompts in run 1) than everyone else (~0 for the rest of the
sample), regardless of how well-optimized their markup was.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Why I'm not overselling this
&lt;/h2&gt;

&lt;p&gt;n=41-44 is small. This is an internal validation exercise for our own&lt;br&gt;
product, not a peer-reviewed study, and I don't want it read as one.&lt;br&gt;
"No significant correlation found" is not the same claim as "we proved&lt;br&gt;
there is no relationship" — at this sample size we can't prove a negative.&lt;br&gt;
Specific caveats:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Two engines so far (Claude, Gemini). Citation behavior differs meaningfully
across ChatGPT, Grok, and Perplexity — we haven't run the same check
across all of them yet.&lt;/li&gt;
&lt;li&gt;One time window, no longitudinal before/after. We didn't take a domain,
improve its score, and watch citations change over months. That's the
actually convincing experiment and we haven't run it yet.&lt;/li&gt;
&lt;li&gt;Prompt-domain matching wasn't blind. Some prompts were picked because a
domain plausibly related to that topic, which likely biases toward
domains that would get mentioned anyway.&lt;/li&gt;
&lt;li&gt;"Brand prominence" is a fuzzy variable that probably absorbs some real
content-quality signal we're not capturing separately. We can't fully rule
out that what looks like "brand wins" is partly "genuinely better/more
authoritative content wins," which on-site markup scoring doesn't measure.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  What we still think is true, with more confidence
&lt;/h2&gt;

&lt;p&gt;Some things aren't correlational guesses — they're closer to mechanical&lt;br&gt;
facts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;robots.txt blocking is binary.&lt;/strong&gt; If &lt;code&gt;GPTBot&lt;/code&gt;, &lt;code&gt;ClaudeBot&lt;/code&gt;, or similar are
disallowed, that engine cites you zero times, by construction. About 89%
of sites we've scanned block at least one AI crawler by default, usually
by accident (a blanket &lt;code&gt;Disallow: /&lt;/code&gt; that predates AI bots existing).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FAQ schema changes extraction, not inclusion.&lt;/strong&gt; For content that's
already in an engine's consideration set, structuring it as self-contained
Q&amp;amp;A chunks seems to affect whether it gets pulled into a RAG-style
citation — this lines up with published research on chunking behavior. But
that's a "how you're cited" lever, not a "whether you're cited" lever.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Where that leaves the product
&lt;/h2&gt;

&lt;p&gt;We're rewriting our own copy to say what the score actually measures:&lt;br&gt;
AI-crawler readiness and machine-readability, not citation probability. No&lt;br&gt;
tool — ours included — can promise the second one. If your on-site work is&lt;br&gt;
mostly aimed at "getting cited more," the more binding constraint for most&lt;br&gt;
sites is probably brand/mentions elsewhere, not another Schema.org type.&lt;/p&gt;

&lt;p&gt;The scoring engine and fix generator are open source (MIT) if you want to&lt;br&gt;
see the logic or run it on your own site without touching our SaaS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;causabi-geo
geo-optimizer analyze https://yourdomain.com
geo-optimizer fix https://yourdomain.com &lt;span class="nt"&gt;--output&lt;/span&gt; ./out
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Repo: &lt;a href="https://github.com/SHADRINMMM/causabi-geo" rel="noopener noreferrer"&gt;https://github.com/SHADRINMMM/causabi-geo&lt;/a&gt;&lt;br&gt;
Site (hosted version + monitoring): &lt;a href="https://causabi.com" rel="noopener noreferrer"&gt;https://causabi.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>ai</category>
      <category>opensource</category>
      <category>datascience</category>
    </item>
  </channel>
</rss>
