<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: RankCLI</title>
    <description>The latest articles on DEV Community by RankCLI (@rankcli).</description>
    <link>https://dev.to/rankcli</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122693%2F8fec9cfb-7727-4284-9333-9b581c760ce8.png</url>
      <title>DEV Community: RankCLI</title>
      <link>https://dev.to/rankcli</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rankcli"/>
    <language>en</language>
    <item>
      <title>Cloudflare May Be Blocking Googlebot On Your Site, And robots.txt Will Not Tell You</title>
      <dc:creator>RankCLI</dc:creator>
      <pubDate>Wed, 30 Sep 2026 18:48:59 +0000</pubDate>
      <link>https://dev.to/rankcli/cloudflare-may-be-blocking-googlebot-on-your-site-and-robotstxt-will-not-tell-you-3i5h</link>
      <guid>https://dev.to/rankcli/cloudflare-may-be-blocking-googlebot-on-your-site-and-robotstxt-will-not-tell-you-3i5h</guid>
      <description>&lt;p&gt;On 2026-09-15 Cloudflare changed what "blocked" means for AI traffic. New domains, new sites on existing accounts, and free-tier accounts that had not changed the setting now start with &lt;strong&gt;AI Training crawlers disallowed and AI Agents blocked on ad-bearing pages&lt;/strong&gt;. Search was supposed to stay open.&lt;/p&gt;

&lt;p&gt;It did not, for a lot of people.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that catches everyone: mixed-use crawlers
&lt;/h2&gt;

&lt;p&gt;Googlebot crawls for Search &lt;strong&gt;and&lt;/strong&gt; collects data that feeds AI training. Cloudflare classes it as &lt;em&gt;mixed-use&lt;/em&gt;, and a mixed-use crawler is treated under whichever rule is most restrictive.&lt;/p&gt;

&lt;p&gt;So if Training is blocked, including through the older "Block AI Bots" toggle, the block applies to &lt;strong&gt;all&lt;/strong&gt; of Googlebot's functions on ad-supported pages. Googlebot starts getting 403 on pages and on sitemap fetches. The same applies to Bingbot and Applebot.&lt;/p&gt;

&lt;p&gt;Your robots.txt still says &lt;code&gt;Allow&lt;/code&gt;. Nothing in your repository changed. Search Console reports fetch failures and nobody can find the cause, because everyone looks at robots.txt first.&lt;/p&gt;

&lt;p&gt;Cloudflare's own network data puts Googlebot's crawl share at &lt;strong&gt;27.49% in Q2 2026, down from 57.20%&lt;/strong&gt; a year earlier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why robots.txt checks cannot see this
&lt;/h2&gt;

&lt;p&gt;This is the part worth internalising, because it applies well beyond Cloudflare.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;robots.txt is a request. A CDN rule is enforcement.&lt;/strong&gt; They live in different layers and they fail in different places. robots.txt is served by your origin and read by a crawler that chooses to obey it. A bot rule at the edge refuses the connection before your origin is ever consulted.&lt;/p&gt;

&lt;p&gt;A crawler can be explicitly allowed in robots.txt and still receive a 403. Every "AI visibility" checker I know of, including ours until this week, stops at robots.txt and reports the site as open.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking it by hand
&lt;/h2&gt;

&lt;p&gt;Two curl requests, one as a browser and one as Googlebot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s2"&gt;"%{http_code}&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s2"&gt;n"&lt;/span&gt; &lt;span class="se"&gt;\\&lt;/span&gt;
  &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="s2"&gt;"Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"&lt;/span&gt; &lt;span class="se"&gt;\\&lt;/span&gt;
  https://your-site.dev/

curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s2"&gt;"%{http_code}&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s2"&gt;n"&lt;/span&gt; &lt;span class="se"&gt;\\&lt;/span&gt;
  &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="s2"&gt;"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 Chrome/131.0"&lt;/span&gt; &lt;span class="se"&gt;\\&lt;/span&gt;
  https://your-site.dev/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Browser 200 and Googlebot 403 means an edge rule, not robots.txt. Check the sitemap the same way, because that is often where it shows first.&lt;/p&gt;

&lt;p&gt;One honest caveat: real crawlers are verified by IP as well as user-agent, so a spoofed agent is indicative rather than conclusive. Confirm with Search Console's URL Inspection, which fetches as the verified Googlebot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Or in one command
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @rankcli/cli@latest audit &lt;span class="nt"&gt;-u&lt;/span&gt; https://your-site.dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It asks as Googlebot, bingbot, GPTBot, ClaudeBot and PerplexityBot, compares each against a browser request, and reports two separate findings:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;SEARCH_CRAWLER_BLOCKED_AT_EDGE&lt;/code&gt;&lt;/strong&gt;, an error, when a search crawler is refused while a browser is served&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;AI_CRAWLER_BLOCKED_AT_EDGE&lt;/code&gt;&lt;/strong&gt;, a warning, for the AI crawlers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Separate because the costs are different. Losing Googlebot costs you the index. Losing GPTBot costs you citations in AI answers. One is an emergency, the other is a decision you might have made deliberately.&lt;/p&gt;

&lt;p&gt;No signup, nothing leaves your machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it will not do
&lt;/h2&gt;

&lt;p&gt;It will not report anything if a browser cannot fetch your site either. A site that is simply down is not a bot-rule problem, and turning one outage into six confident findings about rules that do not exist is exactly the kind of false alarm that teaches people to ignore their tools.&lt;/p&gt;

&lt;p&gt;Tested against reddit.com, which refuses Googlebot, bingbot and GPTBot while serving browsers: it reports all three, split correctly. Tested against nytimes.com, which refuses browsers too: it reports nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix, if you are affected
&lt;/h2&gt;

&lt;p&gt;It is in your CDN, not your repository. On Cloudflare: zone &lt;strong&gt;Security → Bots&lt;/strong&gt;, and any "Block AI Bots" or AI Crawler Control toggle. Because Googlebot is mixed-use, you need to allow Search crawlers explicitly or opt out of the default on ad-supported pages.&lt;/p&gt;

&lt;p&gt;Existing paid customers were notified in advance and can opt out in zone security settings. New domains and free-tier accounts got the default without asking for it, which is why this keeps surprising people who never changed a setting.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://rankcli.dev/blog/is-cloudflare-blocking-googlebot-on-your-site" rel="noopener noreferrer"&gt;rankcli.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>cloudflare</category>
      <category>googlebot</category>
    </item>
    <item>
      <title>Your SEO Fixes Now Land Where Your Framework Actually Serves Them</title>
      <dc:creator>RankCLI</dc:creator>
      <pubDate>Mon, 28 Sep 2026 21:01:08 +0000</pubDate>
      <link>https://dev.to/rankcli/your-seo-fixes-now-land-where-your-framework-actually-serves-them-2h7</link>
      <guid>https://dev.to/rankcli/your-seo-fixes-now-land-where-your-framework-actually-serves-them-2h7</guid>
      <description>&lt;p&gt;Five releases went out over two days. Here is what changed, and why each one exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  Detection rebuilt: 57% to 94%
&lt;/h2&gt;

&lt;p&gt;A &lt;code&gt;robots.txt&lt;/code&gt; written to the wrong directory is worse than no &lt;code&gt;robots.txt&lt;/code&gt; at all. The audit stops reporting it missing, while crawlers still never see it. So before an auto-fix can write a file, it has to know what it is looking at.&lt;/p&gt;

&lt;p&gt;The old detector returned the first signal that matched. That is fast and wrong: signals conflict, and the first one to fire is not the most reliable one. It also matched on short substrings, which produced answers like classifying &lt;strong&gt;nginx as Gin&lt;/strong&gt; because &lt;code&gt;server&lt;/code&gt; contained "gin", and classifying any page with &lt;code&gt;style="width: 50%"&lt;/code&gt; as Spring Boot because a Thymeleaf pattern matched &lt;code&gt;th:&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Detection now weighs every signal on the page. Each contributes evidence at one of three strengths, the highest total wins, and a score below the floor returns &lt;code&gt;unknown&lt;/code&gt; rather than a guess. On a 51-site corpus that moved accuracy from &lt;strong&gt;57% to 94%&lt;/strong&gt; — with no incorrect claims. When it names a framework, it is right; when it is unsure, it says so.&lt;/p&gt;

&lt;p&gt;It recognises &lt;strong&gt;32 frameworks&lt;/strong&gt; now, including Docusaurus and Sphinx.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixes land in the right directory
&lt;/h2&gt;

&lt;p&gt;Knowing the framework is only useful if you then act on it. The served directory is a property of the framework, not of whether you happen to have created it yet:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Framework&lt;/th&gt;
&lt;th&gt;Static files go to&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;React, Vue, Next.js, Nuxt, Astro, Remix, Rails, Laravel&lt;/td&gt;
&lt;td&gt;&lt;code&gt;public/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SvelteKit, Hugo, Gatsby, Docusaurus, Fresh&lt;/td&gt;
&lt;td&gt;&lt;code&gt;static/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ASP.NET Core&lt;/td&gt;
&lt;td&gt;&lt;code&gt;wwwroot/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Phoenix&lt;/td&gt;
&lt;td&gt;&lt;code&gt;priv/static/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spring Boot&lt;/td&gt;
&lt;td&gt;&lt;code&gt;src/main/resources/static/&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jekyll, Pelican&lt;/td&gt;
&lt;td&gt;the app root itself&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;An earlier version looked for an existing &lt;code&gt;public/&lt;/code&gt; directory in your repository. That is right for a project that has one and wrong for every project that does not yet — a fresh Next.js app, a fresh Nuxt app and a Docusaurus site all resolved to the repository root, because the directory they serve from had not been created.&lt;/p&gt;

&lt;p&gt;Monorepos work too: the app root is found by locating an &lt;code&gt;index.html&lt;/code&gt; beside a &lt;code&gt;package.json&lt;/code&gt;, so fixes land in the right package rather than at the top of the tree.&lt;/p&gt;

&lt;h2&gt;
  
  
  A safety layer that refuses to commit
&lt;/h2&gt;

&lt;p&gt;Generated file edits are now checked before they become a commit, and anything carrying a concern is dropped rather than committed. Among what it refuses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A replacement that lost a &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; tag, a stylesheet, or a search-console verification tag&lt;/li&gt;
&lt;li&gt;Invalid JSON, XML or JSON-LD&lt;/li&gt;
&lt;li&gt;A sitemap that came back shorter, or a &lt;code&gt;lastmod&lt;/code&gt; moved backwards&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Invented content&lt;/strong&gt; — a phone number in the reserved 555 range, "123 Main Street, Anytown", an unfilled &lt;code&gt;your_...&lt;/code&gt; placeholder&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A canonical added to a single-page-app shell&lt;/strong&gt;, which would make every route declare itself a duplicate of one URL&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those last two are there because a real generated pull request proposed both. We &lt;a href="https://rankcli.dev/blog/we-found-five-bugs-in-our-own-seo-tool" rel="noopener noreferrer"&gt;wrote that one up separately&lt;/a&gt; — it is the more interesting story.&lt;/p&gt;

&lt;p&gt;Files RankCLI owns from a template — &lt;code&gt;robots.txt&lt;/code&gt;, &lt;code&gt;sitemap.xml&lt;/code&gt;, &lt;code&gt;llms.txt&lt;/code&gt; — may be created but never rewritten by a model. A template cannot hallucinate a date; a model can, and did.&lt;/p&gt;

&lt;h2&gt;
  
  
  One-click install for Claude, Cursor and VS Code
&lt;/h2&gt;

&lt;p&gt;The MCP server is the lowest-friction thing we ship: no signup, no API key, nothing leaves your machine. Every path we documented for it involved hand-editing a JSON config file, which is a strange way to introduce the easy option.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Desktop&lt;/strong&gt; now has a downloadable extension bundle — grab the &lt;code&gt;.mcpb&lt;/code&gt; from the &lt;a href="https://github.com/integrallis/rankcli-cli/releases/latest" rel="noopener noreferrer"&gt;latest release&lt;/a&gt; and double-click it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt; is one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add rankcli &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @rankcli/mcp-server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Cursor and VS Code&lt;/strong&gt; have one-click install links on &lt;a href="https://rankcli.dev/docs" rel="noopener noreferrer"&gt;the docs page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;All of it drives the same &lt;strong&gt;15 tools&lt;/strong&gt; — full audits, AI-crawler access checks, structured data validation, Core Web Vitals, security headers, and generators for the fixes.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI readiness that discriminates
&lt;/h2&gt;

&lt;p&gt;The AI-readiness sub-score used to read near zero for essentially every site, including ones that do it well. A score that says zero for everyone looks like a harsh grader and is actually a broken instrument — it carried no information and could not be acted on.&lt;/p&gt;

&lt;p&gt;It is recalibrated. Real sites now spread across a usable range and rank in a defensible order, so the number means something when it moves.&lt;/p&gt;

&lt;h2&gt;
  
  
  A tier ladder with no flat rung
&lt;/h2&gt;

&lt;p&gt;Auto-fix PR limits now scale properly across plans. The cap on how many pull requests one audit report can open was a flat 3 on every plan — which made the top tier identical to the free one on the limit its customers actually hit, working through a single large report across many sites. It is 3 / 3 / 5 / 10 now, and a test asserts it keeps rising: a rung that does not rise is a rung that means nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @rankcli/cli audit &lt;span class="nt"&gt;-u&lt;/span&gt; https://your-site.dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No signup, no account, 280+ checks. The CLI, the MCP server and the GitHub Action are free forever.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://rankcli.dev/blog/framework-aware-seo-fixes-and-one-click-mcp" rel="noopener noreferrer"&gt;rankcli.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>devops</category>
      <category>ai</category>
    </item>
    <item>
      <title>We Ran Our SEO Tool On Ourselves And Found Five Bugs. One Would Have Deindexed Our Site.</title>
      <dc:creator>RankCLI</dc:creator>
      <pubDate>Mon, 28 Sep 2026 20:45:48 +0000</pubDate>
      <link>https://dev.to/rankcli/we-ran-our-seo-tool-on-ourselves-and-found-five-bugs-one-would-have-deindexed-our-site-p5m</link>
      <guid>https://dev.to/rankcli/we-ran-our-seo-tool-on-ourselves-and-found-five-bugs-one-would-have-deindexed-our-site-p5m</guid>
      <description>&lt;p&gt;We publish rankcli.dev's own audit score every week, whether it went up or not. This week we did something slightly different: instead of reading the number, we read every finding and asked whether it was actually true.&lt;/p&gt;

&lt;p&gt;Five of them were not. Here they are, because the failures are more useful than the score.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. A page that asked not to be indexed, faulted for not being indexed
&lt;/h2&gt;

&lt;p&gt;For the same &lt;code&gt;/login&lt;/code&gt; page, in the same run, the audit reported two things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;NOINDEX_TAG&lt;/code&gt; - correct, it carries &lt;code&gt;&amp;lt;meta name="robots" content="noindex, nofollow"&amp;gt;&lt;/code&gt; on purpose&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;REACT_CSR_DETECTED&lt;/code&gt;, severity &lt;strong&gt;error&lt;/strong&gt;: &lt;em&gt;"Critical - your content is likely invisible to Google"&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It knew the page had opted out of search, then raised a critical error because search could not see it.&lt;/p&gt;

&lt;p&gt;Every SaaS has a noindex &lt;code&gt;/login&lt;/code&gt;, &lt;code&gt;/signup&lt;/code&gt; and &lt;code&gt;/account&lt;/code&gt;. They are all client-rendered shells. So every one of our users with a login page was collecting a critical SEO error that nobody should ever act on.&lt;/p&gt;

&lt;p&gt;Framework checks now carry an &lt;code&gt;indexable_only&lt;/code&gt; flag, and 19 of them across 9 frameworks are marked with it. Deliberately not marked: render-blocking scripts, which is about performance, and hydration mismatches, which break the page for real users too. Those still fire on a noindex page, because they are still true there.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Every dark-themed page failed colour contrast
&lt;/h2&gt;

&lt;p&gt;The contrast checker read only an element's own inline style. If it found no background, it assumed white.&lt;/p&gt;

&lt;p&gt;Our &lt;code&gt;/login&lt;/code&gt; fallback is &lt;code&gt;#a1a1aa&lt;/code&gt; text painted on &lt;code&gt;#0a0a0a&lt;/code&gt;. That is about &lt;strong&gt;9:1&lt;/strong&gt; - comfortably past WCAG AA. We reported &lt;strong&gt;2.56:1&lt;/strong&gt; and called it critical.&lt;/p&gt;

&lt;p&gt;It now reads the background from the nearest ancestor that paints one, and when nothing up the tree paints anything it records the element as undetermined rather than measuring against a guess. Inline styles only, deliberately: resolving a real CSS cascade without a browser is a much larger job, and half-doing it would reintroduce exactly the confident guess we were removing.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Our AI readiness score was 1%. So was everyone's.
&lt;/h2&gt;

&lt;p&gt;This is the one that stung, because AI readiness is the thing people come to us for.&lt;/p&gt;

&lt;p&gt;We assumed our content was bad. Then we ran the same check against sites that ought to score well:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Site&lt;/th&gt;
&lt;th&gt;AI Readiness (before)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;stripe.com&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vercel.com&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cloudflare.com&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;docs.astro.build&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rankcli.dev&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A documentation site - the ideal target for AI citation - scored zero. The metric was saturated. It carried no information at all.&lt;/p&gt;

&lt;p&gt;The cause was a category multiplier copied from categories with a tenth of the checks. Between 15 and 27 AI-readiness checks fire on a typical page, most of them aspirational content notices: no comparison table, no testimonials, no original data. At one point per notice and three per warning, multiplied by three, any real page blew past 100 points of deduction before the interesting checks were counted.&lt;/p&gt;

&lt;p&gt;The same five pages now spread &lt;strong&gt;52-71&lt;/strong&gt;, and rank in a defensible order.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Whitespace could fail your build
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;--fail-on error&lt;/code&gt; exits 1. So an error in this tool is not a line in a report - it is somebody's broken build.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;DOM_SIZE_EXCESSIVE&lt;/code&gt; counted every node the parser returns, while its thresholds (1,500 and 3,000) are Lighthouse's - and Lighthouse counts &lt;strong&gt;elements&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;On tailwindcss.com that was &lt;strong&gt;3,231 counted against 2,230 real elements&lt;/strong&gt;. The difference: 796 whitespace nodes and 205 comments. Enough to push a warning-sized DOM over the error line.&lt;/p&gt;

&lt;p&gt;Which means re-indenting a template, or leaving comments in your output, could have failed your build. It counts elements now, and a test asserts that formatting a page cannot change its severity.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Missing image dimensions was rated "your site is broken"
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;IMAGES_MISSING_DIMENSIONS&lt;/code&gt; fired at &lt;strong&gt;error&lt;/strong&gt; for three or more images. That failed svelte.dev, tailwindcss.com, rust-lang.org, laravel.com and fastapi.tiangolo.com - five of the ten well-built sites we swept.&lt;/p&gt;

&lt;p&gt;Missing width and height is a real CLS cost and we still report it. It is not "your site is broken". Two other places in our own codebase already treated the same condition as a notice and as info; error was the outlier, not the consensus.&lt;/p&gt;

&lt;h2&gt;
  
  
  And then the auto-fix PR tried to deindex us
&lt;/h2&gt;

&lt;p&gt;While reviewing an open auto-fix pull request against our own repo before merging it, we found this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;"canonical"&lt;/span&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;"https://rankcli.dev/"&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;added to the Vite &lt;code&gt;index.html&lt;/code&gt; - the shared shell that serves every route. The file carries a comment three lines above the insertion point explaining precisely why that must never happen: Helmet cannot remove static tags, so every page would inherit it.&lt;/p&gt;

&lt;p&gt;Merged, every page on the site - pricing, docs, blog - would have declared itself a duplicate of the homepage.&lt;/p&gt;

&lt;p&gt;The same PR proposed Organization structured data with a phone number of &lt;code&gt;+1-800-555-0199&lt;/code&gt;, an address of "123 Main Street, Anytown, CA 90210", and a social profile of &lt;code&gt;youtube.com/channel/your_youtube_channel&lt;/code&gt;. None of it real. It also linked &lt;code&gt;/images/og-image.png&lt;/code&gt; and &lt;code&gt;/search&lt;/code&gt;, both of which 404.&lt;/p&gt;

&lt;p&gt;We closed it. Then we fixed the generator: our safety layer now refuses &lt;strong&gt;placeholder content&lt;/strong&gt; and &lt;strong&gt;a canonical added to a single-page-app shell&lt;/strong&gt;, and any file carrying a safety concern is dropped rather than committed. We tested the new guards against that pull request's actual diff - all six problems are caught, while a real phone number in ordinary prose still passes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we actually learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A confident wrong "critical" is worse than a missing check.&lt;/strong&gt; A missing check costs you one finding. A false critical teaches people that the red ones are noise - and then the real one goes unread.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Saturated metrics are invisible failures.&lt;/strong&gt; A score that reads zero for everyone looks like a harsh grader, not a broken instrument. The only way we caught it was running the same check against sites we already believed were good, and being surprised.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generated fixes need a human read.&lt;/strong&gt; Ours carry a disclaimer saying so. This week that disclaimer stopped being boilerplate.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, today
&lt;/h2&gt;

&lt;p&gt;rankcli.dev scores &lt;strong&gt;75/100&lt;/strong&gt; with &lt;strong&gt;0 errors&lt;/strong&gt; and &lt;strong&gt;AI Readiness 67&lt;/strong&gt;, measured with the published CLI. Not because the site improved this week - it barely changed. Because the tool stopped lying about it.&lt;/p&gt;

&lt;p&gt;You can run exactly the same audit, with no signup and no account:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @rankcli/cli audit &lt;span class="nt"&gt;-u&lt;/span&gt; https://your-site.dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it tells you something that is not true, we would genuinely like to know. That is how all five of these were found.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://rankcli.dev/blog/we-found-five-bugs-in-our-own-seo-tool" rel="noopener noreferrer"&gt;rankcli.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>testing</category>
      <category>showdev</category>
    </item>
    <item>
      <title>RankCLI weekly dogfood report — 2026-09-15</title>
      <dc:creator>RankCLI</dc:creator>
      <pubDate>Tue, 15 Sep 2026 10:03:50 +0000</pubDate>
      <link>https://dev.to/rankcli/rankcli-weekly-dogfood-report-2026-09-15-58e7</link>
      <guid>https://dev.to/rankcli/rankcli-weekly-dogfood-report-2026-09-15-58e7</guid>
      <description>&lt;p&gt;🤖 Weekly dogfood: rankcli.dev scored 88/100 this week (11 open issues) — audited by the same tool we ship. rankcli.dev&lt;/p&gt;

&lt;p&gt;Read more or run your own audit: &lt;a href="https://rankcli.dev" rel="noopener noreferrer"&gt;rankcli.dev&lt;/a&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Is Your Site Even Visible to ChatGPT?</title>
      <dc:creator>RankCLI</dc:creator>
      <pubDate>Sun, 13 Sep 2026 13:02:00 +0000</pubDate>
      <link>https://dev.to/rankcli/is-your-site-even-visible-to-chatgpt-281f</link>
      <guid>https://dev.to/rankcli/is-your-site-even-visible-to-chatgpt-281f</guid>
      <description>&lt;p&gt;You've optimized for Google. Sitemap's clean, Core Web Vitals are green, structured data validates. But when someone asks ChatGPT or Perplexity about your product, you're nowhere in the answer — and a competitor with a worse site is.&lt;/p&gt;

&lt;p&gt;The likely reason: the crawlers that feed AI answers aren't Googlebot, and they don't see your site the way Googlebot does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meet the other crawlers
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPTBot&lt;/strong&gt; (OpenAI) — crawls for model training and, via a separate retrieval path, for real-time browsing/citation in ChatGPT answers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ClaudeBot&lt;/strong&gt; (Anthropic) — same idea, powers Claude's web search and citations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PerplexityBot&lt;/strong&gt; — crawls specifically to answer live queries with citations; this is the one most directly tied to "will I get mentioned."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google-Extended&lt;/strong&gt; — a separate opt-in signal from regular Googlebot, controls whether Google's AI features can use your content.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each respects robots.txt, but as a distinct user-agent — blocking Googlebot doesn't block these, and blocking one doesn't block the others. Most sites have never explicitly considered any of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a site that ranks fine can still be invisible
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. robots.txt blocks it, often by accident.&lt;/strong&gt; Some site generators and "AI protection" plugins add blanket Disallow rules for GPTBot/CCBot/etc. without the owner realizing it also kills citation eligibility, not just training-data scraping. Check yours:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://yoursite.com/robots.txt | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-A2&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"gptbot&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="s2"&gt;claudebot&lt;/span&gt;&lt;span class="se"&gt;\|&lt;/span&gt;&lt;span class="s2"&gt;perplexitybot"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. The content is client-rendered only.&lt;/strong&gt; This is the big one. Traditional SEO crawlers (Googlebot) execute JavaScript and wait for your SPA to render. Several AI crawlers fetch raw HTML and move on — no JS execution. If your content only exists after a React/Vue mount, an AI crawler may see an empty &lt;code&gt;&amp;lt;div id="root"&amp;gt;&lt;/code&gt; and nothing else.&lt;/p&gt;

&lt;p&gt;Test what a crawler actually receives, vs. what your browser renders:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="s2"&gt;"GPTBot"&lt;/span&gt; https://yoursite.com/ | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s1"&gt;'&amp;lt;title&amp;gt;.*&amp;lt;/title&amp;gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that title tag (or the body content you care about) is missing or generic, but your browser shows it fine — that's your JS-rendering gap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. No structured data, or broken structured data.&lt;/strong&gt; AI answer engines lean on schema.org JSON-LD to disambiguate what a page is actually about — product, article, FAQ, organization — rather than inferring it purely from prose. A page with solid content but zero structured data is harder for a retrieval system to confidently cite.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. No llms.txt.&lt;/strong&gt; An emerging, not-yet-universal convention — a plain-text file at /llms.txt giving a curated, LLM-readable summary of what your site is and where the important pages are, the same spirit as sitemap.xml but written for language models instead of search indexers. Not every crawler uses it yet, but it costs almost nothing to add and directly addresses the "can a model even summarize what we do" problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Stale noindex leftovers.&lt;/strong&gt; Staging-environment meta tags (&lt;code&gt;&amp;lt;meta name="robots" content="noindex"&amp;gt;&lt;/code&gt;) that never got removed after launch block everything — traditional and AI crawlers alike.&lt;/p&gt;

&lt;h2&gt;
  
  
  A five-minute self-check
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;curl -A "GPTBot" yoursite.com/robots.txt&lt;/code&gt; — confirm you're not accidentally blocked.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;curl -A "GPTBot" yoursite.com/&lt;/code&gt; and diff it against what your browser renders — confirm the content is actually in the raw HTML.&lt;/li&gt;
&lt;li&gt;Validate your structured data (Google's Rich Results Test works fine for this even though it's Google-branded — it's just a schema.org validator).&lt;/li&gt;
&lt;li&gt;Check if /llms.txt exists. If not, it's a 20-minute add.&lt;/li&gt;
&lt;li&gt;Grep your codebase for leftover noindex tags.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Automating this
&lt;/h2&gt;

&lt;p&gt;We ran into this exact problem building &lt;a href="https://rankcli.dev" rel="noopener noreferrer"&gt;RankCLI&lt;/a&gt; — the open MCP server (&lt;code&gt;npx @rankcli/mcp-server&lt;/code&gt;, no signup) runs GEO checks for all of the above alongside the usual technical-SEO audit, so it's one pass instead of five manual curl commands every time you ship. Free tier is genuinely free; the hosted layer just adds scheduling and auto-fix PRs on top.&lt;/p&gt;

&lt;p&gt;But even without any tool, the five checks above take less time than writing this sentence took, and most sites have never run them once.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>ai</category>
      <category>webdev</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
