<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 500wango</title>
    <description>The latest articles on DEV Community by 500wango (@500wango).</description>
    <link>https://dev.to/500wango</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4105642%2Ff5ee14a0-2f9e-47b7-8f90-c656376a12fe.png</url>
      <title>DEV Community: 500wango</title>
      <link>https://dev.to/500wango</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/500wango"/>
    <language>en</language>
    <item>
      <title>llms.txt best practices: facts, canonical URLs, what to leave out</title>
      <dc:creator>500wango</dc:creator>
      <pubDate>Sun, 04 Oct 2026 14:36:54 +0000</pubDate>
      <link>https://dev.to/500wango/llmstxt-best-practices-facts-canonical-urls-what-to-leave-out-1dn2</link>
      <guid>https://dev.to/500wango/llmstxt-best-practices-facts-canonical-urls-what-to-leave-out-1dn2</guid>
      <description>&lt;p&gt;With AI search engines (Perplexity, ChatGPT Search, Claude) shifting how content is discovered, the proposed llms.txt standard (from llmstxt.org) has gained massive attention.          &lt;/p&gt;

&lt;p&gt;Think of llms.txt as a Markdown-based parallel to robots.txt and sitemap.xml. Instead of telling web spiders where to crawl, it serves clean, token-efficient, machine-readable facts and&lt;br&gt;
  canonical documentation directly to LLMs.                                                                                                                                                &lt;/p&gt;

&lt;p&gt;However, after inspecting hundreds of newly deployed llms.txt files, we noticed developers making common mistakes that actually harm how models interpret their sites.                   &lt;/p&gt;

&lt;p&gt;Here are the practical best practices for structuring your discovery file.                                                                                                               &lt;/p&gt;

&lt;p&gt;### 1. Where and How to Serve It                                                                                                                                                         &lt;/p&gt;

&lt;p&gt;• Path: Must be served at the root of your domain: &lt;a href="https://yourdomain.com/llms.txt" rel="noopener noreferrer"&gt;https://yourdomain.com/llms.txt&lt;/a&gt;&lt;br&gt;&lt;br&gt;
  • MIME type: text/plain; charset=utf-8 (Do not serve it as text/html or application/octet-stream).&lt;br&gt;&lt;br&gt;
  • HTTP status: Must return 200 OK without requiring authentication, cookies, or redirect chains.                                                                                         &lt;/p&gt;

&lt;p&gt;### 2. Standard Syntax Anatomy                                                                                                                                                           &lt;/p&gt;

&lt;p&gt;An effective llms.txt follows a simple, hierarchical Markdown structure:                                                                                                                 &lt;/p&gt;

&lt;p&gt;Your Brand / Project Name                                                                                                                                                               &lt;/p&gt;

&lt;p&gt;│ A concise 1-2 sentence description of what the project actually does. Avoid buzzwords and marketing fluff.                                                                             &lt;/p&gt;

&lt;p&gt;## Core Documentation                                                                                                                                                                    &lt;/p&gt;

&lt;p&gt;• Quickstart Guide &lt;a href="https://yourdomain.com/docs/quickstart:" rel="noopener noreferrer"&gt;https://yourdomain.com/docs/quickstart:&lt;/a&gt; Step-by-step setup in under 5 minutes.&lt;br&gt;&lt;br&gt;
  • API Reference &lt;a href="https://yourdomain.com/docs/api:" rel="noopener noreferrer"&gt;https://yourdomain.com/docs/api:&lt;/a&gt; REST API endpoints and authentication rules.&lt;br&gt;&lt;br&gt;
  • Pricing &lt;a href="https://yourdomain.com/pricing:" rel="noopener noreferrer"&gt;https://yourdomain.com/pricing:&lt;/a&gt; Official tiers, free limits, and billing terms.                                                                                                &lt;/p&gt;

&lt;p&gt;## Key Facts                                                                                                                                                                             &lt;/p&gt;

&lt;p&gt;• Founded: 2024&lt;br&gt;&lt;br&gt;
  • Pricing Model: Freemium ($0 / $29 / $99 per month)&lt;br&gt;
  • Supported Frameworks: Node.js, Python, Go&lt;/p&gt;

&lt;p&gt;### 3. What NOT to Include (The Top 3 Pitfalls)&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Stale Pricing or Claims:
Never hardcode pricing numbers into llms.txt that you don't update religiously. If your live page says $49/mo but your llms.txt says $29/mo, AI engines flag this contradiction as a
trust failure. Always link to your live pricing page.&lt;/li&gt;
&lt;li&gt;Marketing Slogans:
Phrases like "The world's leading revolutionary AI platform" take up valuable context window tokens and offer zero factual grounding. Stick to strict technical capabilities and
verifiable facts.&lt;/li&gt;
&lt;li&gt;Private or Gated Routes:
Do not list staging URLs, internal admin panels, or authenticated dashboard links. Crawlers will only fetch HTTP 200 public resources.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;### 4. Optional: Full Context via llms-full.txt&lt;/p&gt;

&lt;p&gt;If your project has comprehensive API docs or whitepapers that fit within a ~100k token context, you can also publish a companion file at /llms-full.txt containing the full concatenated&lt;br&gt;
  text. In your main llms.txt, link to it at the bottom:&lt;/p&gt;

&lt;p&gt;• Full Documentation /llms-full.txt: Complete combined documentation for offline context.&lt;/p&gt;

&lt;p&gt;### Want to Generate or Validate Your Draft?&lt;/p&gt;

&lt;p&gt;If you want to quickly scaffold a clean llms.txt file from your homepage metadata or validate your existing syntax:&lt;/p&gt;

&lt;p&gt;👉 Free llms.txt Generator &amp;amp; Validator &lt;a href="https://citeaura.com/llms-txt-tool" rel="noopener noreferrer"&gt;https://citeaura.com/llms-txt-tool&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Are you already experimenting with llms.txt on your projects? Do you see genuine crawler hits on it in your server logs yet?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>seo</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Is GPTBot Silently Blocked by Your Cloudflare WAF? How to Test and Fix AI Crawler Access</title>
      <dc:creator>500wango</dc:creator>
      <pubDate>Sun, 04 Oct 2026 14:08:12 +0000</pubDate>
      <link>https://dev.to/500wango/is-gptbot-silently-blocked-by-your-cloudflare-waf-how-to-test-and-fix-ai-crawler-3p5i</link>
      <guid>https://dev.to/500wango/is-gptbot-silently-blocked-by-your-cloudflare-waf-how-to-test-and-fix-ai-crawler-3p5i</guid>
      <description>&lt;p&gt;You launch a site, set up your &lt;code&gt;robots.txt&lt;/code&gt; to welcome AI search bots, and move on:                                                                                                    &lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User-agent: GPTBot                                                                                                                                                                     
Allow: /                                                                                                                                                                               
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;A few weeks later, you check your server access logs or wonder why your content never gets cited in ChatGPT Search or Perplexity. You find zero crawler hits.                            &lt;/p&gt;

&lt;p&gt;The culprit is almost never your robots.txt. It is usually an edge security rule or web application firewall (WAF) blocking the crawler before the request ever touches your origin server.                                                                                                                                                                                  &lt;/p&gt;

&lt;p&gt;Here is what is happening under the hood and how to test it.&lt;br&gt;&lt;br&gt;
  ──────                                                                                                                                                                                   &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Edge Challenge Problem
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most modern web architectures sit behind Cloudflare, Fastly, or AWS CloudFront.                                                                                                          &lt;/p&gt;

&lt;p&gt;When you enable features like Cloudflare Super Bot Fight Mode or aggressive rate-limiting:                                                                                               &lt;/p&gt;

&lt;p&gt;• The edge evaluates the incoming HTTP request.&lt;br&gt;&lt;br&gt;
  • OpenAI, Anthropic, and Perplexity use dynamic IP subnets that change frequently.&lt;br&gt;&lt;br&gt;
  • If the WAF cannot instantly verify the crawler via reverse DNS or trusted ASN matching, it serves an HTTP 403 Forbidden or a Cloudflare JavaScript Challenge (Turnstile).              &lt;/p&gt;

&lt;p&gt;Because an automated crawler cannot solve an interactive browser challenge, it simply fails and drops the page from its index.                                                           &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How to Test Your Live Site with cURL
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You do not need to wait weeks to know if you are affected. Run this command from your terminal:                                                                                          &lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;curl -I -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)" https://yourdomain.com                                          
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Look at the HTTP status header:                                                                                                                                                          &lt;/p&gt;

&lt;p&gt;• HTTP/2 200 OK: You are in the clear. The crawler can read your HTML.&lt;br&gt;&lt;br&gt;
  • HTTP/2 403 Forbidden: Your WAF or hosting provider is actively blocking OpenAI.&lt;br&gt;&lt;br&gt;
  • cf-mitigated: challenge: Cloudflare is intercepting the crawler with an interactive bot challenge.&lt;br&gt;&lt;br&gt;
  ──────                                                                                                                                                                                   &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How to Allow AI Crawlers Safely
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you find that GPTBot is being blocked, do not turn off your entire WAF. Instead, create a targeted bypass rule in Cloudflare:                                                         &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Go to Security -&amp;gt; WAF -&amp;gt; Custom Rules.
&lt;/li&gt;
&lt;li&gt;Create a rule named Allow Verified AI Crawlers.
&lt;/li&gt;
&lt;li&gt;Set the condition:
  • (cf.client.bot and http.user_agent contains "GPTBot")
&lt;/li&gt;
&lt;li&gt;Set Action to Skip:
  • Check All remaining custom rules
  • Check Super Bot Fight Mode (or Bot Management)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This ensures legitimate OpenAI search crawlers can fetch public content while your protected admin and API endpoints stay secure.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Need an Instant Sanity Check?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you don't have terminal access or want to check multiple bots (GPTBot, ClaudeBot, PerplexityBot, ByteSpider) in 5 seconds, we built a lightweight, no-signup checker that runs these&lt;br&gt;&lt;br&gt;
  cURL validations for you:&lt;/p&gt;

&lt;p&gt;Free AI Crawler Access Checker &lt;a href="https://citeaura.com/crawler-check" rel="noopener noreferrer"&gt;https://citeaura.com/crawler-check&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;How are you handling AI search crawlers in your current infra? Are you whitelisting them or keeping them blocked by default?&lt;/p&gt;

</description>
      <category>devops</category>
      <category>webdev</category>
      <category>cloudflare</category>
      <category>ai</category>
    </item>
    <item>
      <title>Is GPTBot Silently Blocked by Your Cloudflare WAF? How to Test and Fix AI Crawler Access</title>
      <dc:creator>500wango</dc:creator>
      <pubDate>Sun, 04 Oct 2026 13:58:18 +0000</pubDate>
      <link>https://dev.to/500wango/check-if-your-site-or-cloudflare-waf-is-silently-blocking-gptbot-3img</link>
      <guid>https://dev.to/500wango/check-if-your-site-or-cloudflare-waf-is-silently-blocking-gptbot-3img</guid>
      <description>&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You launch a site, set up your `robots.txt` to welcome AI search bots, and move on:                                                                                                    
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
txt                                                                                                                                                                                 
    User-agent: GPTBot                                                                                                                                                                     
    Allow: /                                                                                                                                                                               

  A few weeks later, you check your server access logs or wonder why your content never gets cited in ChatGPT Search or Perplexity. You find zero crawler hits.                            

  The culprit is almost never your robots.txt. It is usually an edge security rule or web application firewall (WAF) blocking the crawler before the request ever touches your origin      
  server.                                                                                                                                                                                  

  Here is what is happening under the hood and how to test it.                                                                                                                             
  ──────                                                                                                                                                                                   
  ### 1. The Edge Challenge Problem                                                                                                                                                        

  Most modern web architectures sit behind Cloudflare, Fastly, or AWS CloudFront.                                                                                                          

  When you enable features like Cloudflare Super Bot Fight Mode or aggressive rate-limiting:                                                                                               

  • The edge evaluates the incoming HTTP request.                                                                                                                                          
  • OpenAI, Anthropic, and Perplexity use dynamic IP subnets that change frequently.                                                                                                       
  • If the WAF cannot instantly verify the crawler via reverse DNS or trusted ASN matching, it serves an HTTP 403 Forbidden or a Cloudflare JavaScript Challenge (Turnstile).              

  Because an automated crawler cannot solve an interactive browser challenge, it simply fails and drops the page from its index.                                                           
  ──────                                                                                                                                                                                   
  ### 2. How to Test Your Live Site with cURL                                                                                                                                              

  You do not need to wait weeks to know if you are affected. Run this command from your terminal:                                                                                          

    curl -I -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)" https://yourdomain.com                                          

  Look at the HTTP status header:                                                                                                                                                          

  • HTTP/2 200 OK: You are in the clear. The crawler can read your HTML.                                                                                                                   
  • HTTP/2 403 Forbidden: Your WAF or hosting provider is actively blocking OpenAI.                                                                                                        
  • cf-mitigated: challenge: Cloudflare is intercepting the crawler with an interactive bot challenge.                                                                                     
  ──────                                                                                                                                                                                   
  ### 3. How to Allow AI Crawlers Safely                                                                                                                                                   

  If you find that GPTBot is being blocked, do not turn off your entire WAF. Instead, create a targeted bypass rule in Cloudflare:                                                         

  1. Go to Security -&amp;gt; WAF -&amp;gt; Custom Rules.                                                                                                                                                
  2. Create a rule named Allow Verified AI Crawlers.                                                                                                                                       
  3. Set the condition:
      • (cf.client.bot and http.user_agent contains "GPTBot")
  4. Set Action to Skip:
      • Check All remaining custom rules
      • Check Super Bot Fight Mode (or Bot Management)


  This ensures legitimate OpenAI search crawlers can fetch public content while your protected admin and API endpoints stay secure.
  ──────
  ### 4. Need an Instant Sanity Check?

  If you don't have terminal access or want to check multiple bots (GPTBot, ClaudeBot, PerplexityBot, ByteSpider) in 5 seconds, we built a lightweight, no-signup checker that runs these  
  cURL validations for you:

  👉 Free AI Crawler Access Checker https://citeaura.com/crawler-check

  How are you handling AI search crawlers in your current infra? Are you whitelisting them or keeping them blocked by default?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>devops</category>
      <category>webdev</category>
      <category>cloudflare</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
