<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: CiteScore</title>
    <description>The latest articles on DEV Community by CiteScore (@citescore).</description>
    <link>https://dev.to/citescore</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4175459%2F3ad2cce1-b06a-4151-882d-2de2bbb1b81d.png</url>
      <title>DEV Community: CiteScore</title>
      <link>https://dev.to/citescore</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/citescore"/>
    <language>en</language>
    <item>
      <title>How to write an llms.txt file (with a free generator)</title>
      <dc:creator>CiteScore</dc:creator>
      <pubDate>Sat, 10 Oct 2026 17:37:05 +0000</pubDate>
      <link>https://dev.to/citescore/how-to-write-an-llmstxt-file-with-a-free-generator-1f04</link>
      <guid>https://dev.to/citescore/how-to-write-an-llmstxt-file-with-a-free-generator-1f04</guid>
      <description>&lt;p&gt;Large language models increasingly read websites to answer questions, but a typical page is full of navigation menus, scripts, cookie banners, and footers. An llms.txt file is a short Markdown file that gives those models a clean map of your site: what it is, and which pages matter most. This guide covers what the file is, where it goes, how to format it, what to avoid, and how it differs from robots.txt.&lt;/p&gt;

&lt;h2&gt;
  
  
  What llms.txt is
&lt;/h2&gt;

&lt;p&gt;llms.txt is a proposed convention, described at llmstxt.org since 2024. It is not an official standard from a standards body, and no AI provider has promised to read it in every case. Treat it as low-cost documentation for tools that choose to use it. Its main appeal is that it is easy to write and easy to maintain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it goes
&lt;/h2&gt;

&lt;p&gt;Place the file at the root of your domain, so it is available at &lt;a href="https://example.com/llms.txt" rel="noopener noreferrer"&gt;https://example.com/llms.txt&lt;/a&gt;. A file in a subdirectory such as /docs/llms.txt will not be found by convention. Serve it as plain text with a 200 status code, and make sure it is not behind a login or a redirect chain.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Markdown format
&lt;/h2&gt;

&lt;p&gt;The format is deliberately simple. It has four parts, in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;An H1 with the name of the site or project.&lt;/strong&gt; This is the only required element.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A blockquote summary.&lt;/strong&gt; One or two sentences that say what the site is and who it is for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optional paragraphs of context.&lt;/strong&gt; Plain text that explains anything the links cannot, such as scope or how the documentation is organized.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;H2 sections containing link lists.&lt;/strong&gt; Each entry is a Markdown link, a colon, and a short description.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The skeleton looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Site Name&lt;/span&gt;
&lt;span class="gt"&gt;
&amp;gt; One or two sentences that say what the site is and who it is for.&lt;/span&gt;

Optional context paragraph.

&lt;span class="gu"&gt;## Section name&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Page title&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/page&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Short description of what the page covers.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  A complete example
&lt;/h2&gt;

&lt;p&gt;Here is a file for a fictional note-taking app called Quillpad:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Quillpad&lt;/span&gt;
&lt;span class="gt"&gt;
&amp;gt; Quillpad is a note-taking app for teams that stores notes as plain Markdown and syncs them across devices. Setup guides, API docs, and pricing are linked below.&lt;/span&gt;

Quillpad saves every note as a standard .md file. The API is versioned, and the current version is v2.

&lt;span class="gu"&gt;## Documentation&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Getting started&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/getting-started&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Install Quillpad and create your first workspace.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Markdown syntax&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/markdown&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Supported formatting, including tables and task lists.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;API reference&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/api&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Endpoints, authentication, and rate limits for v2.

&lt;span class="gu"&gt;## Product&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Pricing&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/pricing&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Plans, seat limits, and what each tier includes.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Security&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/security&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Encryption, data residency, and audit reports.

&lt;span class="gu"&gt;## Optional&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Changelog&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/changelog&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Release notes by date.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Careers&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/careers&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Open roles.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Optional section has a specific meaning in the convention. It marks secondary content that a tool can skip when it has limited context space.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A missing or wrong H1.&lt;/strong&gt; The H1 anchors the file. Without it, the file does not follow the format.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Links with no descriptions.&lt;/strong&gt; A bare list of URLs tells a model very little. Each link needs a sentence about what is there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Serving HTML instead of Markdown.&lt;/strong&gt; A catch-all route can return your HTML error page at /llms.txt. Check the response with curl or your browser's network tab.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dead or redirecting links.&lt;/strong&gt; Use final URLs, and check them whenever you restructure the site.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Listing everything.&lt;/strong&gt; A 2,000-line file is less useful than a focused 30-line one. Prioritize the pages you most want people and tools to find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Marketing copy in the summary.&lt;/strong&gt; The blockquote should describe the site, not sell it. Phrases like "the best" or "revolutionary" add nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Letting it go stale.&lt;/strong&gt; A file that points to removed pages is worse than no file at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blocking it.&lt;/strong&gt; Do not disallow /llms.txt in robots.txt, and do not require authentication to read it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How it differs from robots.txt
&lt;/h2&gt;

&lt;p&gt;These two files do different jobs, and most sites benefit from having both.&lt;/p&gt;

&lt;p&gt;robots.txt tells crawlers what they may request. It holds allow and disallow rules for each user agent and follows the Robots Exclusion Protocol (RFC 9309). It controls crawling, and compliance is voluntary, though major crawlers generally honor it.&lt;/p&gt;

&lt;p&gt;llms.txt does not restrict access. It is a curated map that points a model to the content you consider most useful and briefly explains each part. If you want to keep a crawler away from pages, use robots.txt or real access controls. If you want to point a model toward your best documentation, write an llms.txt.&lt;/p&gt;

&lt;h2&gt;
  
  
  A short checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Put the file at /llms.txt on your root domain.&lt;/li&gt;
&lt;li&gt;Start with an H1 and a one- or two-sentence blockquote.&lt;/li&gt;
&lt;li&gt;Use H2 sections with full URLs and one-line descriptions.&lt;/li&gt;
&lt;li&gt;Move secondary material under "Optional."&lt;/li&gt;
&lt;li&gt;Confirm that every link returns a 200 status.&lt;/li&gt;
&lt;li&gt;Review the file whenever you make major changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you would rather not write the file by hand, the free &lt;a href="https://citescore.vercel.app/llms-txt-generator" rel="noopener noreferrer"&gt;llms.txt generator&lt;/a&gt; can produce a starting draft. Edit it, then check the result against the rules above before you publish.&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>ai</category>
      <category>webdev</category>
      <category>seo</category>
    </item>
    <item>
      <title>Which AI crawlers should you allow in robots.txt? (GPTBot, PerplexityBot, Google-Extended and more)</title>
      <dc:creator>CiteScore</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:44:15 +0000</pubDate>
      <link>https://dev.to/citescore/which-ai-crawlers-should-you-allow-in-robotstxt-gptbot-perplexitybot-google-extended-and-more-5gdk</link>
      <guid>https://dev.to/citescore/which-ai-crawlers-should-you-allow-in-robotstxt-gptbot-perplexitybot-google-extended-and-more-5gdk</guid>
      <description>&lt;p&gt;AI assistants find and cite websites in two different ways. Some crawlers collect pages to train future models. Others fetch pages so an assistant can search the web and link to your site in an answer. Your robots.txt file can treat these differently, but only if you know what each user-agent does.&lt;/p&gt;

&lt;p&gt;One caveat first: robots.txt is a request, not a lock. Reputable crawlers honor it. Anything else needs server-level or firewall blocking.&lt;/p&gt;

&lt;h2&gt;
  
  
  The major AI crawlers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;User-agent&lt;/th&gt;
&lt;th&gt;Owner&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPTBot&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;Collects content for training OpenAI models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OAI-SearchBot&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;Indexes pages for ChatGPT search&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT-User&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;Fetches a page when a ChatGPT user asks for it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ClaudeBot&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Collects content for training Anthropic models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude-SearchBot&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Indexes pages for Claude's search features&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude-User&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;Fetches a page when a Claude user asks for it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PerplexityBot&lt;/td&gt;
&lt;td&gt;Perplexity&lt;/td&gt;
&lt;td&gt;Indexes pages for Perplexity answers and citations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Perplexity-User&lt;/td&gt;
&lt;td&gt;Perplexity&lt;/td&gt;
&lt;td&gt;Fetches a page when a Perplexity user asks for it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CCBot&lt;/td&gt;
&lt;td&gt;Common Crawl&lt;/td&gt;
&lt;td&gt;Builds a public web archive that many AI teams use for training&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google-Extended&lt;/td&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;A control token, not a crawler. Governs use of content for Gemini training and grounding&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Applebot-Extended&lt;/td&gt;
&lt;td&gt;Apple&lt;/td&gt;
&lt;td&gt;A control token. Governs use of Applebot-crawled content for Apple's AI training&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Google-Extended and Applebot-Extended are not separate crawlers. They are opt-out tokens that apply to content already crawled by Googlebot or Applebot. Blocking them does not remove you from Google Search, and AI Overviews follow Googlebot and standard directives such as &lt;code&gt;nosnippet&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  A recommended robots.txt block
&lt;/h2&gt;

&lt;p&gt;If you have no AI rules today, your site is already open to all of these bots. The block below opts out of training while staying visible in AI search:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="c"&gt;# Training crawlers: opt out
&lt;/span&gt;&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;GPTBot&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;ClaudeBot&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;CCBot&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Google&lt;/span&gt;-&lt;span class="n"&gt;Extended&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Applebot&lt;/span&gt;-&lt;span class="n"&gt;Extended&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="c"&gt;# Search and answer crawlers: allow
&lt;/span&gt;&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;OAI&lt;/span&gt;-&lt;span class="n"&gt;SearchBot&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;ChatGPT&lt;/span&gt;-&lt;span class="n"&gt;User&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Claude&lt;/span&gt;-&lt;span class="n"&gt;SearchBot&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Claude&lt;/span&gt;-&lt;span class="n"&gt;User&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;PerplexityBot&lt;/span&gt;
&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Perplexity&lt;/span&gt;-&lt;span class="n"&gt;User&lt;/span&gt;
&lt;span class="n"&gt;Allow&lt;/span&gt;: /
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep your existing rules for Googlebot and &lt;code&gt;*&lt;/code&gt; in their own groups.&lt;/p&gt;

&lt;h2&gt;
  
  
  Training vs. answer crawlers: the trade-off
&lt;/h2&gt;

&lt;p&gt;Training crawlers feed future model versions. Blocking them signals that you do not want your content used for future training. It does not remove anything a model has already learned, and it is hard to measure its effect on your traffic.&lt;/p&gt;

&lt;p&gt;Answer crawlers decide whether you appear in live answers. If you block PerplexityBot or OAI-SearchBot, your pages are less likely to be cited when people ask a relevant question. For most sites that want visibility, that is the bigger cost.&lt;/p&gt;

&lt;p&gt;The block above is a reasonable default. Adjust it if you have licensing concerns, paywalled content or private material. Vendors also note that user-triggered fetchers such as ChatGPT-User may not apply robots.txt rules, so treat those rules as best effort.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to verify your rules
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Open &lt;code&gt;https://yoursite.com/robots.txt&lt;/code&gt; in a browser. It should return 200 and plain text. A 404 means no rules, so everything is allowed. A 5xx error can stop crawlers entirely.&lt;/li&gt;
&lt;li&gt;Test the rules with Python's standard library:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;urllib.robotparser&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;RobotFileParser&lt;/span&gt;

&lt;span class="n"&gt;rp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RobotFileParser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://yoursite.com/robots.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;rp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;bot&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GPTBot&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OAI-SearchBot&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PerplexityBot&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;can_fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bot&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://yoursite.com/blog/post&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Search your server logs or analytics for these user-agent strings. Confirm that the bots you allowed reach your pages, and that the bots you blocked request only &lt;code&gt;robots.txt&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Check your CDN or WAF. A bot-protection rule that returns 403 blocks crawlers regardless of robots.txt, and it can block the bots you meant to allow.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Python's parser does not handle every wildcard the way Google does, so test important paths with more than one tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Blocking Googlebot by accident.&lt;/strong&gt; Google-Extended is not Googlebot. Blocking Googlebot removes you from Google Search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assuming a &lt;code&gt;*&lt;/code&gt; rule covers AI bots.&lt;/strong&gt; A crawler follows the most specific group that matches its name and ignores &lt;code&gt;*&lt;/code&gt;. If GPTBot has its own group, that group must contain every rule you want for GPTBot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Using robots.txt for privacy.&lt;/strong&gt; The file is public and it is only a request. Protect private pages with authentication.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mixing up Disallow and Allow.&lt;/strong&gt; &lt;code&gt;Disallow: /&lt;/code&gt; blocks the whole site, while an empty &lt;code&gt;Disallow:&lt;/code&gt; allows everything. A misplaced slash can hide your entire site from search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ignoring the CDN.&lt;/strong&gt; Bot-protection rules can override what your robots.txt says.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expecting instant results.&lt;/strong&gt; Crawlers cache robots.txt, and removal from AI answers takes time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Test your robots.txt AI-crawler rules for free at &lt;a href="https://citescore.vercel.app/free-check" rel="noopener noreferrer"&gt;https://citescore.vercel.app/free-check&lt;/a&gt;&lt;/p&gt;

</description>
      <category>beginners</category>
      <category>seo</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How to write an llms.txt file (with a copy-paste template)</title>
      <dc:creator>CiteScore</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:42:06 +0000</pubDate>
      <link>https://dev.to/citescore/how-to-write-an-llmstxt-file-with-a-copy-paste-template-56fh</link>
      <guid>https://dev.to/citescore/how-to-write-an-llmstxt-file-with-a-copy-paste-template-56fh</guid>
      <description>&lt;h1&gt;
  
  
  How to write an llms.txt file (with a copy-paste template)
&lt;/h1&gt;

&lt;p&gt;If you want AI assistants and large language models to understand your site accurately, you can give them a short, curated map. That map is an llms.txt file. This guide explains what it is, where it goes, how to format it, and how to test it before you rely on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What llms.txt is
&lt;/h2&gt;

&lt;p&gt;llms.txt is a proposed standard, introduced by Jeremy Howard in 2024 and documented at llmstxt.org. It gives language models a plain Markdown file that summarises your website: what it is, what it offers, and which pages matter most.&lt;/p&gt;

&lt;p&gt;It works alongside two files you probably already have. robots.txt tells crawlers what they may access. A sitemap.xml lists every URL. llms.txt does something different: it selects the most important content and explains it in a format a model can read within a limited context window.&lt;/p&gt;

&lt;p&gt;Be realistic about adoption. The standard is still informal, and not every AI company has confirmed that it reads llms.txt. Treat it as low-cost documentation that helps the tools that do use it, not as a guaranteed ranking lever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it goes
&lt;/h2&gt;

&lt;p&gt;The file must live at the root of your domain:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;https://example.com/llms.txt&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Serve it with a 200 status as plain text or Markdown. It should not redirect to a login page or depend on JavaScript to render. If you have subdomains such as &lt;code&gt;docs.example.com&lt;/code&gt; and want them covered, publish a separate llms.txt on each one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The format
&lt;/h2&gt;

&lt;p&gt;The structure is deliberately simple. Include these parts, in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;An H1 title.&lt;/strong&gt; The name of the project or site. This is the only required element.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A blockquote summary.&lt;/strong&gt; One or two sentences that explain what the site is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optional paragraphs.&lt;/strong&gt; Extra context in plain prose, such as who the site is for or what is out of scope.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;H2 sections.&lt;/strong&gt; Groups of links, each written as a Markdown list item: &lt;code&gt;- [Link title](URL): short note&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An Optional section.&lt;/strong&gt; Links a model can skip when context is tight.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Use absolute URLs and point them at clean, public pages that load directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Full copy-paste template
&lt;/h2&gt;

&lt;p&gt;Replace the placeholders and delete any section you do not need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Example Co&lt;/span&gt;
&lt;span class="gt"&gt;
&amp;gt; Example Co makes invoicing software for freelancers. The product covers quotes, recurring billing and tax exports for the US and UK.&lt;/span&gt;

Pricing changes from time to time, so check the pricing page for current numbers. Documentation covers the public REST API and the Zapier integration.

&lt;span class="gu"&gt;## Docs&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Getting started&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/getting-started&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Create an account and send your first invoice
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;API reference&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/api&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Endpoints, authentication and rate limits
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Webhooks&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/webhooks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Event types and payload examples

&lt;span class="gu"&gt;## Product&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Features&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/features&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Full list of invoicing and tax features
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Pricing&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/pricing&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Plans and per-seat costs

&lt;span class="gu"&gt;## Optional&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Blog&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/blog&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Product updates and tutorials
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Careers&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/careers&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Open roles
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Common mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Putting the file in the wrong place.&lt;/strong&gt; &lt;code&gt;/docs/llms.txt&lt;/code&gt; is not the root. Crawlers look for &lt;code&gt;/llms.txt&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dead or redirected links.&lt;/strong&gt; Every URL should return 200. A list full of 404s signals an unmaintained file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Listing everything.&lt;/strong&gt; Pasting 300 URLs defeats the purpose. Pick the 10 to 30 pages that answer the most common questions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vague descriptions.&lt;/strong&gt; "Our amazing solutions" tells a model nothing. Say what the page actually contains.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Marketing copy in the summary.&lt;/strong&gt; Keep the blockquote factual. Plain statements are easier to use than slogans.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forgetting the H1.&lt;/strong&gt; Without a title the file does not follow the format, and parsers may fail on it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blocking the file.&lt;/strong&gt; If a firewall, CDN rule or robots.txt rule blocks &lt;code&gt;/llms.txt&lt;/code&gt;, nothing else you write matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Letting it go stale.&lt;/strong&gt; Update the file when you launch, rename or remove a major page.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to test it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check the status and redirects.&lt;/strong&gt; Run &lt;code&gt;curl -I https://example.com/llms.txt&lt;/code&gt; and confirm a 200 response with no redirect chain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check the content type.&lt;/strong&gt; It should be &lt;code&gt;text/plain&lt;/code&gt; or &lt;code&gt;text/markdown&lt;/code&gt;, not &lt;code&gt;text/html&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read it as a stranger would.&lt;/strong&gt; Open the file and ask whether someone who has never seen the site can tell what it offers from the first 20 lines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check every link.&lt;/strong&gt; Extract the URLs from the file and request each one with curl. Flag anything that is not a 200.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confirm the structure.&lt;/strong&gt; The first line should be a single H1, followed by a blockquote, then H2 sections that contain list items.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check robots.txt.&lt;/strong&gt; Make sure the pages you link are not disallowed for the AI user agents you want to reach.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Repeat these checks after every deploy that touches the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;A good llms.txt is short, accurate and maintained. Start with ten strong links, a factual summary and a clean root URL, then expand as your documentation grows.&lt;/p&gt;

&lt;p&gt;You can check your site's AI-crawler setup for free at &lt;a href="https://citescore.vercel.app/free-check" rel="noopener noreferrer"&gt;https://citescore.vercel.app/free-check&lt;/a&gt;&lt;/p&gt;

</description>
      <category>tutorial</category>
      <category>seo</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
