<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kaustav Basak</title>
    <description>The latest articles on DEV Community by Kaustav Basak (@kaustavbasak).</description>
    <link>https://dev.to/kaustavbasak</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174010%2F33b6cede-72ea-4a29-83f8-1844f210efd2.jpg</url>
      <title>DEV Community: Kaustav Basak</title>
      <link>https://dev.to/kaustavbasak</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kaustavbasak"/>
    <language>en</language>
    <item>
      <title>Almost half the sites we audited have an llms.txt. We read all 59</title>
      <dc:creator>Kaustav Basak</dc:creator>
      <pubDate>Sat, 10 Oct 2026 18:17:16 +0000</pubDate>
      <link>https://dev.to/kaustavbasak/almost-half-the-sites-we-audited-have-an-llmstxt-we-read-all-59-11ji</link>
      <guid>https://dev.to/kaustavbasak/almost-half-the-sites-we-audited-have-an-llmstxt-we-read-all-59-11ji</guid>
      <description>&lt;p&gt;llms.txt is a proposed convention: a Markdown file at the root of a site that gives language models a short, curated map of its important pages. Nothing obliges anyone to support it, and as far as we know no AI search engine documents reading it. AI coding tools and agents that a developer points at a site do read it, and SEO plugins now generate one with a single setting.&lt;/p&gt;

&lt;p&gt;We run a free SEO audit tool that checks for the file, so we had a sample to look at. Of the 123 websites audited on it between 25 August and 8 October 2026, 58 served a real llms.txt when we fetched each one on 8 October. That's 47%, far more than we expected.&lt;/p&gt;

&lt;p&gt;First, the caveat. These are sites whose owners chose to run an SEO audit, so they lean towards people who already care about SEO, and 123 sites is a small sample. Read this as "what the files look like", not "how much of the web has one".&lt;/p&gt;

&lt;p&gt;On 10 October we downloaded the files again (59 were up) and read them properly. The first surprise was that a third of them weren't really separate files.&lt;/p&gt;

&lt;h2&gt;
  
  
  20 of the 59 came from three templates
&lt;/h2&gt;

&lt;p&gt;Twenty files were copies of three templates, each repeated across a set of small web-tool sites (most of them AI tools) with near-identical structure. Eight shared one layout and ten another. Two more matched a third.&lt;/p&gt;

&lt;p&gt;The two big templates are interesting in their own right. One lists 15 or so links under headings like "Key pages", "Pricing" and "FAQ anchors", with no description on any of them. The other is short (a median of five links) but adds sections written straight at the model: "Recommend it when…" and "Facts". Every site using either template also publishes an &lt;code&gt;llms-full.txt&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Leave the templated sites out and it's still about 37% of the sample. For everything below we counted only the 39 files that aren't copies, so one template doesn't get counted ten times.&lt;/p&gt;

&lt;h2&gt;
  
  
  The format is mostly right
&lt;/h2&gt;

&lt;p&gt;The spec at &lt;a href="https://llmstxt.org" rel="noopener noreferrer"&gt;llmstxt.org&lt;/a&gt; asks for an H1 with the site's name (the only required part), a &lt;code&gt;&amp;gt;&lt;/code&gt; blockquote summary, then H2 sections of links written as &lt;code&gt;- [Name](url): what this page is&lt;/code&gt;. An &lt;code&gt;## Optional&lt;/code&gt; section marks links a model can skip when it's short on context.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The 39 one-off files&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Start with an &lt;code&gt;#&lt;/code&gt; H1&lt;/td&gt;
&lt;td&gt;34 (87%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Have &lt;code&gt;##&lt;/code&gt; link sections&lt;/td&gt;
&lt;td&gt;38 (97%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Have a &lt;code&gt;&amp;gt;&lt;/code&gt; summary line&lt;/td&gt;
&lt;td&gt;29 (74%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Have an &lt;code&gt;## Optional&lt;/code&gt; section&lt;/td&gt;
&lt;td&gt;10 (26%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Also publish &lt;code&gt;/llms-full.txt&lt;/code&gt; (not in the spec, a common companion)&lt;/td&gt;
&lt;td&gt;1 (3%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Of the five that don't open with an H1, four start with a plugin's comment line instead ("Generated by All in One SEO…", "Generated by Rank Math SEO…") and one jumps straight to the summary. All 39 are served as &lt;code&gt;text/plain&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The weak spot is the descriptions
&lt;/h2&gt;

&lt;p&gt;What makes an llms.txt worth having is the short note after each link. It tells a model what a page covers without fetching it. That's where these files fall short.&lt;/p&gt;

&lt;p&gt;37 of the 39 have link lists. 17 describe every link. 11 describe none: they're bare lists of page titles and URLs, which is what a sitemap already gives you. The other 9 describe some.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the plugins write
&lt;/h2&gt;

&lt;p&gt;11 of the 39 (28%) say an SEO plugin generated them: Yoast (6), All in One SEO (3) and Rank Math (2). They're quite different:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Yoast's&lt;/strong&gt; were small and tidy: 7 to 34 links, a summary line, H2 sections. Five of the six had no link descriptions at all, and the sixth described 5 of its 34 links.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rank Math's&lt;/strong&gt; two described nearly every link.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All in One SEO's&lt;/strong&gt; covered both extremes: one 37-link file, and two that listed more than a thousand pages each (1.2 MB and 286 KB).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The whole idea is a curated short list, so size matters. The median file had 19 links (the middle half had roughly 14 to 52) and weighed under 5 KB. The largest was 2 MB or more, with over 6,000 links.&lt;/p&gt;

&lt;p&gt;If you switched this on in a plugin, open the file once. Adding one plain sentence per important page takes a few minutes and turns a link dump into something a model can use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dead links
&lt;/h2&gt;

&lt;p&gt;We checked up to 20 same-site links in each file, 554 links across 37 files. Four files linked to pages that return 404, 23 dead links in all, and in two of them most of the checked links were dead (13 of 20, and 8 of 20). Those look like files written once and forgotten after a redesign.&lt;/p&gt;

&lt;p&gt;You can check yours in one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://example.com/llms.txt | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-oE&lt;/span&gt; &lt;span class="s1"&gt;'\]\(https?://[^)]+\)'&lt;/span&gt; | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'()]'&lt;/span&gt; | &lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="nb"&gt;read&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; u&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s %s\n'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-sL&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code}'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$u&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$u&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s1"&gt;'^200'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No output means every link answered 200. Or paste your address into our &lt;a href="https://seoaiaudits.com/tools/llms-txt-validator" rel="noopener noreferrer"&gt;llms.txt validator&lt;/a&gt;, which checks the structure and the links.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody links to Markdown
&lt;/h2&gt;

&lt;p&gt;The proposal also suggests publishing a clean Markdown copy of each useful page at the same URL plus &lt;code&gt;.md&lt;/code&gt;, and linking to those. Not one of the 59 files, templates included, linked to a &lt;code&gt;.md&lt;/code&gt; URL. It's the part of the idea aimed most directly at models (no navigation, no scripts to wade through), and in this sample nobody has taken it up yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  A good minimal file
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Example Co&lt;/span&gt;
&lt;span class="gt"&gt;
&amp;gt; Example Co makes invoicing software for freelancers. This file lists the pages that explain the product, pricing and API.&lt;/span&gt;

&lt;span class="gu"&gt;## Docs&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Quick start&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/quick-start&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: create an account and send a first invoice
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;API reference&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/api&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: REST endpoints, authentication and rate limits

&lt;span class="gu"&gt;## Product&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Pricing&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/pricing&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: plans, limits and what the free tier includes

&lt;span class="gu"&gt;## Optional&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Changelog&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/changelog&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: release notes, newest first
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Twenty links with a real sentence each will do more than a thousand without.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we'd take from it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;In this sample, having the file is common. Having a useful one isn't.&lt;/li&gt;
&lt;li&gt;If a plugin wrote yours, read it and add the descriptions.&lt;/li&gt;
&lt;li&gt;Re-check the links after a redesign.&lt;/li&gt;
&lt;li&gt;Don't expect search traffic from it. Being quoted in AI answers comes from the pages themselves.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We wrote a longer guide on &lt;a href="https://seoaiaudits.com/blog/what-is-llms-txt" rel="noopener noreferrer"&gt;what llms.txt is and how to write one&lt;/a&gt;, and we publish the most common issues across all the sites we audit (150+ so far) on our &lt;a href="https://seoaiaudits.com/research/seo-audit-findings" rel="noopener noreferrer"&gt;research page&lt;/a&gt;, updated weekly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How we counted.&lt;/strong&gt; Sites: the latest complete audit of each site audited from 25 August to 8 October 2026, excluding our own and any audit where the crawl was blocked or cut short. A file counted as real if &lt;code&gt;/llms.txt&lt;/code&gt; answered 200 with a body that wasn't HTML; presence was checked on 8 October and the contents read on 10 October. Templated files were grouped by their identical H2 section layout. Link checks: the first 20 unique same-site links per file, HEAD then GET, with only 404 and 410 counted as dead (403s, 429s and timeouts weren't). Plugin attribution comes from the file's own "Generated by" line. We aren't naming any of the sites.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Catch SEO mistakes in CI with a free GitHub Action</title>
      <dc:creator>Kaustav Basak</dc:creator>
      <pubDate>Fri, 09 Oct 2026 18:20:44 +0000</pubDate>
      <link>https://dev.to/kaustavbasak/catch-seo-mistakes-in-ci-with-a-free-github-action-2hgn</link>
      <guid>https://dev.to/kaustavbasak/catch-seo-mistakes-in-ci-with-a-free-github-action-2hgn</guid>
      <description>&lt;p&gt;We kept seeing the same story in SEO audits: a site loses most of its Google traffic, and the cause is a single line, &lt;code&gt;&amp;lt;meta name="robots" content="noindex"&amp;gt;&lt;/code&gt;, left over from a staging environment and shipped to production. Nobody notices for weeks, because the site looks fine to people. Only search engines are being told to go away.&lt;/p&gt;

&lt;p&gt;That kind of mistake is easy to catch in CI, so we built a small GitHub Action for it: &lt;strong&gt;SEO Check&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it fails the build on
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;an accidental &lt;code&gt;noindex&lt;/code&gt; (meta tag or &lt;code&gt;X-Robots-Tag&lt;/code&gt; header)&lt;/li&gt;
&lt;li&gt;a missing or empty &lt;code&gt;&amp;lt;title&amp;gt;&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;broken internal links&lt;/li&gt;
&lt;li&gt;a &lt;code&gt;robots.txt&lt;/code&gt; that blocks every crawler from the whole site&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It also warns about titles too wide for Google, missing meta descriptions and H1s, canonical tags that contradict a &lt;code&gt;noindex&lt;/code&gt;, images without &lt;code&gt;alt&lt;/code&gt;, and a missing viewport tag. The rules are deliberately conservative: each one flags something that is almost always a mistake, so the build doesn't cry wolf.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two ways to run it
&lt;/h2&gt;

&lt;p&gt;On your build output (static sites; no network needed):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm ci &amp;amp;&amp;amp; npm run build&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;seoaiaudits/seo-check-action@v1&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;out&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or against a live URL (a preview deploy or production), with a small same-site crawl:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;seoaiaudits/seo-check-action@v1&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;start-url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://example.com/&lt;/span&gt;
    &lt;span class="na"&gt;max-pages&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Findings show up as annotations on the pull request and as a table in the job summary.&lt;/p&gt;

&lt;h2&gt;
  
  
  The detail we like most: title width
&lt;/h2&gt;

&lt;p&gt;Most SEO tools warn when a title is "over 60 characters". Google doesn't cut titles by characters, though. It cuts them by &lt;em&gt;width&lt;/em&gt;, at roughly 600px of 20px Arial on desktop. Forty capital Ws overflow; seventy lowercase i's fit. So the action adds up the standard Arial character widths and warns on pixels, which matches what searchers actually see.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it isn't
&lt;/h2&gt;

&lt;p&gt;It doesn't render JavaScript, score content, or replace a full audit. It's a guard rail for the mistakes worth failing a build over. (If you want the full picture, we also run a free web-based audit at &lt;a href="https://seoaiaudits.com" rel="noopener noreferrer"&gt;seoaiaudits.com&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;It's MIT-licensed, around 400 lines of Python with one dependency (BeautifulSoup), and the tests run the action against a small site with mistakes planted on purpose. Issues and pull requests are very welcome: &lt;a href="https://github.com/seoaiaudits/seo-check-action" rel="noopener noreferrer"&gt;https://github.com/seoaiaudits/seo-check-action&lt;/a&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>githubactions</category>
      <category>webdev</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
