<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Lucas Fernández Díaz</title>
    <description>The latest articles on DEV Community by Lucas Fernández Díaz (@lucasfernandezdiaz).</description>
    <link>https://dev.to/lucasfernandezdiaz</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4128336%2F5d2bc3aa-dc65-452e-88f9-76c2bed1461e.png</url>
      <title>DEV Community: Lucas Fernández Díaz</title>
      <link>https://dev.to/lucasfernandezdiaz</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lucasfernandezdiaz"/>
    <language>en</language>
    <item>
      <title>We Tested 56 Therapy Websites for AI Access</title>
      <dc:creator>Lucas Fernández Díaz</dc:creator>
      <pubDate>Wed, 07 Oct 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/lucasfernandezdiaz/we-tested-56-therapy-websites-for-ai-access-2mlk</link>
      <guid>https://dev.to/lucasfernandezdiaz/we-tested-56-therapy-websites-for-ai-access-2mlk</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://contento.solutions/academy/therapy-websites-ai-access" rel="noopener noreferrer"&gt;Contento Solutions&lt;/a&gt;, where I write about SEO and AI visibility for service businesses.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We tested 12 AI crawlers against 56 private practice therapy websites. Eight refuse the crawlers behind ChatGPT, Claude or Perplexity, and seven of those eight never wrote that rule: a server or a CDN did. The bigger number, two thirds of the sample, we threw away, and this explains why.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number we threw away, and why
&lt;/h2&gt;

&lt;p&gt;The first pass said 37 of 56 sites, two thirds of the sample, block at least one AI crawler. That headline was available, it was technically true, and it was useless.&lt;/p&gt;

&lt;p&gt;Thirty five of those 37 cases were the same crawler: a bulk collector that many hosts block by default because it requests aggressively. Blocking it is a decision about server load. It is not a decision about whether a person asking an assistant for a therapist will ever see your practice.&lt;/p&gt;

&lt;p&gt;So we split the twelve crawlers into two groups and asked two different questions. The first group is the assistants people actually use: the crawlers behind ChatGPT, Claude, Perplexity and Google's AI surfaces. Blocking those has a consequence you can describe in one sentence. The second group collects at scale for training. Blocking those is defensible and says nothing about visibility.&lt;/p&gt;

&lt;p&gt;Everything below counts the first group only.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two layers, because one of them is invisible
&lt;/h2&gt;

&lt;p&gt;Two places decide whether a crawler can read your site, and only one of them is public.&lt;/p&gt;

&lt;p&gt;The first is robots.txt, a text file at the root of your domain that names which crawlers may request which paths. Anyone can read it, including you, right now. It is where every audit stops.&lt;/p&gt;

&lt;p&gt;The second is the response itself. A site can hand a page to a browser and refuse the identical page to a named crawler, with a 403, at the edge, before any of your website's own code runs. That block appears nowhere in robots.txt. It arrives from a CDN setting, a security plugin, or a managed rule somebody enabled once and never revisited.&lt;/p&gt;

&lt;p&gt;We checked both layers for each of the 12 crawlers: the longest matching rule in robots.txt, and then a real request carrying that crawler's user agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two controls, because without them the test accuses innocent sites
&lt;/h2&gt;

&lt;p&gt;The first control is a generic bot. Before recording a block, we requested the same page with a bot user agent that has nothing to do with AI. A site that refuses every non browser is not making a decision about AI, it is making a decision about bots, and it cannot count. Two sites failed this control, so we dropped them from the sample entirely.&lt;/p&gt;

&lt;p&gt;The second control is patience. We treated every 429 as our own rate limit rather than as their policy: we waited, then asked again more slowly. Five sites produced a 429 at some point, and we counted none of them as blocking.&lt;/p&gt;

&lt;p&gt;That is why the denominator below is 56 and not 58.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we found
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assistant crawler&lt;/th&gt;
&lt;th&gt;Sites that refuse it&lt;/th&gt;
&lt;th&gt;Of&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ClaudeBot&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPTBot&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OAI-SearchBot&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT-User&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PerplexityBot&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude-SearchBot&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Perplexity-User&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google-Extended&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Eight distinct practices out of 56 close the door to at least one assistant crawler. That is roughly one in seven.&lt;/p&gt;

&lt;p&gt;The shape of those eight matters more than the count. Five of them refuse seven of the eight assistant crawlers at once, which is not what a considered policy looks like. A practice that decided to keep ChatGPT out would write one line. Refusing almost every assistant simultaneously, with nothing in robots.txt, is the signature of a single switch somewhere in front of the website.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Seven of the eight never declared it.&lt;/strong&gt; Only one site has the rule written in its own robots.txt, where the owner could find it. In the other seven the block is invisible from inside the website.&lt;/p&gt;

&lt;h2&gt;
  
  
  The contrast nobody expects
&lt;/h2&gt;

&lt;p&gt;Fourteen of the 56 practices publish an llms.txt, the newer file meant to help assistants read a site. That is almost twice as many sites doing the optional new thing as sites accidentally locked out of the assistants entirely.&lt;/p&gt;

&lt;p&gt;Read that in order. Fourteen practices went looking for the current best practice and added a file. Eight have a door closed that they did not close and cannot see. The work went to the part that is visible in a blog post, not to the part that decides whether a crawler gets a page at all.&lt;/p&gt;

&lt;p&gt;One more measurement, because it corrects something we expected to find: not a single site in the sample uses the noai meta tag. Whatever is happening here, almost nobody chose it deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to check your own, in four minutes
&lt;/h2&gt;

&lt;p&gt;None of this needs a tool or a consultant.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, open your robots.txt
&lt;/h2&gt;

&lt;p&gt;Type your domain followed by /robots.txt in a browser. If you see rules naming GPTBot, ClaudeBot, PerplexityBot or Google-Extended with Disallow, that is a declared block. If the file is empty or missing, that layer is clear, which is not the same as open.&lt;/p&gt;

&lt;h2&gt;
  
  
  Second, look at what sits in front of your site
&lt;/h2&gt;

&lt;p&gt;Most invisible blocks come from a CDN or a security layer, not from your website builder. If your domain runs through a service that filters traffic, find its bot settings. Rules that promise to stop scrapers or AI bots do exactly that, including the assistants you want.&lt;/p&gt;

&lt;h2&gt;
  
  
  Third, ask an assistant for your own page
&lt;/h2&gt;

&lt;p&gt;Give ChatGPT or Claude the URL of your practice page and ask it to tell you what the page says. If it cannot fetch it while the page loads fine in your browser, you have the second layer problem, and you just proved it in thirty seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fourth, ask your host the one useful question
&lt;/h2&gt;

&lt;p&gt;The question is not whether they block AI. It is: does anything in front of my site return 403 to a named crawler? A support agent can answer that in a sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does not prove, and it matters
&lt;/h2&gt;

&lt;p&gt;Being readable is not being cited. Forty eight of these 56 practices can be fetched perfectly well by every assistant crawler we tested, and most of them still will not come up when someone asks for a therapist in their city. Access is the floor, not the outcome. A crawler that can read a page nobody would ever quote has changed nothing.&lt;/p&gt;

&lt;p&gt;What the measurement does prove is narrower and worth knowing: &lt;strong&gt;for about one in seven practices, the work of being chosen has not started, because the page cannot be read at all.&lt;/strong&gt; And in seven of those eight cases, nobody made that decision. It was made for them, by something they are paying for.&lt;/p&gt;

&lt;p&gt;If you want the rest of the picture, the part that decides whether an assistant quotes you once it can read you, that is what &lt;a href="https://contento.solutions/method" rel="noopener noreferrer"&gt;our method&lt;/a&gt; measures, and &lt;a href="https://contento.solutions/academy/make-website-citable-ai-search" rel="noopener noreferrer"&gt;how to make a page worth citing&lt;/a&gt; is the practical version. The file everyone is adding is covered in &lt;a href="https://contento.solutions/academy/llms-txt-what-to-put-in-it" rel="noopener noreferrer"&gt;what to actually put in llms.txt&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;We did not name any of the 56 practices. They did not ask to be measured, and seven of the eight blocked sites are not doing anything wrong: they are the ones with the least visibility into their own setup. Naming them would punish the people this piece is meant to help.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions we get about this
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does blocking an AI crawler hurt my Google ranking?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No. These are separate crawlers from the one that indexes your site for classic search results. Blocking GPTBot does nothing to your position in Google's blue links. What it changes is whether an assistant can read your page when someone asks it a question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My robots.txt is empty, so I am fine, right?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Not necessarily. In seven of the eight blocked practices we found, robots.txt was clean and the refusal came from a layer in front of the website. An empty robots.txt means you did not write a block. It does not mean nobody else did.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should I add an llms.txt file?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
It does not hurt, and fourteen of these 56 practices already have one. Just be aware of the order: a file that tells an assistant what your site contains is useless if the assistant cannot fetch your pages in the first place. Check the door before you add the sign.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is being readable enough to get recommended?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No, and this is the honest limit of everything above. Forty eight of these 56 practices are perfectly readable and most still will not be named when someone asks for a therapist in their city. Access is the floor, not the outcome.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why did you not name the practices you tested?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
They did not ask to be measured, and seven of the eight blocked sites are not doing anything wrong. They are the ones with the least visibility into their own setup, and naming them would punish exactly the people the piece is meant to help.&lt;/p&gt;

&lt;p&gt;If you want to know which of the two problems your own site has, the door or the citation, that is a reasonable thing to bring to a call. What the practice does and what gets reported is on the &lt;a href="https://contento.solutions/services" rel="noopener noreferrer"&gt;services page&lt;/a&gt;, with more detail on the &lt;a href="https://contento.solutions/faq" rel="noopener noreferrer"&gt;FAQ page&lt;/a&gt;. The call is &lt;a href="https://contento.solutions/book-a-call" rel="noopener noreferrer"&gt;thirty minutes&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Have you ever found a block on your own site that nobody on the team wrote? Tell me in the comments where it turned out to live: the CDN, a security plugin or the host. I would like to know which layer shows up most often outside this niche.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>ai</category>
      <category>webdev</category>
      <category>datascience</category>
    </item>
    <item>
      <title>How I run a one-person company's website like a production system</title>
      <dc:creator>Lucas Fernández Díaz</dc:creator>
      <pubDate>Tue, 06 Oct 2026 13:02:20 +0000</pubDate>
      <link>https://dev.to/lucasfernandezdiaz/how-i-run-a-one-person-companys-website-like-a-production-system-2j0l</link>
      <guid>https://dev.to/lucasfernandezdiaz/how-i-run-a-one-person-companys-website-like-a-production-system-2j0l</guid>
      <description>&lt;p&gt;&lt;em&gt;A version of this story first appeared on &lt;a href="https://contento.solutions/blog/how-i-built-contento-as-a-system" rel="noopener noreferrer"&gt;Contento Solutions&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;My company's website runs on Squarespace. It also has 33 checks that run against it every morning, 118 scripts, a git history of 684 commits since September 1, two independent backups a day and a decision log with more than 900 rows.&lt;/p&gt;

&lt;p&gt;I'm the only person in the company. That's exactly why it's built this way.&lt;/p&gt;

&lt;p&gt;This is the architecture, what each part is for, and the two places where the platform said no and I had to find another way in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why treat a small site like production?
&lt;/h2&gt;

&lt;p&gt;Because nobody else is going to notice when it breaks. On a team, someone sees the broken button on their phone. When you're alone, a regression can sit on the site for weeks. The fix isn't paying more attention. It's building something that pays attention for you.&lt;/p&gt;

&lt;p&gt;The site has 41 public pages: an Academy with 11 guides, a Blog with 7 posts, six landing pages for specific fields, plus the usual home, services, pricing and so on. Every page carries custom code. That's a lot of surface for one person to eyeball.&lt;/p&gt;

&lt;h2&gt;
  
  
  The front: Squarespace plus custom code
&lt;/h2&gt;

&lt;p&gt;Squarespace handles hosting, the CMS and the editor. Everything that makes the site look like itself is mine: custom HTML blocks and per-page code injections, all wrapped in one root class so the CSS is scoped and never leaks into the platform's own styles. One palette, three typefaces, the same button sizes everywhere.&lt;/p&gt;

&lt;p&gt;Writing to Squarespace is where it got interesting. There's no public API for page content, so I watched what the editor's own Save button sends and replayed those same requests from a logged-in session. Two rules came out of that:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Read the object, change one key, send the whole thing back.&lt;/strong&gt; Rebuilding a payload by hand, or cloning a section with fresh IDs, gets you a 400 with no error body. The same request on the object you just read gets a 200.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The error message is the documentation.&lt;/strong&gt; When the layout grid rejects a change, the 400 tells you the exact limit it wanted. I raised the values until it stopped complaining.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The code injection editor was the other wall. Setting its value programmatically looks perfect: the editor updates, the Save button lights up, you save. Reload the page and the change is gone. It only persists what arrives as real input. So the process is: put the cursor where it belongs, enter the text as a person would, and never call anything saved until a cold reload outside the editor proves it.&lt;/p&gt;

&lt;h2&gt;
  
  
  site.json: one source for every script
&lt;/h2&gt;

&lt;p&gt;The scripts don't keep their own lists. Everything structural lives in one machine-readable file, &lt;code&gt;site.json&lt;/code&gt;: the list of public URLs, the site IDs, every surface where custom code lives (with a SHA-256 of its last known content), and the baseline value for every gate.&lt;/p&gt;

&lt;p&gt;That rule exists because of a bug. At one point the list of URLs had been copied into six different scripts, and one gate spent several sprints checking 29 of the 33 pages that existed then. Nothing failed. It just wasn't looking. Now there's one list, and every script imports it.&lt;/p&gt;

&lt;p&gt;The baselines get rewritten only on purpose, with an explicit update mode, when a change is intended. Never as a side effect of a run.&lt;/p&gt;

&lt;h2&gt;
  
  
  33 checks before coffee
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ct-gates.py&lt;/code&gt; runs all the gates against the public URLs, anonymously, with no session. Each check compares a number against its baseline in &lt;code&gt;site.json&lt;/code&gt; and prints PASS or FAIL. The exit code is 0 if everything passes, 1 if something fails, 2 if the network did.&lt;/p&gt;

&lt;p&gt;The daily run reads the served HTML. Its 33 checks are grouped into families:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Family&lt;/th&gt;
&lt;th&gt;What it catches&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;HTTP and sitemap&lt;/td&gt;
&lt;td&gt;URLs that don't answer, public pages missing from the list&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structured data&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ld+json&lt;/code&gt; blocks that don't parse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Links&lt;/td&gt;
&lt;td&gt;broken links, orphan pages that aren't orphans on purpose&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typography&lt;/td&gt;
&lt;td&gt;a banned font sneaking back in, the wrong family on eyebrows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Page weight&lt;/td&gt;
&lt;td&gt;pages outside their normalized size band&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Footer and CTAs&lt;/td&gt;
&lt;td&gt;missing global footer, drift in call to action labels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FAQ&lt;/td&gt;
&lt;td&gt;questions on a page without matching FAQ schema&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llms.txt&lt;/td&gt;
&lt;td&gt;site pieces missing from the file AI assistants read&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Once a week, a heavier run renders the pages in headless Chrome and checks what HTML can't tell you: computed visibility after JavaScript, real keyboard focus order through the DevTools protocol, and alignment at three screen widths.&lt;/p&gt;

&lt;p&gt;Most of these families exist because something got through once. A banned font lived on the site for months in 183 declarations until I spotted it by accident on a form. Now it's a gate, and it can't come back without a FAIL.&lt;/p&gt;

&lt;h2&gt;
  
  
  The watcher: did something change that I didn't touch?
&lt;/h2&gt;

&lt;p&gt;A gate tells you when a number leaves its band. It doesn't tell you when a number moves inside the band, or when someone updated the baseline without looking. So a second script, the watcher, runs the gates on a schedule and compares every result against the &lt;strong&gt;previous run&lt;/strong&gt;, not only against the baseline. It also records the git HEAD of each run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;number changed and HEAD changed: I did that. Report it, don't alarm.&lt;/li&gt;
&lt;li&gt;number changed and HEAD is the same: it moved on its own. That's the one to look at.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That logic caught the strangest bug on the site. One page appeared to double in weight in a few days, with nothing published. I saved both versions and diffed them. The content was identical byte for byte. The difference was runs of blank space the platform injects depending on which cached copy you get, more than half the page. Raw page weight on Squarespace turned out to be useless, so the weight gate now measures the page with whitespace collapsed.&lt;/p&gt;

&lt;p&gt;The watcher runs from Windows Task Scheduler at 09:00 every day, with the full rendered run on Sundays. The runner holds a lock so two scheduled tasks firing at once after a cold boot don't overwrite each other's log, and it retries once after 90 seconds, because the first request after waking a laptop fails often.&lt;/p&gt;

&lt;p&gt;The honest gap: if the laptop is off at 09:00, nothing runs. Moving the watcher to a small always-on server is next on the list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Backups that don't depend on remembering
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ct-respaldo.py&lt;/code&gt; runs at the end of the same morning round and makes two copies of every repo:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A push to a private GitHub remote. It never forces, and it never commits human work. The only files it commits on its own are the ones the checkers write.&lt;/li&gt;
&lt;li&gt;A &lt;code&gt;git bundle&lt;/code&gt; with the full history in a cloud-synced folder: one latest copy that gets overwritten and one per ISO week, keeping the last two.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Anything it couldn't back up, uncommitted changes or unpushed commits, gets written to a status file, and the run exits 1. Restoring is one command: clone the bundle.&lt;/p&gt;

&lt;h2&gt;
  
  
  A decision log instead of a memory
&lt;/h2&gt;

&lt;p&gt;Every decision goes into a markdown table: ID, evidence, decision, status. It's past 900 rows. The ones that went wrong stay in, with what was learned.&lt;/p&gt;

&lt;p&gt;It's grepped by ID, never read whole. That matters for the agents too: they read the same record, so they don't propose something I already rejected last week. A daily brief, &lt;code&gt;HOY.md&lt;/code&gt;, is composed by a script from the gates, the backups and the queue. No AI writes it by hand, and every number in it says where it came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  llms.txt and IndexNow, kept in sync by scripts
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;llms.txt&lt;/code&gt; was written by hand once, and a few weeks later an audit found it listed 7 of 12 pieces because nobody remembered to add the new ones. Now &lt;code&gt;ct-llms.py&lt;/code&gt; regenerates it from the sitemap. It keeps the hand-written sections intact, rewrites only the Academy and Blog lists in sitemap order, keeps the curated line for any piece already there, and builds new lines from each page's title and meta description. A check mode exits 1 if the live file is missing anything, and that check is one of the daily gates.&lt;/p&gt;

&lt;p&gt;For IndexNow, the problem was that Squarespace won't let you put a file at the domain root, and that's where the key has to live. The workaround: upload the key as a site file and add a URL mapping that redirects the root path to it. &lt;code&gt;ct-indexnow.py&lt;/code&gt; fetches the key from the root before every ping and refuses to send anything if it doesn't match, then notifies the IndexNow endpoints and reports which ones accepted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents split by function
&lt;/h2&gt;

&lt;p&gt;I work with AI agents, and I split them the way you'd split a team. Several do research. One evaluates what all the others produce. One documents. One keeps improving the system itself. Others work on branding, others on code. Some days 22 run at once.&lt;/p&gt;

&lt;p&gt;What they share is the system above: the same &lt;code&gt;site.json&lt;/code&gt;, the same gates, the same decision log. The AI executes. The judgment stays with me.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell anyone running a site alone
&lt;/h2&gt;

&lt;p&gt;Write the decision down before you build. Put every list in one file. Make each bug you find into a check that runs forever. Compare against yesterday, not only against the rule. And never call anything saved until a cold reload says so.&lt;/p&gt;

&lt;p&gt;This is what I build for businesses: the site, the measurement, the content and the automation behind it, end to end, with a strategy and a plan.&lt;/p&gt;

&lt;p&gt;Bring me a digital problem.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>python</category>
      <category>automation</category>
      <category>devops</category>
    </item>
    <item>
      <title>Why a Perfect Consistency Score Is a Warning Sign</title>
      <dc:creator>Lucas Fernández Díaz</dc:creator>
      <pubDate>Wed, 30 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/lucasfernandezdiaz/why-a-perfect-consistency-score-is-a-warning-sign-289</link>
      <guid>https://dev.to/lucasfernandezdiaz/why-a-perfect-consistency-score-is-a-warning-sign-289</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://contento.solutions/academy/perfect-consistency-score-warning-sign" rel="noopener noreferrer"&gt;Contento Solutions&lt;/a&gt;, where I write about SEO and AI visibility for service businesses.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Consistency is the signal almost nobody fails. Fifty one of the 54 therapy websites we measured scored 85 or better, and 17 scored a perfect 100. Then we split the cohort by that score and it ran backwards: the flawless sites were lower on all five of the other signals we measured, and lower on the index overall.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we actually measure when we say consistency
&lt;/h2&gt;

&lt;p&gt;The word covers a lot of ground in marketing writing, so here is the narrow version we score. Consistency asks whether a site contradicts itself in the places a machine reads. Not whether the brand feels coherent. Whether the served HTML says two different things.&lt;/p&gt;

&lt;p&gt;Three checks, and they start at 100 and come down from there. We deduct 15 points if any page title drops the brand name that the rest of the titles carry. We deduct 25 if the word placeholder is sitting in text a visitor can read, which is the fingerprint of a template that was filled in halfway. And we count the distinct prices published across the sampled pages, which we report as data and do not penalise at all.&lt;/p&gt;

&lt;p&gt;That last decision was a correction to our own tooling. An earlier version subtracted 20 points from any site publishing five or more price points, and that punishes a practice with seven services for having seven services. The real contradiction is the same item carrying two different prices, and this method cannot tell those apart from the outside, so it does not claim to. The &lt;a href="https://contento.solutions/academy/how-to-build-an-ai-visibility-index" rel="noopener noreferrer"&gt;method piece&lt;/a&gt; covers the other four times our instrument produced a confident wrong number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fifty one sites out of 54 pass it
&lt;/h2&gt;

&lt;p&gt;Here is the whole distribution. Seventeen sites at 100, thirty four at 85, one at 75, two at 60. Nothing below 60 in the entire cohort.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What loses points&lt;/th&gt;
&lt;th&gt;Deduction&lt;/th&gt;
&lt;th&gt;Sites affected&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A page title that drops the brand the other titles carry&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;36 of 54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The word placeholder visible in served text&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;3 of 54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Several different prices published across pages&lt;/td&gt;
&lt;td&gt;0, reported as data&lt;/td&gt;
&lt;td&gt;34 publish at least one price&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Compare that to the rest of the index. Citability averages 50.5 across the same 54 sites. Verifiable identity, 57.8. Decision served, 60.5. Consistency sits at 88.6, which makes it the second highest of the six signals and the highest one that anybody has to work at, since retrievability at 97 is mostly the hosting platform doing its job.&lt;/p&gt;

&lt;p&gt;A signal that 94 percent of a niche passes cannot separate anyone inside that niche. That was the finding we expected to write up, and it would have been a short and slightly boring piece. Then we split the cohort.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then we split the cohort, and the score ran backwards
&lt;/h2&gt;

&lt;p&gt;We grouped the 54 sites by their consistency score and averaged everything else. We expected the two groups to look roughly the same, because a signal nobody fails should not predict anything in either direction.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;17 sites scoring 100&lt;/th&gt;
&lt;th&gt;34 sites scoring 85&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Retrievability&lt;/td&gt;
&lt;td&gt;92.9&lt;/td&gt;
&lt;td&gt;98.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Citability&lt;/td&gt;
&lt;td&gt;38.4&lt;/td&gt;
&lt;td&gt;56.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Verifiable identity&lt;/td&gt;
&lt;td&gt;48.2&lt;/td&gt;
&lt;td&gt;62.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decision served&lt;/td&gt;
&lt;td&gt;57.1&lt;/td&gt;
&lt;td&gt;62.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Friction to contact&lt;/td&gt;
&lt;td&gt;68.2&lt;/td&gt;
&lt;td&gt;77.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overall index score&lt;/td&gt;
&lt;td&gt;67.5&lt;/td&gt;
&lt;td&gt;73.9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every row goes the wrong way. The sites that scored perfectly on consistency came out 6.4 points lower on the index than the sites that lost 15 points to a stray page title, and the gap is widest on citability, at 18.5 points.&lt;/p&gt;

&lt;p&gt;Our first move was to suspect the measurement rather than the finding, which is the habit that piece eight is mostly about. Three of the 17 perfect scorers had fewer than eight pages read, so we dropped them and ran it again on the 50 sites where the full sample came back. The gap survived: 63.2 against 71.7 on the five other signals. Smaller, still there, still pointing the same way.&lt;/p&gt;

&lt;p&gt;Say the size of this honestly. It is 17 sites against 34, in one cohort, in one niche, with no significance test run on it. That is a pattern worth explaining, not a law worth quoting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a thin site cannot contradict itself
&lt;/h2&gt;

&lt;p&gt;The mechanism falls out of the deduction, once you look at what it takes to trigger it.&lt;/p&gt;

&lt;p&gt;Losing those 15 points requires having a page whose title says something the other titles do not. A site with eight distinct pages, each about a specific problem, each titled for that problem, will eventually have a title that carries the topic and not the practice name. A site with four pages titled Home, About, Services and Contact, each with the brand pasted on the end, cannot possibly trip the check. There is nothing there to be inconsistent with.&lt;/p&gt;

&lt;p&gt;Citability is the row that makes this concrete, because it counts the opposite thing: headings, questions, lists, dated content, answers near the top of the page. A site built out of specific pages scores well there and is exposed to the title deduction. A site built out of four template slots scores 100 on consistency and 38 on citability, and those two numbers are the same fact described from two directions.&lt;/p&gt;

&lt;p&gt;So the perfect score is not measuring discipline. In this cohort it is mostly measuring how little the site says. That is a hypothesis that fits the data rather than something we proved, and it is the one we would test first on the next niche.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this instrument cannot see
&lt;/h2&gt;

&lt;p&gt;The uncomfortable part is that consistency is also the signal where our measurement is furthest from the thing it is named after.&lt;/p&gt;

&lt;p&gt;Real inconsistency is a fee that says 150 on the services page and 180 on the FAQ. It is an intake process described one way in the about page and another way in the booking flow. It is a therapist listed as accepting a payer that the contact page says is not accepted. None of that is visible to a script reading eight pages, because catching it requires knowing which two statements are supposed to match.&lt;/p&gt;

&lt;p&gt;What we can see are two proxies, brand in the title and template leftovers, plus a count of prices we deliberately do not score. Calling that combination consistency is convenient shorthand. It is also the weakest name to score mapping in the whole index, and we would rather write that down than let the 88.6 sit there looking like good news.&lt;/p&gt;

&lt;h2&gt;
  
  
  Twenty sites publish no price at all
&lt;/h2&gt;

&lt;p&gt;One number from this signal is worth pulling out on its own, even though it costs nobody a point. Twenty of the 54 sites publish no price anywhere in the sampled pages. No fee, no range, no sliding scale.&lt;/p&gt;

&lt;p&gt;Those sites cannot be caught contradicting themselves on price, because they never say anything to contradict. Their consistency score is clean by absence. Their decision served score, which is the signal that asks whether a visitor can work out fit and cost without contacting anyone, is where that silence gets paid for, and the cohort average there is 60.5.&lt;/p&gt;

&lt;p&gt;This is the pattern the whole piece keeps circling. Publishing less is the cheapest way to look consistent, and every signal that measures usefulness moves in the other direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with a signal that runs backwards
&lt;/h2&gt;

&lt;p&gt;Three things, and only the first one is work.&lt;/p&gt;

&lt;p&gt;Check the template leftovers, because that one is real and it is fast. Three sites out of 54 are serving the word placeholder to visitors, and there is no argument for leaving it there. Search your own site for it, along with lorem ipsum and any stock copy your platform shipped with. It takes ten minutes and it is the only part of this signal we would ask anyone to fix on a deadline.&lt;/p&gt;

&lt;p&gt;Do not chase the other 15 points. If your consistency score is 85 because a page title carries a topic rather than your practice name, that title is very likely doing its job. Rewriting eight titles to include the brand will lift one signal out of six by 15 points and will not touch citability, identity or whether a visitor can decide. Spending an afternoon there is spending it on the scoreboard.&lt;/p&gt;

&lt;p&gt;And treat a perfect score as a prompt rather than a result. If you scored 100, the question is not what you did right. It is whether there is enough on the site to be inconsistent about. That question is answered by the citability and decision served numbers sitting next to it, which is the argument for scoring the three axes separately instead of blending them, and it is the whole reason the &lt;a href="https://contento.solutions/method" rel="noopener noreferrer"&gt;method&lt;/a&gt; is built the way it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does a perfect consistency score mean my site is bad?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No. It means the score on its own tells you almost nothing, so read it next to citability and decision served. In this cohort the sites at 100 averaged 38.4 on citability, and that is the number that carries the information.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should I put my brand name in every page title?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Only if it fits. Our check looks for the brand somewhere in the title rather than at the end, and dropping it on one page costs 15 points on one signal of six. A title that describes what the page is about is worth more than the 15 points.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why do you report prices without scoring them?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Because publishing seven prices is what a practice with seven services looks like, not a contradiction. The contradiction we would want to catch is the same service priced twice, and reading eight pages from the outside cannot distinguish the two, so we report the count and leave the judgement to the reader.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this finding strong enough to act on?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Act on the template leftovers. Treat the inversion as a reason not to spend time on the remaining 15 points, which is a decision about where not to work rather than a claim about what will happen if you do. It is 17 sites against 34 in a single niche with no significance test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How would I check this on my own site?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Read the served HTML of your eight most important pages, not the rendered page. Count how many carry your brand in the title, search all eight for placeholder text, and list every price you find. The &lt;a href="https://contento.solutions/academy/how-to-build-an-ai-visibility-index" rel="noopener noreferrer"&gt;method piece&lt;/a&gt; has the sampling and exclusion rules that make the result repeatable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that transfers
&lt;/h2&gt;

&lt;p&gt;A signal everybody passes is not a standard, it is a floor, and a floor tells you who fell through rather than who is doing well. The useful move is to check which of your signals separate the field and which of them only separate the broken from everyone else. In this cohort, retrievability at 97 and consistency at 88.6 are floors. Citability at 50.5 is where the field actually spreads out.&lt;/p&gt;

&lt;p&gt;Run the same split on your own index and expect at least one signal to run backwards. Ours did, on the axis we would have been happiest to report, which is roughly how these things go. The full cohort results are in &lt;a href="https://contento.solutions/academy/therapy-websites-ai-visibility-index" rel="noopener noreferrer"&gt;the index itself&lt;/a&gt;, and we should say plainly that we sell into the niche we measured.&lt;/p&gt;

&lt;p&gt;If you want a read on what a machine can currently retrieve, trust and quote from your own site, that is a reasonable thing to bring to a call. What the practice does and what gets reported is on the &lt;a href="https://contento.solutions/services" rel="noopener noreferrer"&gt;services page&lt;/a&gt;, with more detail on the &lt;a href="https://contento.solutions/faq" rel="noopener noreferrer"&gt;FAQ page&lt;/a&gt;. The call is &lt;a href="https://contento.solutions/book-a-call" rel="noopener noreferrer"&gt;thirty minutes&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you score sites, or your own site, on more than one signal: which one would you expect to run backwards if you split by the signal everybody passes? Tell me in the comments. I would like to know whether ours is a quirk of one niche or something that shows up everywhere.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>ai</category>
      <category>datascience</category>
    </item>
    <item>
      <title>What Actually Goes In an llms.txt</title>
      <dc:creator>Lucas Fernández Díaz</dc:creator>
      <pubDate>Wed, 16 Sep 2026 20:05:37 +0000</pubDate>
      <link>https://dev.to/lucasfernandezdiaz/what-actually-goes-in-an-llmstxt-2me5</link>
      <guid>https://dev.to/lucasfernandezdiaz/what-actually-goes-in-an-llmstxt-2me5</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://contento.solutions/academy/llms-txt-what-to-put-in-it" rel="noopener noreferrer"&gt;Contento Solutions&lt;/a&gt;, where I write about SEO and AI visibility for service businesses.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;An llms.txt file is a plain text summary of your site. It is written for language models, not for crawlers. It is also a proposal rather than a standard (&lt;a href="https://llmstxt.org/" rel="noopener noreferrer"&gt;Jeremy Howard published it&lt;/a&gt; in September 2024), and no major AI company has confirmed that it reads one. Write it anyway. The reason is not the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest version of what this file is
&lt;/h2&gt;

&lt;p&gt;Somebody proposes a convention. A few thousand sites adopt it. Nobody who runs a large model says whether they use it. That is roughly where llms.txt sits today, and any article that tells you otherwise is selling something.&lt;/p&gt;

&lt;p&gt;So the useful question is not whether assistants read it. It is whether writing one is worth an hour of your time given that they might not.&lt;/p&gt;

&lt;p&gt;It is. Here is the part that gets left out of most explanations: the file is a forcing function. To write it you have to state, in fewer than a thousand words and without marketing cover, what your business does, who it is for, who it is not for, what it charges, and what it refuses to claim. Most businesses have never written that down anywhere. The ones that have tend to have a clearer website, because the same clarity leaks into every page.&lt;/p&gt;

&lt;p&gt;The file is the cheap part. The decision is the asset.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it sits next to the files you already have
&lt;/h2&gt;

&lt;p&gt;Four files describe your site to machines, and they answer four different questions. Confusing them is the most common mistake in this area.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Who reads it&lt;/th&gt;
&lt;th&gt;The question it answers&lt;/th&gt;
&lt;th&gt;Honored?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;robots.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Any crawler that chooses to&lt;/td&gt;
&lt;td&gt;May you fetch this path?&lt;/td&gt;
&lt;td&gt;Yes, by long convention&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sitemap.xml&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Search crawlers&lt;/td&gt;
&lt;td&gt;What exists here, and when did it change?&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Schema, as JSON-LD&lt;/td&gt;
&lt;td&gt;Search engines and assistants&lt;/td&gt;
&lt;td&gt;What is this page, expressed as data?&lt;/td&gt;
&lt;td&gt;Yes, where the type is supported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;llms.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Proposed for language models&lt;/td&gt;
&lt;td&gt;What is this business, in prose?&lt;/td&gt;
&lt;td&gt;Unconfirmed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notice that only one of the four is prose. That is the whole idea. Schema is excellent at facts with a fixed shape, and useless for the things that do not have one: judgement, scope, refusals, the difference between two services that look identical on a price list. An &lt;a href="https://contento.solutions/academy/entity-map-small-brands-content-architecture" rel="noopener noreferrer"&gt;entity map&lt;/a&gt; handles the structured half. This file handles the half that only survives in sentences.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually goes in it
&lt;/h2&gt;

&lt;p&gt;The convention asks for Markdown at &lt;code&gt;/llms.txt&lt;/code&gt;, opening with an H1 that names the business and a blockquote that summarises it. After that the format is loose, which is where most files go wrong. Loose does not mean write whatever you want. It means the discipline has to come from you.&lt;/p&gt;

&lt;p&gt;Six things earn their place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One sentence that survives paraphrase.&lt;/strong&gt; Assume a model will compress everything you wrote into two lines and hand those to somebody who has never heard of you. Write that compression yourself. If your summary only makes sense with the rest of the page attached, it is not a summary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The boundary of what you do.&lt;/strong&gt; Not the list of services. The edge. A sentence like "this is not a general agency and does not sell rankings, packages or volume" tells a model more than six bullet points of capability, because it can be used to rule you out. Being ruled out correctly is not a loss. It is the only way being ruled in means anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who the buyer is, in their words.&lt;/strong&gt; Sector, size, situation. A therapist opening a second practice is a different reader from a firm of nine accountants, and if your file cannot tell them apart, neither can anything reading it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Names, spelled the way people get them wrong.&lt;/strong&gt; If the founder's name carries an accent, write the unaccented spelling too and say it refers to the same person. This sounds small. It is the difference between a model resolving you to one entity and resolving you to two half entities that never accumulate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your citation terms.&lt;/strong&gt; Say plainly whether an assistant may quote you, at what length, and with what attribution. Most sites are silent here, and silence is the least useful answer. If you want to be cited, saying so costs nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A contact route that is not a form.&lt;/strong&gt; An address a system can read and repeat. Forms are drawn by JavaScript on most platforms, which means a machine reading your page may find no way to reach you at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to leave out, and this is the harder half
&lt;/h2&gt;

&lt;p&gt;The failure mode is not omission. It is writing the file as a pitch.&lt;/p&gt;

&lt;p&gt;Leave out adjectives that cannot fail. Leading, innovative, trusted, results driven. Every one of them is true of every competitor, which means none of them separate you, which means a model summarising ten businesses in your category will produce ten identical sentences and pick on some other basis.&lt;/p&gt;

&lt;p&gt;Leave out numbers you cannot stand behind in public. A percentage with no method attached is worse than no percentage, because the first person to ask where it came from finds out that nobody knows. If a number matters, say how it was measured. If it cannot be measured honestly with the tools you have, say that instead. We do exactly this on our own file, and it is not modesty, it is that the alternative is a claim we would have to defend.&lt;/p&gt;

&lt;p&gt;Leave out anything the site itself contradicts. A file that says pricing is transparent, on a site where pricing is on request, has told a reader something more damaging than either statement alone. This is the cheapest audit in the whole exercise: read the file, then read the site, and fix whichever one is lying.&lt;/p&gt;

&lt;p&gt;And leave out the entire sitemap. A list of every URL is what &lt;code&gt;sitemap.xml&lt;/code&gt; is for. Six or eight pages that matter, each with a line saying why it exists, is worth more than eighty links with no explanation.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to check it did anything
&lt;/h2&gt;

&lt;p&gt;You mostly cannot, and pretending otherwise is how this field earns its reputation.&lt;/p&gt;

&lt;p&gt;There is no report for this. No console panel, no impression count, no line in analytics that says an assistant read your file. Anyone offering you a dashboard for llms.txt performance has invented the numbers on it.&lt;/p&gt;

&lt;p&gt;What you can do is sample. Ask four or five assistants the questions a buyer would actually ask, in a fresh session with no history: who does this kind of work for this kind of business, what does it cost, who is it not for. Record the answers verbatim, with the date. Repeat in a month. You are not measuring a metric, you are watching a description of your business drift toward or away from the one you wrote. That is a slower and less satisfying instrument than a chart, and it is the honest one. The same sampling discipline is what makes &lt;a href="https://contento.solutions/academy/search-console-ai-overviews-visibility-measurement" rel="noopener noreferrer"&gt;Search Console useful after AI Overviews&lt;/a&gt;, where the temptation to read a number that does not mean what it looks like is even stronger.&lt;/p&gt;

&lt;p&gt;One more check, and it takes a minute. Fetch your own file with a plain request and read it end to end as a stranger would. If you finish it and could not say who this business turns away, the file is not finished.&lt;/p&gt;

&lt;h2&gt;
  
  
  The order to do this in
&lt;/h2&gt;

&lt;p&gt;Write the file last, not first.&lt;/p&gt;

&lt;p&gt;An llms.txt that describes a site a model cannot fetch, or cannot parse, or finds contradicts itself page to page, is a well written summary of a problem. The sequence that works is unglamorous: make the pages readable without JavaScript, make the structured data say what the page says, make the site agree with itself, and then write the file that summarises all of it. We wrote about the first three as &lt;a href="https://contento.solutions/academy/make-website-citable-ai-search" rel="noopener noreferrer"&gt;making a site citable&lt;/a&gt;, and the order matters more than any individual step.&lt;/p&gt;

&lt;p&gt;Then the hour you spend on the file is spent describing something true, which is the only version of this work that survives contact with a reader who checks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions people ask about this
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Do ChatGPT, Claude, Gemini or Perplexity actually read llms.txt?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
None has confirmed it. Treat any claim that they do, or that they do not, as unverified. The case for writing one rests on the clarity it forces and its near zero cost, not on a confirmed behaviour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is llms.txt a replacement for schema markup?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No. They do different jobs. Schema states facts in a form machines parse reliably. llms.txt states judgement, scope and boundaries in prose, which schema has no way to express. A site that needs one usually needs both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does adding llms.txt affect Google rankings?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
There is no evidence that it does, and no reason to expect it. It is not a ranking file. Anybody selling it as one is describing a mechanism that has never been shown to exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long should it be?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Long enough to state what you do, who for, who not for, and how to reach you. Ours runs about seven thousand characters and that is on the generous side. Under a thousand words is a reasonable ceiling for most businesses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where does the file go, and how do I add one on a hosted platform?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
It belongs at the root, as &lt;code&gt;yourdomain.com/llms.txt&lt;/code&gt;. Platforms that do not let you place arbitrary files at the root make this awkward, and the workaround varies by platform. If yours will not allow it, the schema and page level work still stands on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;Write the file because writing it makes you decide what your business is. Keep it true, keep it specific, keep it consistent with the site it describes. If assistants turn out to read it, you were ready. If they never do, you still ended up with the clearest paragraph anyone in your company has written about what you sell.&lt;/p&gt;

&lt;p&gt;If you want a second read on whether your site says the same thing to a person, a crawler and an assistant, that is a reasonable thing to bring to a call. How this practice works and what gets reported is on the &lt;a href="https://contento.solutions/services" rel="noopener noreferrer"&gt;services page&lt;/a&gt;, with more detail on the &lt;a href="https://contento.solutions/faq" rel="noopener noreferrer"&gt;FAQ page&lt;/a&gt;. The call is &lt;a href="https://contento.solutions/book-a-call" rel="noopener noreferrer"&gt;thirty minutes&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you already have an llms.txt, drop the URL in the comments and tell me what you decided to leave out. That choice says more about a business than anything that made it into the file.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>ai</category>
      <category>webdev</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
