<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: WebPixie</title>
    <description>The latest articles on DEV Community by WebPixie (@webpixie).</description>
    <link>https://dev.to/webpixie</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4066463%2Ff0bd83d1-25f9-4721-8f05-fcc9e19d0379.png</url>
      <title>DEV Community: WebPixie</title>
      <link>https://dev.to/webpixie</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/webpixie"/>
    <language>en</language>
    <item>
      <title>GEO vs SEO: What actually changes in your content</title>
      <dc:creator>WebPixie</dc:creator>
      <pubDate>Sat, 29 Aug 2026 12:13:51 +0000</pubDate>
      <link>https://dev.to/webpixie/geo-vs-seo-what-actually-changes-in-your-content-8ld</link>
      <guid>https://dev.to/webpixie/geo-vs-seo-what-actually-changes-in-your-content-8ld</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;GEO is SEO with one extra job: getting cited inside AI assistant responses. The 2023 paper that coined the term tested 9 content interventions. Four worked strongly, three gave moderate gains, and two flatlined or backfired, including classic keyword stuffing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;GEO is SEO with one extra job: getting cited inside AI assistant responses, not only ranked on a Google results page. It is not a new discipline. The original GEO paper (Aggarwal et al., 2023) tested nine content interventions against generative engines. Four moved the citation needle reliably, three gave moderate gains, and two flatlined or actively hurt visibility, including the classic SEO move of keyword stuffing.&lt;/p&gt;

&lt;p&gt;Most online comparisons treat GEO as a brand-new craft requiring a brand-new workflow. The empirical evidence says the opposite: most of what works for GEO is what a good editor would do anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  What GEO actually is
&lt;/h2&gt;

&lt;p&gt;GEO stands for Generative Engine Optimization, a term introduced by Aggarwal et al. in a 2023 arXiv paper that was accepted to KDD 2024. The paper formalises the question that came up the moment ChatGPT, Perplexity, Claude, and Gemini started answering queries with citations: how do you raise the probability that your page is one of the cited sources?&lt;/p&gt;

&lt;p&gt;The paper measures two visibility outcomes. Position-Adjusted Word Count (PAWC) is roughly 'how prominently your content appears in the answer, weighted by where it appears.' Subjective Impression (SI) is a model-rated assessment of how much the answer relies on your source. Both are imperfect proxies, but together they give a measurable target.&lt;/p&gt;

&lt;p&gt;The headline result: the paper's composite GEO approach boosted visibility by up to 40 percent across the test set. Per-intervention results from the paper's Table 1 are more useful than the composite, because they show which specific moves matter and which do not. The numbers below come from that table.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four strongest interventions
&lt;/h2&gt;

&lt;p&gt;These four moves produced the largest gains on both metrics in the paper's evaluation.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Quotation Addition: +42.6% PAWC, +23.3% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Adding short, properly attributed quotations from authoritative sources had the largest effect of any intervention. Quotations from credible publications, papers, or official documentation are treated by generative engines as high-signal evidence.&lt;/p&gt;

&lt;p&gt;Practical version: when you can paraphrase a source or quote them directly, quote them. Pick the sentence that says the thing most precisely. Attribute inline with the source name and date.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Statistics Addition: +32.8% PAWC, +16.6% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Replacing qualitative language with quantitative claims tied to specific data sources came second. The effect is strongest in factual queries and comparison content.&lt;/p&gt;

&lt;p&gt;In practice, every claim that can take a number should take a number, and every number should name its source. Replace 'a lot of users complain' with the actual percentage from the actual survey, named. Generic 'studies show' phrasing is the opposite of what works here.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Fluency Optimization: +28.7% PAWC, +9.3% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Improving the readability and linguistic quality of the page came third. The takeaway is not 'use simpler words.' It is that messy, hard-to-parse prose is less likely to be selected by the model than clean prose that says the same thing. The Subjective Impression gain is much smaller than PAWC, suggesting fluency helps the model find and include your content more than it changes how the model judges your authority.&lt;/p&gt;

&lt;p&gt;Working version: shorter paragraphs, active voice, sentences that resolve on a clear claim rather than trailing into hedging. Most editors already know how to do this; GEO is the empirical confirmation that it pays off.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Cite Sources: +27.7% PAWC, +10.9% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Adding hyperlinked citations to credible sources came fourth, narrowly behind Fluency. The model treats outbound links to credible publications as a quality signal for the citing page.&lt;/p&gt;

&lt;p&gt;What this looks like: when you make a non-trivial claim, link to where you got it. The anchor text should describe the destination, not say 'click here.' If you cannot find a credible source for a claim, downgrade the claim, do not link to a weaker source.&lt;/p&gt;

&lt;h2&gt;
  
  
  The moderate-gain interventions
&lt;/h2&gt;

&lt;p&gt;These three gave smaller but still meaningful gains, with effects that vary by content area.&lt;/p&gt;

&lt;h3&gt;
  
  
  Technical Terms: +18.5% PAWC, +8.3% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Adding domain-specific terminology gave a larger boost than the paper's narrative framing implied. It works because generative engines index by entity; precise technical vocabulary helps the model match the page to the query. The lift is real but smaller than the top four, and the catch is that forcing technical terms into a general-audience piece probably reverses the gain by making the prose harder to parse.&lt;/p&gt;

&lt;h3&gt;
  
  
  Easy-to-Understand: +13.8% PAWC, +4.7% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Simplifying language while keeping accuracy improved PAWC but barely moved Subjective Impression. The model rewards clarity for content selection but does not particularly judge simpler content as more reliable. Effect is intertwined with Fluency, and the paper found it hard to fully separate the two.&lt;/p&gt;

&lt;h3&gt;
  
  
  Authoritative: +11.8% PAWC, +15.5% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Switching to a more confident, definitive voice stands out because it is the only intervention where Subjective Impression gains exceed PAWC gains. Definitive prose changes how the model judges your reliability more than it changes how often the model picks your content. Position-taking matters, but the lift is moderate, not transformational.&lt;/p&gt;

&lt;p&gt;Across all three moderate-gain interventions, the move is the same: take a clear position in entity-precise prose, written for the audience the topic implies. The bigger lift still comes from quotes, statistics, and citations; tone and clarity are smaller multipliers on top.&lt;/p&gt;

&lt;h2&gt;
  
  
  The interventions that flatlined or hurt
&lt;/h2&gt;

&lt;p&gt;Two interventions produced near-baseline or actively negative results.&lt;/p&gt;

&lt;h3&gt;
  
  
  Unique Words: +6.2% PAWC, +6.2% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;Sprinkling rare vocabulary across the page produced near-baseline gains on both metrics. The model is not impressed by lexical diversity for its own sake. Skip this move entirely.&lt;/p&gt;

&lt;h3&gt;
  
  
  Keyword Stuffing: -8.7% PAWC, +2.6% Subjective Impression
&lt;/h3&gt;

&lt;p&gt;The classic SEO move of increasing search-keyword density actually decreased Position-Adjusted Word Count by 8.7 percent. Subjective Impression saw a small +2.6 percent bump, but the headline finding is clear: keyword density hurts citation visibility, not helps it. Generative engines do not rank documents the way a keyword-based search index does; they read for meaning and citation-worthiness. Keyword density is not a citation signal, and aggressive stuffing is a citation anti-signal.&lt;/p&gt;

&lt;p&gt;In short, do not write GEO content like 2012 SEO content. Density is not depth; in the paper, density was a penalty.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where GEO and SEO overlap and where they diverge
&lt;/h2&gt;

&lt;p&gt;Most of the empirically supported GEO interventions are also SEO best practice. Citing sources, using statistics, writing clearly, taking positions: these have been in the SEO playbook for over a decade, even if the rationale was different.&lt;/p&gt;

&lt;p&gt;The divergences are smaller than the GEO marketing suggests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Keyword density actively hurts in GEO. Classical SEO rewards keyword presence and proximity to natural phrasing. GEO penalises high density. A page that says the right thing once with good context outperforms a page that repeats the same thing five times with awkward stuffing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Outbound links matter more, not less. Classical SEO worried about 'link juice' leaking through outbound links. GEO rewards citations to credible sources at roughly +27.7 percent PAWC. Outbound links to authoritative work are a feature, not a leak.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Quote density matters. SEO does not care whether you paraphrase or quote your source. GEO rewards direct quotation with attribution at the largest effect size in the paper (+42.6 percent PAWC). This is the single largest divergence between the two playbooks.&lt;br&gt;
Structural signals carry similar weight across both. Headings, scannable structure, direct-answer paragraphs near the top help in both regimes. The recent rise of llms.txt is a separate access-layer concern, not a content-layer one; see whether llms.txt actually moves the needle for that side of the picture.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Net effect: GEO does not require you to throw out the SEO workflow. It requires you to dial up two practices (quotation and citation) and dial down one (keyword stuffing, which the paper shows is now actively harmful). The rest is recognisable.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical GEO checklist
&lt;/h2&gt;

&lt;p&gt;Distil the empirical findings into something you can hold next to a draft:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Quote your sources directly when paraphrasing weakens the claim. Pick the sentence that says it most precisely. Attribute inline with source name and date. This is the single highest-leverage move per the paper.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Convert vague claims to specific numbers with sourced provenance. Replace 'most users' with the actual percentage from the actual survey, named.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Cite credible sources with descriptive anchor text. If you cannot find a credible source for a claim, downgrade the claim instead of citing a weaker source.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Take a position. Hedging across every claim signals to the model that the page is not the best source on the topic.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Write entity-precise prose. Use named tools, exact version numbers, official acronyms ('TLS 1.3', not 'modern encryption'). Generative engines index by entity.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Keep the prose clean. Short paragraphs, active voice, sentences that resolve. The 'fluency' effect in the paper is partly about not making the model work harder than necessary.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Skip the SEO-era keyword stuffing entirely. The paper shows it actively hurts citation visibility, not just fails to help.&lt;br&gt;
The 2023 GEO paper is the empirical baseline, not the final word.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For now, the practical workflow is short. Quote, cite, count, write clearly, take a position. None of that requires a new content team or a new editorial calendar. It requires an editor who treats every paragraph as a possible answer in someone else's AI search.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://webpixie.io/blog/post/geo-vs-seo-content-changes" rel="noopener noreferrer"&gt;Blog Link&lt;/a&gt;&lt;br&gt;
&lt;a href="https://webpixie.io/" rel="noopener noreferrer"&gt;Site Link&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Should you block AI crawlers? Honest answer for 2026</title>
      <dc:creator>WebPixie</dc:creator>
      <pubDate>Tue, 18 Aug 2026 12:28:20 +0000</pubDate>
      <link>https://dev.to/webpixie/should-you-block-ai-crawlers-honest-answer-for-2026-3ii5</link>
      <guid>https://dev.to/webpixie/should-you-block-ai-crawlers-honest-answer-for-2026-3ii5</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Most sites should allow AI training crawlers in 2026: invisibility in AI assistant answers now costs more than uncrawled-for-training saves. Here is the per-provider breakdown (GPTBot, ClaudeBot, Google-Extended) and the robots.txt for each realistic decision.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Most sites should allow the major AI training crawlers in 2026, not block them. The cost of being invisible inside AI assistant answers is now larger than the upside of being uncrawled for training. The exceptions are real but narrow: IP-sensitive content, regulated data, content where commercial moats depend on scarcity. For everyone else, blocking is the lazy default that quietly costs you discovery.&lt;/p&gt;

&lt;p&gt;Most existing posts on this topic hedge to a 'balanced hybrid approach' without explaining what the actual tradeoff looks like per provider. This post takes a clearer line: walk through what each AI provider's bots actually do, separate training crawl from live retrieval, and give the concrete robots.txt for each realistic decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three providers and their bots in 2026
&lt;/h2&gt;

&lt;p&gt;Every major AI provider now runs more than one crawler. The names are confusing, the purposes overlap, and most robots.txt advice on the internet treats them as one thing. They are not.&lt;/p&gt;

&lt;h3&gt;
  
  
  OpenAI: three bots, three jobs
&lt;/h3&gt;

&lt;p&gt;OpenAI documents three bots in OpenAI's bots reference:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;GPTBot collects content used to train future OpenAI models. Verifiable IP list at OpenAI's GPTBot IP list.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;OAI-SearchBot crawls the web for ChatGPT search results. Verifiable at OpenAI's OAI-SearchBot IP list.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;ChatGPT-User fetches pages on demand when a ChatGPT user asks a question that requires browsing. Verifiable at OpenAI's ChatGPT-User IP list.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Blocking GPTBot stops your content from feeding new model training runs. Blocking OAI-SearchBot stops you from being a citable source in ChatGPT search results. Blocking ChatGPT-User stops live retrieval when a user explicitly asks ChatGPT about your site. The three have very different consequences, and most blanket 'block ChatGPT' advice treats them as one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Anthropic: three bots, three jobs
&lt;/h3&gt;

&lt;p&gt;Anthropic documents the same structural pattern in Anthropic's crawler FAQ:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ClaudeBot collects content for training Anthropic's Claude models.&lt;/li&gt;
&lt;li&gt;Claude-User performs live retrieval when a Claude user directs it to fetch a page.&lt;/li&gt;
&lt;li&gt;Claude-SearchBot analyses content to improve Claude's search results.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;IP ranges are published at Anthropic's crawler IP list for verification. Anthropic notes explicitly that their bots respect robots.txt 'do not crawl' directives.&lt;/p&gt;

&lt;h3&gt;
  
  
  Google: one control token, no separate bot
&lt;/h3&gt;

&lt;p&gt;Google takes a different approach, documented in Google's common crawlers documentation.&lt;/p&gt;

&lt;p&gt;Google-Extended is not a separate crawler with its own user-agent string. It is a robots.txt control token. The actual fetching is done by Googlebot; Google-Extended tells Google whether your content can be used for training Gemini and the Gemini API. The critical clarification from Google's own docs: 'Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.' Blocking Google-Extended does not affect Search.&lt;/p&gt;

&lt;h2&gt;
  
  
  What allowing actually buys you
&lt;/h2&gt;

&lt;p&gt;The honest case for allowing training crawlers comes down to citation visibility. It is stronger in 2026 than it was in 2023.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Citation share in AI answers. ChatGPT, Claude, and Gemini increasingly answer questions by citing or summarising specific pages. If your content is not in the training set and not in the live retrieval surface, you are not in the answer.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Long-tail discovery. AI assistants surface niche content that traditional search undervalues. A small blog with deep expertise on one topic can earn AI citations it could never earn against high-DA SEO competitors on Google. Blocking gives up that lane entirely.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Trust signals during sales cycles. If your buyers research vendors through ChatGPT or Claude before making contact, a site blocked from AI retrieval can surface as "unable to verify" in answers, which reads as a credibility gap.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Compounding inertia. Sites that established AI-citation footprints early benefit from retrieval models that learned them as canonical sources. Catching up later is harder than not blocking now.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are not hypothetical benefits any more. Anecdotal reports from publishers suggest some AI-driven referral traffic, though hard numbers are scarce and vary widely by site type.&lt;/p&gt;

&lt;h2&gt;
  
  
  What blocking actually costs you
&lt;/h2&gt;

&lt;p&gt;The cost of blocking has shifted since 2023. Three points worth noting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Invisibility in AI answers. The most direct cost. Competitors who allow crawlers appear in answers where you do not. For SEO-equivalent visibility purposes, this is now a real lane to compete in.&lt;/li&gt;
&lt;li&gt;No exit clause from existing training sets. Blocking GPTBot today does not remove your content from models already trained on it. Major training datasets sampled the public web years ago and content from those scrapes is already baked in. Blocking is forward-looking, not retroactive.&lt;/li&gt;
&lt;li&gt;Operational complexity. Maintaining an accurate bot blocklist requires keeping up with new user-agent strings as providers add them. Anthropic added Claude-SearchBot after the initial ClaudeBot; OpenAI added OAI-SearchBot and ChatGPT-User after the initial GPTBot. Blocklists that aged out of date allow exactly what they were trying to prevent.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The per-provider decision
&lt;/h2&gt;

&lt;p&gt;The three providers behave differently. The decision is not one-size-fits-all.&lt;/p&gt;

&lt;h3&gt;
  
  
  OpenAI
&lt;/h3&gt;

&lt;p&gt;Most sites should allow ChatGPT-User and OAI-SearchBot, because those are the surfaces where ChatGPT actually cites you. The training crawl (GPTBot) is the one with the lowest direct upside (training-set inclusion has the longest delay between crawl and benefit) and the loudest objections (IP, compensation). Reasonable position: allow ChatGPT-User and OAI-SearchBot, block GPTBot. More permissive: allow all three.&lt;/p&gt;

&lt;h3&gt;
  
  
  Anthropic
&lt;/h3&gt;

&lt;p&gt;Same pattern. Allow Claude-User and Claude-SearchBot so you can be cited in Claude responses and Claude-powered search. Block ClaudeBot training if you object to training-set inclusion, allow it if you do not. The middle ground is well supported by Anthropic's own bot taxonomy.&lt;/p&gt;

&lt;h3&gt;
  
  
  Google
&lt;/h3&gt;

&lt;p&gt;Special case. There is no separate Google-Extended bot to block crawling-wise; it only controls whether your content trains Gemini. Blocking has no Search-visibility downside per Google's own documentation. The decision here is purely about training-set inclusion, with no AI-citation tradeoff (Gemini's citation behaviour is governed by separate Google products, not by Google-Extended). For most sites, the choice is roughly neutral; for IP-sensitive content, blocking Google-Extended is the cheapest privacy improvement available.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hidden middle ground: block training, allow retrieval
&lt;/h2&gt;

&lt;p&gt;Most online advice treats the decision as binary: block all AI crawlers, or allow all AI crawlers. The bot taxonomies above make a third position available.&lt;/p&gt;

&lt;p&gt;Block training crawl, allow live retrieval and search. In concrete terms: Disallow GPTBot, allow ChatGPT-User and OAI-SearchBot. Disallow ClaudeBot, allow Claude-User and Claude-SearchBot. Disallow Google-Extended, leave Googlebot alone (which you would not have blocked anyway).&lt;/p&gt;

&lt;p&gt;This middle ground says: 'do not train on me without consent or compensation, but do cite me when a user asks about me directly.' It preserves discovery and citation upside while objecting to training-set extraction. It is the most defensible default for an editorial content business: you keep the citation lane open and only give up the training-set inclusion you were most ambivalent about.&lt;/p&gt;

&lt;p&gt;The position has tradeoffs. Some critics argue the distinction is naive: training data drives the same models that do live retrieval, so blocking training while allowing retrieval is morally inconsistent. Others argue the distinction is exactly the point: training is a permanent extraction, retrieval is an in-the-moment citation. Pick a side, but pick it knowingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to actually configure robots.txt
&lt;/h2&gt;

&lt;p&gt;Three concrete configurations cover almost every realistic scenario. Pick the one that matches your position and paste it into your /robots.txt.&lt;/p&gt;

&lt;h3&gt;
  
  
  Allow everything (default for most public sites)
&lt;/h3&gt;

&lt;p&gt;Do nothing. The default of no rules is the most permissive option. AI crawlers respect robots.txt; if you have no rules excluding them, they fetch by default. For most content businesses interested in citation, this is the right baseline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Block training only, allow retrieval and search
&lt;/h3&gt;

&lt;p&gt;The middle-ground configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;GPTBot&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;ClaudeBot&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Google&lt;/span&gt;-&lt;span class="n"&gt;Extended&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="c"&gt;# ChatGPT-User, OAI-SearchBot, Claude-User, Claude-SearchBot
# are allowed by default (no Disallow rules above)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This blocks the three training crawlers while leaving the live-retrieval and search bots free to fetch. It is the configuration that matches the block-training, allow-retrieval position above.&lt;/p&gt;

&lt;h3&gt;
  
  
  Block everything AI-related
&lt;/h3&gt;

&lt;p&gt;The maximum-blocking configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;GPTBot&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;OAI&lt;/span&gt;-&lt;span class="n"&gt;SearchBot&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;ChatGPT&lt;/span&gt;-&lt;span class="n"&gt;User&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;ClaudeBot&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Claude&lt;/span&gt;-&lt;span class="n"&gt;User&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Claude&lt;/span&gt;-&lt;span class="n"&gt;SearchBot&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;Google&lt;/span&gt;-&lt;span class="n"&gt;Extended&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use this only if you have a specific reason to be invisible to AI products. Most sites do not.&lt;/p&gt;

&lt;h2&gt;
  
  
  When blocking actually makes sense
&lt;/h2&gt;

&lt;p&gt;Blocking has real use cases. The honest list is short:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;IP-sensitive content. Original research, paywalled journalism, proprietary datasets, internal documentation accidentally exposed. The cost of inclusion in a training set may exceed the benefit of being cited.&lt;/li&gt;
&lt;li&gt;Regulated content. Healthcare, legal, financial advice content that comes with compliance obligations. Regulators are paying increasing attention to AI training data sourcing; blocking is the conservative posture.&lt;/li&gt;
&lt;li&gt;Commercial moats. If your value proposition depends on customers needing to come to your site to access the answer, citation by AI assistants is a direct threat, not an opportunity. Few sites are actually in this position, but those that are tend to know it.&lt;/li&gt;
&lt;li&gt;Contractual constraints. Some B2B content contracts require that content not be used for training third-party models. Honour the contract; configure the blocklist accordingly.
If none of these applies to your site, blocking AI training is reflex, not strategy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The default decision should be 'allow', with the option to block training-specific bots if you have a real reason. The hidden middle ground (block training, allow retrieval) is the most coherent position for editorial content businesses. The maximalist block-everything posture is rarely justified and usually costs more than it saves.&lt;/p&gt;

&lt;p&gt;For background on the access-layer file these decisions get expressed alongside, see what llms.txt is and whether crawlers read it. The bots respect robots.txt as the gate; llms.txt is the brief on what to read once they are inside.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://webpixie.io/" rel="noopener noreferrer"&gt;See More&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>seo</category>
    </item>
    <item>
      <title>Website downtime cost: deconstructing the famous numbers</title>
      <dc:creator>WebPixie</dc:creator>
      <pubDate>Sat, 15 Aug 2026 10:46:56 +0000</pubDate>
      <link>https://dev.to/webpixie/website-downtime-cost-deconstructing-the-famous-numbers-40mf</link>
      <guid>https://dev.to/webpixie/website-downtime-cost-deconstructing-the-famous-numbers-40mf</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;The famous downtime cost numbers (Gartner's $5,600/min, IDC's $1M/hour) come from enterprise IT surveys that rarely apply to most sites. Here is where each one comes from, why it overstates the cost for most businesses, and a calculation framework that does apply.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Direct answer: for most websites, one hour of downtime costs in the low thousands of dollars, not the $336,000 implied by Gartner's $5,600 per minute. The famous figures come from 2014 surveys of Fortune 1000 IT decision-makers and rarely fit smaller sites.&lt;/p&gt;

&lt;p&gt;The famous downtime cost numbers, Gartner's $5,600 per minute, IDC's $1 million per hour, Amazon's $13 million per hour, are real citations. They are also answers to a question almost nobody asking you for a monitoring budget is actually asking. The honest cost of your downtime depends on your traffic, your revenue model, and what your customers can substitute for your service. For most sites the real number is far smaller than the headline figures, and the real risk is qualitative.&lt;/p&gt;

&lt;p&gt;The headline numbers are not wrong. They are misapplied. This post traces where each one came from, explains why it rarely fits your site, and gives a calculation framework you can actually use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the famous downtime numbers come from
&lt;/h2&gt;

&lt;p&gt;Three numbers dominate the citation graph for downtime cost. Each has a real source, and each one has been quoted without context for over a decade.&lt;/p&gt;

&lt;h3&gt;
  
  
  $5,600 per minute (Gartner)
&lt;/h3&gt;

&lt;p&gt;The most-cited figure traces to a 2014 Gartner blog post by Andrew Lerner titled "The Cost of Downtime", published July 16, 2014, that estimated the average cost of IT downtime at $5,600 per minute. Lerner himself noted in the same post that this was an average with a large degree of variance: one study he referenced put the actual range at $140,000 to $540,000 per hour, equivalent to roughly $2,300 to $9,000 per minute. The $5,600 figure is the midpoint of that range, not a universal constant. The original Gartner blog URL has since been retired, so the number is now widely cited through secondary copies; treat the precise figure as a directional benchmark rather than a confirmable primary source.&lt;/p&gt;

&lt;p&gt;Two things to notice. The survey measured IT downtime across full enterprise stacks, including ERP outages, internal application failures, and back-office systems. Web-facing downtime is a subset, not the whole. And the surveyed organisations were medium-to-large enterprises, not the long tail of online businesses.&lt;/p&gt;

&lt;h3&gt;
  
  
  $1 million per hour (IDC, ITIC)
&lt;/h3&gt;

&lt;p&gt;IDC and ITIC, two separate research firms often conflated in citations, publish annual surveys of enterprise IT reliability. The $1 million per hour figure most commonly traces to ITIC's Hourly Cost of Downtime survey. The 2024 edition reports that 41 percent of surveyed enterprises put hourly downtime cost in the $1 million to $5 million-plus range, with the top tier (financial services, energy, telecommunications) clustered at the high end.&lt;/p&gt;

&lt;p&gt;The exact headline ITIC uses for the 2024 report is unambiguous about what is being measured:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;41% of Enterprises Say Hourly Downtime Costs $1 Million to Over $5 Million&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same caveat as before. The respondents are typically Fortune 1000 IT operations, not online businesses in general. The number reflects the high end of a heavy-tail distribution, not a median, and the survey deliberately samples the segment with the largest cost exposure.&lt;/p&gt;

&lt;h3&gt;
  
  
  ~$13 million per hour (Amazon)
&lt;/h3&gt;

&lt;p&gt;The Amazon figure is derived arithmetic, not a survey. It takes Amazon's annual e-commerce revenue, divides by the number of seconds in a year, and produces a per-second revenue figure that is sometimes annualised back to a per-hour cost of downtime. Variants of this calculation circulated between 2019 and 2021. A 2021 Amazon outage was widely reported in the business press as costing tens of millions of dollars in lost sales over roughly one hour; precise dollar estimates varied across outlets and the underlying primary source is not independently verifiable, so treat the order of magnitude rather than any specific figure as the takeaway. Even at the lower end of those estimates, the per-minute cost runs two orders of magnitude above the Gartner average.&lt;/p&gt;

&lt;p&gt;The Amazon number applies to Amazon. It applies to almost no other site.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why these numbers rarely apply to your site
&lt;/h2&gt;

&lt;p&gt;There are four common ways the famous figures get misused:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Sample skew. The surveys behind these numbers measure enterprises with large IT estates and complex internal dependencies. A small SaaS, a marketing site, or a single-product DTC store has a smaller blast radius and a smaller cost surface.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Scope mismatch. Survey 'downtime' usually means any IT system being unavailable, including internal tools, ERP, payroll, and back-end databases. Web outage is a subset. Quoting an IT downtime figure to budget for web monitoring inflates the case.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Indirect-cost inclusion. The headline numbers often include lost employee productivity (engineers, support, sales unable to work) and contractual penalties. Those are real costs but they show up on different lines of the P&amp;amp;L than your hosting bill.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Median versus mean. Heavy-tail distributions have a small number of large outliers (a payment processor going down) dominating the mean. Quoting the mean as if it applies to the median site is statistically misleading.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this means downtime is free. It means your number is different from the famous number, and it is your number that matters when you decide what to spend preventing outages.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually calculate
&lt;/h2&gt;

&lt;p&gt;Four categories add up to your real downtime cost. They are easy to estimate badly and harder to estimate well, but even a rough estimate is more useful than borrowing someone else's headline.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Direct revenue lost during the outage&lt;br&gt;
For transactional sites, this is the clearest line. Take your average revenue per hour during the outage window and multiply by the duration. Cleaner if you split by hour of day; revenue at 4 AM and 4 PM are not the same.&lt;br&gt;
Watch out for substitution. If your site goes down and customers come back ten minutes later to complete the order, the revenue is delayed, not lost. If they go to a competitor instead, it is gone. Cart abandonment data from your own analytics is the best proxy.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Internal team cost during the incident&lt;br&gt;
Engineers stop shipping. Customer success answers tickets instead of doing their job. The on-call rotation absorbs hours that come out of feature work or sleep. Multiply the fully-loaded hourly cost of everyone involved by the duration of the response, including post-incident review and remediation. This number is rarely tiny. A four-engineer incident lasting two hours, plus a one-hour post-mortem, at a fully-loaded $150 per hour, runs $1,800 per incident before anyone counts revenue.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Recovery cost&lt;br&gt;
Refunds and chargebacks where applicable. Goodwill credits issued to keep customers from leaving. Support ticket surge during and after the outage. Re-engagement marketing if churn risk spikes. These costs land in the days following the outage, not during it, which is why they often get missed in the first estimate.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Trust and reputation damage&lt;br&gt;
This category is real and is also the hardest to quantify honestly. You can proxy it by tracking churn rate before and after notable outages, by tracking NPS changes, by watching social media sentiment, or by measuring sales-cycle elongation for B2B customers asking pointed reliability questions. None of these proxies is perfect, but ignoring the category gives you a number that is too low.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Worked examples by business type
&lt;/h2&gt;

&lt;p&gt;These are hypothetical scenarios meant to illustrate the framework, not specific customer figures. Plug in your own numbers when you do this for real.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hypothetical: a DTC ecommerce store, $5 million ARR
&lt;/h3&gt;

&lt;p&gt;Average revenue per hour: roughly $570 (revenue divided by 8,760 hours per year). Peak Thursday-evening or weekend hours might run 3 to 5 times that. A one-hour outage during a peak slot loses an estimated $1,700 to $2,800 in direct revenue, with maybe 40 percent recovered by returning customers, leaving net direct loss around $1,000 to $1,700. Add internal team cost ($1,800 for a typical incident response), recovery support ($400 to $800 in support tickets and goodwill credits), and the total lands at roughly $3,500 to $4,500 per peak-hour outage. Not the Gartner $5,600 per minute, but not nothing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hypothetical: a small B2B SaaS, $1 million ARR
&lt;/h3&gt;

&lt;p&gt;Average revenue per hour: roughly $114. Revenue is not the right line for SaaS downtime. The real costs are SLA credit exposure (often 10 to 30 percent of monthly revenue for severe breaches, per typical SaaS SLA terms), churn risk (a small fraction of customers per outage, multiplied by lifetime value), and internal team response. A two-hour outage with an SLA breach might trigger $5,000 to $15,000 in service credits, plus the team cost, plus a hard-to-quantify churn increment. SaaS downtime cost is dominated by contractual and trust factors, not by per-minute revenue.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hypothetical: a content site with ad revenue, $300K annual
&lt;/h3&gt;

&lt;p&gt;Daily ad revenue is roughly $820. Each hour of downtime during organic-traffic hours costs around $35 to $100 in lost impressions, plus a small but real SEO impact if the outage coincides with crawler activity. A two-hour outage in this profile is closer to $100 to $200 in immediate revenue loss, plus the qualitative impact on returning-reader retention. The famous numbers grossly overstate this case.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hypothetical: a paid API service, $2 million ARR
&lt;/h3&gt;

&lt;p&gt;API revenue is more linear with uptime than any other model. Customers integrate your API into their products; your downtime cascades into their downtime, which means their angry customers become your angry customers. Direct revenue loss is straightforward (revenue per hour times outage duration), but the indirect cost is larger and shows up in renewal conversations months later. SLA penalty clauses in B2B API contracts are often more punitive than SaaS dashboard SLAs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hypothetical: a brochure marketing site
&lt;/h3&gt;

&lt;p&gt;Direct revenue loss during an outage: $0. Internal team cost during incident response: real, but often nominal because the team rarely needs to wake up at night for a brochure site. Recovery cost: low. Trust damage: real but slow, and mostly relevant if a prospect happens to visit during the outage. The honest downtime cost for a brochure site is dominated by qualitative impressions, not currency. Spending a top-tier monitoring budget on a brochure site is hard to justify on cost arithmetic alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the cost spikes
&lt;/h2&gt;

&lt;p&gt;Downtime cost is not constant across the clock. A few factors that shift it materially:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Peak versus off-peak hours. Most consumer sites see 3 to 10 times the traffic during peak hours compared to overnight. A one-hour outage at 8 PM and a one-hour outage at 4 AM are not comparable losses.&lt;/li&gt;
&lt;li&gt;Campaign windows. Paid advertising drives traffic to specific landing pages. An outage during a paid campaign window wastes ad spend in real time, on top of losing the conversions the campaign was meant to drive.&lt;/li&gt;
&lt;li&gt;Sales seasons. Black Friday, Cyber Monday, end-of-quarter B2B closing windows, holiday gift-giving periods. The same outage that costs $1,000 in March can cost ten times that in late November.&lt;/li&gt;
&lt;li&gt;Customer time zones. A US-only B2C site cares about evenings and weekends. A B2B SaaS serving global customers cares about Monday-morning hours rotating across timezones. Tailor monitoring intensity to when the cost actually concentrates.&lt;/li&gt;
&lt;li&gt;Outage clustering. The first outage of the month is expensive. The fourth outage of the month is exponentially more expensive in churn risk and trust damage, even if each individual outage is short. Cumulative impact compounds.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What the famous numbers do get right
&lt;/h2&gt;

&lt;p&gt;The famous numbers exaggerate for most sites, but they are not wrong about everything. Three observations they capture correctly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Downtime cost is non-linear. A two-minute outage is cheaper per minute than a two-hour outage. Customer abandonment, social-media amplification, and trust erosion accelerate as outage duration grows. The Gartner and ITIC surveys implicitly capture this by reporting on incidents large enough to be measurable; short blips do not show up in their data because they do not break things badly enough to count.&lt;/li&gt;
&lt;li&gt;The indirect costs are real, even when they are uncomfortable to measure. Productivity loss from incident response, opportunity cost of delayed feature work, sales-cycle elongation. All of these show up somewhere on the P&amp;amp;L even when nobody books them as a downtime cost. The big surveys at least try to count them.&lt;/li&gt;
&lt;li&gt;Recurring outages carry a multiplier on top of the per-incident cost. Customers tolerate one bad day. Three bad days in a quarter triggers renegotiation, churn, and lost deals that none of the single-incident calculations capture. The enterprise IT operations behind the famous numbers absolutely have data on this multiplier effect, even when their headline figure obscures it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Hold both ideas at once. The famous figures are wrong for you in magnitude. They are right in shape.&lt;/p&gt;

&lt;p&gt;The famous numbers are not lies; they are answers to a different question. The right calculation for you uses your traffic, your conversion math, your team's hourly cost during incidents, and your honest read on customer trust. Most sites land far below the headline numbers, and a few sites land far above them. Find which one you are before deciding what to fund.&lt;/p&gt;

&lt;p&gt;Or, more usefully: invest enough in uptime monitoring and incident management to make the calculation irrelevant most months. Then the only number you really need is zero.&lt;br&gt;
&lt;a href="https://webpixie.io/" rel="noopener noreferrer"&gt;See More&lt;/a&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>monitoring</category>
      <category>incident</category>
      <category>uptime</category>
    </item>
    <item>
      <title>robots.txt vs llms.txt vs sitemap.xml: what each is for</title>
      <dc:creator>WebPixie</dc:creator>
      <pubDate>Sat, 15 Aug 2026 10:12:25 +0000</pubDate>
      <link>https://dev.to/webpixie/robotstxt-vs-llmstxt-vs-sitemapxml-what-each-is-for-8c1</link>
      <guid>https://dev.to/webpixie/robotstxt-vs-llmstxt-vs-sitemapxml-what-each-is-for-8c1</guid>
      <description>&lt;p&gt;Three small files with three different jobs. robots.txt controls who crawls, sitemap.xml maps what to crawl, and llms.txt curates what AI assistants read first. Get the mental model right with minimum-viable examples and the mistakes that confuse them.&lt;br&gt;
robots.txt is a fence, sitemap.xml is a map, llms.txt is a brief. Three small files, three different jobs. robots.txt controls who is allowed to crawl your site. sitemap.xml tells search crawlers what is worth crawling. llms.txt is meant to brief AI assistants on what to read first about you. The files are not interchangeable, and most online comparisons blur the boundary between them.&lt;/p&gt;

&lt;p&gt;Get the metaphors right and the rest falls into place.&lt;/p&gt;
&lt;h2&gt;
  
  
  robots.txt is a fence
&lt;/h2&gt;

&lt;p&gt;robots.txt sits at the root of your site at /robots.txt and tells well-behaved bots which paths they are allowed to fetch. It is an access policy expressed as a plain-text allowlist or denylist. It is voluntary; malicious crawlers ignore it. Major search engines and the public AI crawlers respect it. The protocol was standardized as RFC 9309 in 2022, formalizing decades of de-facto convention.&lt;/p&gt;

&lt;p&gt;A minimum-viable robots.txt looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: *
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /&lt;span class="n"&gt;admin&lt;/span&gt;/
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /&lt;span class="n"&gt;api&lt;/span&gt;/&lt;span class="n"&gt;internal&lt;/span&gt;/

&lt;span class="n"&gt;User&lt;/span&gt;-&lt;span class="n"&gt;agent&lt;/span&gt;: &lt;span class="n"&gt;GPTBot&lt;/span&gt;
&lt;span class="n"&gt;Disallow&lt;/span&gt;: /

&lt;span class="n"&gt;Sitemap&lt;/span&gt;: &lt;span class="n"&gt;https&lt;/span&gt;://&lt;span class="n"&gt;example&lt;/span&gt;.&lt;span class="n"&gt;com&lt;/span&gt;/&lt;span class="n"&gt;sitemap&lt;/span&gt;.&lt;span class="n"&gt;xml&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This file allows every bot except into /admin/ and /api/internal/, blocks OpenAI GPTBot entirely, and points crawlers at the sitemap. The Sitemap directive is the one bridge between robots.txt and sitemap.xml. The two files are otherwise separate concerns.&lt;/p&gt;

&lt;p&gt;What robots.txt does not do:&lt;/p&gt;

&lt;p&gt;It does not remove pages from search indexes. Disallow blocks crawling; it does not deindex content already known to a search engine. Use the noindex meta tag or HTTP header for that.&lt;br&gt;
It does not protect private content. The file is publicly readable. Anyone can fetch /robots.txt and read the paths you are hiding. Authentication is for protection; robots.txt is for crawl politeness.&lt;br&gt;
It does not bind crawlers that ignore it. Scrapers, spam bots, and many AI-training crawlers fetch your content regardless. Server-side blocking by user-agent or IP is the only enforcement layer.&lt;br&gt;
It does not control rendering safely. Disallowing the path of a JavaScript file used by indexed pages can confuse search engines about page content. Disallow paths carefully on JS-heavy sites.&lt;br&gt;
It does not control AI training as a single policy stance. Blocking GPTBot stops the OpenAI training crawler, but ChatGPT can still cite your site via a separate live-browsing bot and other AI products use different bot names. The relationship between training and retrieval is per-vendor.&lt;/p&gt;
&lt;h2&gt;
  
  
  sitemap.xml is a map
&lt;/h2&gt;

&lt;p&gt;sitemap.xml lives anywhere on your domain (commonly /sitemap.xml) and tells search crawlers which URLs you want them to consider. Each entry can include a last-modified date, a priority hint, and a change-frequency hint. The format is defined by the sitemap protocol at sitemaps.org. It is the canonical way to surface URLs that crawlers might not find through internal linking.&lt;/p&gt;

&lt;p&gt;A minimum-viable sitemap looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?xml version="1.0" encoding="UTF-8"?&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;urlset&lt;/span&gt; &lt;span class="na"&gt;xmlns=&lt;/span&gt;&lt;span class="s"&gt;"http://www.sitemaps.org/schemas/sitemap/0.9"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;url&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;loc&amp;gt;&lt;/span&gt;https://example.com/&lt;span class="nt"&gt;&amp;lt;/loc&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;lastmod&amp;gt;&lt;/span&gt;2026-05-21&lt;span class="nt"&gt;&amp;lt;/lastmod&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/url&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;url&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;loc&amp;gt;&lt;/span&gt;https://example.com/blog/welcome&lt;span class="nt"&gt;&amp;lt;/loc&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;lastmod&amp;gt;&lt;/span&gt;2026-05-10&lt;span class="nt"&gt;&amp;lt;/lastmod&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/url&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/urlset&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What sitemap.xml does not do:&lt;/p&gt;

&lt;p&gt;It does not force indexing. Listing a URL in a sitemap is a suggestion, not a command. Search engines decide what to index based on content quality, internal linking, and many other signals.&lt;br&gt;
It does not replace internal linking. Pages that nothing on your site links to are weak candidates for indexing even when they appear in the sitemap. Site structure carries more weight than a sitemap entry.&lt;br&gt;
It does not honor lastmod lies. Updating every lastmod to today on every crawl makes the file untrustworthy and is widely discounted by major search engines. Accurate dates only.&lt;br&gt;
It does not target AI assistants in any specific way. Some AI crawlers use sitemaps to discover URLs, but the format was designed for search-engine indexing, not for LLM retrieval.&lt;/p&gt;
&lt;h2&gt;
  
  
  llms.txt is a brief
&lt;/h2&gt;

&lt;p&gt;llms.txt is the newest of the three, introduced in 2024 with the spec hosted at llmstxt.org. It is a single markdown file at /llms.txt that gives AI assistants a curated, structured view of your most important content. The point is to save the LLM from crawling and synthesizing the entire site every time someone asks about you, by pointing it at a short list of pages that already say what matters.&lt;/p&gt;

&lt;p&gt;A minimum-viable llms.txt looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Example.com&lt;/span&gt;
&lt;span class="gt"&gt;
&amp;gt; A short paragraph describing what Example.com is, the audience it serves, and the kind of questions an AI assistant should be able to answer from this site.&lt;/span&gt;

&lt;span class="gu"&gt;## Docs&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Getting started&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/start.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: The 5-minute onboarding guide.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;API reference&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/docs/api.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: The full HTTP API surface.

&lt;span class="gu"&gt;## Blog&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Why we added llms.txt&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/blog/llms-txt.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Our reasoning and implementation.

&lt;span class="gu"&gt;## Optional&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Detailed pricing&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://example.com/pricing.md&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;: Plan limits and feature matrix.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sections under H2 headings are required (Docs, Blog, etc.). The Optional section is treated as lower priority and may be skipped by an LLM with limited context. Linked targets are usually served as markdown variants (page.md alongside page.html) for cleaner LLM ingestion, although the standard does not require it.&lt;/p&gt;

&lt;p&gt;What llms.txt does not do:&lt;/p&gt;

&lt;p&gt;It is not robots.txt. robots.txt controls access; llms.txt curates what to read. An LLM that respects llms.txt still respects robots.txt for crawl permission.&lt;br&gt;
It is not a sitemap. A sitemap aims for completeness for indexers; llms.txt aims for curation for synthesizers. Listing every page defeats the point.&lt;br&gt;
It does not bind any specific model. Adoption is voluntary and uneven. Some AI products use llms.txt as a primary source, others ignore it entirely. Treat it as an opportunity to be cited, not a guarantee.&lt;br&gt;
It does not replace the linked pages. The LLM still fetches the targets you list. Make those destinations high-signal; a broken or thin linked page wastes the brief.&lt;br&gt;
For a closer look at llms.txt on its own, what the file is, the format, and whether AI crawlers actually read it, see llms.txt and AI crawler indexability.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the three files interact in practice
&lt;/h2&gt;

&lt;p&gt;Crawlers and AI assistants read combinations of the files in different orders, depending on what they are doing:&lt;/p&gt;

&lt;p&gt;A search crawler (Googlebot, Bingbot) fetches robots.txt first to check permission, then sitemap.xml to discover URLs, then crawls the discovered URLs. robots.txt is the gate; sitemap.xml is the hint sheet.&lt;br&gt;
An AI training crawler (GPTBot, ClaudeBot, Google-Extended) fetches robots.txt to check whether crawling is allowed at all. If allowed, it typically does not use llms.txt; it crawls from the homepage or sitemap.&lt;br&gt;
An AI assistant doing live retrieval (the ChatGPT browser, Perplexity, Gemini grounding) may read llms.txt when it exists, but adoption is voluntary and uneven, so most retrieval still works from the sitemap or live search. The file is designed to be read first; whether an assistant honors that is up to the vendor.&lt;br&gt;
Conflicts are rare because the files address different layers, but a few are worth watching for:&lt;/p&gt;

&lt;p&gt;robots.txt blocks a path that llms.txt links to. The AI assistant has the brief but cannot fetch the target. Audit your llms.txt against your robots.txt before shipping.&lt;br&gt;
sitemap.xml includes URLs that robots.txt disallows. Major search engines treat the disallow as authoritative; the sitemap entry is wasted. Remove or update one to match the other.&lt;br&gt;
llms.txt lists pages that no longer exist. Stale 404s in llms.txt train an LLM that your site is unreliable. Generate llms.txt from a source of truth, not by hand at file-rotation time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which to ship first
&lt;/h2&gt;

&lt;p&gt;Most sites already have a robots.txt; if you do not, ship one. It is the smallest, oldest, and least optional of the three. The second priority depends on what you want.&lt;/p&gt;

&lt;p&gt;If discoverability in search is the goal, prioritize sitemap.xml. It is well-supported, low-risk, and search-engine-friendly. llms.txt can come later.&lt;br&gt;
If citation by AI assistants is part of your strategy, the page fundamentals and an accurate sitemap do more than llms.txt; treat the file as optional rather than a citation lever.&lt;br&gt;
If you have the time, ship both. They are not mutually exclusive, and the marginal cost of generating one when you already have the other is small. Most static-site generators have plugins for both.&lt;br&gt;
Three files, three jobs, one mental model. robots.txt fences off what crawlers should not touch. sitemap.xml maps what they should consider. llms.txt briefs AI assistants on what matters most. Ship them in that order, keep them in sync, and they do their work without anyone reading another comparison post on the subject.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://webpixie.io" rel="noopener noreferrer"&gt;See More&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Why Cloudflare Can Make Your Uptime Monitor Lie to You</title>
      <dc:creator>WebPixie</dc:creator>
      <pubDate>Fri, 07 Aug 2026 21:10:28 +0000</pubDate>
      <link>https://dev.to/webpixie/why-cloudflare-can-make-your-uptime-monitor-lie-to-you-d7g</link>
      <guid>https://dev.to/webpixie/why-cloudflare-can-make-your-uptime-monitor-lie-to-you-d7g</guid>
      <description>&lt;p&gt;A “website down” alert does not always mean your server is actually down.&lt;/p&gt;

&lt;p&gt;If your site is behind Cloudflare, your monitoring request can fail because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cloudflare blocks the monitoring bot&lt;/li&gt;
&lt;li&gt;A WAF rule returns 403&lt;/li&gt;
&lt;li&gt;Rate limiting returns 429&lt;/li&gt;
&lt;li&gt;Cloudflare itself is experiencing an outage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Meanwhile, your origin server may still be perfectly healthy.&lt;/p&gt;

&lt;p&gt;One useful signal is the &lt;code&gt;cf-mitigated&lt;/code&gt; response header. If it is present, Cloudflare intervened before the request reached your server.&lt;/p&gt;

&lt;p&gt;The more reliable approach is to distinguish between:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Cloudflare blocking the monitor&lt;/li&gt;
&lt;li&gt;Cloudflare having an outage&lt;/li&gt;
&lt;li&gt;Your origin actually being down&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We broke down how to diagnose each case — and three ways to reduce false downtime alerts:&lt;/p&gt;

&lt;p&gt;👉 &lt;a href="https://webpixie.io/blog/post/cloudflare-false-downtime-alerts" rel="noopener noreferrer"&gt;https://webpixie.io/blog/post/cloudflare-false-downtime-alerts&lt;/a&gt;&lt;/p&gt;

</description>
      <category>cloudflarechallenge</category>
      <category>devops</category>
      <category>monitoring</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How Often Should You Check Website Uptime?</title>
      <dc:creator>WebPixie</dc:creator>
      <pubDate>Thu, 06 Aug 2026 20:50:13 +0000</pubDate>
      <link>https://dev.to/webpixie/how-often-should-you-check-website-uptime-bg6</link>
      <guid>https://dev.to/webpixie/how-often-should-you-check-website-uptime-bg6</guid>
      <description>&lt;p&gt;Checking a website every 30 seconds sounds better than checking it every 5 minutes.&lt;/p&gt;

&lt;p&gt;But the fastest interval is not always the right one.&lt;/p&gt;

&lt;p&gt;A checkout endpoint, a SaaS application, a marketing website, an SSL certificate, and a domain registration all have different failure patterns and levels of impact.&lt;/p&gt;

&lt;p&gt;The right monitoring frequency depends on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How quickly the failure affects users&lt;/li&gt;
&lt;li&gt;How much downtime the service can tolerate&lt;/li&gt;
&lt;li&gt;How retries and verification work&lt;/li&gt;
&lt;li&gt;How quickly the team can respond&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We explored practical monitoring intervals for websites, APIs, checkout flows, SSL, DNS, and domains in our latest guide:&lt;/p&gt;

&lt;p&gt;👉 Blog: &lt;a href="https://webpixie.io/blog/post/uptime-monitoring-frequency" rel="noopener noreferrer"&gt;https://webpixie.io/blog/post/uptime-monitoring-frequency&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;👉 Free Uptime &amp;amp; Downtime Calculator:&lt;a href="https://webpixie.io/free-tools/uptime-calculator" rel="noopener noreferrer"&gt;https://webpixie.io/free-tools/uptime-calculator&lt;/a&gt;&lt;/p&gt;

</description>
      <category>monitoring</category>
      <category>dns</category>
      <category>ssl</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
