<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Edward Chapman</title>
    <description>The latest articles on DEV Community by Edward Chapman (@edchapman).</description>
    <link>https://dev.to/edchapman</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156848%2F7d04bf9c-7217-4ee3-a954-71e2752d79c8.png</url>
      <title>DEV Community: Edward Chapman</title>
      <link>https://dev.to/edchapman</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/edchapman"/>
    <language>en</language>
    <item>
      <title>Robots.txt allows OAI-SearchBot. That does not prove ChatGPT can fetch your page.</title>
      <dc:creator>Edward Chapman</dc:creator>
      <pubDate>Sat, 03 Oct 2026 07:37:20 +0000</pubDate>
      <link>https://dev.to/edchapman/robotstxt-allows-oai-searchbot-that-does-not-prove-chatgpt-can-fetch-your-page-2e8l</link>
      <guid>https://dev.to/edchapman/robotstxt-allows-oai-searchbot-that-does-not-prove-chatgpt-can-fetch-your-page-2e8l</guid>
      <description>&lt;p&gt;A line in robots.txt that allows OAI-SearchBot grants permission to a compliant crawler. It does not tell you whether your CDN lets the request through, whether the response contains your content, or whether the page appears in ChatGPT answers. Each question needs its own evidence.&lt;/p&gt;

&lt;p&gt;Two OpenAI user agents, two decisions&lt;/p&gt;

&lt;p&gt;OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for training. Blocking GPTBot alone does not automatically block OAI-SearchBot. To allow search crawling and disallow training crawling, use separate groups:&lt;/p&gt;

&lt;p&gt;User-agent: OAI-SearchBot&lt;br&gt;
Allow: /&lt;/p&gt;

&lt;p&gt;User-agent: GPTBot&lt;br&gt;
Disallow: /&lt;/p&gt;

&lt;p&gt;These rules state crawler permissions. They do not override firewall rules or guarantee training exclusion in every context. Check OpenAI's documentation for the current scope of each agent:&lt;br&gt;
&lt;a href="https://platform.openai.com/docs/bots" rel="noopener noreferrer"&gt;https://platform.openai.com/docs/bots&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Check the rule that applies&lt;/p&gt;

&lt;p&gt;Start with the hostname. The file at example.com/robots.txt does not govern &lt;a href="http://www.example.com" rel="noopener noreferrer"&gt;www.example.com&lt;/a&gt; or shop.example.com.&lt;/p&gt;

&lt;p&gt;Then check the exact path. Plain path rules match by prefix: Disallow: /pricing also matches /pricing-old and /pricing/eu.&lt;/p&gt;

&lt;p&gt;Check the user-agent group too. Under standard matching, a crawler that finds a group naming it follows that group rather than falling back to User-agent: *. Review all groups for the named agent.&lt;/p&gt;

&lt;p&gt;Permission to crawl is not indexing or citation. Robots.txt manages crawler traffic; it is not authentication or access control. Anyone can read the file and ignore it, so confidential data needs proper access controls.&lt;/p&gt;

&lt;p&gt;Google explains the scope and limitations here:&lt;br&gt;
&lt;a href="https://developers.google.com/search/docs/crawling-indexing/robots/intro" rel="noopener noreferrer"&gt;https://developers.google.com/search/docs/crawling-indexing/robots/intro&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The edge can refuse first&lt;/p&gt;

&lt;p&gt;A CDN or firewall can block or challenge a permitted request before it reaches your server. Your origin logs may show nothing about that request.&lt;/p&gt;

&lt;p&gt;Review edge security events and origin logs separately. Record the hostname, path, time, action or status, and the rule that matched. If a rule catches traffic you meant to allow, adjust that specific rule. Do not disable site-wide protection to troubleshoot one crawler.&lt;/p&gt;

&lt;p&gt;Cloudflare's documentation, checked on 2 October 2026, separates Search, Agent and Training categories. An Allow setting in the AI bot policy does not remove other WAF controls. Published defaults do not tell you how an existing zone is configured. Check your own settings:&lt;br&gt;
&lt;a href="https://developers.cloudflare.com/bots/additional-configurations/block-ai-bots/" rel="noopener noreferrer"&gt;https://developers.cloudflare.com/bots/additional-configurations/block-ai-bots/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A copied user agent tests your request, not theirs&lt;/p&gt;

&lt;p&gt;Running curl with OAI-SearchBot's user-agent string sends a request from your machine and IP address. If the edge uses IP ranges or verified-bot status, it may respond differently to the real crawler.&lt;/p&gt;

&lt;p&gt;A 200 from your laptop does not prove the crawler gets a 200. A 403 does not prove it gets a 403. The User-Agent header comes from the client, so it does not verify identity.&lt;/p&gt;

&lt;p&gt;Use the verification information OpenAI currently publishes and compare it with your edge request records. If no verified requests appear in the period you checked, report that they were not observed. That does not establish that they were blocked.&lt;/p&gt;

&lt;p&gt;A 200 is not necessarily your content&lt;/p&gt;

&lt;p&gt;A 200 response can contain a login page, a challenge interstitial, or a JavaScript shell with little text. Where response evidence is available for a verified crawler request, compare what was served with the expected public page.&lt;/p&gt;

&lt;p&gt;Status and response size are clues, not proof that the intended text was delivered. Standard access logs usually do not contain response bodies. Even confirmed delivery does not prove the provider ingested the page.&lt;/p&gt;

&lt;p&gt;Keep four findings separate&lt;/p&gt;

&lt;p&gt;Declared permission: the matching robots.txt group and path rule. This does not show that a request arrived.&lt;/p&gt;

&lt;p&gt;Observed access: edge or origin evidence of a verified crawler request and its response. This does not show that the content was indexed or used.&lt;/p&gt;

&lt;p&gt;Provider indexing: signals the provider publishes, where available. This does not promise a citation.&lt;/p&gt;

&lt;p&gt;Cited answer: an actual answer linking the page. This does not show that the citation will recur or appear for other queries.&lt;/p&gt;

&lt;p&gt;A practical checklist&lt;/p&gt;

&lt;p&gt;Read robots.txt on the exact hostname and find the group and path rule for OAI-SearchBot.&lt;/p&gt;

&lt;p&gt;Check GPTBot separately against your training decision.&lt;/p&gt;

&lt;p&gt;Review edge security events for that hostname and path, then origin logs.&lt;/p&gt;

&lt;p&gt;Verify crawler identity using provider documentation, not the User-Agent string alone.&lt;/p&gt;

&lt;p&gt;For verified requests, record the action, status, matching rule, and available evidence of the content served.&lt;/p&gt;

&lt;p&gt;Report anything you could not confirm as unknown. Access evidence alone cannot establish that ChatGPT cited a page.&lt;/p&gt;

&lt;p&gt;An optional check for the robots.txt part&lt;/p&gt;

&lt;p&gt;I'm with Firm Beacon. Our free browser checker compares robots.txt permissions for search and training agents on a specific URL. You can enter a URL or paste robots.txt without creating an account or giving an email address:&lt;br&gt;
&lt;a href="https://www.firmbeacon.co.uk/tools/ai-crawler-check?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=ai_crawler_access" rel="noopener noreferrer"&gt;https://www.firmbeacon.co.uk/tools/ai-crawler-check?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=ai_crawler_access&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It checks the declared rules, not your firewall. It does not verify indexing, citations or rankings. Use it for the first step, then check access evidence separately.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
    </item>
    <item>
      <title>What to Check Before You Spend More Time on SEO</title>
      <dc:creator>Edward Chapman</dc:creator>
      <pubDate>Sat, 03 Oct 2026 03:27:45 +0000</pubDate>
      <link>https://dev.to/edchapman/what-to-check-before-you-spend-more-time-on-seo-2dpi</link>
      <guid>https://dev.to/edchapman/what-to-check-before-you-spend-more-time-on-seo-2dpi</guid>
      <description>&lt;p&gt;When a page does not perform in search, it is tempting to write more content or start building links. Those may be useful later, but first check whether the site is sending clear technical signals to crawlers and browsers.&lt;/p&gt;

&lt;p&gt;A small audit will not explain every ranking change. It can show whether basic page information is present, whether important files are reachable, and whether the sample of pages you checked is consistent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the site-level files
&lt;/h2&gt;

&lt;p&gt;Begin with HTTPS, &lt;code&gt;robots.txt&lt;/code&gt;, and the XML sitemap.&lt;/p&gt;

&lt;p&gt;HTTPS is the expected protocol for a public website. Check that the address you use is the HTTPS version and that it loads as a public page.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;robots.txt&lt;/code&gt; contains crawler rules. A rule can block a path that you intended to make public, so read the file instead of assuming it is correct. The XML sitemap is another useful signal to inspect. It should be available at the address declared by the site and should represent the URLs you actually want discovered.&lt;/p&gt;

&lt;p&gt;These files have different jobs. &lt;code&gt;robots.txt&lt;/code&gt; controls crawler access to paths. A sitemap lists URLs for discovery. Neither one proves that a page is indexed or will rank.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the pages that matter
&lt;/h2&gt;

&lt;p&gt;The homepage is a sensible starting point, but it is not enough. Include a product or service page, an important article, and a page that has changed recently. If the site is larger, use a sample rather than assuming that every template behaves the same way.&lt;/p&gt;

&lt;p&gt;The free &lt;a href="https://www.firmbeacon.co.uk/tools/website-audit?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=website_audit_workflow" rel="noopener noreferrer"&gt;Firm Beacon Website Audit&lt;/a&gt; accepts a public website address and checks up to ten public HTML pages. It follows public links on the same site, so the result is a sample of what the audit can reach from the starting page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review the page signals
&lt;/h2&gt;

&lt;p&gt;For each page, look at the title, meta description, and H1 heading. These elements should describe the page that the visitor sees. A missing value is a reason to inspect the template or CMS settings. It is not, by itself, proof of a ranking penalty.&lt;/p&gt;

&lt;p&gt;Check the canonical URL as well. It should point to the version of the page that you want search engines to treat as canonical. Pay attention to the hostname and protocol. A page using &lt;code&gt;https://www.example.com&lt;/code&gt; and a canonical pointing to a different version deserves a closer look.&lt;/p&gt;

&lt;p&gt;The audit also reports the robots meta directive and the &lt;code&gt;X-Robots-Tag&lt;/code&gt; response header. These can contain &lt;code&gt;noindex&lt;/code&gt; instructions. If a page is meant to appear in search, confirm that an exclusion was intentional.&lt;/p&gt;

&lt;p&gt;Language and mobile viewport tags are easy to overlook. The language value gives a page a declared language, while the viewport tag helps a browser size the page for a mobile screen. Open Graph tags control how a page is represented when someone shares it on a social platform. They are useful checks even though they do not establish search rankings.&lt;/p&gt;

&lt;p&gt;Internal links matter for navigation and discovery. A small audit can show the links present in the public HTML it reads. Follow up on important pages that are hard to reach from the pages you checked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the result in context
&lt;/h2&gt;

&lt;p&gt;A technical HTML audit has clear limits. It does not verify Google indexing, rankings, backlinks, Core Web Vitals, firewall access, or JavaScript-rendered content. A page can contain a title and canonical and still have a separate indexing problem. A page can also lack one optional signal without needing an urgent rewrite.&lt;/p&gt;

&lt;p&gt;Treat each finding as a question:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the value missing on purpose?&lt;/li&gt;
&lt;li&gt;Does it match the page's purpose and preferred URL?&lt;/li&gt;
&lt;li&gt;Does the same issue appear in the other templates you sampled?&lt;/li&gt;
&lt;li&gt;Can you confirm the result in the CMS, response headers, or Search Console?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This keeps a short audit from turning into a list of automatic fixes. The right change depends on the page and the site's publishing setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  A short checklist
&lt;/h2&gt;

&lt;p&gt;Before changing content or starting outreach, check:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The public address loads over HTTPS.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;robots.txt&lt;/code&gt; does not block a path you need crawled.&lt;/li&gt;
&lt;li&gt;The sitemap is reachable and contains intended public URLs.&lt;/li&gt;
&lt;li&gt;Important pages have useful titles, descriptions, and H1 headings.&lt;/li&gt;
&lt;li&gt;Canonical URLs use the intended hostname and protocol.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;noindex&lt;/code&gt; directives are deliberate.&lt;/li&gt;
&lt;li&gt;The pages declare language, viewport, and sharing information where needed.&lt;/li&gt;
&lt;li&gt;Internal links reach the pages you want visitors and crawlers to find.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can run this first pass with the free &lt;a href="https://www.firmbeacon.co.uk/tools/website-audit?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=website_audit_workflow" rel="noopener noreferrer"&gt;Website Audit&lt;/a&gt;. It requires no account or email and reports the signals found in up to ten public HTML pages. Use the result to choose what to investigate next, then confirm important changes with the tools that cover indexing, performance, or server access.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How to create an XML sitemap from a URL list</title>
      <dc:creator>Edward Chapman</dc:creator>
      <pubDate>Fri, 02 Oct 2026 17:52:32 +0000</pubDate>
      <link>https://dev.to/edchapman/how-to-create-an-xml-sitemap-from-a-url-list-2o0n</link>
      <guid>https://dev.to/edchapman/how-to-create-an-xml-sitemap-from-a-url-list-2o0n</guid>
      <description>&lt;p&gt;An XML sitemap is a list of public URLs that you want search engines to discover. It can help with discovery, but it does not guarantee crawling, indexing or rankings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the sitemap your site already has
&lt;/h2&gt;

&lt;p&gt;If your site uses a CMS, check whether it already generates a sitemap. WordPress, Wix and other platforms can maintain one as pages are added or removed. Use that sitemap rather than creating a second file that will go stale.&lt;/p&gt;

&lt;p&gt;A manual sitemap makes sense when you have a small site, a curated URL list, or a build process that produces the final URLs separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the URLs
&lt;/h2&gt;

&lt;p&gt;Include canonical, public pages that you want search engines to consider. Do not add:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;login pages or other private URLs;&lt;/li&gt;
&lt;li&gt;URLs that redirect somewhere else;&lt;/li&gt;
&lt;li&gt;duplicate versions of the same page;&lt;/li&gt;
&lt;li&gt;pages you deliberately exclude from search;&lt;/li&gt;
&lt;li&gt;broken or temporary URLs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use absolute URLs, and keep the hostname consistent. For example, decide whether the site uses &lt;code&gt;https://example.com&lt;/code&gt; or &lt;code&gt;https://www.example.com&lt;/code&gt; and do not mix both versions in one sitemap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write the XML
&lt;/h2&gt;

&lt;p&gt;The basic sitemap format is small. Each page goes inside a &lt;code&gt;&amp;lt;url&amp;gt;&lt;/code&gt; element, and its address goes inside &lt;code&gt;&amp;lt;loc&amp;gt;&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="cp"&gt;&amp;lt;?xml version=1.0 encoding=UTF-8?&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;urlset&lt;/span&gt; &lt;span class="na"&gt;xmlns=&lt;/span&gt;&lt;span class="s"&gt;http://www.sitemaps.org/schemas/sitemap/0.9&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;url&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;loc&amp;gt;&lt;/span&gt;https://example.com/&lt;span class="nt"&gt;&amp;lt;/loc&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/url&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;url&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;loc&amp;gt;&lt;/span&gt;https://example.com/about&lt;span class="nt"&gt;&amp;lt;/loc&amp;gt;&lt;/span&gt;
  &lt;span class="nt"&gt;&amp;lt;/url&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/urlset&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Escape characters that have a special meaning in XML. For example, an ampersand in a query string becomes &lt;code&gt;&amp;amp;amp;&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;loc&amp;gt;&lt;/span&gt;https://example.com/search?a=1&lt;span class="ni"&gt;&amp;amp;amp;&lt;/span&gt;b=2&lt;span class="nt"&gt;&amp;lt;/loc&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You do not need to add &lt;code&gt;priority&lt;/code&gt; or &lt;code&gt;changefreq&lt;/code&gt;. Google ignores those fields. Add &lt;code&gt;lastmod&lt;/code&gt; only when the date reflects a real, significant update to that URL. Invented dates make the file less useful, so it is better to leave the field out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the limits
&lt;/h2&gt;

&lt;p&gt;A sitemap can contain up to 50,000 URLs or 50 MB when uncompressed. Larger sites can split the URLs across several sitemap files and reference them from a sitemap index.&lt;/p&gt;

&lt;p&gt;For a small, already curated list, a browser-local converter is enough. The &lt;a href="https://www.firmbeacon.co.uk/tools/sitemap-generator?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=sitemap_guide" rel="noopener noreferrer"&gt;free XML sitemap generator&lt;/a&gt; removes duplicate URLs and fragments, escapes XML characters and downloads &lt;code&gt;sitemap.xml&lt;/code&gt;. It does not crawl the site or upload the pasted list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Publish and submit it
&lt;/h2&gt;

&lt;p&gt;Upload the file to the public location you choose, commonly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://example.com/sitemap.xml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then check the URL in a browser or with an HTTP client. It should return the XML file, not an HTML error page or a login screen.&lt;/p&gt;

&lt;p&gt;You can also declare the location in &lt;code&gt;robots.txt&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Sitemap: https://example.com/sitemap.xml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a site you manage, submit the sitemap URL in Google Search Console. Keep the sitemap current when public URLs are added, removed or replaced.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a sitemap cannot tell you
&lt;/h2&gt;

&lt;p&gt;A successful upload does not prove that Google has crawled or indexed every listed page. It also does not improve a weak page, override a canonical decision, or guarantee rankings. Review indexing evidence separately in Search Console and check the pages themselves for useful content, access, canonical URLs and appropriate robots directives.&lt;/p&gt;

&lt;p&gt;The useful test is simple: include the right public URLs, publish valid XML at a stable address, and then verify what search engines actually did with it.&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
      <category>javascript</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
