<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anosh</title>
    <description>The latest articles on DEV Community by Anosh (@marketingwithanosh).</description>
    <link>https://dev.to/marketingwithanosh</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167173%2F26de9bad-3ca2-4860-ba3c-6cfcdac86cb6.png</url>
      <title>DEV Community: Anosh</title>
      <link>https://dev.to/marketingwithanosh</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/marketingwithanosh"/>
    <language>en</language>
    <item>
      <title>Structured Data: How Schema Markup Actually Helps Search Engines Understand Pages</title>
      <dc:creator>Anosh</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:11:00 +0000</pubDate>
      <link>https://dev.to/marketingwithanosh/structured-data-how-schema-markup-actually-helps-search-engines-understand-pages-37pl</link>
      <guid>https://dev.to/marketingwithanosh/structured-data-how-schema-markup-actually-helps-search-engines-understand-pages-37pl</guid>
      <description>&lt;p&gt;"Add schema to get rich snippets" is the most common way structured data gets explained, and it's backwards. Rich results are a possible side effect. The real job of schema is simpler: it helps a machine understand what a page is about.&lt;/p&gt;

&lt;p&gt;Search engines are good at reading text, but text is ambiguous. Is "Apple" a fruit or a company? Is "Jordan" a country or a person? Is "$49" a price, a fee or a random number? Schema removes that guesswork.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model
&lt;/h2&gt;

&lt;p&gt;Think of it as a pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Page content
    ↓
Structured entities (things)
    ↓
Machine-readable relationships
    ↓
Search engine understanding
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Page content:&lt;/strong&gt; the words, images and numbers a human reads&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Entities:&lt;/strong&gt; the distinct things the page is about, like a company, a person, a product or an event&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relationships:&lt;/strong&gt; how those things connect (this article was written by this person, who works for this organization)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Understanding:&lt;/strong&gt; the search engine can place your page in its knowledge of the world, instead of guessing from keywords&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Schema.org and JSON-LD
&lt;/h2&gt;

&lt;p&gt;These two get mixed up constantly, but they're different things.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Schema.org&lt;/strong&gt; is the shared vocabulary. It defines types (&lt;code&gt;Product&lt;/code&gt;, &lt;code&gt;Event&lt;/code&gt;) and properties (&lt;code&gt;price&lt;/code&gt;, &lt;code&gt;startDate&lt;/code&gt;). Think of it as the dictionary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JSON-LD&lt;/strong&gt; is the format you write it in. Think of it as the sentence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Google recommends JSON-LD because it lives in a separate &lt;code&gt;&amp;lt;script&amp;gt;&lt;/code&gt; block, so it doesn't tangle with your visible HTML:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;script &lt;/span&gt;&lt;span class="na"&gt;type=&lt;/span&gt;&lt;span class="s"&gt;"application/ld+json"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@context&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://schema.org&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Organization&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;name&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Example Co&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;url&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://example.com&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;logo&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://example.com/logo.png&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;sameAs&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://www.linkedin.com/company/example&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
  &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/script&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few core ideas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;@context&lt;/code&gt; says which vocabulary you're using&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;@type&lt;/code&gt; says what kind of thing this is&lt;/li&gt;
&lt;li&gt;Properties describe it&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;@id&lt;/code&gt; lets you give an entity a stable identifier, so other blocks can refer to it&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The types you'll actually use
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Organization
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes the business or brand behind the site&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;logo&lt;/code&gt;, &lt;code&gt;sameAs&lt;/code&gt; (links to official profiles)&lt;/li&gt;
&lt;li&gt;It tells search engines which entity your site belongs to, and connects your profiles together&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Person
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes an individual, such as an author or founder&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;jobTitle&lt;/code&gt;, &lt;code&gt;sameAs&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Helpful for tying content to a real author&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  WebSite
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes the site as a whole&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Helps clarify your site name. Google retired the old sitelinks search box feature, so don't add it expecting that result.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Article
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes a blog post or news piece&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;headline&lt;/code&gt;, &lt;code&gt;author&lt;/code&gt;, &lt;code&gt;datePublished&lt;/code&gt;, &lt;code&gt;dateModified&lt;/code&gt;, &lt;code&gt;image&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Tells the engine who wrote it and when, and when it was last updated&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Product
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes something for sale&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;image&lt;/code&gt;, &lt;code&gt;description&lt;/code&gt;, &lt;code&gt;offers&lt;/code&gt; (price, currency, availability), &lt;code&gt;aggregateRating&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Turns a product page into clean, comparable data: what it is, what it costs, whether it's in stock&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  BreadcrumbList
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes where a page sits in your site's hierarchy&lt;/li&gt;
&lt;li&gt;Key properties: a list of items with &lt;code&gt;position&lt;/code&gt;, &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;item&lt;/code&gt; (URL)&lt;/li&gt;
&lt;li&gt;Shows how your content is organized, which helps both understanding and navigation&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  LocalBusiness
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes a business with a physical presence&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;address&lt;/code&gt;, &lt;code&gt;telephone&lt;/code&gt;, &lt;code&gt;openingHours&lt;/code&gt;, &lt;code&gt;geo&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Connects the page to a real-world location, which supports local search&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  FAQPage
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes a page of questions and answers&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;mainEntity&lt;/code&gt; with &lt;code&gt;Question&lt;/code&gt; and &lt;code&gt;acceptedAnswer&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Marks up the question-and-answer structure clearly. Note that Google now shows FAQ rich results only for a small set of well-known authoritative sites, so the value for most sites is clarity, not appearance.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Review
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes a review of something, including the rating and reviewer&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;reviewRating&lt;/code&gt;, &lt;code&gt;author&lt;/code&gt;, &lt;code&gt;itemReviewed&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Make sure the review is about something real and visible on the page. Self-serving reviews of your own business, placed on your own site, aren't eligible for review stars.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Event
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Describes something happening at a specific time and place&lt;/li&gt;
&lt;li&gt;Key properties: &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;startDate&lt;/code&gt;, &lt;code&gt;location&lt;/code&gt;, &lt;code&gt;eventAttendanceMode&lt;/code&gt;, &lt;code&gt;offers&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Gives the engine exact dates, venues and ticket info, which is hard to pull reliably from prose&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where relationships come in
&lt;/h2&gt;

&lt;p&gt;Single blocks are useful, but the real power is connecting entities. Here, an article points to its author and publisher using &lt;code&gt;@id&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"@context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://schema.org"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"@graph"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"@type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Organization"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"@id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/#org"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Example Co"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"@type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Person"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"@id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/#author"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Jane Doe"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"worksFor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"@id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/#org"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"@type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Article"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"headline"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"How Schema Works"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"author"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"@id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/#author"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"publisher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"@id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com/#org"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"datePublished"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-01"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the engine sees a small graph: an article, written by a person, who works for an organization. That's understanding, not just labelling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Valid schema does not equal a guaranteed rich result
&lt;/h2&gt;

&lt;p&gt;This is where expectations go wrong. Valid markup means your code follows the rules. It doesn't mean Google will do anything visible with it.&lt;/p&gt;

&lt;p&gt;Reasons valid schema may produce no rich result:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Eligibility.&lt;/strong&gt; Not every type has a rich result, and some have been restricted or retired.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality and trust.&lt;/strong&gt; Google decides whether your page and site deserve the enhancement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content mismatch.&lt;/strong&gt; Markup must reflect what users can see on the page. Marking up hidden or invented content breaks guidelines and can lead to a manual action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Missing required properties.&lt;/strong&gt; The markup might validate as schema but lack what a specific rich result needs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google's discretion.&lt;/strong&gt; Even eligible pages don't always get the feature on every search.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So treat the two things separately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Valid:&lt;/strong&gt; passes syntax and vocabulary checks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eligible:&lt;/strong&gt; meets Google's documented requirements for a given feature&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Displayed:&lt;/strong&gt; Google chooses to show it, sometimes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Your control stops at the first two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Marking up content that isn't visible on the page&lt;/li&gt;
&lt;li&gt;Copying a template without changing the values&lt;/li&gt;
&lt;li&gt;Duplicate or conflicting blocks from both a plugin and the theme&lt;/li&gt;
&lt;li&gt;Adding &lt;code&gt;aggregateRating&lt;/code&gt; with invented numbers&lt;/li&gt;
&lt;li&gt;Using the wrong type because it "gets stars"&lt;/li&gt;
&lt;li&gt;Treating schema as a ranking factor. It helps understanding. It isn't a shortcut to higher rankings.&lt;/li&gt;
&lt;li&gt;Setting it up once and never updating prices, dates or availability&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to test it
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Run the page through the &lt;strong&gt;Rich Results Test&lt;/strong&gt; to see which features it's eligible for&lt;/li&gt;
&lt;li&gt;Use the &lt;strong&gt;Schema Markup Validator&lt;/strong&gt; to check general schema.org syntax&lt;/li&gt;
&lt;li&gt;Watch the &lt;strong&gt;Enhancements&lt;/strong&gt; reports in Search Console for errors and warnings&lt;/li&gt;
&lt;li&gt;Compare the markup against the visible page and ask whether they say the same thing&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Schema is a translation layer between how humans read a page and how machines parse it. Done well, it makes your pages clearer, your entities more consistent and your content easier to connect to the wider web of information. Sometimes that earns a rich result, and sometimes it simply means a search engine understood you correctly.&lt;/p&gt;

&lt;p&gt;Start with the basics, which are Organization, WebSite, Article or Product depending on the page, and BreadcrumbList. Keep it accurate and keep it honest.&lt;/p&gt;

&lt;p&gt;What's the schema type that gave you the most trouble? Tell me in the comments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Suggested tags:&lt;/strong&gt; &lt;code&gt;seo&lt;/code&gt;, &lt;code&gt;webdev&lt;/code&gt;, &lt;code&gt;json&lt;/code&gt;, &lt;code&gt;beginners&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Visit my website: &lt;a href="https://anoshbb.com" rel="noopener noreferrer"&gt;anoshbb.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>html</category>
      <category>seo</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The Complete Guide to HTTP Status Codes for SEO</title>
      <dc:creator>Anosh</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:06:32 +0000</pubDate>
      <link>https://dev.to/marketingwithanosh/the-complete-guide-to-http-status-codes-for-seo-2fk7</link>
      <guid>https://dev.to/marketingwithanosh/the-complete-guide-to-http-status-codes-for-seo-2fk7</guid>
      <description>&lt;p&gt;Every time Googlebot requests a URL, your server answers with a three-digit number before anything else happens. That number decides whether the page gets indexed, replaced, dropped or crawled less often.&lt;/p&gt;

&lt;p&gt;This is a reference you can come back to. For each code, you get what it means, the SEO implication, when to use it, and the mistake people commonly make.&lt;/p&gt;

&lt;h2&gt;
  
  
  2xx: Success
&lt;/h2&gt;

&lt;h3&gt;
  
  
  200 OK
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The request worked and the content was returned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; This is the only status that makes a page eligible for indexing. Google still decides whether it's worth keeping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; The page exists and should be seen by users and crawlers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Returning 200 for pages that are really errors, like empty search results or "product not found" pages. Google calls these &lt;strong&gt;soft 404s&lt;/strong&gt; and treats them as missing anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  3xx: Redirects
&lt;/h2&gt;

&lt;h3&gt;
  
  
  301 Moved Permanently
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The URL has changed for good. Go to the new one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Google treats the destination as the new canonical home. Ranking signals pass through, and the old URL eventually drops out of the index.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A URL changes&lt;/li&gt;
&lt;li&gt;You migrate to a new domain&lt;/li&gt;
&lt;li&gt;You move from HTTP to HTTPS&lt;/li&gt;
&lt;li&gt;You restructure your URL paths&lt;/li&gt;
&lt;li&gt;You merge duplicate pages&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Redirect chains (A → B → C → D) and redirecting a batch of old URLs to the homepage. Redirect each page to its closest equivalent, in a single hop.&lt;/p&gt;

&lt;h3&gt;
  
  
  302 Found
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The URL has moved temporarily. Keep using the original.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Google generally keeps the original URL indexed. If a 302 stays in place for a long time, Google may start treating it like a permanent move anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A/B testing&lt;/li&gt;
&lt;li&gt;Short-term promotions or seasonal pages&lt;/li&gt;
&lt;li&gt;Sending users to a regional or device-specific version for a while&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Using 302 for a permanent change because it's the default in many frameworks and plugins. Check what your redirect actually returns.&lt;/p&gt;

&lt;h3&gt;
  
  
  303 See Other
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The response is somewhere else, and the client should fetch it with a GET request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Rarely relevant for crawled pages. Google treats it as a temporary redirect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; A form submission or POST request should send the user to a confirmation page, so refreshing doesn't resubmit the form.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Using 303 for regular page moves. It's meant for the post-submission flow, not for content that changed address.&lt;/p&gt;

&lt;h3&gt;
  
  
  307 Temporary Redirect
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; A temporary redirect that preserves the HTTP method. A POST stays a POST.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Treated like a 302 for search. Google keeps the original URL.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; The move is temporary and the request method matters, such as in APIs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Assuming a 307 means "permanent but safer." You'll often see 307s when a site has HSTS, where the browser itself redirects HTTP to HTTPS. That's a browser behavior, not what your server sent.&lt;/p&gt;

&lt;h3&gt;
  
  
  308 Permanent Redirect
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; A permanent redirect that preserves the HTTP method.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Treated like a 301. Signals consolidate on the new URL.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; The move is permanent and you want the method preserved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Thinking 308 is a different SEO outcome from 301. For regular page navigation they behave the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  301 vs 302 vs 307 (and 308)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;301:&lt;/strong&gt; permanent, new URL gets indexed, old one fades out&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;302:&lt;/strong&gt; temporary, original URL usually stays indexed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;307:&lt;/strong&gt; temporary like 302, but the HTTP method is preserved&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;308:&lt;/strong&gt; permanent like 301, but the HTTP method is preserved&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical rule is simple. Ask yourself whether the old URL is coming back. If &lt;strong&gt;no&lt;/strong&gt;, use 301 (or 308). If &lt;strong&gt;yes&lt;/strong&gt;, use 302 (or 307). Both kinds pass signals, so the old line that "302s don't pass PageRank" doesn't hold up. What differs is which URL Google chooses to keep in the index.&lt;/p&gt;

&lt;h2&gt;
  
  
  4xx: Client errors
&lt;/h2&gt;

&lt;h3&gt;
  
  
  404 Not Found
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; Nothing is here, and the server doesn't say whether it will return.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; The URL drops out of the index over time. A handful of 404s is normal and doesn't hurt your site overall.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; A page doesn't exist, or was removed with no replacement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Panicking over every 404 in Search Console. Fix the ones that have backlinks, internal links or real traffic, and ignore the junk.&lt;/p&gt;

&lt;h3&gt;
  
  
  410 Gone
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The page existed, was removed on purpose, and isn't coming back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; A slightly stronger signal than 404. Google may drop the URL a little faster, though the difference is small.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; You deliberately delete content permanently, like expired listings or retired campaigns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Expecting 410 to give a big SEO advantage. It's a clean way to say "gone," not a ranking trick.&lt;/p&gt;

&lt;h3&gt;
  
  
  429 Too Many Requests
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The client is sending requests too quickly. Slow down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Google treats 429 like a server error and reduces its crawl rate. If it persists, indexing can suffer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; You rate-limit abusive bots or scrapers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Rate-limiting so aggressively that Googlebot gets caught, usually through a security plugin or firewall. Check server logs if crawl stats suddenly drop.&lt;/p&gt;

&lt;h2&gt;
  
  
  5xx: Server errors
&lt;/h2&gt;

&lt;h3&gt;
  
  
  500 Internal Server Error
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; Something broke on the server, and it can't say what.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Google retries and slows its crawling. If the error continues for days, URLs can start falling out of the index.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; You shouldn't be using it on purpose. It means a bug.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Ignoring intermittent 500s because "the site loads for me." Check Search Console's crawl stats and your logs.&lt;/p&gt;

&lt;h3&gt;
  
  
  502 Bad Gateway
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; A server acting as a gateway or proxy got an invalid response from the server behind it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Same as other 5xx errors: crawl slowdown, and de-indexing if prolonged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; It's generated by infrastructure, not chosen.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Blaming the application when the issue is the proxy, load balancer or CDN configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  503 Service Unavailable
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The server is temporarily unable to handle the request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; This is the right code for planned downtime. Google understands it's temporary and will come back, especially when a &lt;code&gt;Retry-After&lt;/code&gt; header is present.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Maintenance windows&lt;/li&gt;
&lt;li&gt;Planned deployments&lt;/li&gt;
&lt;li&gt;Short-term overload&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Showing a "we'll be back soon" page with a 200 status. Google may index the maintenance page. Return 503 instead. And don't leave it up for days.&lt;/p&gt;

&lt;h3&gt;
  
  
  504 Gateway Timeout
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;What it means:&lt;/strong&gt; The gateway waited too long for the upstream server.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEO implication:&lt;/strong&gt; Treated as a server error. Slow backends lead to repeated timeouts and reduced crawling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use it when:&lt;/strong&gt; Again, infrastructure produces it, not you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Common mistake:&lt;/strong&gt; Treating it as a one-off. Repeated 504s usually point to slow database queries or overloaded servers that also hurt your Core Web Vitals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick reference
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;200:&lt;/strong&gt; indexable, but confirm it isn't a soft 404&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;301 / 308:&lt;/strong&gt; permanent move, signals consolidate on the new URL&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;302 / 307:&lt;/strong&gt; temporary move, original URL usually stays&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;303:&lt;/strong&gt; post-form redirect, rarely an SEO concern&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;404:&lt;/strong&gt; not found, drops out over time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;410:&lt;/strong&gt; deliberately gone, drops out slightly faster&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;429:&lt;/strong&gt; slow down, Google reduces crawling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;500 / 502 / 504:&lt;/strong&gt; server trouble, fix fast&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;503:&lt;/strong&gt; temporary downtime, the correct code for maintenance&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to check your own status codes
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-I&lt;/span&gt; https://example.com/old-page
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at the first line for the status and the &lt;code&gt;location&lt;/code&gt; header for redirect targets. To follow a whole chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-IL&lt;/span&gt; https://example.com/old-page
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Also use Search Console's Page indexing and Crawl stats reports, and a crawler like Screaming Frog to catch chains, loops and soft 404s across the whole site.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Status codes are the cheapest part of technical SEO to get right and one of the costliest to get wrong. A wrong redirect type, a maintenance page returning 200, or a firewall quietly serving 429s to Googlebot can undo months of content work.&lt;/p&gt;

&lt;p&gt;Pick the code that matches what's really happening, and keep your signals consistent: the redirect, the canonical, the sitemap and your internal links should all agree.&lt;/p&gt;

&lt;p&gt;Which status code has caused you the most confusion? Tell me in the comments.&lt;/p&gt;

&lt;p&gt;Also do visit my website : &lt;a href="https://anoshbb.com/" rel="noopener noreferrer"&gt;anoshbb.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Suggested tags:&lt;/strong&gt; &lt;code&gt;seo&lt;/code&gt;, &lt;code&gt;webdev&lt;/code&gt;, &lt;code&gt;http&lt;/code&gt;, &lt;code&gt;beginners&lt;/code&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>tutorial</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Core Web Vitals: What They Actually Mean Technically</title>
      <dc:creator>Anosh</dc:creator>
      <pubDate>Tue, 06 Oct 2026 19:02:02 +0000</pubDate>
      <link>https://dev.to/marketingwithanosh/core-web-vitals-what-they-actually-mean-technically-17m0</link>
      <guid>https://dev.to/marketingwithanosh/core-web-vitals-what-they-actually-mean-technically-17m0</guid>
      <description>&lt;p&gt;"Core Web Vitals are important for SEO" is the kind of sentence that's true and useless at the same time. It tells you to care, but not what to open in your editor on Monday morning.&lt;/p&gt;

&lt;p&gt;So let's skip the pep talk and look at what each metric measures, what causes it to go bad, and what a developer can actually change.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three metrics, quickly
&lt;/h2&gt;

&lt;p&gt;Core Web Vitals are three field metrics, measured on real users and judged at the &lt;strong&gt;75th percentile&lt;/strong&gt; of page loads:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LCP (Largest Contentful Paint):&lt;/strong&gt; how long until the biggest visible element renders. Good is &lt;strong&gt;2.5 seconds or less&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;INP (Interaction to Next Paint):&lt;/strong&gt; how quickly the page responds visually after a user interaction. Good is &lt;strong&gt;200 milliseconds or less&lt;/strong&gt;. It replaced FID in 2024.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CLS (Cumulative Layout Shift):&lt;/strong&gt; how much visible content jumps around unexpectedly. Good is &lt;strong&gt;0.1 or less&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One thing worth saying early: these are &lt;em&gt;field&lt;/em&gt; numbers. Lighthouse in your own browser is a lab test on your fast laptop. Real users on mid-range phones and patchy connections are the ones Google measures, so use PageSpeed Insights (the CrUX data at the top) or Search Console to see what they experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  LCP: why the biggest thing loads late
&lt;/h2&gt;

&lt;p&gt;LCP is usually a hero image, a large heading, or a big text block. The element is whatever the browser decides is largest in the viewport.&lt;/p&gt;

&lt;p&gt;The useful mental model is that LCP is a chain of four phases:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Time to First Byte (TTFB):&lt;/strong&gt; how long the server takes to respond&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resource load delay:&lt;/strong&gt; how long before the browser &lt;em&gt;starts&lt;/em&gt; fetching the LCP resource&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resource load duration:&lt;/strong&gt; how long the download takes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Element render delay:&lt;/strong&gt; how long between "downloaded" and "painted"&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When LCP is bad, one of those phases is almost always the culprit. Find which one before changing anything.&lt;/p&gt;

&lt;h3&gt;
  
  
  What causes slow LCP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Slow server response.&lt;/strong&gt; If TTFB is already 1.5 seconds, you've spent most of your budget before the browser sees any HTML. Common causes are no caching, slow database queries, distant servers and no CDN.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hero images that are too heavy.&lt;/strong&gt; A 2 MB unoptimized JPEG as your banner is the classic offender.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hero images discovered late.&lt;/strong&gt; If the image is a CSS background or gets injected by JavaScript, the browser can't find it until much later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lazy-loading the LCP image.&lt;/strong&gt; &lt;code&gt;loading="lazy"&lt;/code&gt; on your hero tells the browser to deprioritize the exact thing it should hurry up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Render-blocking CSS.&lt;/strong&gt; The browser won't paint until it has the stylesheets in the &lt;code&gt;&amp;lt;head&amp;gt;&lt;/code&gt;. A giant CSS bundle delays everything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Render-blocking JavaScript.&lt;/strong&gt; Synchronous scripts in the head pause HTML parsing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Web fonts.&lt;/strong&gt; If your LCP is a headline, it may sit invisible while the font file downloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Client-side rendering.&lt;/strong&gt; If the page is an empty &lt;code&gt;&amp;lt;div id="root"&amp;gt;&lt;/code&gt; until a JS bundle runs, the LCP element doesn't exist until all that work finishes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What you can change
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Fix the server first, if TTFB is the problem:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cache HTML where you can (page caching, edge caching)&lt;/li&gt;
&lt;li&gt;Put a CDN in front of the site&lt;/li&gt;
&lt;li&gt;Optimize slow queries and heavy backend work&lt;/li&gt;
&lt;li&gt;Avoid redirect chains before the page even loads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Make the hero image fast and discoverable:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;img&lt;/span&gt;
  &lt;span class="na"&gt;src=&lt;/span&gt;&lt;span class="s"&gt;"/hero.avif"&lt;/span&gt;
  &lt;span class="na"&gt;width=&lt;/span&gt;&lt;span class="s"&gt;"1200"&lt;/span&gt;
  &lt;span class="na"&gt;height=&lt;/span&gt;&lt;span class="s"&gt;"600"&lt;/span&gt;
  &lt;span class="na"&gt;fetchpriority=&lt;/span&gt;&lt;span class="s"&gt;"high"&lt;/span&gt;
  &lt;span class="na"&gt;alt=&lt;/span&gt;&lt;span class="s"&gt;"Product hero image"&lt;/span&gt;
&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Serve modern formats like WebP or AVIF&lt;/li&gt;
&lt;li&gt;Use responsive images (&lt;code&gt;srcset&lt;/code&gt; and &lt;code&gt;sizes&lt;/code&gt;) so phones don't download desktop-sized files&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;fetchpriority="high"&lt;/code&gt; to the LCP image&lt;/li&gt;
&lt;li&gt;Never lazy-load it&lt;/li&gt;
&lt;li&gt;Keep it in the HTML as a real &lt;code&gt;&amp;lt;img&amp;gt;&lt;/code&gt;, not a CSS background&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If it must be discovered early and can't be in plain HTML, preload it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;"preload"&lt;/span&gt; &lt;span class="na"&gt;as=&lt;/span&gt;&lt;span class="s"&gt;"image"&lt;/span&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;"/hero.avif"&lt;/span&gt; &lt;span class="na"&gt;fetchpriority=&lt;/span&gt;&lt;span class="s"&gt;"high"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Deal with render-blocking resources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Inline the small amount of critical CSS needed for above-the-fold content&lt;/li&gt;
&lt;li&gt;Load the rest of the CSS without blocking&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;defer&lt;/code&gt; (or &lt;code&gt;async&lt;/code&gt;, where order doesn't matter) to scripts&lt;/li&gt;
&lt;li&gt;Remove unused CSS and JS instead of just minifying them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Handle fonts sensibly:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use &lt;code&gt;font-display: swap&lt;/code&gt; so text shows immediately in a fallback&lt;/li&gt;
&lt;li&gt;Preload your one or two critical font files&lt;/li&gt;
&lt;li&gt;Self-host fonts instead of waiting on a third-party domain&lt;/li&gt;
&lt;li&gt;Subset fonts so you're not shipping characters you never use&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Consider rendering strategy.&lt;/strong&gt; If you're on a JS framework, server-side rendering or static generation gets real HTML to the browser faster than shipping an empty shell.&lt;/p&gt;

&lt;h2&gt;
  
  
  INP: why the page feels sluggish
&lt;/h2&gt;

&lt;p&gt;INP measures the delay between a click, tap or keypress and the next frame the browser paints in response. It looks at interactions across the whole visit and reports roughly the worst one, which is why one bad interaction can hurt you.&lt;/p&gt;

&lt;p&gt;Each interaction has three parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input delay:&lt;/strong&gt; the wait before the handler can even start (the main thread is busy)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Processing time:&lt;/strong&gt; how long your event handlers run&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Presentation delay:&lt;/strong&gt; the time to recalculate layout and paint the result&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What causes poor INP
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Long tasks.&lt;/strong&gt; Any task that holds the main thread for more than 50 ms can block the browser from responding. Several in a row make a page feel frozen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Main-thread blocking in general.&lt;/strong&gt; The main thread handles JavaScript, style calculation, layout and painting. If JS hogs it, nothing else happens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heavy event handlers.&lt;/strong&gt; A click handler that filters 5,000 items, updates a huge state tree and re-renders half the page will feel slow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Third-party scripts.&lt;/strong&gt; Chat widgets, tag managers, analytics, A/B testing tools and ad scripts all compete for the same main thread, and you don't control their code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Large DOM size.&lt;/strong&gt; Updating a page with tens of thousands of nodes costs more at every step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expensive rendering work.&lt;/strong&gt; Forcing layout repeatedly (layout thrashing) or re-rendering too much UI after each interaction.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What you can change
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Break up long tasks.&lt;/strong&gt; Yield back to the main thread so the browser can respond to the user between chunks of work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;processItems&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;heavyWork&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="c1"&gt;// Let the browser handle pending input and paint&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;resolve&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Newer APIs like &lt;code&gt;scheduler.yield()&lt;/code&gt; do this more cleanly where supported, but the idea is the same: don't hold the thread for 300 ms in one go.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep event handlers light:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do the minimum needed to show visual feedback first, then defer the heavy part&lt;/li&gt;
&lt;li&gt;Debounce or throttle handlers for input, scroll and resize&lt;/li&gt;
&lt;li&gt;Avoid synchronous work you don't need before the next paint&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Move heavy work off the main thread.&lt;/strong&gt; Web Workers are good for parsing, calculations and data crunching that doesn't touch the DOM.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cut down the JavaScript you ship:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Code-split so each page loads only what it needs&lt;/li&gt;
&lt;li&gt;Remove unused libraries (that date library you imported for one function)&lt;/li&gt;
&lt;li&gt;Lazy-load features that appear only after interaction&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Tame third-party scripts:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Audit every one and delete what nobody uses&lt;/li&gt;
&lt;li&gt;Load non-essential scripts after the page is interactive, or on user interaction&lt;/li&gt;
&lt;li&gt;Prefer lighter alternatives where they exist&lt;/li&gt;
&lt;li&gt;Measure the impact of each one, because the cost is often a surprise&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Reduce rendering cost:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Keep the DOM small and avoid deeply nested structures&lt;/li&gt;
&lt;li&gt;Virtualize very long lists&lt;/li&gt;
&lt;li&gt;Batch DOM reads and writes to avoid layout thrashing&lt;/li&gt;
&lt;li&gt;In React-style frameworks, avoid unnecessary re-renders with sensible state placement and memoization&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  CLS: why things jump around
&lt;/h2&gt;

&lt;p&gt;CLS adds up unexpected layout shifts. A shift happens when a visible element moves between frames without the user causing it. The score combines how much of the viewport moved and how far it moved.&lt;/p&gt;

&lt;p&gt;The painful version is familiar: you go to tap a button, an ad loads above it, everything slides down, and you tap the wrong thing.&lt;/p&gt;

&lt;h3&gt;
  
  
  What causes layout shifts
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Images and videos without dimensions.&lt;/strong&gt; The browser doesn't know how much space to reserve, so content below gets pushed when the media loads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ads and embeds.&lt;/strong&gt; Ad slots that resize after loading, or iframes (maps, videos, social posts) with no reserved space.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic content.&lt;/strong&gt; Banners, cookie notices, "related items" blocks or personalized content inserted above existing content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Web fonts.&lt;/strong&gt; When the fallback font swaps to the real one, text can reflow if the two have different metrics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Injected elements.&lt;/strong&gt; Anything added to the DOM above content the user is already looking at, like a promo bar sliding in after load.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Late-loading CSS.&lt;/strong&gt; Styles that arrive after first paint can change the layout.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Note that shifts caused directly by user input (like clicking to expand an accordion) are excluded. It's the &lt;em&gt;unexpected&lt;/em&gt; ones that count.&lt;/p&gt;

&lt;h3&gt;
  
  
  What you can change
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Always give media dimensions:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;img&lt;/span&gt; &lt;span class="na"&gt;src=&lt;/span&gt;&lt;span class="s"&gt;"/photo.webp"&lt;/span&gt; &lt;span class="na"&gt;width=&lt;/span&gt;&lt;span class="s"&gt;"800"&lt;/span&gt; &lt;span class="na"&gt;height=&lt;/span&gt;&lt;span class="s"&gt;"450"&lt;/span&gt; &lt;span class="na"&gt;alt=&lt;/span&gt;&lt;span class="s"&gt;"..."&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The browser uses the width and height to calculate the aspect ratio and reserve space before the file arrives. You can also set &lt;code&gt;aspect-ratio&lt;/code&gt; in CSS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight css"&gt;&lt;code&gt;&lt;span class="nc"&gt;.video-wrapper&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="py"&gt;aspect-ratio&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16&lt;/span&gt; &lt;span class="p"&gt;/&lt;/span&gt; &lt;span class="m"&gt;9&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Reserve space for ads and embeds:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Give ad slots a &lt;code&gt;min-height&lt;/code&gt; that matches the most common ad size&lt;/li&gt;
&lt;li&gt;Put placeholders around iframes and embedded widgets&lt;/li&gt;
&lt;li&gt;Avoid inserting ads above content that's already in view&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Be careful with dynamic content:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Insert new content &lt;em&gt;below&lt;/em&gt; the viewport or in space you've already reserved&lt;/li&gt;
&lt;li&gt;For banners and cookie notices, use fixed or overlay positioning, or reserve their height from the start&lt;/li&gt;
&lt;li&gt;Show skeleton placeholders that match the final size of the content&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Reduce font-swap shifts:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Preload critical fonts&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;font-display: optional&lt;/code&gt; or &lt;code&gt;swap&lt;/code&gt; depending on how much a brief flash matters to you&lt;/li&gt;
&lt;li&gt;Match the fallback font's metrics to the web font with &lt;code&gt;size-adjust&lt;/code&gt;, &lt;code&gt;ascent-override&lt;/code&gt; and similar descriptors, so the swap barely changes the layout&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Animate with transforms.&lt;/strong&gt; Animations that change &lt;code&gt;top&lt;/code&gt;, &lt;code&gt;height&lt;/code&gt; or &lt;code&gt;margin&lt;/code&gt; trigger layout and can cause shifts. &lt;code&gt;transform&lt;/code&gt; and &lt;code&gt;opacity&lt;/code&gt; don't.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to find what's actually wrong
&lt;/h2&gt;

&lt;p&gt;Fixing things blindly is how teams waste a sprint. A better workflow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start with field data.&lt;/strong&gt; Check Search Console's Core Web Vitals report and the CrUX section in PageSpeed Insights to see which metric fails and on which pages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproduce in the lab.&lt;/strong&gt; Use Lighthouse and the Chrome DevTools Performance panel, ideally with CPU and network throttling to mimic a mid-range phone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Find the specific cause.&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;For LCP, look at which element it is and which of the four phases is longest&lt;/li&gt;
&lt;li&gt;For INP, record an interaction in the Performance panel and look for long tasks&lt;/li&gt;
&lt;li&gt;For CLS, use the Layout Shift regions in DevTools to see what moved&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fix one thing, then measure again.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wait for field data.&lt;/strong&gt; CrUX uses a rolling 28-day window, so improvements take a few weeks to show in your real-user numbers.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Quick reference
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LCP too high:&lt;/strong&gt; slow server, heavy or late-discovered hero image, render-blocking CSS and JS, fonts, client-side rendering. Fix with caching, image optimization, &lt;code&gt;fetchpriority&lt;/code&gt;, preload, deferred scripts and SSR.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;INP too high:&lt;/strong&gt; long tasks, heavy handlers, third-party scripts, big DOM. Fix by splitting work, using workers, trimming JS and auditing third parties.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CLS too high:&lt;/strong&gt; media without dimensions, ads, injected content, font swaps. Fix with explicit sizes, reserved space, stable fonts and transform-based animation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Core Web Vitals look intimidating until you notice that each one is really a simple complaint from a real user:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LCP: "this is taking forever to show up"&lt;/li&gt;
&lt;li&gt;INP: "I tapped it and nothing happened"&lt;/li&gt;
&lt;li&gt;CLS: "everything moved right as I tried to tap"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every technical fix above comes from taking one of those complaints seriously. And it's worth remembering that these metrics are one signal among many. A fast page won't rescue weak content, but a slow, jumpy one makes everything else harder.&lt;/p&gt;

&lt;p&gt;If you've got a Core Web Vitals problem that refuses to go away, share it in the comments. The stubborn ones are usually the most educational.&lt;/p&gt;

&lt;p&gt;Also check out my website : &lt;a href="https://anoshbb.com/" rel="noopener noreferrer"&gt;anoshbb.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Suggested tags:&lt;/strong&gt; &lt;code&gt;webdev&lt;/code&gt;, &lt;code&gt;performance&lt;/code&gt;, &lt;code&gt;seo&lt;/code&gt;, &lt;code&gt;javascript&lt;/code&gt;&lt;/p&gt;

</description>
      <category>frontend</category>
      <category>performance</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Robots.txt Deep Dive: What It Can and Cannot Do</title>
      <dc:creator>Anosh</dc:creator>
      <pubDate>Tue, 06 Oct 2026 18:53:46 +0000</pubDate>
      <link>https://dev.to/marketingwithanosh/robotstxt-deep-dive-what-it-can-and-cannot-do-1bn7</link>
      <guid>https://dev.to/marketingwithanosh/robotstxt-deep-dive-what-it-can-and-cannot-do-1bn7</guid>
      <description>&lt;p&gt;Robots.txt is one of the smallest files on your site and one of the most misunderstood. It's a plain text file, and a single wrong line can hide your whole site from search engines. Even more often, people expect it to do something it was never built for.&lt;/p&gt;

&lt;p&gt;Let's sort out what it actually does.&lt;/p&gt;

&lt;h2&gt;
  
  
  What robots.txt controls
&lt;/h2&gt;

&lt;p&gt;Robots.txt controls &lt;strong&gt;crawling&lt;/strong&gt;. That's all. It tells compliant bots which URLs they may or may not request.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It lives at the root of a host: &lt;code&gt;https://example.com/robots.txt&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Each host, protocol and port needs its own file (&lt;code&gt;blog.example.com&lt;/code&gt; doesn't inherit from &lt;code&gt;example.com&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;It's a polite request, not a lock. Reputable crawlers like Googlebot and Bingbot follow it. Malicious bots ignore it&lt;/li&gt;
&lt;li&gt;It's public. Anyone can read it, so never use it to hide private pages&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does &lt;strong&gt;not&lt;/strong&gt; control indexing, and it does not provide security.&lt;/p&gt;

&lt;h2&gt;
  
  
  The basic syntax
&lt;/h2&gt;

&lt;p&gt;A robots.txt file is made of groups. Each group starts with a user-agent and is followed by rules.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight apache"&gt;&lt;code&gt;&lt;span class="nc"&gt;User&lt;/span&gt;-agent: *
Disallow: /admin/
&lt;span class="nc"&gt;Allow&lt;/span&gt;: /admin/help/

Sitemap: https://example.com/sitemap.xml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  User-agent
&lt;/h3&gt;

&lt;p&gt;This names the bot a group applies to.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;*&lt;/code&gt; means all bots&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Googlebot&lt;/code&gt;, &lt;code&gt;Bingbot&lt;/code&gt; and other names target specific crawlers&lt;/li&gt;
&lt;li&gt;A bot follows only the &lt;strong&gt;most specific group&lt;/strong&gt; that matches it. It does not combine that group with the &lt;code&gt;*&lt;/code&gt; group&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point trips people up. If you write a &lt;code&gt;Googlebot&lt;/code&gt; group, Googlebot ignores the &lt;code&gt;*&lt;/code&gt; group completely, so repeat any shared rules inside it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Disallow
&lt;/h3&gt;

&lt;p&gt;Blocks a path from being crawled.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Disallow: /private/&lt;/code&gt; blocks everything under that folder&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Disallow: /&lt;/code&gt; blocks the entire site&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Disallow:&lt;/code&gt; (empty) blocks nothing&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Allow
&lt;/h3&gt;

&lt;p&gt;Makes an exception inside a blocked area.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Useful for opening one subfolder inside a blocked directory&lt;/li&gt;
&lt;li&gt;When rules conflict, Google follows the &lt;strong&gt;most specific (longest) match&lt;/strong&gt;. If it's a tie, the less restrictive rule wins&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Sitemap
&lt;/h3&gt;

&lt;p&gt;Points crawlers to your XML sitemap.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use a full absolute URL&lt;/li&gt;
&lt;li&gt;It isn't tied to any user-agent group, so it can sit anywhere in the file&lt;/li&gt;
&lt;li&gt;You can list more than one&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Wildcards
&lt;/h2&gt;

&lt;p&gt;Google and Bing support two special characters:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;*&lt;/code&gt; matches any sequence of characters&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;$&lt;/code&gt; marks the end of a URL
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight apache"&gt;&lt;code&gt;&lt;span class="nc"&gt;User&lt;/span&gt;-agent: *
Disallow: /*.pdf$
Disallow: /search*
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first rule blocks URLs ending in &lt;code&gt;.pdf&lt;/code&gt;. Without the &lt;code&gt;$&lt;/code&gt;, it would also block something like &lt;code&gt;/file.pdf?download=1&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Directory blocking
&lt;/h2&gt;

&lt;p&gt;Be careful with trailing slashes. They change the meaning.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Disallow: /blog/&lt;/code&gt; blocks the folder and everything inside it&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Disallow: /blog&lt;/code&gt; blocks anything that &lt;strong&gt;starts with&lt;/strong&gt; &lt;code&gt;/blog&lt;/code&gt;, including &lt;code&gt;/blog-news&lt;/code&gt; and &lt;code&gt;/blogger-tools&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Paths are also &lt;strong&gt;case-sensitive&lt;/strong&gt;. &lt;code&gt;/Admin/&lt;/code&gt; and &lt;code&gt;/admin/&lt;/code&gt; are different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  Query parameters
&lt;/h2&gt;

&lt;p&gt;Parameters can create thousands of near-identical URLs: sorting, filters, session IDs, tracking tags. Robots.txt is a common way to stop crawlers from wasting time on them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight apache"&gt;&lt;code&gt;&lt;span class="nc"&gt;User&lt;/span&gt;-agent: *
Disallow: /*?sort=
Disallow: /*?*sessionid=
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few cautions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Block only parameters that truly produce duplicate or useless pages&lt;/li&gt;
&lt;li&gt;If a parameter changes the actual content (like &lt;code&gt;?page=2&lt;/code&gt; on a real listing), blocking it can hide products or articles from crawlers&lt;/li&gt;
&lt;li&gt;For duplicates you want consolidated rather than ignored, a canonical tag is often the better tool&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Bot-specific rules
&lt;/h2&gt;

&lt;p&gt;You can give different bots different instructions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight apache"&gt;&lt;code&gt;&lt;span class="nc"&gt;User&lt;/span&gt;-agent: Googlebot
Disallow: /internal-search/

&lt;span class="nc"&gt;User&lt;/span&gt;-agent: GPTBot
Disallow: /

&lt;span class="nc"&gt;User&lt;/span&gt;-agent: *
Disallow: /admin/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;This is how people opt out of certain AI crawlers while staying open to search engines&lt;/li&gt;
&lt;li&gt;Remember that each bot reads only its own group&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Crawl-delay&lt;/code&gt; is ignored by Google. Bing supports it. Don't rely on it for Googlebot&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why robots.txt is not an indexing directive
&lt;/h2&gt;

&lt;p&gt;This is the big misunderstanding.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Disallow&lt;/code&gt; stops Google from &lt;strong&gt;fetching&lt;/strong&gt; a page&lt;/li&gt;
&lt;li&gt;It doesn't stop Google from &lt;strong&gt;knowing the URL exists&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;If other pages link to a blocked URL, Google can still index the bare address and show it in results with no description&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So a blocked page can still appear in search. It just appears without content, because Google was never allowed to read it.&lt;/p&gt;

&lt;p&gt;Google also no longer supports a &lt;code&gt;noindex&lt;/code&gt; line inside robots.txt. Any guide telling you to write &lt;code&gt;Noindex: /page/&lt;/code&gt; in the file is out of date.&lt;/p&gt;

&lt;h2&gt;
  
  
  Blocking crawling vs noindex
&lt;/h2&gt;

&lt;p&gt;These two tools solve different problems.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Disallow in robots.txt:&lt;/strong&gt; "Don't fetch this URL"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;noindex (meta tag or &lt;code&gt;X-Robots-Tag&lt;/code&gt; header):&lt;/strong&gt; "You may fetch this, but don't keep it in the index"
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;meta&lt;/span&gt; &lt;span class="na"&gt;name=&lt;/span&gt;&lt;span class="s"&gt;"robots"&lt;/span&gt; &lt;span class="na"&gt;content=&lt;/span&gt;&lt;span class="s"&gt;"noindex"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The trap is combining them. If you block a page in robots.txt &lt;strong&gt;and&lt;/strong&gt; add a &lt;code&gt;noindex&lt;/code&gt; tag, Google never fetches the page, so it never sees the &lt;code&gt;noindex&lt;/code&gt;. The URL can stay indexed for a long time.&lt;/p&gt;

&lt;p&gt;To remove a page from search properly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Let Google crawl it (no &lt;code&gt;Disallow&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Serve a &lt;code&gt;noindex&lt;/code&gt; tag or header&lt;/li&gt;
&lt;li&gt;Wait for Google to recrawl and drop it&lt;/li&gt;
&lt;li&gt;Only then consider blocking it, if you still want to&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For pages that are truly private, use authentication. Neither robots.txt nor &lt;code&gt;noindex&lt;/code&gt; is real protection.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common syntax mistakes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Leaving &lt;code&gt;Disallow: /&lt;/code&gt; live after launch.&lt;/strong&gt; The classic staging-site disaster. Check it after every migration&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blocking CSS and JavaScript.&lt;/strong&gt; Google needs these to render your pages properly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wrong case or missing slash.&lt;/strong&gt; &lt;code&gt;/Blog/&lt;/code&gt; won't match &lt;code&gt;/blog/&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Putting rules before any &lt;code&gt;User-agent&lt;/code&gt; line.&lt;/strong&gt; Rules must belong to a group&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relative sitemap URLs.&lt;/strong&gt; Use the full &lt;code&gt;https://&lt;/code&gt; address&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Using robots.txt to hide sensitive content.&lt;/strong&gt; The file is public, so it can actually point people to what you want hidden&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assuming &lt;code&gt;*&lt;/code&gt; rules apply to every group.&lt;/strong&gt; They don't, as covered above&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wrong location or format.&lt;/strong&gt; It must sit at the root, be plain text, and return a 200 status. If it returns a server error for a long time, Google may treat the site cautiously&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A sensible starter file
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight apache"&gt;&lt;code&gt;&lt;span class="nc"&gt;User&lt;/span&gt;-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /*?sessionid=

Sitemap: https://example.com/sitemap.xml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Short, clear and easy to audit. Add rules only when you have a reason, and write a comment (&lt;code&gt;# reason&lt;/code&gt;) so future you remembers why.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to test it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Use the robots.txt report in Google Search Console to see how Google reads your file&lt;/li&gt;
&lt;li&gt;Run the URL Inspection tool on important pages and confirm they aren't blocked&lt;/li&gt;
&lt;li&gt;Open &lt;code&gt;/robots.txt&lt;/code&gt; in your browser after every deployment&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Think of robots.txt as a traffic sign for crawlers, not a vault door and not a deletion button. It manages where bots spend their time. If you want something kept out of search results, use &lt;code&gt;noindex&lt;/code&gt;. If you want something truly private, use a login.&lt;/p&gt;

&lt;p&gt;Have you ever found a rogue &lt;code&gt;Disallow: /&lt;/code&gt; in production? Tell me about it in the comments.&lt;/p&gt;

&lt;p&gt;Also , if you're interested, do visit my website - &lt;a href="https://anoshbb.com/" rel="noopener noreferrer"&gt;anoshbb.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Crawlability vs Indexability: A Decision Tree for Figuring Out Why a Page Isn't in Google</title>
      <dc:creator>Anosh</dc:creator>
      <pubDate>Tue, 06 Oct 2026 18:48:16 +0000</pubDate>
      <link>https://dev.to/marketingwithanosh/crawlability-vs-indexability-a-decision-tree-for-figuring-out-why-a-page-isnt-in-google-4opf</link>
      <guid>https://dev.to/marketingwithanosh/crawlability-vs-indexability-a-decision-tree-for-figuring-out-why-a-page-isnt-in-google-4opf</guid>
      <description>&lt;p&gt;A page missing from Google is one of the most common problems in SEO. It's also one of the most misdiagnosed, because most people treat "not indexed" as a single problem. It's at least five different problems that look identical from the outside.&lt;/p&gt;

&lt;p&gt;The mix-up usually starts with two words that get used interchangeably: crawlability and indexability. They are not the same thing, and fixing the wrong one is how people lose a whole afternoon.&lt;/p&gt;

&lt;p&gt;Two different questions&lt;br&gt;
Crawlability: Can Googlebot access this URL and fetch its content?&lt;br&gt;
Indexability: Once Google has the content, is it allowed (and willing) to store the page in its index and show it in results?&lt;/p&gt;

&lt;p&gt;An analogy: crawling is a librarian walking into your shop and reading your book. Indexing is deciding whether the book goes on the shelf. The librarian can read it and still leave it off the shelf.&lt;/p&gt;

&lt;p&gt;That gives you four possible states:&lt;/p&gt;

&lt;p&gt;Crawlable and indexable: the goal.&lt;br&gt;
Crawlable but not indexable: Google reads the page but is told, or decides, not to keep it. noindex, canonicals pointing elsewhere and quality problems live here.&lt;br&gt;
Not crawlable but still indexed: yes, this happens. A URL blocked in robots.txt can appear in results with no description if other pages link to it.&lt;br&gt;
Not crawlable and not indexed: the usual "my page doesn't exist to Google" case.&lt;/p&gt;

&lt;p&gt;Most beginners assume robots.txt is the tool for keeping pages out of Google. It isn't. It controls crawling, not indexing. That one misunderstanding explains a surprising number of broken sites.&lt;/p&gt;

&lt;p&gt;The decision tree&lt;/p&gt;

&lt;p&gt;Here's the flow I use whenever someone says "my page isn't showing up":&lt;/p&gt;

&lt;p&gt;Page doesn't appear in Google&lt;br&gt;
        ↓&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Can Googlebot crawl it?
(status code, server, redirects, login walls)
    ↓&lt;/li&gt;
&lt;li&gt;Is crawling blocked?
(robots.txt)
    ↓&lt;/li&gt;
&lt;li&gt;Is indexing blocked?
(noindex meta tag / X-Robots-Tag header)
    ↓&lt;/li&gt;
&lt;li&gt;Is there a canonical pointing elsewhere?
    ↓&lt;/li&gt;
&lt;li&gt;Is Google choosing not to index it?
(discovered vs crawled, duplicates, soft 404)
    ↓&lt;/li&gt;
&lt;li&gt;Is the page actually valuable enough?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Work through it in order. The first three steps are technical and yes/no. The last three are judgment calls, and they're harder to fix. Don't skip ahead to "content quality" when a stray noindex is sitting in your template.&lt;/p&gt;

&lt;p&gt;Step 1: Can Googlebot even reach the page?&lt;/p&gt;

&lt;p&gt;Start with the boring stuff. Check what your server actually returns.&lt;/p&gt;

&lt;p&gt;HTTP status codes&lt;/p&gt;

&lt;p&gt;200: the page loaded. This is what you want, but it's not proof the page is fine (see soft 404s below).&lt;br&gt;
301 / 308: permanent redirect. Google follows it and generally treats the destination as the page that matters.&lt;br&gt;
302 / 307: temporary redirect. Google may keep the original URL as the one it shows.&lt;br&gt;
404 / 410: not found / gone. These URLs drop out of the index over time. A 410 is a slightly stronger signal that it's gone for good.&lt;br&gt;
401 / 403: access denied. Googlebot can't get in, so it can't index the content.&lt;br&gt;
429 / 5xx: the server is overloaded or erroring. If this keeps happening, Google slows its crawl rate, and pages can get dropped.&lt;/p&gt;

&lt;p&gt;A quick way to see this yourself:&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
curl -I &lt;a href="https://example.com/your-page" rel="noopener noreferrer"&gt;https://example.com/your-page&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Look at the first line for the status code, and look for x-robots-tag and location headers while you're there. Both come up again below.&lt;/p&gt;

&lt;p&gt;Redirects&lt;/p&gt;

&lt;p&gt;Redirects deserve their own attention because they fail quietly:&lt;/p&gt;

&lt;p&gt;Chains: A → B → C → D. Google follows a limited number of hops (around 10), but every extra hop wastes crawl effort and slows things down. Keep it to one.&lt;br&gt;
Loops: A → B → A. Googlebot gives up.&lt;br&gt;
Redirecting everything to the homepage: Google often treats this like a soft 404, because the destination has nothing to do with the original page.&lt;br&gt;
JavaScript or meta-refresh redirects: these can work, but they're slower and less reliable than a server-side redirect.&lt;/p&gt;

&lt;p&gt;Other access problems&lt;/p&gt;

&lt;p&gt;Login walls or paywalls with no accessible content for the crawler&lt;br&gt;
Firewalls or bot protection accidentally blocking Googlebot (a surprisingly common one with aggressive security plugins)&lt;br&gt;
Pages that only appear after a user action, like a button click or infinite scroll&lt;br&gt;
Content that only exists after heavy JavaScript runs and fails to render&lt;br&gt;
Step 2: Is crawling blocked by robots.txt?&lt;/p&gt;

&lt;p&gt;Open yourdomain.com/robots.txt and read it with fresh eyes. A single stray line can hide a whole section:&lt;/p&gt;

&lt;p&gt;User-agent: *&lt;br&gt;
Disallow: /blog/&lt;/p&gt;

&lt;p&gt;That tells every crawler not to fetch anything under /blog/.&lt;/p&gt;

&lt;p&gt;Things to know:&lt;/p&gt;

&lt;p&gt;Disallow stops Google from fetching the page. It does not tell Google to remove the page from the index.&lt;br&gt;
If a blocked URL has links pointing at it from elsewhere, Google can still index the bare URL. You'll see it in results with a note that no description is available.&lt;br&gt;
Blocking CSS or JavaScript files can stop Google from rendering your page properly, which affects how it understands your content.&lt;br&gt;
A staging Disallow: / that survives a site launch is the classic disaster. Check this first after any migration.&lt;/p&gt;

&lt;p&gt;The trap that catches people: if you want a page out of the index, you can't block it in robots.txt and also put noindex on it. Google never fetches the page, so it never sees the noindex. The page can stay indexed indefinitely. To remove something, let Google crawl it and see the noindex.&lt;/p&gt;

&lt;p&gt;Step 3: Is indexing blocked?&lt;/p&gt;

&lt;p&gt;If Google can crawl the page, the next question is whether you've told it not to index. There are two places to check.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The meta robots tag in the HTML head:&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;html&lt;br&gt;
&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The X-Robots-Tag HTTP header:&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;X-Robots-Tag: noindex&lt;/p&gt;

&lt;p&gt;The header version is easy to miss because it never shows up in "view source." It's also the standard way to noindex non-HTML files like PDFs. If you've checked the HTML and found nothing, check the headers with curl -I.&lt;/p&gt;

&lt;p&gt;Common ways a noindex ends up where it shouldn't be:&lt;/p&gt;

&lt;p&gt;A CMS setting like "Discourage search engines from indexing this site" left on after launch&lt;br&gt;
An SEO plugin applying noindex to a whole post type, category or tag archive&lt;br&gt;
A staging environment's config copied to production&lt;br&gt;
A template-level tag that gets inherited by every page&lt;/p&gt;

&lt;p&gt;If Search Console says "Excluded by 'noindex' tag," it's telling you the truth. The only real question is whether you meant it.&lt;/p&gt;

&lt;p&gt;Step 4: Is there a canonical pointing somewhere else?&lt;/p&gt;

&lt;p&gt;A canonical tag tells Google which version of a page you consider the main one:&lt;/p&gt;

&lt;p&gt;html&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;Used correctly, it consolidates duplicate or near-duplicate URLs. Used incorrectly, it can quietly remove a page from the index.&lt;/p&gt;

&lt;p&gt;Things to check:&lt;/p&gt;

&lt;p&gt;Does the canonical point to a different URL? If your page says its canonical is another page, you've told Google to index that other one instead.&lt;br&gt;
Is the canonical pointing at a redirect, a 404 or a noindexed page? Mixed signals like this make Google ignore the tag.&lt;br&gt;
Is the same canonical on every page? A template bug that canonicalizes everything to the homepage is more common than you'd think.&lt;br&gt;
Do the canonical, sitemap entry and internal links all agree on one URL? Conflicting signals make Google choose for itself.&lt;/p&gt;

&lt;p&gt;Remember that a canonical is a hint, not a command. Google can override it. In Search Console, two statuses matter here:&lt;/p&gt;

&lt;p&gt;"Alternate page with proper canonical tag": usually fine. That page is a deliberate duplicate.&lt;br&gt;
"Duplicate, Google chose different canonical than user": Google disagreed with your tag and picked another URL. Look at why the two pages seem the same.&lt;br&gt;
Step 5: Is Google choosing not to index it?&lt;/p&gt;

&lt;p&gt;Now we get into the harder territory. Everything technical is clean, and the page still isn't indexed. Search Console usually gives you one of two statuses that look similar but mean very different things.&lt;/p&gt;

&lt;p&gt;Discovered, currently not indexed&lt;/p&gt;

&lt;p&gt;Google knows the URL exists (from a sitemap or a link) but hasn't crawled it yet.&lt;br&gt;
It often points to crawl priority: a large site, a slow server, or little perceived value in the section.&lt;br&gt;
Fixes: improve internal linking to the page, make sure the server is fast and stable, and prune low-value URLs competing for attention.&lt;/p&gt;

&lt;p&gt;Crawled, currently not indexed&lt;/p&gt;

&lt;p&gt;Google fetched the page and decided not to keep it, for now.&lt;br&gt;
This is almost always a quality or duplication signal, not a technical one.&lt;br&gt;
Fixes: make the page more useful, more distinct and better connected to the rest of your site.&lt;/p&gt;

&lt;p&gt;Telling these two apart saves a lot of guessing. "Discovered" is mainly a crawling story. "Crawled" is mainly an indexing story.&lt;/p&gt;

&lt;p&gt;Duplicate content&lt;/p&gt;

&lt;p&gt;Duplicate content isn't a penalty. Google groups similar pages into a cluster and picks one to show. The usual sources are:&lt;/p&gt;

&lt;p&gt;URL parameters (?sort=price, tracking parameters)&lt;br&gt;
HTTP vs HTTPS and www vs non-www versions&lt;br&gt;
Trailing slash vs no trailing slash&lt;br&gt;
Printer-friendly pages&lt;br&gt;
Boilerplate-heavy pages where only a sentence or two differs, like location pages with swapped city names&lt;/p&gt;

&lt;p&gt;If your page is in a cluster and loses, it shows up as excluded, even though nothing is technically wrong with it.&lt;/p&gt;

&lt;p&gt;Soft 404s&lt;/p&gt;

&lt;p&gt;A soft 404 is a page that returns a 200 OK status but looks like an error or an empty page to Google:&lt;/p&gt;

&lt;p&gt;A "no results found" page that returns 200&lt;br&gt;
A product page that says "out of stock" with no other content&lt;br&gt;
A near-empty template with almost no text&lt;br&gt;
Redirecting deleted pages to an irrelevant page&lt;/p&gt;

&lt;p&gt;Google decides these are effectively not-found pages and leaves them out. The fix is to be honest with your status codes: return a real 404 or 410 for pages that are gone, or add real content to pages that should exist.&lt;/p&gt;

&lt;p&gt;Step 6: Is the page valuable enough?&lt;/p&gt;

&lt;p&gt;If you've cleared every step above, you're left with the uncomfortable question. Google has no reason to reject the page technically. It just doesn't think the page earns a spot.&lt;/p&gt;

&lt;p&gt;Signals worth examining:&lt;/p&gt;

&lt;p&gt;Thin content: a few generic sentences, no original information or depth.&lt;br&gt;
Rehashed content: the same points as ten other pages, with nothing new.&lt;br&gt;
Weak internal linking: orphan pages (nothing links to them) signal that you don't value them either.&lt;br&gt;
Site-wide quality: if much of your site is thin, Google may be more hesitant about everything on it.&lt;br&gt;
Mismatch with intent: the page doesn't clearly answer a question anyone is asking.&lt;br&gt;
Little to no external signals: a brand-new site or page with no links or mentions takes longer to earn trust.&lt;/p&gt;

&lt;p&gt;Practical things that help:&lt;/p&gt;

&lt;p&gt;Add information the other results don't have: examples, data, your own experience&lt;br&gt;
Link to the page from relevant, already-indexed pages&lt;br&gt;
Merge several weak pages into one strong one&lt;br&gt;
Remove or noindex pages that exist only to fill space&lt;br&gt;
Make the page's purpose obvious from the title, headings and first paragraph&lt;/p&gt;

&lt;p&gt;There's no switch to flip here. Quality problems get fixed by making the page better, and then waiting for Google to re-evaluate.&lt;/p&gt;

&lt;p&gt;Putting it into practice&lt;/p&gt;

&lt;p&gt;When a page is missing, I run this in about ten minutes:&lt;/p&gt;

&lt;p&gt;Search Console URL Inspection: paste the URL. It tells you whether the URL is on Google, when it was last crawled, and which canonical Google chose.&lt;br&gt;
Run the live test: it shows what Google sees right now, including whether the page is blocked and what the rendered HTML looks like.&lt;br&gt;
Check robots.txt: read it for any rule matching the path.&lt;br&gt;
curl -I the URL: confirm the status code and look for X-Robots-Tag and redirects.&lt;br&gt;
View source: search for noindex and canonical.&lt;br&gt;
Check the sitemap: is the URL listed, and is it the exact version you want indexed?&lt;br&gt;
Open the Page indexing report: find the exact status for the page. That label usually tells you which branch of the tree you're on.&lt;/p&gt;

&lt;p&gt;One more tip: the site:yourdomain.com/page search is a quick sanity check, but it isn't a reliable way to confirm indexing. Trust Search Console over it.&lt;/p&gt;

&lt;p&gt;Quick reference&lt;br&gt;
Blocked in robots.txt: crawling is blocked. The URL might still be indexed. Remove the rule.&lt;br&gt;
noindex present: indexing is blocked. Remove the tag.&lt;br&gt;
Canonical to another URL: you told Google to prefer something else. Fix or confirm the tag.&lt;br&gt;
404 / 410: the page is gone. Restore it or leave it gone on purpose.&lt;br&gt;
Soft 404: the page looks empty. Add real content or return a real error code.&lt;br&gt;
Redirect issue: shorten the chain and point to a relevant destination.&lt;br&gt;
Discovered, not indexed: crawl priority issue. Improve linking and server health.&lt;br&gt;
Crawled, not indexed: quality issue. Improve the page.&lt;br&gt;
Final thoughts&lt;/p&gt;

&lt;p&gt;The big shift is to stop asking "why won't Google index my page?" and start asking "which question in the chain failed?" Technical blocks answer yes or no, and you can fix them in minutes. The quality questions take longer, but they're the ones that decide whether a page deserves to rank at all.&lt;/p&gt;

&lt;p&gt;Next time a page is missing, walk the tree from top to bottom before changing anything. Most of the time the answer shows up in the first three steps.&lt;/p&gt;

&lt;p&gt;If you've got a weird indexing case that doesn't fit this flow, drop it in the comments. Those are usually the interesting ones.&lt;/p&gt;

&lt;p&gt;Also if you're interested , feel free to check out my website - &lt;a href="https://anoshbb.com/" rel="noopener noreferrer"&gt;anoshbb.com&lt;/a&gt;&lt;/p&gt;

</description>
      <category>google</category>
      <category>howto</category>
      <category>seo</category>
    </item>
  </channel>
</rss>
