DEV Community

Anosh
Anosh

Posted on

Crawlability vs Indexability: A Decision Tree for Figuring Out Why a Page Isn't in Google

A page missing from Google is one of the most common problems in SEO. It's also one of the most misdiagnosed, because most people treat "not indexed" as a single problem. It's at least five different problems that look identical from the outside.

The mix-up usually starts with two words that get used interchangeably: crawlability and indexability. They are not the same thing, and fixing the wrong one is how people lose a whole afternoon.

Two different questions
Crawlability: Can Googlebot access this URL and fetch its content?
Indexability: Once Google has the content, is it allowed (and willing) to store the page in its index and show it in results?

An analogy: crawling is a librarian walking into your shop and reading your book. Indexing is deciding whether the book goes on the shelf. The librarian can read it and still leave it off the shelf.

That gives you four possible states:

Crawlable and indexable: the goal.
Crawlable but not indexable: Google reads the page but is told, or decides, not to keep it. noindex, canonicals pointing elsewhere and quality problems live here.
Not crawlable but still indexed: yes, this happens. A URL blocked in robots.txt can appear in results with no description if other pages link to it.
Not crawlable and not indexed: the usual "my page doesn't exist to Google" case.

Most beginners assume robots.txt is the tool for keeping pages out of Google. It isn't. It controls crawling, not indexing. That one misunderstanding explains a surprising number of broken sites.

The decision tree

Here's the flow I use whenever someone says "my page isn't showing up":

Page doesn't appear in Google
↓

  1. Can Googlebot crawl it? (status code, server, redirects, login walls) ↓
  2. Is crawling blocked? (robots.txt) ↓
  3. Is indexing blocked? (noindex meta tag / X-Robots-Tag header) ↓
  4. Is there a canonical pointing elsewhere? ↓
  5. Is Google choosing not to index it? (discovered vs crawled, duplicates, soft 404) ↓
  6. Is the page actually valuable enough?

Work through it in order. The first three steps are technical and yes/no. The last three are judgment calls, and they're harder to fix. Don't skip ahead to "content quality" when a stray noindex is sitting in your template.

Step 1: Can Googlebot even reach the page?

Start with the boring stuff. Check what your server actually returns.

HTTP status codes

200: the page loaded. This is what you want, but it's not proof the page is fine (see soft 404s below).
301 / 308: permanent redirect. Google follows it and generally treats the destination as the page that matters.
302 / 307: temporary redirect. Google may keep the original URL as the one it shows.
404 / 410: not found / gone. These URLs drop out of the index over time. A 410 is a slightly stronger signal that it's gone for good.
401 / 403: access denied. Googlebot can't get in, so it can't index the content.
429 / 5xx: the server is overloaded or erroring. If this keeps happening, Google slows its crawl rate, and pages can get dropped.

A quick way to see this yourself:

bash
curl -I https://example.com/your-page

Look at the first line for the status code, and look for x-robots-tag and location headers while you're there. Both come up again below.

Redirects

Redirects deserve their own attention because they fail quietly:

Chains: A → B → C → D. Google follows a limited number of hops (around 10), but every extra hop wastes crawl effort and slows things down. Keep it to one.
Loops: A → B → A. Googlebot gives up.
Redirecting everything to the homepage: Google often treats this like a soft 404, because the destination has nothing to do with the original page.
JavaScript or meta-refresh redirects: these can work, but they're slower and less reliable than a server-side redirect.

Other access problems

Login walls or paywalls with no accessible content for the crawler
Firewalls or bot protection accidentally blocking Googlebot (a surprisingly common one with aggressive security plugins)
Pages that only appear after a user action, like a button click or infinite scroll
Content that only exists after heavy JavaScript runs and fails to render
Step 2: Is crawling blocked by robots.txt?

Open yourdomain.com/robots.txt and read it with fresh eyes. A single stray line can hide a whole section:

User-agent: *
Disallow: /blog/

That tells every crawler not to fetch anything under /blog/.

Things to know:

Disallow stops Google from fetching the page. It does not tell Google to remove the page from the index.
If a blocked URL has links pointing at it from elsewhere, Google can still index the bare URL. You'll see it in results with a note that no description is available.
Blocking CSS or JavaScript files can stop Google from rendering your page properly, which affects how it understands your content.
A staging Disallow: / that survives a site launch is the classic disaster. Check this first after any migration.

The trap that catches people: if you want a page out of the index, you can't block it in robots.txt and also put noindex on it. Google never fetches the page, so it never sees the noindex. The page can stay indexed indefinitely. To remove something, let Google crawl it and see the noindex.

Step 3: Is indexing blocked?

If Google can crawl the page, the next question is whether you've told it not to index. There are two places to check.

  1. The meta robots tag in the HTML head:

html

  1. The X-Robots-Tag HTTP header:

X-Robots-Tag: noindex

The header version is easy to miss because it never shows up in "view source." It's also the standard way to noindex non-HTML files like PDFs. If you've checked the HTML and found nothing, check the headers with curl -I.

Common ways a noindex ends up where it shouldn't be:

A CMS setting like "Discourage search engines from indexing this site" left on after launch
An SEO plugin applying noindex to a whole post type, category or tag archive
A staging environment's config copied to production
A template-level tag that gets inherited by every page

If Search Console says "Excluded by 'noindex' tag," it's telling you the truth. The only real question is whether you meant it.

Step 4: Is there a canonical pointing somewhere else?

A canonical tag tells Google which version of a page you consider the main one:

html

Used correctly, it consolidates duplicate or near-duplicate URLs. Used incorrectly, it can quietly remove a page from the index.

Things to check:

Does the canonical point to a different URL? If your page says its canonical is another page, you've told Google to index that other one instead.
Is the canonical pointing at a redirect, a 404 or a noindexed page? Mixed signals like this make Google ignore the tag.
Is the same canonical on every page? A template bug that canonicalizes everything to the homepage is more common than you'd think.
Do the canonical, sitemap entry and internal links all agree on one URL? Conflicting signals make Google choose for itself.

Remember that a canonical is a hint, not a command. Google can override it. In Search Console, two statuses matter here:

"Alternate page with proper canonical tag": usually fine. That page is a deliberate duplicate.
"Duplicate, Google chose different canonical than user": Google disagreed with your tag and picked another URL. Look at why the two pages seem the same.
Step 5: Is Google choosing not to index it?

Now we get into the harder territory. Everything technical is clean, and the page still isn't indexed. Search Console usually gives you one of two statuses that look similar but mean very different things.

Discovered, currently not indexed

Google knows the URL exists (from a sitemap or a link) but hasn't crawled it yet.
It often points to crawl priority: a large site, a slow server, or little perceived value in the section.
Fixes: improve internal linking to the page, make sure the server is fast and stable, and prune low-value URLs competing for attention.

Crawled, currently not indexed

Google fetched the page and decided not to keep it, for now.
This is almost always a quality or duplication signal, not a technical one.
Fixes: make the page more useful, more distinct and better connected to the rest of your site.

Telling these two apart saves a lot of guessing. "Discovered" is mainly a crawling story. "Crawled" is mainly an indexing story.

Duplicate content

Duplicate content isn't a penalty. Google groups similar pages into a cluster and picks one to show. The usual sources are:

URL parameters (?sort=price, tracking parameters)
HTTP vs HTTPS and www vs non-www versions
Trailing slash vs no trailing slash
Printer-friendly pages
Boilerplate-heavy pages where only a sentence or two differs, like location pages with swapped city names

If your page is in a cluster and loses, it shows up as excluded, even though nothing is technically wrong with it.

Soft 404s

A soft 404 is a page that returns a 200 OK status but looks like an error or an empty page to Google:

A "no results found" page that returns 200
A product page that says "out of stock" with no other content
A near-empty template with almost no text
Redirecting deleted pages to an irrelevant page

Google decides these are effectively not-found pages and leaves them out. The fix is to be honest with your status codes: return a real 404 or 410 for pages that are gone, or add real content to pages that should exist.

Step 6: Is the page valuable enough?

If you've cleared every step above, you're left with the uncomfortable question. Google has no reason to reject the page technically. It just doesn't think the page earns a spot.

Signals worth examining:

Thin content: a few generic sentences, no original information or depth.
Rehashed content: the same points as ten other pages, with nothing new.
Weak internal linking: orphan pages (nothing links to them) signal that you don't value them either.
Site-wide quality: if much of your site is thin, Google may be more hesitant about everything on it.
Mismatch with intent: the page doesn't clearly answer a question anyone is asking.
Little to no external signals: a brand-new site or page with no links or mentions takes longer to earn trust.

Practical things that help:

Add information the other results don't have: examples, data, your own experience
Link to the page from relevant, already-indexed pages
Merge several weak pages into one strong one
Remove or noindex pages that exist only to fill space
Make the page's purpose obvious from the title, headings and first paragraph

There's no switch to flip here. Quality problems get fixed by making the page better, and then waiting for Google to re-evaluate.

Putting it into practice

When a page is missing, I run this in about ten minutes:

Search Console URL Inspection: paste the URL. It tells you whether the URL is on Google, when it was last crawled, and which canonical Google chose.
Run the live test: it shows what Google sees right now, including whether the page is blocked and what the rendered HTML looks like.
Check robots.txt: read it for any rule matching the path.
curl -I the URL: confirm the status code and look for X-Robots-Tag and redirects.
View source: search for noindex and canonical.
Check the sitemap: is the URL listed, and is it the exact version you want indexed?
Open the Page indexing report: find the exact status for the page. That label usually tells you which branch of the tree you're on.

One more tip: the site:yourdomain.com/page search is a quick sanity check, but it isn't a reliable way to confirm indexing. Trust Search Console over it.

Quick reference
Blocked in robots.txt: crawling is blocked. The URL might still be indexed. Remove the rule.
noindex present: indexing is blocked. Remove the tag.
Canonical to another URL: you told Google to prefer something else. Fix or confirm the tag.
404 / 410: the page is gone. Restore it or leave it gone on purpose.
Soft 404: the page looks empty. Add real content or return a real error code.
Redirect issue: shorten the chain and point to a relevant destination.
Discovered, not indexed: crawl priority issue. Improve linking and server health.
Crawled, not indexed: quality issue. Improve the page.
Final thoughts

The big shift is to stop asking "why won't Google index my page?" and start asking "which question in the chain failed?" Technical blocks answer yes or no, and you can fix them in minutes. The quality questions take longer, but they're the ones that decide whether a page deserves to rank at all.

Next time a page is missing, walk the tree from top to bottom before changing anything. Most of the time the answer shows up in the first three steps.

If you've got a weird indexing case that doesn't fit this flow, drop it in the comments. Those are usually the interesting ones.

Also if you're interested , feel free to check out my website - anoshbb.com

Top comments (0)