DEV Community

Max
Max

Posted on Originally published at storecanary.io

Why is my Shopify product page unavailable in Merchant Center?

TL;DR

  • "Product page unavailable" means Google's Shopping crawler (Storebot-Google) failed to fetch that product URL on its last attempt, it does not mean the page itself is broken.
  • This check runs separately from Search Console's indexing report, on a different schedule and sometimes a different crawler, so a page can be fully indexed and still fail this check.
  • On Shopify the cause is almost always something in front of the app: a redirect, an edited robots.txt, a proxy/WAF/bot-protection layer, or DNS, not the theme or the product data itself.

If you build or maintain Shopify storefronts for clients, you'll eventually get a message like "Google says our products are unavailable, but the site works fine." The client opens the product URL, it loads. You open it, it loads. Merchant Center still shows "Product page unavailable" on some or all products, and running a manual recrawl only clears it for a few days before it comes back.

This isn't a theme bug, and you won't find it by inspecting the product in admin. Merchant Center's own error list for this issue is almost entirely about the request Google made, not about the page's content: redirects, DNS, a robots.txt Google couldn't reach, a server that didn't respond in time. That means the fix usually lives in whatever sits between the domain and Shopify, a layer a developer configured or an app installed, not in the storefront code.

Here's what the check actually verifies, why it can disagree with Search Console, and how to track down the real cause in a client's stack instead of guessing at the product.

What does "Product page unavailable" actually check?

Google's help page for this issue is specific: "For one or more of your products, you submitted a landing page ... that can't be accessed by Google and may also not be accessible to Google users."

The listed causes are almost all request failures, not content problems:

  • landing page "was not found on the server" (404)
  • "Your landing page redirects too many times"
  • "the server couldn't process the request"
  • robots.txt "couldn't be reached"
  • "We weren't able to resolve the hostname"

Four out of five are infrastructure. On platforms where merchants hand-write feed URLs, "you submitted a bad link" is a reasonable first guess. On Shopify the product link is generated from the handle, the same one the sitemap and theme use, so a wrong URL is the fastest thing to rule out, not the most likely cause.

Why does Search Console say the page is fine while Merchant Center flags it?

Because they're different checks, on different schedules, and not necessarily the same crawler. Google's crawler docs list Storebot-Google (user agent token Storebot-Google) as the crawler for "all surfaces of Google Shopping," distinct from Googlebot, which serves Search. Merchant Center's product crawl issues page tells merchants to allow both "Googlebot" (used for landing pages) and "Googlebot-image."

The matching rule in robots.txt matters here. Google states: "you need to match only one crawler token for a rule to apply." A named User-agent: Storebot-Google or User-agent: AdsBot-Google group pulls that crawler out of the global * group entirely. For AdsBot specifically, Google says "the global user agent (*) is ignored." A single robots.txt file can enforce three different crawl policies at once without anyone writing three on purpose.

On stock Shopify stores this usually isn't the problem: checking seven live Shopify storefronts' public robots.txt files, six were readable, all six had a global User-agent: * group plus a named adsbot-google group, and none named Storebot-Google anywhere. Shopping's crawler falls under the global (permissive) group by default. If a client's store has had robots.txt edited since (an app, a developer, an SEO consultant), that default no longer applies. Worth checking against what a blocked-by-robots.txt issue looks like if you find a suspicious custom group.

Why doesn't "it opens fine in my browser" prove the page is reachable?

Because reachability isn't a property of a URL, it's a property of one request, from one caller, at one moment. Bot-protection and WAF layers exist specifically to answer different callers differently. Google's own troubleshooting page names this directly: "Your website may be behind a firewall, have geo-restrictions, or use a private IP that blocks our crawlers," and recommends allowing crawling from the US and allowlisting "the Googlebot user agent" at the server or firewall level.

A quick test run on 19 September 2026 makes the point concretely: of seven live Shopify storefronts, four served product pages identically (same 200 status, near-identical body) to both a Googlebot and a Storebot-Google user agent string. A fifth returned HTTP 403 on every single request, homepage, robots.txt, sitemap, regardless of user agent, from an edge/bot-protection provider, serving an interstitial page marked noindex,nofollow. The storefront was presumably fine; a shopper on a residential connection would never see that response. The test requests came from datacenter IPs, exactly the kind of caller these systems are tuned to challenge, and it's a reasonable proxy for what happens when Google's own crawler IP ranges hit the same rule.

Two practical consequences follow. Testing with a spoofed user-agent header from your own laptop tells you almost nothing, since the requesting address is what's being judged, not the header string. And a failure like this is intermittent by construction, since rate limits and reputation scores shift on their own, which is exactly why a manual recrawl can "fix" it for a few days and then it comes back.

How do you find the actual cause in a client's stack?

Work through these in order. Each step rules out more than the one after it, for less effort:

  1. Count the affected products. One product flagged points at that product's link specifically. The whole catalog flagged points at something in front of the domain; no per-product edit will touch that.
  2. Diff the submitted URL against what the store actually serves. Does it redirect? Does the host match the primary domain? An apex domain on a store that serves www, or a stale secondary domain, adds a redirect hop to every check, and "redirects too many times" is a documented cause.
  3. Request the URL from outside the client's own network (a server, a different connection) and compare the status code to what you get locally. This is the only step that can surface a rule aimed at a caller other than you.
  4. Read every User-agent block in /robots.txt, and confirm the file returns 200 (an unreachable robots.txt is its own documented cause). Check whether a named group has pulled a crawler out of the default global rule.
  5. Identify anything sitting between the domain and Shopify. Check response headers for a proxy or CDN that isn't Shopify, and check DNS for a provider that proxies traffic rather than only resolving it. This layer answers before Shopify ever sees the request, so nothing in Shopify admin will surface it.
  6. Trigger a website recrawl in Merchant Center and log the date. Google says this can take "up to 12 hours." If the flag clears with zero changes on your end, the page was never the problem; the request was.

Which causes are worth checking first on Shopify?

Wrong domain in the feed. Cheap to rule out first. If the feed carries an apex URL and the store redirects to www (or references a domain the client migrated off), every check burns a redirect hop.

Unreachable robots.txt. A crawler that can't read your rules doesn't assume permission, it stops entirely, taking every product page down with it even though none of them was directly requested.

An edited robots.txt.liquid. Shopify's docs say the default robots.txt is already "optimal for Search Engine Optimization," and warn that customizing it (blocking crawlers, adding crawl-delay rules) can result in "loss of all traffic" if done incorrectly. A crawl-delay added months ago to slow an aggressive scraper, or a named group copied from a forum thread, keeps applying long after anyone remembers adding it, and it lives in a theme file, invisible from admin. Worth checking whether any app, developer, or SEO tool has ever touched this file.

A proxy/WAF/bot-protection layer refusing the caller. The most likely cause for a catalog-wide, recurring flag. Same family of causes as a 403 in Search Console: a CDN rule, a geo-restriction, a third-party service mounted on the domain.

An allowlist keyed on user-agent string instead of IP range. The common "fix," allowing the Googlebot user agent, doesn't hold up on its own: a user-agent string is trivially spoofable, so a protection layer that trusts it has a reason not to, and one that ignores it entirely won't budge no matter what string you add. The durable fix is allowlisting Google's published crawler IP ranges and re-testing from outside the network.

A slow server response under load. Google's list includes "the server couldn't process the request." A request that times out during a large import or a traffic spike leaves no trace afterward, the page is fast again by the time anyone checks. If flags cluster around a known load event, that correlation is the only evidence you'll get.

Why does the flag come back after a recrawl clears it?

Because a recrawl re-asks the question, it doesn't fix anything. If the original failure came from something that varies by request (a rate limit, a bot-protection score, a momentary DNS hiccup), asking again later gets a better answer and the flag clears. Nothing was actually resolved, so the next unlucky request re-triggers it, one product at a time, which reads as a slow drift rather than an outage.

That's the expensive part. Google says a fixed store can take "up to 12 hours" to review, and Shopping visibility recovers gradually after that. Every relapse costs days of visibility on a product that was in stock and correctly configured the entire time. For a developer or agency managing several client stores, this is worth monitoring continuously rather than closing the ticket after one fix, since the trigger is usually outside your code: a hosting provider tightening a default, a security app installed for an unrelated reason, a DNS migration, a stricter fraud rule during a busy week. Shopify admin keeps showing the product as active and published the whole time, because from Shopify's point of view, the request never arrived.

Have you run into this on a client's store, and did the actual cause turn out to be robots.txt, a WAF/proxy layer, or something else entirely?

The full version of this article — with screenshots and ongoing updates — lives on the StoreCanary blog.

Top comments (0)