Correction (originally published as "Bing indexed 0 pages of a 133-page site"). The bug I describe below is real and the fix is verified. The headline conclusion was wrong, and the diagnostic I recommended was invalid. Both are corrected here, in place, with the evidence.
What was actually broken
pureip.app is an IP diagnostics toolbox: Cloudflare Workers plus a single-page app. This line in wrangler.toml:
[assets]
not_found_handling = "single-page-application"
is the standard Workers SPA setting. It returns index.html with status 200 for any unmatched path. So:
/wp-admin → 200 (returns the homepage)
/this-page-does-not-exist → 200 (returns the homepage)
/phpmyadmin → 200 (returns the homepage)
/anything-you-type → 200 (returns the homepage)
An unlimited supply of identical 200-status pages, from a crawler's point of view. That was a real defect and it is now fixed: those paths return 404. Verified against the live site, not from memory:
/wp-admin → 404
/phpmyadmin → 404
/this-does-not-exist-12345 → 404
/ip → 200
/guides → 200
/ → 200
The fix is a top-level path allowlist in the Worker; anything not on it gets a real 404:
const KNOWN_TOP = new Set(["", "ai", "ai-check", "ai-unlock", "all", "api",
"articles", "assets", "browser", "cdn", "claude", "compare", /* ... 45 segments */]);
if (!/\.[a-z0-9]{1,5}$/i.test(P) && !KNOWN_TOP.has(P.split("/")[1].toLowerCase())) {
return new Response("Not found", { status: 404 });
}
The hard part is completing the allowlist. Miss one legitimate path and you de-index pages that were already indexed. I cross-checked five sources: the full sitemap (134 URLs across 8 sitemaps, all HTTP 200), the site's own /all index page, the App.tsx route table, legacyRoutes (22 entries), and the path prefixes the Worker handles itself. Then a two-way regression, re-run for this correction: 151 / 151 legitimate paths return 200 (the sitemap union plus the route table) and 19 / 19 junk paths return 404.
The conclusion that was wrong
I claimed Bing had indexed zero pages of 134, and that the soft 404 was why. Then I recommended a diagnostic for other people to use:
Run a Bing
site:query. If it returns unrelated domains, that means no page from your domain is in the index at all.
That diagnostic is invalid, and I can prove it with a control I should have run the first time. Running site:cloudflare.com the same way returned pages from arsenal.com — a football club. Cloudflare is one of the most thoroughly indexed domains on the web. If the operator returns unrelated domains for site:cloudflare.com, then "unrelated domains came back" tells you nothing about any domain's indexing status.
The site: operator was returning filler, not a filter result. What I read as evidence of absence was evidence of a broken tool.
Repeated measurements made the failure mode clearer. The same query, minutes apart, returned three completely different things: unrelated domains, then a full page of results, then nothing at all. And the "nothing at all" state applied to control queries too — site:cloudflare.com and site:stripe.com both returned zero results at the same moment site:pureip.app did. When your control goes to zero, the measurement is broken, not the thing you are measuring.
What I actually know now
Via DuckDuckGo, which serves Bing's index, I got a clean reading once:
site:pureip.app → 7 results
pureip.app/research/ip-audit-2026-09/en/
pureip.app/research/vps-audit-2026-09/en/
pureip.app
pureip.app/ai-unlock
pureip.app/guides/alibaba-cloud-ip-check
pureip.app/guides/greencloudvps-ip-check
pureip.app/guides/bandwagonhost-ip-check
So Bing had indexed this domain. Not zero. I could not repeat that measurement afterwards — every endpoint degraded into rate-limited empty responses — so I cannot tell you the number, and I am not going to guess it. The honest state of knowledge is: more than zero, less than everything, and I have no reliable instrument for the exact figure from this environment.
The genuinely useful takeaway is smaller and more boring than the original post implied. I found a real soft-404 defect, fixed it, and verified the fix on both sides. I did not establish that Bing was refusing the domain, and I should not have said I had.
Why this happens
Scraped search results are the worst possible instrument for an indexing decision, because the failure mode is silent and looks exactly like the answer you were hoping for.
- A blocked or throttled endpoint returns an empty result set. Regex-counting a scraped SERP reports "0 indexed" whenever the HTML changes, the request is throttled, or a bot check appears. It never reports "I don't know."
-
site:is not a lookup, it is a query. It goes through ranking and can be ignored, substituted, or filled with unrelated results under load. -
Datacenter IPs get a degraded internet. Google served this machine's requests a stripped page with no result count and no result titles. Bing served local shopping results for commercial queries and unrelated domains for
site:queries. - Absence of evidence reads as evidence of absence when the thing you are measuring is a count that can legitimately be zero.
Always run a positive control. Query a domain you know is indexed, with the identical code path, in the same minute. If the control fails, you have measured nothing. I skipped this and published a number.
Reusable lessons
1. A control is not optional when the expected answer is zero. If your instrument can silently return "0", you cannot distinguish a real zero from a broken instrument without a known-positive comparison.
2. Never publish an indexing number you cannot reproduce. I had one reading and a large pile of unreliable ones. That is not a measurement, and the headline should not have carried it.
3. "All checks pass" is still a clue. robots.txt clean, sitemaps 200, crawler visiting daily, submission returning 200 — if those are all green and coverage looks bad, your measurement deserves re-checking before your infrastructure does.
4. Fix both directions. Verifying "junk now 404s" is not enough. Killing a legitimate page costs more than the junk was worth. 151 legitimate paths verified 200, zero collateral damage.
5. Wait for propagation. During edge rollout the same path alternates between 200 and 404. Do not conclude anything from a single probe.
Bugs fixed along the way
Legacy status pages were being killed. /gpt/status.html and /claude/status.html are legitimate old links in legacyRoutes that should 301 to /status/openai. Because they carry an .html extension, the earlier extension rule 404-ed them. After the fix the suite is 142 pass / 0 fail (re-run to confirm).
A brittle assertion. One test read tree.props.children[1].props.children[0].length. A later change added a <HomePageSearch /> component which took over children[1], shifting every index. The functionality was fine — the test was positional. Rewritten to find nodes by predicate, plus the missing negative case:
assert.equal(merged, 1, "same IP should collapse to 1 card");
assert.equal(distinct, 2, "different IPs should produce 2 cards");
The old test only verified "same IP merges". It never verified "different IPs produce two cards", so if deduplication had swallowed everything it would still have passed. Mutation-tested after the rewrite: delete the dedupe line and the assertion fails immediately.
Tool: PureIP — check an IP before you buy a VPS: datacenter vs residential, blocklists, reverse DNS, AI-service reachability. Free, no signup.
Every figure in this post was re-measured against the live site and the local source tree on the date of this correction. The Bing indexing figure is the one thing I could not re-measure, and it is labelled as such rather than stated.
Top comments (0)