DEV Community

Hammad Shams Uddin
Hammad Shams Uddin

Posted on

My monitoring said 26 pages were broken. Every one of them was fine.

This morning's report from my own daily checker opened like this:

NEEDS YOU — 27 thing(s)
  * / returns HTTP 403 but had 33 impressions
  * /es/tools/text-summarizer returns HTTP 403 but had 11 impressions
  * /tools/text-to-slug returns HTTP 520 but had 11 impressions
  … 23 more …
  * sitemap.xml returned nothing usable
Enter fullscreen mode Exit fullscreen mode

Twenty-six pages and the sitemap. The homepage among them. My first thought was that the host had suspended the account — SSH to the box was timing out at the same moment, which fitted.

Every one of those pages was serving normally.

What the site was actually doing

Before touching anything I measured the same URL through different clients. Forty requests, alternating, same minute:

client result
curl sending a Chrome user agent 20 of 20 → 403
curl sending Googlebot's user agent 20 of 20 → 200
real headless Chrome 5 of 5 → 200, full content
curl sending curl/8.0 mixed, mostly 200

Read that table twice. The requests that claimed to be a browser were refused. The request that claimed to be Googlebot was served. An actual browser was served. So the site was fine for humans, fine for Google — and hostile to exactly one population: things that say "Chrome" without being Chrome.

That is not a fault. That is bot protection at the edge doing its job. Cloudflare-style protection compares what a client claims against how it behaves, and a curl request wearing a Chrome user agent is the oldest impersonation in the book. My checker had been making that exact request every morning for months, and the day the protection tightened, my checker became a bot to be blocked.

The 520s were the same story from a different angle: a proxy failing to get a clean answer from the origin, on requests that were being refused anyway.

The bug is not the block. It is the report.

The block is somebody else's defence working. The failure that cost me an hour is that my own script turned I was refused into the page is broken — twenty-six times, in a list headed NEEDS YOU.

Four changes, and they are all the same change:

1. Ask whether you can see, before saying what you see.

function seoCanSee(): array
{
    // fetch the homepage, then:
    if ($code === 200 && stripos($body, SEO_MARKER) !== false) {
        return ['ok' => true, 'code' => $code, 'why' => ''];
    }
    return ['ok' => false, 'code' => $code,
            'why' => 'the homepage answered HTTP ' . $code . ' to this script'];
}
Enter fullscreen mode Exit fullscreen mode

If that fails, nothing downstream is allowed to report on individual pages at all. One line goes out — this script cannot reach the site; page, sitemap and robots checks did NOT run; a browser and Googlebot may well be served normally — and the per-page loop never runs.

2. Look at the body, not only at the status. A 403 from my application would carry my page. A 403 from the edge carries the edge's page. Same integer, different fact:

if (($st['code'] === 403 || $st['code'] === 520) && !$st['ours']) { $edgeBlocked++; continue; }
Enter fullscreen mode Exit fullscreen mode

3. Report one fact once. The blockage is upstream of every page, so it is one line — 3 pages could not be checked: our own edge answered 403 to this script — not one alarm per page. Twenty-six alarms for one cause is how a NEEDS YOU list becomes wallpaper.

4. Check what your loudest alarm does when the probe is blocked. Mine compares the homepage as Googlebot sees it against the homepage as a browser sees it, and shouts CLOAKING if they differ — the alarm that means your site is hacked. With browser-UA requests being refused, the "browser" side was a 403 page, whose <title> is 403 Forbidden. Two different titles. That alarm was one edge-rule away from telling me I had been hacked again, three weeks after I actually had been. Now a response that is not my page is not a comparison — it is "did not see it".

After all four: 27 items became 4, and the four are true.

If you are chasing the same ghost

Three things worth doing before you believe your own monitoring:

  • Load the page in a real browser. Not curl with a browser string — an actual browser. If it renders, your users are fine and your problem is your client.
  • Look at what your client is really sending. Your checker, your CI job, your uptime probe: each one has a user agent, and it may not be what you think. See exactly what a client sends, then make your script say something honest rather than dressing up as Chrome.
  • Read the whole response, not the code. The headers a URL actually returns will usually name the thing that answered — server: cloudflare, x-turbo-charged-by: LiteSpeed, a cf-ray id — and that tells you whether you reached your application or something standing in front of it.

The uncomfortable part is that none of this was a site problem for a single minute. Every visitor and every crawler was served the whole time. The only thing that was down was my ability to tell.

Top comments (0)