Four days ago I gave my site its own crawler log, because the host keeps access logs somewhere a script cannot read them and Search Console answers two days late. Here is the first full read of it.
bingbot 9,629 hits bytespider 357
googlebot 183 hits yandexbot 94
And the status codes those two were served:
| day | bingbot 200 | bingbot 410 | googlebot 200 | googlebot other |
|---|---|---|---|---|
| 29 Sep (from 15:51) | 70 | 982 | 45 | — |
| 30 Sep | 198 | 3,914 | 66 | — |
| 1 Oct | 172 | 2,487 | 54 | 1× 404 |
| 2 Oct (to 16:03) | 123 | 1,653 | 17 | — |
9,036 of bingbot's 9,629 requests — 94% — were to URLs that do not exist. 9,015 of them were distinct. Every one has the same shape: a dictionary word, a slash, a word with digits pushed between its letters.
They are not Bing's fault. In September this site was compromised, and the malware handed every crawler that asked for a page roughly 291 invented URLs. It did that for six days. The malware has been gone for weeks; the queue it created is still being worked through, at about two thousand requests a day.
Googlebot, over the same four days: 183 requests, 135 distinct real pages, and not one spam URL. Two crawlers, the same site, the same week, completely different pictures. If I had only Search Console I would have had neither.
What changed in those four days
On the 29th those dead URLs were answering 404. A 404 means not here right now, which is an invitation to come back. I switched the cloaker's URL shape to 410 Gone — stop asking — after checking the pattern against every URL in my sitemap (9 of 9 known spam paths matched; 0 of 6,240 real ones did).
The daily 410 count since: 3,914 → 2,487 → 1,653. The last day is two-thirds of a day, so it is not yet a trend I would bet on, but it is pointing the right way and it is the first number I have had at all.
The rule itself lives inside the 404 handler, not at the top of the router. That ordering is the whole safety argument: a request only reaches notFound() once it has failed to match anything real, so a live route cannot be caught by the pattern — and one of mine, /s/{code}, has exactly the same shape as the spam.
"Unverified" was hiding two different facts
I had the log check each crawler's IP against the published address ranges — Google's googlebot.json, Bing's bingbot.json — and report verified against unverified. A reader of the first write-up, who runs the same check over nginx logs, pointed out that my split quietly merged two things that are not alike:
- a request that failed the list: claims to be Googlebot, is not in Google's ranges — suspicious;
- a request with no list to check: Anthropic publishes no ranges at all, so a ClaudeBot hit can never be more than unchecked — not suspicious, just unknowable.
Collapsing those is the same error as calling an unreachable sitemap "empty", which this project has made more than once. So there are three counters now, and the digest prints what it actually knows:
googlebot 183 hits (178 verified, 5 FAILED the list)
yandexbot 94 hits (94 no published list)
claudebot 39 hits (39 no published list)
google-inspection 2 hits (2 FAILED the list)
The same reader made a second point I have taken: these lists move, so a verification you made today will not reproduce against next month's list. The cache now keeps a dated snapshot of every list it fetches, sixty days of them, which costs a few hundred kilobytes and makes an old answer checkable.
Those 5 Googlebot hits that failed the list are all mine, and I had to look at the addresses to know it: two came from my laptop testing with a spoofed agent, and three from the server itself — my own daily cloaking check, which asks for the homepage as Googlebot and compares it against the homepage as a browser. My monitoring impersonates Googlebot for a living, so of course it fails a Googlebot address check. The log recorded all five exactly as it was told, which is the point: a user agent is a claim, and a claim is not an identity.
If you want the same visibility
You do not need a log pipeline. Roughly 120 lines of PHP: match the user agent against a list of crawler tokens, append one line at shutdown — time, crawler, status served, path, IP — and never write a line for a human. Then:
- Count distinct paths, not requests. 9,629 hits sounded like a crawl-rate problem; 9,015 distinct dead URLs is a completely different diagnosis. Any list of lines will collapse to its unique entries in a second, and the gap between the two numbers is usually where the story is.
-
Check what you are telling crawlers to skip. Mine already blocks the training crawlers I want blocked and nothing else — worth confirming rather than assuming, because a
Disallowin the wrong user-agent group reads as harmless and is not. - Log the status you served, not the path requested. The interesting line is never "Googlebot asked for X". It is "Googlebot was given a 404 for X".
Three weeks after the site was clean, most of one search engine's attention was still going to the attacker's imagination. Nothing in Search Console would have told me that, and nothing in my code was telling me either — until the site started writing it down.
Top comments (0)