DEV Community

Hammad Shams Uddin
Hammad Shams Uddin

Posted on

I gave my site an access log. Within a minute it showed Bing crawling a hacker''s URLs.

My host keeps access logs in a control panel I have to click through, and nowhere a script can read them. So after September's compromise I could not answer the one question that mattered — are the crawlers coming back, and what are they being served? — except by waiting two days for Search Console to round it into a weekly shape.

So I gave the site its own log. About 120 lines of PHP: if the user agent names a known crawler, append one line at shutdown with the time, the crawler, the status code it was actually served, the path and the IP. Humans are never written to it.

The first day's output, in full, from the first minute:

15:51:01  bingbot  200  /calculators/pace-calculator          40.77.167.123
15:51:34  bingbot  404  /stimulated/a0d3o2b2o25378994         40.77.167.123
15:51:36  bingbot  404  /harshly/p0o9s1i9t9i3o7n38906         40.77.167.123
15:51:38  bingbot  404  /disservice/u1n6g2u8e5s4670872        40.77.167.123
15:51:40  bingbot  404  /dictation/g0o6g3g2l3e3e1y3e0264      40.77.167.123
15:51:41  bingbot  404  /coadjust/v1i4m1i5n4a2r1i8a6529       40.77.167.123
15:51:48  bingbot  404  /selfapproving/p1r3e2l1e9c2t1e3d9648  40.77.167.123
…
Enter fullscreen mode Exit fullscreen mode

By the end of the day: 24 of 26 crawler requests were to URLs a cloaker invented three weeks earlier.

Those paths are not random junk. They are the exact shape the malware generated — a dictionary word, a slash, a second word with digits pushed between its letters — and it handed about 291 fresh ones to every crawler that asked for a page, for six days, before the site was cleaned. The malware has been gone for weeks. The queue it created has not.

404 was the wrong answer

Every one of those URLs was returning 404, which is correct in the sense that they are not there. But 404 means not here right now, and crawlers treat it as a reason to come back and check again later. That is why they are still working through a list from three weeks ago.

410 means gone; stop asking. Both Google and Bing drop a 410 faster than a 404, and the difference is exactly the behaviour I want.

The risk in a rule like this is obvious: match too broadly and you 410 something real. So I wrote the shape, then measured it before it went anywhere near the site:

$rule = function (string $path): bool {
    $seg = explode('/', trim($path, '/'));
    if (count($seg) !== 2) return false;
    $s = $seg[1];
    if (strpos($s, '-') !== false) return false;     // our slugs are hyphenated
    if (strlen($s) < 10) return false;
    return preg_match_all('~[0-9]~', $s) >= 4
        && preg_match_all('~[a-z]~i', $s) >= 4;
};
Enter fullscreen mode Exit fullscreen mode

Result: 9 of 9 known spam paths matched, 0 of 6,240 real URLs did.

One more precaution, and it is the part I would get wrong if I were rushing. The rule does not run at the start of routing. It runs at the end, inside the 404 handler — the place a request only reaches once it has failed to match anything real:

function notFound(): void
{
    // …if the path has the cloaker's shape, 410 instead of 404…
    http_response_code(410);
    header('X-Robots-Tag: noindex');
    exit;
}
Enter fullscreen mode Exit fullscreen mode

That ordering is not decoration. My site has a /s/{code} short-link route: two segments, no hyphen, letters and digits — it matches the shape. Put the rule first and every short link a user shares becomes 410 Gone. Put it last, and a live route can never be caught by it, because a live route never reaches the 404 handler at all.

Verified on the live site afterwards:

/accidence/d0e3p2o6s2i8t6a6r8y548   410      (cloaker shape)
/no-such-page                        404      (ordinary typo)
/s/a1b2c3d4e5                        302      (short link, still works)
/tools/json-formatter                200
Enter fullscreen mode Exit fullscreen mode

The log lied to me once on day one

One line said googlebot 200 /. I was pleased for about ten seconds, until I looked at the IP: 39.34.131.166 — my own network. It was me, testing with a spoofed user agent earlier that afternoon. The log had recorded exactly what it was told, which is all a user agent ever is: a claim anybody can type.

So the log now checks the claim against the crawlers' own published address lists — Google's googlebot.json and special-crawlers.json, Bing's bingbot.json, cached for a day — and reports both numbers:

bingbot     25 hits (25 verified, 0 unverified)
googlebot    1 hit  (0 verified, 1 unverified)
Enter fullscreen mode Exit fullscreen mode

And the alarm that says "Googlebot has not come" now counts verified visits only, because an alarm that can be silenced by anyone typing Googlebot into a user agent string is not an alarm.

If you are running this kind of check yourself: your crawler's status codes are what it believes about your site, so log what was served, not what was requested. And if you end up needing to retire a batch of URLs permanently, an .htaccess rule is the cheaper place to do it than in application code — as long as you have measured the pattern against your real URLs first.

Three weeks after the cleanup, most of my crawl budget was still being spent on the attacker's imagination. I only know that because the site finally started writing things down.

Top comments (1)

Collapse
 
danielecangi profile image
DaC •

The most interesting part is certainly observing what the Internet continues to do as a consequence of what the malware had done