The worst scraping failure I've dealt with didn't throw a single error.
No 403. No captcha. No timeout. The request came back HTTP 200, 884 KB of HTML, in 557 ms. The parser ran fine. And it returned 0 listings.
If nobody is watching, that run gets logged as a success. So does the next one, and the one after that. You find out weeks later, when someone asks why the price chart went flat on the 12th.
Here's what happened, how I fixed it, and the checks I now put on every scraper I run.
WHAT CHANGED
The target was kleinanzeigen.de, a large German classifieds site. One day they migrated their front end to utility CSS classes. Same pages, same URLs, same data on screen. Different markup underneath.
My parser was anchored on class names. I counted how often each selector matched on one results page, before and after:
| Selector | Before | After |
|---|---|---|
.aditem |
27 | 0 |
.price-shipping--price |
27 | 0 |
.ellipsis |
27 | 3 |
Three anchors dead overnight, a fourth barely alive. The page itself was perfectly healthy.
WHAT DIDN'T CHANGE
Then I counted the things that describe the data rather than the design:
| Anchor | After the redesign |
|---|---|
[data-adid] |
27 |
[data-href] |
28 |
<article> |
27 |
| JSON-LD blocks | 28 |
Every one of them survived. That's the whole lesson in one table: CSS classes belong to the designers, data attributes and structured data belong to the product. Designers change styling all the time. The product team rarely touches the IDs that the site's own JavaScript depends on.
THE FIX
I rebuilt the parser around anchors that describe the data instead of the design, with a fallback layer for when those move too.
Result: 0 → 505 listings across 4 requests, 100% field completeness on price, mileage, year and location.
THE TRAP INSIDE THE FIX
One bug cost me more time than the redesign itself.
To read the visible text, you strip the HTML tags. Done the obvious way, it quietly glues neighbouring words together: 150.000 km and EZ 03/2019 become kmEZ. The mileage regex stops matching. No error, just empty fields. A second silent failure inside the fix for the first one.
The correction is one character. Finding it is the hard part, because nothing tells you it's broken.
WHAT I CHECK ON EVERY RUN NOW
Status codes tell you whether the server answered. They say nothing about whether you got the data. So every run gets judged on its output:
- Zero rows on a 200 is a failure, never an empty day.
- Volume against the recent average: a sudden drop means something changed.
- Field completeness: a parser can find every row and still lose a field.
Any of these trips an alert within minutes, instead of someone noticing weeks later.
TAKEAWAYS
- A 200 is not a success. Judge a run by what it extracted.
- Anchor on data, not on design. CSS changes far more often than the data behind it.
- The question isn't whether your scraper will break. It's how fast you'll know.
I build and maintain scrapers for protected and fast-changing sites, with this kind of monitoring built in. If your scraper is quietly returning less than it should, or getting blocked, reach out on Upwork and I'll take a look.
Top comments (0)