DEV Community

Cover image for My scraper returned HTTP 200 and zero rows. Nothing crashed.
Hugo.H
Hugo.H

Posted on

My scraper returned HTTP 200 and zero rows. Nothing crashed.

The worst scraping failure I've dealt with didn't throw a single error.

No 403. No captcha. No timeout. The request came back HTTP 200, 884 KB of HTML, in 557 ms. The parser ran fine. And it returned 0 listings.

If nobody is watching, that run gets logged as a success. So does the next one, and the one after that. You find out weeks later, when someone asks why the price chart went flat on the 12th.

Here's what happened, how I fixed it, and the checks I now put on every scraper I run.

WHAT CHANGED

The target was kleinanzeigen.de, a large German classifieds site. One day they migrated their front end to utility CSS classes. Same pages, same URLs, same data on screen. Different markup underneath.

My parser was anchored on class names. I counted how often each selector matched on one results page, before and after:

Selector Before After
.aditem 27 0
.price-shipping--price 27 0
.ellipsis 27 3

Three anchors dead overnight, a fourth barely alive. The page itself was perfectly healthy.

WHAT DIDN'T CHANGE

Then I counted the things that describe the data rather than the design:

Anchor After the redesign
[data-adid] 27
[data-href] 28
<article> 27
JSON-LD blocks 28

Every one of them survived. That's the whole lesson in one table: CSS classes belong to the designers, data attributes and structured data belong to the product. Designers change styling all the time. The product team rarely touches the IDs that the site's own JavaScript depends on.

THE FIX

I rebuilt the parser around anchors that describe the data instead of the design, with a fallback layer for when those move too.

Result: 0 → 505 listings across 4 requests, 100% field completeness on price, mileage, year and location.

THE TRAP INSIDE THE FIX

One bug cost me more time than the redesign itself.

To read the visible text, you strip the HTML tags. Done the obvious way, it quietly glues neighbouring words together: 150.000 km and EZ 03/2019 become kmEZ. The mileage regex stops matching. No error, just empty fields. A second silent failure inside the fix for the first one.

The correction is one character. Finding it is the hard part, because nothing tells you it's broken.

WHAT I CHECK ON EVERY RUN NOW

Status codes tell you whether the server answered. They say nothing about whether you got the data. So every run gets judged on its output:

  1. Zero rows on a 200 is a failure, never an empty day.
  2. Volume against the recent average: a sudden drop means something changed.
  3. Field completeness: a parser can find every row and still lose a field.

Any of these trips an alert within minutes, instead of someone noticing weeks later.

TAKEAWAYS

  • A 200 is not a success. Judge a run by what it extracted.
  • Anchor on data, not on design. CSS changes far more often than the data behind it.
  • The question isn't whether your scraper will break. It's how fast you'll know.

I build and maintain scrapers for protected and fast-changing sites, with this kind of monitoring built in. If your scraper is quietly returning less than it should, or getting blocked, reach out on Upwork and I'll take a look.

Top comments (0)