DEV Community

Fatih İlhan
Fatih İlhan

Posted on

My congressional trading scraper was quietly losing half its data

I run two small Apify actors that pull U.S. congressional stock trading disclosures (the STOCK Act "PTR" filings) for the House and the Senate and turn them into clean JSON. A handful of people pay for them. Revenue had plateaued, so I did what I should have done months ago: I stopped looking at the marketing and started auditing my own output.

It was worse than I thought.

The number that ruined my evening

I took the 50 most recent House filings and checked how many rows each one produced.

Almost half of them produced zero rows. Not errors. Not warnings anyone would see. Just nothing. The run finished green, the dataset looked healthy, and a big chunk of Congress's trades simply weren't there.

After fixing it, the same 50 filings went from 238 rows to 416.

If you're a paying user, that's the kind of thing you'd never notice. You'd just assume your data was complete. That's what bothered me most.

Why it happened

There wasn't one bug. There were several small ones that all failed the same way: quietly.

1. An amount format I never planned for. House filings report amounts as ranges like $1,001 - $15,000. My regex required a range. Then one filing showed up with an exact amount, $2,722.50, with cents. The regex didn't match, the parser returned an empty array, and the filing vanished.

2. Exchange transactions. Some rows are exchanges (swap one holding for another), not buys or sells. I didn't have a type for them, so they got dropped.

3. Scanned paper filings. About 14% of House filings in the last 90 days aren't digital at all. Someone printed the form, filled it in, and hand-delivered it. The PDF is just an image. My parser saw no text, returned nothing, moved on.

4. The Senate had the same problem, and I had a comment saying it didn't. There was literally a code comment in my Senate actor saying "Senate has no scanned case, source is HTML." Wrong. The Senate also accepts paper filings, they just live at a different URL. In the last 30 days: 3 out of 49.

5. Network failures. When a PDF download failed after retries, I incremented an error counter and... that's it. The filing was known, it was in the index, and it disappeared from the output.

The pattern is obvious in hindsight: every one of these paths ended in return []. And an empty array is the most polite way a program can lie to you.

The fix: every filing leaves a trace

The rule I landed on is simple. Every filing that appears in the official index must show up in the output. Either as real transaction rows, or as a clearly marked placeholder that tells you why there's no data.

So now every row has a parse_status:

parse_status What it means
ok Parsed normally
scanned_unparsed Paper filing, image only, no text layer
parse_failed Text exists, but my parser couldn't extract rows (my bug, not yours)
fetch_failed The source couldn't be downloaded after retries

Placeholders keep the politician, filing date, and a link to the original document, so you can go look yourself. And they're free. Nobody should pay for a row that says "sorry, couldn't read this."

The run output now also includes counters (paper filings, parse failures, fetch failures), so a silent drop can't hide in a green run anymore.

Fixing things broke other things

A few smaller lessons from the same week:

My ticker fix broke Mastercard. Some rows had the ticker only inside the asset name, like Electronic Arts Inc. (EA). I added a fallback to extract it, with a stoplist to avoid false positives like (LLC) or (NY). I put all U.S. state codes on that stoplist. Which meant MA (Mastercard), MS (Morgan Stanley), DE (Deere), and MO (Altria) all quietly became null. Removed the state list, added tests for exactly those tickers.

Not every "bug" was a bug. I was sure duplicate content_hash values meant I was losing data. They didn't. The hash fingerprints the transaction content on purpose, so identical trades in the same filing share it, and the unique id handles the rest. It was already documented. I was the one who hadn't read it.

I tried OCR. I stopped.

Since 9 out of 10 scanned forms I checked were printed (not handwritten), OCR seemed doable. It mostly was.

The forms turned out to be checkbox grids: the amount isn't written as text, it's an X in one of eleven lettered columns. Full-page OCR turned that into garbage. So I built a hybrid: OCR for text cells, pixel density for checkboxes.

It got close. On one 6-page filing, date errors went from 29 of 111 rows down to 4. The remaining 4 were Tesseract reading a printed 2 as 7, consistently, on digits a human reads instantly. 06/02/2026 became 06/02/7026.

I could have "fixed" that with rules like "the year must be 2026, so 7 is probably 2." But that same guess on a month or day turns 02 into 07, and now you have a wrong trade date that looks perfectly real. So the policy is all or nothing: if any row in a filing fails validation, the whole filing stays a scanned_unparsed placeholder.

The OCR code is sitting behind a disabled flag. Maybe a different engine later. For now, an honest placeholder beats a confident wrong answer.

What I'd tell anyone building a data product

  • Count what goes in, not just what comes out. Index size vs output size, per run.
  • return [] in a parser is a smell. Make failure produce something visible.
  • Treat "the run succeeded" and "the data is complete" as two different claims.
  • When in doubt, emit less data with a clear flag rather than more data you can't vouch for. The actors are here if you want to poke at them: House and Senate. If you find a filing that isn't accounted for, I genuinely want to know.

Top comments (0)