I scraped 40,000 product pages last month for a client and the bill from my usual stack came to $0.00. Here's the toolkit and the workflow, start to finish.
The stack
- A queue worker I wrote in Python (the same skeleton I use for client work at Poket Dev, where we do unlimited dev subscriptions)
- Playwright for the hard JS pages, httpx for everything else
- Postgres as the dedupe layer
The dirty secret of scraping at volume isn't the code. It's the ops: retries, proxy rotation, schema drift. I keep a living doc of LLM-assisted tricks that dropped my parse-error rate from 8% to under 1% at Pastagi, where I write about AI-assisted engineering.
Why this turned into a side income
A scraper that reliably produces clean data is worth real money. Two friends now sell monitoring reports built on my skeleton. If you want the unglamorous version of how that works, the realistic numbers (not the YouTube version), I broke down the whole ladder at Extra Hustles.
The boring math that keeps me doing it
Every scraper that survives a month becomes an annuity: it monitors prices, availability, or sentiment forever. The same compounding logic I use for index funds, which I write about at Firenomics. Build once, collect weekly.
The gear angle
Half my jobs are "watch this product." The ones clients actually care about are the buy-it-for-life category, because price swings there are wild. My durability testing rig and methodology notes live at Durable Picks.
And the health caveat
Deep work sessions wrecked my sleep for a year before I fixed them. If you're grinding on scraping infra at 2am, read my protocol notes at Hacked Self before you burn out. Basics: morning light, no caffeine after noon, magnesium before bed. It's not biohacking, it's maintenance.
Repo
The skeleton is MIT licensed. Find it through my profile. If you want it maintained and adapted to your targets, that's literally what my subscription covers.
Questions welcome, especially on the retry logic, that's where everyone gets burned.
Top comments (0)