DEV Community

Ryan Cole
Ryan Cole

Posted on

Web Scraping at Scale Without the $99–$250/Month API Bill

Most "scraping APIs" charge $99–$250/month the moment you need concurrency or a rotating proxy. For an indie project or a B2B lead-gen pipeline, that's often the single biggest avoidable line item in the stack.

Here's the setup I run instead: a modular Python scraper that does retries, rotation, and structured extraction locally, then writes clean JSON you can push straight into a CRM, a spreadsheet, or an n8n workflow.

Why hosted scraping APIs get expensive

The pricing is usually metered by successful request. That sounds fair until you realize a "successful request" for a dynamic page can mean 5–8 internal calls after rendering, retries, and anti-bot backoff. A pipeline that needs 20k rows/day can burn through a mid-tier plan before lunch.

The alternative isn't "scrape more carefully." It's to own the extraction layer.

The architecture that keeps costs at zero

Four pieces, each swappable:

  1. Fetcher — httpx for static pages, Playwright only when JS is required.
  2. Rotator — a pool of user-agents + backoff, so one 429 doesn't kill the run.
  3. Extractor — CSS/XPath selectors or a JSON schema, decoupled from the fetcher.
  4. Sink — write JSONL first, then load to wherever the data needs to live.
import httpx, json, asyncio, random, time

UA = [
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125 Safari/537.36",
    "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17 Safari/605.1.15",
]

async def fetch(client, url, tries=4):
    for i in range(tries):
        try:
            r = await client.get(url, headers={"User-Agent": random.choice(UA)}, timeout=20)
            if r.status_code == 200:
                return r.text
            if r.status_code in (429, 503):
                await asyncio.sleep(2 ** i + random.random())
        except httpx.HTTPError:
            await asyncio.sleep(2 ** i)
    return None

async def run(urls, out="rows.jsonl", conc=8):
    sem = asyncio.Semaphore(conc)
    async with httpx.AsyncClient(follow_redirects=True) as client:
        async def one(u):
            async with sem:
                html = await fetch(client, u)
                if not html:
                    return
                row = extract(html, u)          # your selector/schema logic
                with open(out, "a", encoding="utf-8") as f:
                    f.write(json.dumps(row) + "\n")
        await asyncio.gather(*(one(u) for u in urls))
Enter fullscreen mode Exit fullscreen mode

The key discipline: the extractor never touches the network, and the fetcher never knows the schema. That's what makes it reusable across sites instead of a one-off script per target.

Where people get blocked (and the fixes)

  • 429 / 503 storms → exponential backoff with jitter, plus a smaller concurrency ceiling. Start at 4, not 40.
  • JS-only content → escalate to a headless browser only for those routes; don't pay the render cost for pages that return static HTML.
  • Layout drift → assert on your extracted row shape (required keys, types) and fail loudly instead of writing garbage to the sink.
  • Duplicate rows → dedupe on a stable business key (SKU, URL, email), not on the raw HTML hash.

Turning it into leads, not just rows

If the goal is B2B leads, the pipeline doesn't stop at JSON. The useful shape is:

scrape → normalize → enrich (company, role) → qualify → deliver (CRM / sheet / Telegram)
Enter fullscreen mode Exit fullscreen mode

That last hop is where most "scraper" tutorials stop and where the actual value is. A list of 10k raw rows is a chore; 200 qualified, deduped, enriched leads with a WhatsApp/email column is revenue.

The honest math

A hosted API at $149/mo is $1,788/year. A local pipeline costs you an afternoon of setup and a few dollars of proxy credit if you even need it. For teams doing < 100k pages/month, the local route wins on cost and on control.

If you'd rather skip the build, I packaged the exact modular toolkit I use — retries, rotation, structured extraction, and the CRM-ready output — as Production Python Scraper Toolkit v2.0.

And if you want the whole AI-builder stack — scrapers, n8n lead-automation workflows, and the prompt/agent vaults — the full suite is bundled here: https://ancuboy.gumroad.com/l/ai-developer-suite?utm_source=devto&utm_medium=article&utm_campaign=scrape_cost


Built for indie hackers and small agencies who'd rather own their tooling than rent it. If this saved you a monthly bill, the toolkit is the fastest way to repay the favor.

Top comments (0)