How to Scrape 10,000 Leads in 10 Minutes with Python (HTTPX + Playwright Stealth)
Most developers and agency founders believe web scraping at scale requires paying $99 to $250 every single month for cloud proxy SaaS.
That belief is expensive, fragile, and obsolete.
When you rely on cloud scraping platforms:
- You pay rent on basic HTTP requests.
- Credits quietly expire at the end of the billing cycle.
- The moment a target website updates its DOM or adds client-side React hydration, your metered credits fail and burn money.
In 2026, the economics of data extraction have shifted. By combining asynchronous Python (HTTPX) for blazing directory traversal with headless Playwright for dynamic anti-detection rendering, you can reliably extract 10,000 clean, structured B2B leads locally in under 10 minutes—with $0 monthly cloud bills.
The 3-Layer Architecture: Fast Async + Stealth Browser Fallback
A naive script either uses requests (fast, but instantly dies on JavaScript SPAs and Cloudflare) or runs full Chromium for every page (bulletproof, but eats 16 GB RAM and takes 2 hours for 500 pages).
The production solution is a two-tier adaptive engine:
[ Target URL List ]
│
▼
[ Tier 1: Async HTTPX Worker Pool ] ──(200 OK + Valid Data)──► [ Lead Normalizer ]
│
▼ (403 / Cloudflare Challenge / Dynamic SPA Detected)
[ Tier 2: Pooled Headless Playwright ] ──(Rendered DOM)──────► [ Lead Normalizer ]
│
▼
[ In-Memory Deduplication & Regex Triage ]
│
▼
[ Styled Excel (.xlsx) + UTF-8 CSV Export ]
1. Fast Mode: Async HTTPX with Connection Pooling
For static public directories, sitemaps, and government company registers, spinning up a browser is unnecessary overhead.
import httpx
import asyncio
async def fetch_directory(urls: list[str]):
limits = httpx.Limits(max_keepalive_connections=50, max_connections=100)
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36",
"Accept-Language": "en-US,en;q=0.9"
}
async with httpx.AsyncClient(limits=limits, headers=headers, timeout=10.0) as client:
tasks = [client.get(u) for u in urls]
responses = await asyncio.gather(*tasks, return_exceptions=True)
return [r.text for r in responses if isinstance(r, httpx.Response) and r.status_code == 200]
Throughput: 500 directory records extracted in 11.4 seconds on a standard 8-core CPU.
2. Stealth Fallback: Headless Playwright with Navigator Spoofing
When HTTPX hits a page returning <div id="root"></div> or a bot challenge, the request is automatically promoted to Tier 2:
from playwright.async_api import async_playwright
async def render_stealth(url: str):
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(
user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
viewport={"width": 1920, "height": 1080}
)
page = await context.new_page()
# Evade common webdriver checks
await page.add_init_script("delete Object.getPrototypeOf(navigator).webdriver")
await page.goto(url, wait_until="domcontentloaded")
content = await page.content()
await browser.close()
return content
3. Automated Lead Extraction & Deduplication
Raw HTML is useless noise. The pipeline extracts:
- Direct email addresses (filtered against common image/tracking junk)
- International phone numbers (E.164 normalization)
- WhatsApp direct message URLs
- Verified social links (LinkedIn, X, GitHub)
Duplicate root domains are stripped in-memory using set hashing before disk write.
4. Get the Complete Production CLI Tool
Instead of spending weeks reinventing browser pools, Cloudflare bypass heuristics, and Excel styling libraries:
We have packaged our complete production tool: OmniScraper AI.
Features included:
- ✅ 10 Ready-to-Run Site Presets (B2B directories, e-commerce reviews, seller catalogs)
- ✅ Dual-Engine Auto-Switching (HTTPX fast mode + Playwright stealth)
- ✅ Styled Excel (.xlsx) Exporter with frozen headers, column auto-fit, and active WhatsApp chat links
- ✅ 22 Rotated Desktop & Mobile User-Agents
- ✅ 100% Python Source Code & Permissive Commercial License (Zero Monthly Fees)
👉 Download OmniScraper AI on Gumroad:
https://ancuboy.gumroad.com/l/omniscraper-ai?utm_source=devto&utm_medium=article&utm_campaign=scrape_10k_leads_20260930
(Or explore our complete automation tools & business templates catalog at https://lynk.id/ancudigitalsolution?utm_source=devto&utm_medium=article&utm_campaign=scrape_10k_leads_20260930)
Stop renting what you can run locally.
Top comments (0)