This article was originally published on the Zenrows blog. Read the original here: https://www.zenrows.com/blog/large-scale-web-scraping-zenrows-batch
This guide shows you how to run a large scraping job as a single managed call with Zenrows Batch, from one Cloudflare-protected page to a 100,000-URL job, without building the queue, retry logic, and delivery pipeline yourself. You need Python 3.9+ and a Zenrows API key.
Every example ran against the live Batch API. Full code is in the GitHub repo.
Before you start
- Python 3.9 or later
- Zenrows account for your API key
- The
client.pyhelper from the repo, which wraps auth, submission, polling, results, reruns, webhooks, HMAC, CSV uploads, and scheduling
Why concurrency alone doesn't scale
A scraper that works on 500 URLs breaks at 50,000. The target rate-limits you, 429 fills your logs, open connections eat memory, and failed requests stay failed because nothing retries them.
asyncio, aiohttp, and Scrapy solve concurrency. They don't solve the infrastructure underneath: the queue, retry logic, failure tracking, and result delivery. Every large-scale system ends up needing the same six pieces, and Batch replaces all six with one API call.
Submitting a Cloudflare-protected page
The simplest real job submits one protected page with mode=auto.
# refer to 01_cloudflare_target.py in repo
import client
payload = {
"type": "regular",
"status": "closed",
"zenrows_params": {"mode": "auto"}, # handles anti-bot per target
"tasks": [
{"url": "https://www.scrapingcourse.com/cloudflare-challenge", "external_id": "cloudflare-challenge"}
],
}
submitted = client.submit_job(payload)
job_id = submitted["job_id"]
print(job_id, submitted["latest_run"]["status"])
With mode=auto, the job used dynamic website support and advanced network access, consuming 25 credits. The content came back with the page's confirmation text, "You bypassed the Cloudflare challenge! :D"
Attaching a webhook
A webhook pushes results to your server the moment a run completes, instead of you polling. Signed webhooks use an HMAC key. Generate one first:
# refer to 05_webhook.py in repo
key = client.rotate_hmac_key()
# store key["secret"] now, it's returned exactly once
Attach it to the submission.
payload = {
"type": "regular",
"status": "closed",
"zenrows_params": {"mode": "auto"},
"tasks": [
{"url": "https://www.scrapingcourse.com/ecommerce/product/abominable-hoodie/", "external_id": "good-1"},
],
"webhook": {"url": "https://your-app.com/webhooks/zenrows", "signature": True},
}
submitted = client.submit_job(payload)
Here's the completed run delivery.
{
"schema": "zenrows.webhook/v1",
"event_type": "run.completed",
"job_id": "JOB_ID",
"stats": {"total": 3, "completed": 3, "successful": 3, "failed": 0, "spend": {"credits": 3, "cost": 0.00038997}}
}
Verifying the webhook signature
Verify the delivery came from Zenrows with the X-Signature header and your HMAC secret.
import hmac
import hashlib
import base64
def verify(raw_body: bytes, t: str, v1: str, secret_b64: str) -> bool:
secret = base64.b64decode(secret_b64)
expected = hmac.new(secret, f"{t}.".encode() + raw_body, hashlib.sha256).hexdigest()
return hmac.compare_digest(expected, v1) # constant-time compare
Handling failures and reruns
We submitted five real product pages plus three URLs built to fail, a nonexistent domain, a dead page, and an unroutable IP. Query the results and filter for failures.
# refer to 03_failure_and_rerun.py in repo
results = client.list_results(job_id, status="all")
failed_rows = [r for r in results["results"] if r["status"] == "failed"]
for row in failed_rows:
print(row["external_id"], row["url"], row["error"])
Each failure carries a real error code.
error.code=RESP002 detail="The requested URL page returned a 404 HTTP Status Code..."
error.code=RESP007 detail="The requested target domain could not be resolved..."
error.code=REQS004 title="No ip-based target URLs are allowed"
Rerun only the failures, and it doesn't charge for what already succeeded.
rerun = client.rerun_job(job_id, status="failed")
print(rerun["retried_tasks"], rerun["inherited_tasks"])
# 3 retried, 5 inherited
A URL that's malformed at the syntax level rejects the entire submission with a 400. Fix it and resubmit the whole batch. Only well-formed URLs that fail at scrape time become individual failed tasks you can rerun.
Submitting 100,000 URLs with a credit ceiling
We submitted 100,000 URLs and watched credit use as it ran. The monitor stopped the job at 300,787 credits after 14,105 completed. Here's the watch loop:
# refer to 10_scale_run.py in repo
import time
import client
job_id = "YOUR_JOB_ID"
credit_ceiling = 300_000
while True:
job = client.get_job(job_id)
stats = job["latest_run"]["stats"]
spend = stats.get("spend", {}).get("credits", 0)
if spend >= credit_ceiling: # stop before overspending
client.stop_job(job_id)
print(f"stopped at {spend} credits, {stats['completed']}/{stats['total']} completed")
break
if job["latest_run"]["status"] in {"completed", "stopped"}:
break
time.sleep(5)
CSV upload for large lists
For a list too large to paste inline, upload a CSV and reference it.
# refer to 07_csv_upload.py in repo
slot = client.create_csv_input_slot(fields={"url": "URL", "external_id": "Customer Ref"}, header=True)
client.upload_csv(slot, csv_bytes)
payload = {
"type": "regular",
"status": "closed",
"file_input_id": slot["file_input_id"], # reference the uploaded file
"zenrows_params": {"mode": "auto"},
}
submitted = client.submit_job(payload)
CSV uploads cap at 50MB or 100,000 rows.
Scheduling recurring jobs
A scheduled job stores your URLs as a template and creates a new run each firing.
# refer to 08_scheduling.py in repo
payload = {
"type": "scheduled",
"status": "closed",
"schedule": {"rate": {"every": 24, "unit": "hour"}},
"tasks": [{"url": u, "external_id": f"sched-{i}"} for i, u in enumerate(urls, start=1)],
}
submitted = client.submit_job(payload)
Structured extraction
Batch returns structured JSON with css_extractor passed as a JSON-encoded string, or automatic output with Zenrows Extract and no selector map.
# refer to 04_extraction.py in repo
import json
selectors = {"title": "h1", "price": ".price", "description": ".woocommerce-product-details__short-description"}
payload = {
"type": "regular",
"status": "closed",
"zenrows_params": {
"mode": "auto",
"css_extractor": json.dumps(selectors), # must be JSON-encoded
},
"tasks": [
{"url": "https://www.scrapingcourse.com/ecommerce/product/abominable-hoodie/", "external_id": "product-1"},
],
}
submitted = client.submit_job(payload)
What Batch leaves to you
Batch holds no browser session open between requests, so a target that needs login, clicking through steps, or a form submission before the data appears calls for Browser Sessions.
What's next
- Web data for LLM fine-tuning for running collected data into a training pipeline
-
Zenrows Fetch for the single-URL path with the same
zenrows_params - Full project: GitHub repo
Top comments (0)