DEV Community

Tiioluwani for Zenrows

Posted on Originally published at zenrows.com

How to run large-scale web scraping jobs with Zenrows Batch

This article was originally published on the Zenrows blog. Read the original here: https://www.zenrows.com/blog/large-scale-web-scraping-zenrows-batch

This guide shows you how to run a large scraping job as a single managed call with Zenrows Batch, from one Cloudflare-protected page to a 100,000-URL job, without building the queue, retry logic, and delivery pipeline yourself. You need Python 3.9+ and a Zenrows API key.

Every example ran against the live Batch API. Full code is in the GitHub repo.

Before you start

  • Python 3.9 or later
  • Zenrows account for your API key
  • The client.py helper from the repo, which wraps auth, submission, polling, results, reruns, webhooks, HMAC, CSV uploads, and scheduling

Why concurrency alone doesn't scale

A scraper that works on 500 URLs breaks at 50,000. The target rate-limits you, 429 fills your logs, open connections eat memory, and failed requests stay failed because nothing retries them.

asyncio, aiohttp, and Scrapy solve concurrency. They don't solve the infrastructure underneath: the queue, retry logic, failure tracking, and result delivery. Every large-scale system ends up needing the same six pieces, and Batch replaces all six with one API call.

Submitting a Cloudflare-protected page

The simplest real job submits one protected page with mode=auto.

# refer to 01_cloudflare_target.py in repo
import client

payload = {
    "type": "regular",
    "status": "closed",
    "zenrows_params": {"mode": "auto"},  # handles anti-bot per target
    "tasks": [
        {"url": "https://www.scrapingcourse.com/cloudflare-challenge", "external_id": "cloudflare-challenge"}
    ],
}

submitted = client.submit_job(payload)
job_id = submitted["job_id"]
print(job_id, submitted["latest_run"]["status"])
Enter fullscreen mode Exit fullscreen mode

With mode=auto, the job used dynamic website support and advanced network access, consuming 25 credits. The content came back with the page's confirmation text, "You bypassed the Cloudflare challenge! :D"

Attaching a webhook

A webhook pushes results to your server the moment a run completes, instead of you polling. Signed webhooks use an HMAC key. Generate one first:

# refer to 05_webhook.py in repo
key = client.rotate_hmac_key()
# store key["secret"] now, it's returned exactly once
Enter fullscreen mode Exit fullscreen mode

Attach it to the submission.

payload = {
    "type": "regular",
    "status": "closed",
    "zenrows_params": {"mode": "auto"},
    "tasks": [
        {"url": "https://www.scrapingcourse.com/ecommerce/product/abominable-hoodie/", "external_id": "good-1"},
    ],
    "webhook": {"url": "https://your-app.com/webhooks/zenrows", "signature": True},
}
submitted = client.submit_job(payload)
Enter fullscreen mode Exit fullscreen mode

Here's the completed run delivery.

{
  "schema": "zenrows.webhook/v1",
  "event_type": "run.completed",
  "job_id": "JOB_ID",
  "stats": {"total": 3, "completed": 3, "successful": 3, "failed": 0, "spend": {"credits": 3, "cost": 0.00038997}}
}
Enter fullscreen mode Exit fullscreen mode

Verifying the webhook signature

Verify the delivery came from Zenrows with the X-Signature header and your HMAC secret.

import hmac
import hashlib
import base64

def verify(raw_body: bytes, t: str, v1: str, secret_b64: str) -> bool:
    secret = base64.b64decode(secret_b64)
    expected = hmac.new(secret, f"{t}.".encode() + raw_body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, v1)  # constant-time compare
Enter fullscreen mode Exit fullscreen mode

Handling failures and reruns

We submitted five real product pages plus three URLs built to fail, a nonexistent domain, a dead page, and an unroutable IP. Query the results and filter for failures.

# refer to 03_failure_and_rerun.py in repo
results = client.list_results(job_id, status="all")
failed_rows = [r for r in results["results"] if r["status"] == "failed"]
for row in failed_rows:
    print(row["external_id"], row["url"], row["error"])
Enter fullscreen mode Exit fullscreen mode

Each failure carries a real error code.

error.code=RESP002  detail="The requested URL page returned a 404 HTTP Status Code..."
error.code=RESP007  detail="The requested target domain could not be resolved..."
error.code=REQS004  title="No ip-based target URLs are allowed"
Enter fullscreen mode Exit fullscreen mode

Rerun only the failures, and it doesn't charge for what already succeeded.

rerun = client.rerun_job(job_id, status="failed")
print(rerun["retried_tasks"], rerun["inherited_tasks"])
# 3 retried, 5 inherited
Enter fullscreen mode Exit fullscreen mode

A URL that's malformed at the syntax level rejects the entire submission with a 400. Fix it and resubmit the whole batch. Only well-formed URLs that fail at scrape time become individual failed tasks you can rerun.

Submitting 100,000 URLs with a credit ceiling

We submitted 100,000 URLs and watched credit use as it ran. The monitor stopped the job at 300,787 credits after 14,105 completed. Here's the watch loop:

# refer to 10_scale_run.py in repo
import time
import client

job_id = "YOUR_JOB_ID"
credit_ceiling = 300_000

while True:
    job = client.get_job(job_id)
    stats = job["latest_run"]["stats"]
    spend = stats.get("spend", {}).get("credits", 0)

    if spend >= credit_ceiling:      # stop before overspending
        client.stop_job(job_id)
        print(f"stopped at {spend} credits, {stats['completed']}/{stats['total']} completed")
        break

    if job["latest_run"]["status"] in {"completed", "stopped"}:
        break

    time.sleep(5)
Enter fullscreen mode Exit fullscreen mode

CSV upload for large lists

For a list too large to paste inline, upload a CSV and reference it.

# refer to 07_csv_upload.py in repo
slot = client.create_csv_input_slot(fields={"url": "URL", "external_id": "Customer Ref"}, header=True)
client.upload_csv(slot, csv_bytes)

payload = {
    "type": "regular",
    "status": "closed",
    "file_input_id": slot["file_input_id"],  # reference the uploaded file
    "zenrows_params": {"mode": "auto"},
}
submitted = client.submit_job(payload)
Enter fullscreen mode Exit fullscreen mode

CSV uploads cap at 50MB or 100,000 rows.

Scheduling recurring jobs

A scheduled job stores your URLs as a template and creates a new run each firing.

# refer to 08_scheduling.py in repo
payload = {
    "type": "scheduled",
    "status": "closed",
    "schedule": {"rate": {"every": 24, "unit": "hour"}},
    "tasks": [{"url": u, "external_id": f"sched-{i}"} for i, u in enumerate(urls, start=1)],
}
submitted = client.submit_job(payload)
Enter fullscreen mode Exit fullscreen mode

Structured extraction

Batch returns structured JSON with css_extractor passed as a JSON-encoded string, or automatic output with Zenrows Extract and no selector map.

# refer to 04_extraction.py in repo
import json

selectors = {"title": "h1", "price": ".price", "description": ".woocommerce-product-details__short-description"}

payload = {
    "type": "regular",
    "status": "closed",
    "zenrows_params": {
        "mode": "auto",
        "css_extractor": json.dumps(selectors),  # must be JSON-encoded
    },
    "tasks": [
        {"url": "https://www.scrapingcourse.com/ecommerce/product/abominable-hoodie/", "external_id": "product-1"},
    ],
}
submitted = client.submit_job(payload)
Enter fullscreen mode Exit fullscreen mode

What Batch leaves to you

Batch holds no browser session open between requests, so a target that needs login, clicking through steps, or a form submission before the data appears calls for Browser Sessions.

What's next

Top comments (0)