This article was originally published on the Zenrows blog.
In this tutorial, you'll build a lead generation agent that reads a bot-protected business directory, pulls every company listed on the page, enriches each one from its own website, and scores it against your ideal customer profile (ICP). You'll need Python 3.10 or above, an OpenAI API key, and a Zenrows API key.
The OpenAI Agents SDK runs the agent loop, and its @function_tool decorator turns plain Python functions into tools. Zenrows Fetch handles the step that usually breaks, which is getting the page at all. In the test run, one directory page produced 78 companies, and the agent scored each one in under 20 seconds.
The complete code is on GitHub if you'd like to clone it and follow along.
Prerequisites
- You need Python 3.10 or above.
- You need an OpenAI API key, which the agent, the extraction tool, and the scoring tool all use. Get one from your OpenAI developer dashboard.
- You need a Zenrows API key to retrieve the full contents of web pages. Create one on the signup page.
Install the dependencies.
python3 -m pip install openai-agents openai requests pydantic python-dotenv
The SDK runs the agent, requests makes the Zenrows calls, pydantic defines the tool schemas, and python-dotenv loads your keys. Create a .env file in your project root with your OPENAI_API_KEY and ZENROWS_API_KEY, and add .env to .gitignore so the credentials never reach version control.
How the agent is structured
The work splits into four tools, and each one does a single job.
| Tool | Input | Output | Why it's separate |
|---|---|---|---|
fetch_page |
a URL | the page as Markdown | it's the only job that touches the network, so retries and errors live in one place |
extract_leads |
Markdown plus the source URL |
Lead records |
it runs in chunks because a full directory page exceeds one reliable extraction call |
enrich_lead |
a company and its domain | enrichment signals from the company's site | it makes one fetch per nav link, so it drives the cost and is worth isolating |
score_lead |
enriched lead plus the ICP | a score from 0 to 100 with reasoning | it's pure reasoning with no network calls, so you can rerun it without re-fetching |
Splitting the work keeps retries local. When enrichment fails on one lead, the agent retries that tool and keeps the rest of the run intact. Each job's output also stays small enough to hand to the next one.
The agent only sees two composed tools. discover_leads runs fetch_page and extract_leads, and qualify_lead runs enrich_lead and score_lead. If the agent passed an enriched lead between two separate tools, it would have to reproduce every page excerpt as a tool argument, and that overflows the context window on a large directory.
1. Build the Zenrows Fetch tool
Every code block from here goes into one file called tools.py. Append each block to the same file.
import os
import re
import requests
from urllib.parse import urlparse, parse_qs, unquote
from agents import function_tool
from dotenv import load_dotenv
from openai import OpenAI
from pydantic import BaseModel
load_dotenv()
ZENROWS_API_KEY = os.getenv("ZENROWS_API_KEY")
ZENROWS_ENDPOINT = "https://api.zenrows.com/v1/"
if not ZENROWS_API_KEY:
raise RuntimeError("set ZENROWS_API_KEY in your .env file")
client = OpenAI()
# ===========================================================================
# job 1: fetch a page, protected or not
# ===========================================================================
def _fetch_page(url: str) -> str:
"""fetch logic, importable and testable without the agent."""
params = {
"url": url,
"apikey": ZENROWS_API_KEY,
"mode": "auto", # start cheap, escalate only when the target needs it
# markdown strips nav and boilerplate before the model reads it
"response_type": "markdown"
}
try:
response = requests.get(ZENROWS_ENDPOINT, params=params, timeout=90)
response.raise_for_status()
# return failures as text instead of raising, so one bad url does not end the run.
# each message tells the model whether the failure is worth retrying
except requests.exceptions.Timeout:
return "FETCH_ERROR: timed out after 90s. retry once, then move on."
except requests.exceptions.HTTPError as exc:
status = exc.response.status_code
if status in (401, 403):
return f"FETCH_ERROR: {status}. api key rejected, do not retry."
if status == 429:
return "FETCH_ERROR: 429 concurrency limit. wait, then retry."
if status in (400, 404):
return f"FETCH_ERROR: {status}. this url is not retrievable, do not retry."
return f"FETCH_ERROR: {status}. retry once, then report the failure."
except requests.exceptions.RequestException as exc:
return f"FETCH_ERROR: {exc}. do not retry."
content = response.text.strip()
if not content:
return "FETCH_ERROR: empty response body. do not retry."
return content
# the decorator replaces the function with a FunctionTool object, which is not
# callable. keeping the logic in _fetch_page above leaves it testable
@function_tool
def fetch_page(url: str) -> str:
"""Fetch any web page as markdown, including pages behind anti-bot protection.
Args:
url: full url of the page to fetch
"""
return _fetch_page(url)
# ===========================================================================
# job 2: turn a directory page into leads
# ===========================================================================
# matches the tracking links directories wrap around outbound urls
REDIRECT_LINK = re.compile(r"https?://[a-z0-9.-]*/redirect\?[^\s\)\"]+")
# a full directory page is too long for one reliable extraction call
CHUNK_SIZE = 40000
# slices cut mid-listing, and the extraction prompt is told to skip partial
# entries, so a listing straddling a boundary is dropped by both slices.
# overlapping the window carries each boundary listing whole into one of them;
# _key dedupes the entries the overlap sees twice
CHUNK_OVERLAP = 2000
_fetch_page holds the logic, and fetch_page wraps it with @function_tool so the agent can call it. The decorator replaces the function with a FunctionTool object that you can't call directly, so keeping the logic in a plain function leaves it testable.
Failures come back as text strings, so one bad URL doesn't end the run. Each message also tells the model whether a retry is worth it.
mode=auto lets Zenrows pick the access configuration each target needs. Simple pages resolve on the lightest path, and only protected targets escalate. That matters because protected sites rarely block you with an error. A Cloudflare-protected page can return a 200 status code with a challenge page in the body, and your agent will treat that challenge as a valid page with zero leads on it.
response_type=markdown strips navigation and boilerplate, which gives the model cleaner input for extraction. If your directory is an API-like endpoint, use response_type=json and skip the conversion. The three constants at the bottom of the block belong to the extraction step in the next section.
2. Extract leads from the directory page
Append the extraction code to tools.py.
# strict tool schemas reject bare dicts, so every tool input and output is a model
class Lead(BaseModel):
company: str
name: str
website: str
source_url: str
class LeadList(BaseModel):
leads: list[Lead]
def _unwrap_redirects(markdown: str) -> str:
"""rewrite directory tracking links to the destination domain."""
def replace(match):
url = match.group(0)
# the real destination sits url-encoded in the u parameter
target = parse_qs(urlparse(url).query).get("u", [""])[0]
if not target:
return url
parsed = urlparse(unquote(target))
return f"{parsed.scheme}://{parsed.netloc}"
return REDIRECT_LINK.sub(replace, markdown)
def _reject_source_domain(leads: list[dict], source_url: str) -> list[dict]:
"""blank any website that points back at the directory itself."""
directory = urlparse(source_url).netloc.replace("www.", "")
for lead in leads:
host = urlparse(lead["website"]).netloc.replace("www.", "")
if host == directory:
lead["website"] = ""
return leads
def _key(lead: dict) -> str:
"""stable identity for dedupe across chunks."""
# the model returns the same domain with and without www, so normalise before comparing
site = lead["website"].lower()
for prefix in ("https://", "http://", "www."):
site = site.replace(prefix, "")
site = site.rstrip("/")
# fall back to company name so leads without a domain are not collapsed into one
return site or lead["company"].strip().lower()
def _extract_chunk(chunk: str, source_url: str) -> list[dict]:
"""run one extraction call over a slice of the page."""
response = client.responses.parse(
model="gpt-4o-mini",
max_output_tokens=8000,
input=[
{
"role": "system",
"content": (
"extract every company listed in this page fragment. "
"website is the company's own domain, taken from that listing's "
"visit website link. "
"never use a link from the directory's own domain as the website. "
"name is a person's name and is usually absent from a directory "
"listing, so leave it empty unless a person is actually named. "
"use an empty string for any field the fragment does not state, "
"and never invent a value. "
"the fragment may start or end mid listing, so skip any partial entry."
)
},
{
"role": "user",
"content": f"source_url: {source_url}\n\n{chunk}"
}
],
text_format=LeadList
)
return [lead.model_dump() for lead in response.output_parsed.leads]
def _extract_leads(content: str, source_url: str) -> list[dict]:
"""parse markdown into lead dicts, chunking long pages."""
# a fetch failure is a string, not a page. parsing it would return an empty
# list that looks identical to a genuinely empty directory
if content.startswith("FETCH_ERROR"):
return []
# resolve tracking links before the model reads them
content = _unwrap_redirects(content)
leads, seen = [], set()
for start in range(0, len(content), CHUNK_SIZE - CHUNK_OVERLAP):
chunk = content[start:start + CHUNK_SIZE]
for lead in _extract_chunk(chunk, source_url):
key = _key(lead)
if key in seen:
continue
seen.add(key)
leads.append(lead)
return _reject_source_domain(leads, source_url)
@function_tool
def extract_leads(content: str, source_url: str) -> list[Lead]:
"""Parse a fetched page into a list of leads.
Args:
content: markdown returned by fetch_page
source_url: url the content came from, recorded on every lead
"""
return [Lead(**lead) for lead in _extract_leads(content, source_url)]
# fetch and extract are one job from the agent's point of view. exposing them
# separately means the whole page passes through the model's context to get
# from one tool to the next, which overflows the window on a large directory
@function_tool
def discover_leads(source_url: str) -> list[Lead]:
"""Fetch a directory page and return the companies listed on it.
Args:
source_url: url of the directory or listing page
"""
markdown = _fetch_page(source_url)
if markdown.startswith("FETCH_ERROR"):
return []
return [Lead(**lead) for lead in _extract_leads(markdown, source_url)]
# ===========================================================================
# job 3: enrich each lead from its own website
# ===========================================================================
Strict tool schemas reject bare dicts, so every tool input and output is a Pydantic model. _unwrap_redirects rewrites the directory's tracking links to each company's own domain before the model reads them, and _reject_source_domain blanks any website that points back at the directory.
A full directory page is too long for one reliable extraction call, so _extract_leads splits it into 40,000-character chunks with a 2,000-character overlap. The overlap carries any listing that straddles a boundary into one chunk whole, and _key removes the duplicates the overlap creates.
discover_leads combines fetching and extraction in one tool because the raw output of fetch_page can exceed the context length by itself.
Test both jobs without the agent. Save this script as test_fetch_extract.py.
# test_fetch_extract.py
import json
import os
from tools import _fetch_page, _unwrap_redirects, _extract_leads
SOURCE_URL = "https://clutch.co/it-services"
FIXTURE = "fixtures/directory_page.md"
LEADS_CACHE = "fixtures/leads.json"
# ---- job 1: fetch ---------------------------------------------------------
# cache the page so extraction can be tuned without paying for a fetch each time
if os.path.exists(FIXTURE):
with open(FIXTURE) as f:
markdown = f.read()
print(f"using cached page, {len(markdown)} chars")
else:
markdown = _fetch_page(SOURCE_URL)
if markdown.startswith("FETCH_ERROR"):
raise SystemExit(markdown)
os.makedirs("fixtures", exist_ok=True)
with open(FIXTURE, "w") as f:
f.write(markdown)
print(f"fetched and cached, {len(markdown)} chars")
unwrapped = _unwrap_redirects(markdown)
print(f"{len(unwrapped)} chars after unwrapping redirects, "
f"{len(markdown) - len(unwrapped)} saved\n")
# ---- job 2: extract -------------------------------------------------------
leads = _extract_leads(markdown, SOURCE_URL)
with open(LEADS_CACHE, "w") as f:
json.dump(leads, f, indent=2)
print(f"{len(leads)} leads extracted\n")
for lead in leads[:5]:
print(f"{lead['company']:<32} {lead['website'] or '(none)'}")
missing_site = sum(1 for lead in leads if not lead["website"])
print(f"\n{missing_site} of {len(leads)} leads have no website")
Run it from the project root.
python3 test_fetch_extract.py
Here's the output from a local run against an IT services directory.
using cached page, 543637 chars
457647 chars after unwrapping redirects, 85990 saved
78 leads extracted
Infracore https://infracore.net
Miles IT https://www.milesit.com
Geniusee https://geniusee.com
Peeklogic, LLC https://www.peeklogic.com
TechQuarter LLC https://www.techquarter.io
0 of 78 leads have no website
The page came back at 543,637 characters, and unwrapping the redirects saved 85,990 of them. Counts vary between runs because directory pages change and the model can read the same page differently.
3. Enrich and score each lead
Append the enrichment code to tools.py.
# matches any markdown link, used to read a company's own navigation
MARKDOWN_LINK = re.compile(r"\[([^\]]+)\]\((https?://[^\s\)]+)\)")
# guessing paths costs a fetch per miss, so read the site's nav instead and
# follow only the links it actually has
SIGNAL_KEYWORDS = {
"hiring": ["career", "job", "join", "hiring", "work with us", "we are hiring"],
"services": ["service", "what we do", "solutions", "expertise", "capabilities"],
"portfolio": ["portfolio", "case stud", "our work", "projects", "clients"],
"about": ["about", "who we are", "our story", "team"],
"contact": ["contact", "get in touch", "book a call", "let's talk", "talk to us"],
}
# dict[str, X] is rejected by strict schemas too, which is why signals is a list
class Signal(BaseModel):
name: str
found: bool
url: str
excerpt: str
class EnrichedLead(BaseModel):
company: str
website: str
signals: list[Signal]
def _discover_links(markdown: str, website: str) -> dict:
"""find real urls for each signal by reading the homepage nav."""
host = urlparse(website).netloc.replace("www.", "")
homepage_path = urlparse(website).path.rstrip("/") or "/"
found = {}
for text, url in MARKDOWN_LINK.findall(markdown):
parsed = urlparse(url)
# only follow links on the company's own domain
if parsed.netloc.replace("www.", "") != host:
continue
# an anchor on the homepage is content already in the homepage excerpt
if parsed.fragment and (parsed.path.rstrip("/") or "/") == homepage_path:
continue
# drop the fragment, it never changes what the server returns
clean_url = f"{parsed.scheme}://{parsed.netloc}{parsed.path}"
haystack = f"{text} {url}".lower()
for signal, keywords in SIGNAL_KEYWORDS.items():
if signal not in found and any(k in haystack for k in keywords):
found[signal] = clean_url
return found
def _enrich_lead(company: str, website: str) -> EnrichedLead:
"""collect signals from the company's own site."""
# a lead with no domain still flows through to scoring, judged on what little is known
if not website:
return EnrichedLead(company=company, website="", signals=[])
homepage = _fetch_page(website)
if homepage.startswith("FETCH_ERROR"):
return EnrichedLead(company=company, website=website, signals=[])
# the homepage is the most informative single page, and on a one-page site
# it holds everything the nav links point at
signals = [Signal(name="homepage", found=True, url=website, excerpt=homepage[:6000])]
for name, url in _discover_links(homepage, website).items():
content = _fetch_page(url)
ok = not content.startswith("FETCH_ERROR")
signals.append(Signal(
name=name,
found=ok,
url=url,
# keep excerpts small, scoring reads all of them in one call
excerpt=content[:2000] if ok else ""
))
return EnrichedLead(company=company, website=website, signals=signals)
@function_tool
def enrich_lead(company: str, website: str) -> EnrichedLead:
"""Collect signals from a company's own website.
Args:
company: company name from the directory listing
website: company's own domain
"""
return _enrich_lead(company, website)
# ===========================================================================
# job 4: score each lead against the icp
# ===========================================================================
An extracted lead carries only a company name and a domain, which is too little to judge fit. enrich_lead reads the company's homepage, finds the navigation links the site exposes, and fetches only those pages. Guessing paths like /careers would cost a fetch for every miss. The signals that matter live on the company's own site in its hiring pages, services pages, case studies, and contact pages.
Markdown works well when the model does the reading. If you'd prefer structured fields, Zenrows Extract returns JSON without a selector map.
Next, append the scoring code.
class Score(BaseModel):
score: int
reasoning: str
class ScoredLead(BaseModel):
company: str
website: str
score: int
reasoning: str
def _score_lead(lead: EnrichedLead, icp: str) -> ScoredLead:
"""score an enriched lead against the icp description."""
# no fetching happens here. every page was already retrieved in job 3
summary = "\n\n".join(
f"{s.name} page: {'found at ' + s.url if s.found else 'not found'}\n{s.excerpt}"
for s in lead.signals
)
response = client.responses.parse(
model="gpt-4o-mini",
max_output_tokens=500,
input=[
{
"role": "system",
"content": (
"score how well this company matches the ideal customer profile. "
"return an integer from 0 to 100 and one sentence of reasoning. "
"base the score only on evidence in the signals provided. "
"a missing page is weak evidence, not disqualifying. "
"if the signals are too thin to judge, score below 30 and say so."
)
},
{
"role": "user",
"content": (
f"ideal customer profile:\n{icp}\n\n"
f"company: {lead.company}\n"
f"website: {lead.website}\n\n"
f"signals:\n{summary}"
)
}
],
text_format=Score
)
return ScoredLead(
company=lead.company,
website=lead.website,
score=response.output_parsed.score,
reasoning=response.output_parsed.reasoning
)
@function_tool
def score_lead(lead: EnrichedLead, icp: str) -> ScoredLead:
"""Score an enriched lead against an ICP description.
Args:
lead: enriched lead returned by enrich_lead
icp: plain-language description of the ideal customer
"""
return _score_lead(lead, icp)
# an EnrichedLead carries page excerpts the model has no reason to read. passing
# it between two tools means the agent has to reproduce all of it as an argument,
# and it summarises instead, so scoring reads a summary rather than the pages
@function_tool
def qualify_lead(company: str, website: str, icp: str) -> ScoredLead:
"""Enrich a lead from its own website and score it against an ICP.
Args:
company: company name from the directory listing
website: company's own domain
icp: plain-language description of the ideal customer
"""
enriched = _enrich_lead(company, website)
return _score_lead(enriched, icp)
score_lead reads every signal alongside the ICP and returns a score from 0 to 100 with one sentence of reasoning. The prompt ties the score to evidence, so a lead with only a domain scores lower than one with a careers page and named case studies. No fetching happens in this step, which means you can rerun scoring without paying for the pages again.
qualify_lead pairs enrichment and scoring for the same reason discover_leads pairs fetching and extraction. The bulky intermediate data stays inside one Python call.
Test both halves before you wire the agent, so any later failure belongs to the agent loop. Save this script as test_enrich_score.py.
# test_enrich_score.py
import json
import os
import time
from tools import _fetch_page, _unwrap_redirects, _extract_leads, _enrich_lead, _score_lead
SOURCE_URL = "https://clutch.co/it-services"
FIXTURE = "fixtures/directory_page.md"
LEADS_CACHE = "fixtures/leads.json"
ICP = (
"IT services agencies that build custom software for B2B clients, "
"publish detailed case studies with named clients, and offer ai and cloud services"
)
# ---- job 1: fetch ---------------------------------------------------------
if os.path.exists(FIXTURE):
with open(FIXTURE) as f:
markdown = f.read()
print(f"using cached page, {len(markdown)} chars")
else:
markdown = _fetch_page(SOURCE_URL)
if markdown.startswith("FETCH_ERROR"):
raise SystemExit(markdown)
os.makedirs("fixtures", exist_ok=True)
with open(FIXTURE, "w") as f:
f.write(markdown)
print(f"fetched and cached, {len(markdown)} chars")
unwrapped = _unwrap_redirects(markdown)
print(f"{len(unwrapped)} chars after unwrapping redirects, "
f"{len(markdown) - len(unwrapped)} saved\n")
# ---- job 2: extract -------------------------------------------------------
if os.path.exists(LEADS_CACHE):
with open(LEADS_CACHE) as f:
leads = json.load(f)
print(f"using {len(leads)} cached leads\n")
else:
leads = _extract_leads(markdown, SOURCE_URL)
with open(LEADS_CACHE, "w") as f:
json.dump(leads, f, indent=2)
print(f"{len(leads)} leads extracted\n")
for lead in leads[:5]:
print(f"{lead['company']:<32} {lead['website'] or '(none)'}")
missing_site = sum(1 for lead in leads if not lead["website"])
print(f"\n{missing_site} of {len(leads)} leads have no website\n")
print("-" * 60 + "\n")
# ---- jobs 3 and 4: enrich and score ---------------------------------------
# enrichment is one fetch per nav link found, so start with a few leads
sample = [lead for lead in leads if lead["website"]][:3]
for lead in sample:
start = time.time()
enriched = _enrich_lead(lead["company"], lead["website"])
elapsed = time.time() - start
found = [s.name for s in enriched.signals if s.found]
print(f"{enriched.company} ({elapsed:.1f}s)")
print(f" found: {', '.join(found) if found else 'nothing'}")
scored = _score_lead(enriched, ICP)
print(f" score: {scored.score}")
print(f" {scored.reasoning}\n")
Run it.
python3 test_enrich_score.py
Here are the first two leads from the output.
Infracore (12.8s)
found: homepage, contact, services, portfolio, about, hiring
score: 80
Infracore fits well within the ideal customer profile as they offer IT services, including cloud solutions and have case studies showcasing customized work for various industries, but they may lack detailed named client presentations.
Miles IT (12.1s)
found: homepage, contact, about, services, portfolio, hiring
score: 75
Miles IT provides software development and AI services, aligning well with the ideal customer profile, but lacks detailed case studies with named clients which slightly diminishes the score.
Each lead took around 12 seconds to enrich. The page list reflects what each site exposes in its navigation, so a lead with no careers page has no hiring signal.
4. Wire the agent and run it
Save the agent as agent.py.
import asyncio
from agents import Agent, Runner, set_tracing_disabled
from tools import discover_leads, qualify_lead
set_tracing_disabled(True)
SOURCE_URL = "https://clutch.co/it-services"
ICP = (
"IT services agencies that build custom software for B2B clients, "
"publish detailed case studies with named clients, and offer ai and cloud services"
)
INSTRUCTIONS = """You find and qualify sales leads from web directories.
Follow this sequence exactly:
1. call discover_leads on the source url the user gives you
2. for each lead that has a website, call qualify_lead with its company, website,
and the icp description from the user message
3. return the scored leads as a json array sorted by score descending
rules:
- process only the first 10 leads that have a website, then stop and return them
- skip leads with no website, they cannot be scored fairly
- if a tool returns a string starting with FETCH_ERROR, follow the instruction in
that message. do not retry a call the message tells you not to retry
- never invent a lead, a website, or a score. every value comes from a tool result
"""
agent = Agent(
name="lead generation agent",
instructions=INSTRUCTIONS,
model="gpt-4o-mini", # the reasoning here is scoring, not planning
tools=[discover_leads, qualify_lead],
)
async def main():
result = await Runner.run(
agent,
input=f"source url: {SOURCE_URL}\n\nideal customer profile:\n{ICP}",
max_turns=50, # one qualify call per lead, plus discovery and the reply
)
# the trace shows which tools ran and in what order
for item in result.new_items:
if item.type == "tool_call_item":
print(item.raw_item.name)
print()
print(result.final_output)
if __name__ == "__main__":
asyncio.run(main())
The agent gets the two composed tools and a fixed sequence in its instructions. A small model is enough here because the sequence is already decided and the heavy reasoning happens inside score_lead. The instructions also cap each run at 10 leads, skip leads with no website, and forbid invented values.
Run the agent.
python3 agent.py
The agent prints the ranked leads as JSON. Here are the first two.
[
{
"company": "Geniusee",
"website": "https://geniusee.com",
"score": 85,
"reasoning": "Geniusee offers a range of IT services including custom software development, AI solutions, and cloud services, which align closely with the ideal customer profile; however, the absence of detailed case studies with named clients slightly limits the fit."
},
{
"company": "Sigli",
"website": "https://www.sigli.com",
"score": 85,
"reasoning": "Sigli aligns closely with the ideal customer profile by offering custom software development for B2B clients, specializing in AI and cloud services, and publishing case studies, although the specificity of named clients could not be verified."
}
]
Every value comes from a tool result, so the tool-call trace is the first thing to check when a run misbehaves.
5. Scale across pages with Batch
Save this script as batch.py.
# batch.py
# Scale discovery across several directory pages with Zenrows Batch.
#
# One Fetch call retrieves one page. Batch takes the whole list as a single
# managed job and handles the queue, concurrency and retries, so a directory
# with many pages costs one submission instead of one call per page.
import time
import requests
from agents import function_tool
from tools import ZENROWS_API_KEY, Lead, _extract_leads
BATCH_ENDPOINT = "https://async.api.zenrows.com/v1/jobs"
# batch authenticates by header, unlike fetch which takes apikey as a query param
BATCH_HEADERS = {"X-API-Key": ZENROWS_API_KEY, "Content-Type": "application/json"}
# a job accepts up to 100,000 urls
MAX_URLS_PER_JOB = 100_000
def _submit_batch(urls: list[str]) -> str:
"""submit a url list as one job, return the job id."""
if len(urls) > MAX_URLS_PER_JOB:
raise ValueError(f"a job accepts at most {MAX_URLS_PER_JOB} urls")
payload = {
"type": "regular",
"status": "closed", # run once, accept no further tasks
"zenrows_params": {"mode": "auto", "response_type": "markdown"},
"tasks": [{"url": url} for url in urls],
}
response = requests.post(
BATCH_ENDPOINT, headers=BATCH_HEADERS, json=payload, timeout=30
)
response.raise_for_status()
return response.json()["job_id"]
def _collect_batch(job_id: str, max_attempts: int = 60) -> list[tuple[str, str]]:
"""wait for the job to reach a terminal state, then pull each task's markdown."""
# terminal run states are completed, stopped and deleted. "failed" is a task
# status, not a run status, so polling for it never returns
for _ in range(max_attempts):
time.sleep(5)
status = requests.get(
f"{BATCH_ENDPOINT}/{job_id}", headers=BATCH_HEADERS, timeout=30
)
status.raise_for_status()
run = status.json()["latest_run"]
if run["status"] in ("completed", "stopped", "deleted"):
break
else:
raise TimeoutError(f"job {job_id} did not reach a terminal state in time")
response = requests.get(
f"{BATCH_ENDPOINT}/{job_id}/results", headers=BATCH_HEADERS, timeout=30
)
response.raise_for_status()
pages = []
for task in response.json()["results"]:
if task["status"] != "successful":
continue
# result_url is presigned and valid for 2 hours; re-list the results for a
# fresh link rather than storing this one
markdown = requests.get(task["result_url"], timeout=60).text
pages.append((task["url"], markdown))
return pages
@function_tool
def discover_leads_batch(source_urls: list[str]) -> list[Lead]:
"""Fetch several directory pages as one job and return every company listed.
Args:
source_urls: directory or listing page urls to fetch
"""
job_id = _submit_batch(source_urls)
leads = []
for url, markdown in _collect_batch(job_id):
leads.extend(Lead(**lead) for lead in _extract_leads(markdown, url))
return leads
if __name__ == "__main__":
# two pages of the same directory, submitted as one job
urls = [
"https://clutch.co/it-services",
"https://clutch.co/it-services?page=2",
]
job_id = _submit_batch(urls)
print(f"submitted {job_id}")
for url, markdown in _collect_batch(job_id):
print(f"{url}: {len(markdown)} chars")
One directory page rarely produces enough leads to matter, and a directory usually has many pages. Running each page as its own Fetch call means you manage concurrency and retries yourself. Zenrows Batch takes the whole URL list as one managed job and handles the queue, concurrency, retries, and result delivery. A single job accepts up to 100,000 URLs.
Batch authenticates with an X-API-Key header, while Fetch takes the key as a query parameter. Each successful task returns a presigned result_url that stays valid for two hours.
You can also reach Batch through the Python SDK, the zenrows batch CLI, and the Zenrows MCP server, and none of them need the polling loop. This example uses raw REST so every step stays visible. For the full implementation guide, read running large-scale scraping jobs with Batch.
Wrapping up
You now have an AI lead generation agent that reads a directory with Zenrows Fetch, pulls out every company listed, visits each company's domain, and scores the results against a plain-language ICP. The run in this tutorial pulled 78 companies off a single page, and agent.py scores the first 10 of them per run at under 20 seconds each.
That's automated lead generation with no vendor list underneath it. The agent reads the source, and the source is whichever directory your market uses.
What's next
- Raise or remove the 10-lead cap in
INSTRUCTIONSwhen you want a full pass over the directory. - Put a human review checkpoint between scoring and outreach, then push qualifying leads into a CRM such as HubSpot or Salesforce, an email sequencing tool, or a CSV export.
- Wrap
agent.pyin an HTTP endpoint and trigger it as a webhook step from n8n, Zapier, or Power Automate, passing the source URL and the ICP in the request body. - Skip the wrapper entirely with the Zenrows MCP server, which exposes the same fetch capability to any MCP-compatible agent. This walkthrough of the MCP server covers the setup.
FAQ
How much does it cost to run this agent?
A standard Zenrows request costs one credit. mode=auto selects the configuration each target needs, and a page that requires the full configuration costs up to 25 credits. These weights stay fixed across all plans.
In this tutorial, the directory page resolved at the standard rate. Enrichment drives the cost because it fetches the homepage plus each nav link it finds, which came to five or six fetches per lead. A 10-lead run costs roughly 60 credits. You only pay for successful requests, so a blocked page costs nothing. The Zenrows pricing documentation has the details.
Can I use a different LLM?
You can swap the model behind the agent wrapper because the Zenrows tools are plain Python functions. extract_leads and score_lead depend on OpenAI's structured output API, so you'll need to replace those two calls if you switch providers.
Does this agent work on LinkedIn?
You shouldn't point it at LinkedIn. LinkedIn's user agreement prohibits automated access, and enforcement can lead to account restrictions. Point the agent at sources that permit automated access, such as public directories, business registries, and review platforms.
How does this compare to using Apollo or Clay?
Indexed datasets return the same companies when everyone queries with the same filters, and those filters are limited to the fields the dataset captures. This agent reads the source directly and scores each company on signals from its own site, so two teams with the same ICP can still surface different leads.




Top comments (0)