This article was originally published on the Zenrows blog.
In this tutorial, you'll build a Mastra agent with a custom Zenrows Fetch tool that reads protected and JavaScript-rendered pages, then add a workflow that fetches a list of URLs in parallel. You'll need Node.js 22.13 or later, an OpenAI or Anthropic API key, and a Zenrows API key.
The finished agent compares live Walmart product listings and returns either Markdown to reason over or structured JSON with fields like price and rating. The complete code is on GitHub.
Prerequisites
- You need Node.js 22.13 or later.
- You need an API key for one model provider. This tutorial uses OpenAI by default and shows the Anthropic alternative. Get an OpenAI key from your OpenAI developer dashboard or an Anthropic key from your Anthropic developer dashboard.
- You need a Zenrows API key to retrieve the full contents of web pages. Create one from your Zenrows dashboard.
Why the built-in fetch tool falls short
Mastra's built-in webFetchTool makes plain HTTP requests, doesn't run JavaScript, and returns raw HTML. That's enough for simple pages. On browser-rendered or protected pages, such as a product listing protected by Cloudflare, it breaks quietly.
Cloudflare returns its challenge page with a normal HTTP status code, so the tool raises no error. The agent treats the challenge as the result and reports a missing or wrong price.
The fix belongs in the transport layer, which controls HTTP requests, headers, proxies, and whether JavaScript runs. Zenrows Fetch with mode=auto handles DataDome, Akamai, and Cloudflare in a single call. Wrapping that call in Mastra's createTool keeps typed input and output schemas around the fetch.
1. Set up the Mastra project
npm create mastra@latest
The CLI scaffolds the project structure, config files, and a starter agent, and installs @mastra/core and its dependencies. Follow the prompts to name the project and pick your LLM provider. Then add Zod if the scaffold didn't include it.
npm install zod
Create a .env file in the project root with your keys. You only need the key for the provider you plan to use.
ZENROWS_API_KEY=your_zenrows_key
OPENAI_API_KEY=your_openai_key
ANTHROPIC_API_KEY=your_anthropic_key
Add .env to .gitignore so your credentials stay out of version control. Mastra loads .env automatically, but the standalone test scripts in this tutorial run outside Mastra's loader, so install dotenv for them.
npm install dotenv --save
Your project structure should look similar to this.
src/
└── mastra/
├── agents/
│ └── agent.ts
├── tools/
│ └── schedule-tools.ts
└── index.ts
Start the dev server.
npm run dev
Mastra Studio opens at http://localhost:4111. It reads the registered agents from your Mastra instance and gives you a chat UI to test them.
2. Build the Zenrows Fetch tool
Create src/mastra/tools/zenrows-fetch.ts next to the template's example tools.
import { createTool } from '@mastra/core/tools';
import { z } from 'zod';
export const zenrowsFetchTool = createTool({
id: 'zenrows-fetch',
description:
'Fetches web pages using Zenrows. Use this tool when you need information from a specific URL. Set extractJson to true when you need structured data such as product details, pricing, availability, ratings, or other fields. Leave extractJson false when you need the page content as Markdown.',
inputSchema: z.object({
url: z.string().url(),
extractJson: z.boolean().optional(),
}),
outputSchema: z.object({
content: z.union([z.string(), z.record(z.string(), z.unknown())]),
statusCode: z.number(),
url: z.string(),
}),
execute: async ({ url, extractJson }) => {
const zenrowsUrl = new URL('https://api.zenrows.com/v1/');
zenrowsUrl.searchParams.set('url', url);
zenrowsUrl.searchParams.set(
'apikey',
process.env.ZENROWS_API_KEY!,
);
zenrowsUrl.searchParams.set('mode', 'auto');
if (extractJson) {
zenrowsUrl.searchParams.set('extract', 'auto');
} else {
zenrowsUrl.searchParams.set('response_type', 'markdown');
}
const response = await fetch(zenrowsUrl);
if (!response.ok) {
const error = await response.text();
throw new Error(
`Zenrows request failed with status ${response.status}: ${error}`,
);
}
if (extractJson) {
const body = await response.json();
return {
content: body.parsed,
statusCode: response.status,
url,
};
}
const content = await response.text();
return {
content,
statusCode: response.status,
url,
};
},
});
A createTool definition needs five parts. The id names the tool, and the description is the only signal the model uses to decide whether to call it, so keep it brief and exact. The inputSchema validates inputs, the outputSchema types the response, and execute calls Zenrows Fetch with mode=auto.
z.string().url() rejects malformed URLs before they reach the API. extractJson is optional and defaults to undefined, which falls through to the Markdown path.
Test the tool on its own with src/mastra/tools/test-tool.ts. It doesn't need the dev server or the agent.
import 'dotenv/config';
import { zenrowsFetchTool } from './zenrows-fetch';
if (!zenrowsFetchTool.execute) {
throw new Error('Zenrows tool does not have an execute function');
}
const result = await zenrowsFetchTool.execute(
{
url: 'https://www.walmart.com/search?q=bike',
extractJson: false,
},
{} as any
);
if (!result || 'statusCode' in result === false) {
throw new Error('Zenrows tool returned an unexpected result');
}
console.log(result.statusCode);
console.log(
typeof result.content === 'string'
? result.content.slice(0, 300)
: result.content
);
Run it.
npx tsx src/mastra/tools/test-tool.ts
The output prints the status code, then the first lines of the page as Markdown.
200
[](https://www.walmart.com)
[Skip to Main Content](https://www.walmart.com#maincontent)
[](https://www.walmart.com/all-departments)[](https://www.walmart.com/)
[](https://www.walmart.com/)
If you'd prefer to skip maintaining a custom tool, the Zenrows MCP server is an alternative integration. This walkthrough of the Zenrows MCP server and its implementation covers the setup.
3. Add structured extraction
// part of the previous script
zenrowsUrl.searchParams.set('mode', 'auto');
if (extractJson) {
// server side parsing, returns fields instead of page text
zenrowsUrl.searchParams.set('extract', 'auto');
} else {
zenrowsUrl.searchParams.set('response_type', 'markdown');
}
// parses as JSON
if (extractJson) {
const body = await response.json();
return {
content: body.parsed,
statusCode: response.status,
url,
};
}
const content = await response.text();
return {
content,
statusCode: response.status,
url,
};
These two excerpts from the tool handle the JSON path. When extractJson is true, the request sends extract=auto, and Zenrows Extract parses the page server side into named fields. The response branch then reads the body as JSON and returns body.parsed. When extractJson is false, the tool requests Markdown and reads the body as text.
Markdown suits pages the agent needs to read and reason over. Product listings work better as fields. Test the JSON path with src/mastra/tools/test-tool-extend-extract.ts.
import 'dotenv/config';
import { zenrowsFetchTool } from './zenrows-fetch';
if (!zenrowsFetchTool.execute) {
throw new Error('Zenrows tool does not have an execute function');
}
const result = await zenrowsFetchTool.execute(
{
url: 'https://www.walmart.com/search?q=bike',
extractJson: true,
},
{} as any,
);
if (!result || 'statusCode' in result === false) {
throw new Error('Zenrows tool returned an unexpected result');
}
console.log(result.statusCode);
console.log(
typeof result.content === 'string'
? result.content.slice(0, 300)
: result.content
);
npx tsx src/mastra/tools/test-tool-extend-extract.ts
Here's the parsed response, truncated to two product records.
{
"products": [
{
"badge": null,
"current_price": 149.99,
"fulfillment_label": null,
"is_sponsored": true,
"item_id": "10479873441",
"low_stock": null,
"options_from_price": 149.99,
"original_price": 207.99,
"rating": 3.9,
"review_count": 130,
"title": "Ktaxon 20\" Mountain Bike, 7 Speed Bike with Disc Brakes, White",
"walmart_plus_savings": null
},
{
"badge": null,
"current_price": 227,
"fulfillment_label": null,
"is_sponsored": true,
"item_id": "475992717",
"low_stock": null,
"options_from_price": null,
"original_price": null,
"rating": 4.4,
"review_count": 739,
"title": "Mongoose Rebel X1 BMX Bike, 20-in. Wheels, Kids Ages 7-14 Years, Gray Child Bicycle",
"walmart_plus_savings": null
}
],
"search_query": {
"query": "bike",
"total_results": null
}
}
4. Wire the tool into a Mastra agent
Create the agent at src/mastra/agents/market-research-agent.ts.
import { Agent } from '@mastra/core/agent';
import { zenrowsFetchTool } from '../tools/zenrows-fetch';
export const marketResearchAgent = new Agent({
id: 'market-research-agent',
name: 'Market Research Agent',
description:
'An agent that researches and compares products using live web data.',
instructions: `
You are a product market research agent.
Your job is to research and compare products using the zenrows_fetch tool.
When the user provides one or more product URLs:
1. Fetch each URL using zenrows_fetch.
2. When product information such as name, price, availability, or rating is needed, use structured extraction.
3. Extract the relevant product information from each result.
4. Never invent or assume information that was not returned by the tool.
5. If information is missing, clearly state that it is unavailable.
6. Preserve the currency returned by the source.
7. Do not directly compare prices when the products use different currencies unless a reliable currency conversion is available.
8. When comparing products, present the results clearly in a table when appropriate.
9. After presenting the data, provide useful observations based only on the retrieved information.
Do not use web search when the user asks you to use zenrows_fetch only.
`,
model: 'openai/gpt-5.6-terra',
tools: {
zenrows_fetch: zenrowsFetchTool,
},
});
The instructions need to be specific. A vague prompt gives you a confident answer with a price from training data and no way to tell which fields came from the page. These instructions forbid invented values, require the agent to flag missing data, and tell it when to use structured extraction.
You can switch to any supported model by changing the model line. Make sure the matching key is in your .env file.
// changing model. Replace it in the previous script
// openai
model: 'openai/gpt-5.6-terra',
// anthropic
model: 'anthropic/claude-sonnet-4-5',
Register the agent in src/mastra/index.ts. Registration exposes it to Mastra Studio and the server.
import { Mastra } from '@mastra/core/mastra';
import { LibSQLStore } from '@mastra/libsql';
import { DuckDBStore } from '@mastra/duckdb';
import { MastraCompositeStore } from '@mastra/core/storage';
import {
MastraStorageExporter,
MastraPlatformExporter,
Observability,
SensitiveDataFilter,
} from '@mastra/observability';
import { agent } from './agents/agent';
import { marketResearchAgent } from './agents/market-research-agent';
import { startScheduleTool, stopScheduleTool } from './tools/schedule-tools';
export const mastra = new Mastra({
bundler: {
externals: ['@duckdb/node-bindings'],
},
agents: {
agent,
marketResearchAgent,
},
tools: { startScheduleTool, stopScheduleTool },
storage: new MastraCompositeStore({
id: 'composite-storage',
default: new LibSQLStore({
id: 'mastra-storage',
url: process.env.TURSO_DATABASE_URL || 'file:./mastra.db',
authToken: process.env.TURSO_AUTH_TOKEN || undefined,
}),
domains: {
observability: await new DuckDBStore().getStore('observability'),
},
}),
observability: new Observability({
configs: {
default: {
serviceName: 'mastra',
exporters: [new MastraStorageExporter(), new MastraPlatformExporter()],
spanOutputProcessors: [new SensitiveDataFilter()],
},
},
}),
});
Test the agent with src/mastra/tools/test-agent.ts.
import 'dotenv/config';
import { mastra } from '..';
const agent = mastra.getAgent('marketResearchAgent');
const result = await agent.generate(
'Compare the bikes on https://www.walmart.com/search?q=bike. Give me the five cheapest with their ratings, and flag any that are sponsored.',
);
console.log(JSON.stringify(result.toolCalls, null, 2));
console.log(result.text);
npx tsx src/mastra/tools/test-agent.ts
The script prints the tool call first. The agent chose structured extraction on its own.
[
{
"type": "tool-call",
"runId": "018d9545-eddd-4fde-8cd1-1884316f4212",
"from": "AGENT",
"payload": {
"toolCallId": "call_3H76SowFUi3SRknDFYXrJu7W",
"toolName": "zenrows_fetch",
"args": {
"url": "https://www.walmart.com/search?q=bike",
"extractJson": true
},
"providerMetadata": {
"openai": {
"itemId": "fc_06708f1f9910751e006a84956972bc81a084b7ebc5cc25f70d"
}
}
}
}
]
Then it prints the agent's answer.
The five lowest-priced bikes in the retrieved Walmart search results are:
| Rank | Bike | Price* | Rating | Reviews | Sponsored |
|---:|---|---:|---:|---:|---|
| 1 | Dynacraft 16 Inch Suspect Boys BMX Bike for Child 5-7 Years | 108 | 4.2/5 | 772 | No sponsorship flag returned |
| 2 | 18" Kent Bicycle Abyss Boy's Freestyle BMX Child Bicycle, Blue | 128 | 4.2/5 | 2,325 | No sponsorship flag returned |
| 3 | 20" Kent Tempest BMX Bicycle, Fits Riders 4'2"-5', Black/Aqua, Child, Unisex | 138 | 4.3/5 | 1,496 | No sponsorship flag returned |
| 4 | Ktaxon 24" Women's 7-Speed Cruiser Bike with Basket & Rack, Green | 189.99 | 4.2/5 | 78 | **Sponsored** |
| 5 | Mongoose Rebel X1 BMX Bike, 20-in. Wheels, Kids Ages 7-14, Gray | 227 | 4.4/5 | 740 | **Sponsored** |
The agent ranked the five cheapest bikes from the live page. Where is_sponsored came back null, it said so and didn't guess. This run happened at a different time from the earlier tool test, so some values differ.
5. Test the agent in Mastra Studio
npm run dev
Open http://localhost:4111, select the market research agent, and send it this prompt.
Compare the bikes on Walmart. Give me the five cheapest with their ratings, and flag any that are sponsored.
Studio shows the zenrows_fetch call the agent chose and the ranked table it built from the returned fields.
To add memory, follow the steps in the Mastra memory documentation.
6. Add a workflow for bulk URLs
Create src/mastra/workflow/bulk-fetch.ts.
import { createStep, createWorkflow } from '@mastra/core/workflows';
import { z } from 'zod';
import { zenrowsFetchTool } from '../tools/zenrows-fetch';
const prepareUrlsStep = createStep({
id: 'prepare-urls',
inputSchema: z.object({
urls: z.array(z.string().url()),
}),
outputSchema: z.array(
z.object({
url: z.string().url(),
}),
),
execute: async ({ inputData }) => {
return inputData.urls.map(url => ({
url,
}));
},
});
const fetchProductStep = createStep({
id: 'fetch-product',
inputSchema: z.object({
url: z.string().url(),
}),
outputSchema: z.object({
url: z.string(),
content: z.record(z.string(), z.unknown()),
statusCode: z.number(),
}),
execute: async ({ inputData }) => {
// createTool types execute as optional, so narrow it before calling
if (!zenrowsFetchTool.execute) {
throw new Error('Zenrows tool does not have an execute function');
}
const result = await zenrowsFetchTool.execute(
{
url: inputData.url,
extractJson: true,
},
{} as any,
);
if (!result || 'statusCode' in result === false) {
throw new Error('Zenrows tool returned an unexpected result');
}
return {
url: inputData.url,
// the tool returns a union, this workflow always extracts
content: result.content as Record<string, unknown>,
statusCode: result.statusCode,
};
},
});
export const productResearchWorkflow = createWorkflow({
id: 'product-research-workflow',
inputSchema: z.object({
urls: z.array(z.string().url()),
}),
outputSchema: z.array(
z.object({
url: z.string(),
content: z.record(z.string(), z.unknown()),
statusCode: z.number(),
}),
),
})
.then(prepareUrlsStep)
// concurrency controls how many urls run at once
.foreach(fetchProductStep, { concurrency: 5 })
.commit();
Mastra workflows are deterministic, so they suit a known list of URLs. prepare-urls turns the input list into one task per URL, and foreach runs fetch-product on five URLs at a time with structured extraction on.
createTool types execute as optional and as taking two arguments. The step narrows it before the call and passes both arguments, because calling it with the input alone doesn't compile.
Register the workflow in src/mastra/index.ts.
// add to the imports
import { productResearchWorkflow } from './workflow/bulk-fetch';
// add to the Mastra constructor, alongside agents
workflows: {
productResearchWorkflow,
},
Test it with src/mastra/workflow/test-workflow.ts.
import 'dotenv/config';
import { mastra } from '..';
const workflow = mastra.getWorkflow('productResearchWorkflow');
const run = await workflow.createRun();
const result = await run.start({
inputData: {
urls: [
'https://www.walmart.com/search?q=tanktop',
'https://www.walmart.com/search?q=shirts',
'https://www.walmart.com/browse/home/shop-water-bottles/4044_623679_639999_7751805_4269055_2976621',
],
},
});
console.log(JSON.stringify(result, null, 2));
npx tsx src/mastra/workflow/test-workflow.ts
The truncated output shows two of the three results. Every URL returned status 200, and the results come back in the same order as the input.
{
"result": [
{
"url": "https://www.walmart.com/search?q=tanktop",
"content": {
"products": [
{
"badge": "rollback",
"current_price": 11.98,
"original_price": 29.98,
"fulfillment_label": "Shipping, arrives today",
"is_sponsored": true,
"item_id": "342153793",
"low_stock": true,
"rating": 4.5,
"review_count": 15482,
"title": null
}
// 8 more products
],
"search_query": { "query": "tanktop", "total_results": null }
},
"statusCode": 200
},
{
"url": "https://www.walmart.com/browse/home/shop-water-bottles/...",
"content": {
"products": [
{
"badge": null,
"current_price": 15.95,
"original_price": 17.98,
"is_sponsored": null,
"item_id": "1604723",
"rating": 4.7,
"review_count": 19927,
"title": "Brita Standard Replacement Water Filter (3 Pack)"
}
// 11 more products
],
"search_query": null
},
"statusCode": 200
}
]
}
You can also run the workflow from Mastra Studio.
Wrapping up
You now have a Mastra agent backed by Zenrows that fetches data from protected sites and returns structured results. The Mastra documentation covers more advanced patterns, and the Zenrows MCP server gives you a different integration path.
What's next
- Add memory so the agent remembers earlier comparisons.
- Move large URL lists to Zenrows Batch so the agent stops firing every request at once. The Batch documentation walks through your first job.
- See the same fetch pattern in Python with this guide to an AI lead generation agent with the OpenAI Agents SDK.
FAQ
Why does Mastra's built-in webFetchTool fail on some sites?
It sends HTTP requests without a browser and without anti-bot handling. Cloudflare-protected targets return a challenge page or block page and raise no exception, so the agent reasons over the challenge copy or empty markup and returns an incomplete answer.
Does the Zenrows tool work with both OpenAI and Anthropic models?
It does. Only the agent's model line changes when you switch providers. The tool definition stays the same because Mastra passes your Zod inputSchema to the AI SDK in whatever tool-calling format the provider expects.
How do I handle rate limits when the agent calls the tool frequently?
Handle 429 responses by limiting concurrency, queueing requests, and retrying with exponential backoff and jitter. For large workloads, send the URLs through Zenrows Batch so the agent doesn't fire every request at once.
Can I connect Zenrows through MCP?
You can. Zenrows has an official MCP server, and Mastra's MCPClient works with it like any other MCP server. It supports Streamable HTTP and local transport and gives agents Fetch and browser tools without a wrapper. You still need an API key.






Top comments (0)