An MCP server makes a tool available over a standard protocol, so MCP-compatible clients can call it without a custom integration. For scraping, that tool takes a URL and returns content a model can read. The quickest version wraps requests.get() in a tool decorator and works on static pages. Client-rendered pages are harder: the response can arrive as HTTP 200 with content still inside the page's JavaScript. A reliable tool says when content is missing, rather than letting the model answer from a page it never read. This guide builds one that handles both, in about 90 lines of Python.
TL;DR
- With rendering on, our Web Scraping API cleared 4 or 5 of the 5 protected pages (3 passes), and 2 or 3 without it (5 passes), so headless is the parameter worth setting. Plain httpx cleared 0, a TLS-impersonating client 1.
- Rendering raised one page from 174 characters to 1,567. Both calls reported success, so the server checks content length, not just the status.
- markdown: true turned 652,316 characters of Zillow HTML into 10,010.
- On mcp 2.x the import is from mcp.server.mcpserver_ import MCPServer_. Version 2.0 renamed FastMCP and moved its module, so the mcp.server.fastmcp path in older tutorials no longer exists.
What your agent needs to read a live page
Agents already manage context, call tools, and write code. Reading a live page is the capability worth adding next.
A custom fetch-and-parse tool has 2 problems, and only the first is about scraping. It can break on sites you didn't test. It also lives inside one agent's codebase, so you rebuild it for the next MCP client you adopt.
An MCP server solves the second problem by design. You write the tool once, and any MCP client can call it, including Claude Code, Cursor, and Windsurf.
The first problem is the backend's job.
What you'll need
- Python 3.10 or newer
- A Web Scraping API token from the free plan: run printf 'user:pass' | base64 with your credentials, or [Convert]::ToBase64String([Text.Encoding]::UTF8.GetBytes('user:pass')) in PowerShell
- A virtual environment: run python3 -m venv .venv, then .venv/bin/pip install "mcp>=2,<3" httpx (Windows: python -m venv .venv, then .venv\Scripts\pip.exe and .venv\Scripts\python.exe)
What a scraping backend handles
A plain HTTP client faces 3 failure modes on the live web.
- Pages that load their content in the browser. The quotes on quotes.toscrape.com/js/ exist only inside a var data = [...] array until JavaScript runs, so a client that can't run it sees none of them. The page contains 0 quote elements with rendering off, and 10 with it on. The response is still HTTP 200 and still full of valid markup, so neither the status nor the size of the HTML tells you the content is missing.
- Sites that refuse unfamiliar clients. We tested 5 search pages – jobs, retail, and property – 5 times each. Plain httpx with a Chrome user agent retrieved nothing on any pass. It advertises Chrome while sending Python's TLS signature, and anti-bot systems fingerprint the signature, not the header. curl_cffi, which impersonates Chrome's TLS, retrieved 1 of 5.
- Markup that varies by site. A parser written for one site's markup often won't work on the next one, and raw HTML uses far more context than it needs.
On those same 5 pages, our Web Scraping API retrieved 4 or 5 with rendering on, averaging 4.7 over 3 passes, against 2 or 3 on default settings, averaging 2.4 over 5 passes. Rendering recovered the sites that the default never cleared.
One POST endpoint covers all 3. You send a URL against the universal target, add headless: "html" when a page needs JavaScript, and set geo for locale or device_type for device. We route requests through our standard and premium proxy pools, so most of the blocking problem is handled on our side. Our docs list the current country coverage. No scraper clears every page every time, which is why the tool reports failures instead of hiding them.
With markdown: true set in the Web Scraping API, the HTML becomes Markdown server-side, so a Python server needs no HTML-parsing dependency.
Markdown output is a fraction of the HTML's size:
| Page | Raw HTML | Markdown |
|---|---|---|
| Zillow, Austin listings | 652,316 chars | 10,010 chars |
| Wikipedia, Web scraping | 235,439 chars | 63,166 chars |
| Hacker News front page | 34,275 chars | 10,046 chars |
Every character you don't send is context the model keeps for reasoning.
Building the server
The server's only job is turning a tool call into an API call and formatting the response. One tool does that, and render_js goes in the signature with rendering off by default, so it stays off unless a page needs it.
Save this as server.py:
import os
import httpx
from mcp.server.mcpserver import MCPServer
from mcp.types import ToolAnnotations
API = "https://scraper-api.decodo.com/v2/scrape"
TOKEN = os.environ.get("DECODO_TOKEN")
if not TOKEN:
raise SystemExit("DECODO_TOKEN is unset. Set it to your Decodo user:password in base64.")
ALLOWED = [h for h in os.environ.get("ALLOWED_HOSTS", "").split(",") if h]
MAX_CHARS = 40_000
client = httpx.AsyncClient(timeout=90.0, limits=httpx.Limits(max_connections=10))
server = MCPServer(name="web-scraper", version="1.0.0")
def permitted(url: str) -> bool:
"""Allow everything when ALLOWED_HOSTS is unset, otherwise match the host exactly."""
if not ALLOWED:
return True
try:
return httpx.URL(url).host in ALLOWED
except httpx.InvalidURL:
return False
@server.tool(
description=(
"Fetch a live web page and return it as Markdown. Set render_js=true only when a page "
"builds its content in the browser, because rendering is slower and costs more. "
"country takes an ISO 3166-1 alpha-2 code such as US, DE or JP."
),
annotations=ToolAnnotations(read_only_hint=True, open_world_hint=True),
)
async def scrape_page(url: str, render_js: bool = False, country: str | None = None) -> str:
"""Scrape one URL through the Decodo Web Scraping API and return Markdown."""
if not permitted(url):
return f"Refused: {url} is not in ALLOWED_HOSTS."
payload = {"url": url, "target": "universal", "markdown": True}
if render_js:
payload["headless"] = "html"
if country:
payload["geo"] = country
try:
r = await client.post(API, json=payload, headers={"Authorization": f"Basic {TOKEN}"})
except httpx.RequestError as exc:
return f"Scrape failed: could not reach Decodo ({type(exc).__name__})."
try:
body = r.json()
except ValueError:
return f"Scrape failed: Decodo returned HTTP {r.status_code} and a non-JSON body."
if not isinstance(body, dict):
return f"Scrape failed: Decodo returned HTTP {r.status_code} and an unexpected body."
if r.status_code != 200:
errors = [e for e in (body.get("errors") or []) if isinstance(e, dict)]
named = "; ".join(e.get("message", "") for e in errors)
return f"Scrape failed ({r.status_code}): {named or body.get('message', 'bad request')}"
results = body.get("results")
if not isinstance(results, list) or not results or not isinstance(results[0], dict):
return f"Scrape failed: {body.get('message', 'the target could not be scraped')}"
result = results[0]
landed = result.get("url") or url
if not permitted(landed):
return f"Refused: {url} redirected to {landed}, which is not in ALLOWED_HOSTS."
status = result.get("status_code")
content = str(result.get("content") or "")
if len(content) < 500:
if status != 200:
return f"Scrape failed: {url} returned HTTP {status}."
retry = "" if render_js else " If the page renders in a browser, retry with render_js=true."
return f"{content}\n\n[Only {len(content)} characters came back.{retry}]"
flag = "" if status == 200 else f"\n\n[{url} returned HTTP {status}, so this page may be partial.]"
if len(content) > MAX_CHARS:
return f"{content[:MAX_CHARS]}{flag}\n\n[Truncated at {MAX_CHARS} of {len(content)} chars.]"
return content + flag
if __name__ == "__main__":
server.run(transport="stdio")
MCP 2.0 changed that import line, renaming FastMCP to MCPServer. The migration guide documents the rest, including Tool.inputSchema becoming Tool.input_schema.
The 500-character threshold on the short-content check is a starting point worth adjusting for your sites.
Most of the rest of the file is response handling. The API reports the page's own status inside results, so a 404 page arrives as HTTP 200 with status_code: 404, and the tool checks both codes. On a 400, the errors array carries a message naming the offending field, such as headless or proxy_pool, and the tool returns it. The allowlist is checked twice, on the supplied URL and on results[0].url, because the API follows redirects and one open redirect would defeat a single check.
Connecting it to your client
In the directory holding server.py, check that the server works before you register it:
DECODO_TOKEN=your_base64_token .venv/bin/python -c \
"import asyncio; from server import scrape_page; \
print(asyncio.run(scrape_page(url='https://quotes.toscrape.com'))[:200])"
That tests the token and the dependencies without an MCP client involved. Username invalid means the token isn't base64; Incorrect username or password means it is. Then register the server, using the .venv interpreter, not the system one:
claude mcp add web-scraper -e DECODO_TOKEN=your_base64_token \
-- /full/path/to/project/.venv/bin/python /full/path/to/project/server.py
That command is for Claude Code, and registers the server to the directory you run it in. For Cursor, Windsurf, or Claude Desktop, the equivalent entry goes in that client's MCP config file:
{
"mcpServers": {
"web-scraper": {
"command": "/full/path/to/project/.venv/bin/python",
"args": ["/full/path/to/project/server.py"],
"env": { "DECODO_TOKEN": "your_base64_token" }
}
}
}
Ask for a page in plain language, and the agent chooses the arguments itself:
The agent read render_js from the tool description and set it unprompted. That page is worth seeing, because every quote arrives through JavaScript:
Without rendering, the tool returns 174 characters of navigation links, and not one quote. The hint it appends tells the model to retry with render_js=true rather than summarize what little arrived.
The parameter reference lists the rest of what universal accepts, and is where to look before widening the signature.
Bulk work uses different endpoints. POST /v3/task/batch queues many URLs, and GET /v3/task/{id}/results answers 204 until ready. Pass the URLs under url rather than query. Our 3 queued URLs finished within 10 seconds.
We also publish a ready-made server covering search, eCommerce, and social targets, plus code examples in several languages.
Running it in production
- Reading a near-empty scrape as an API problem. A missing headless parameter and a genuinely broken page both return HTTP 200, so the status code won't distinguish them. When a page returns too little, try rendering first, then the premium pool, which you can reach by adding a proxy_pool parameter. In our runs, rendering was enough, and the premium pool handles sites that need more. Only successful responses are billed, so a failed attempt is free. Rendered requests are priced separately from standard ones – see the plans page.
- Trusting the outer status code. On 1 of 5 passes, Etsy answered HTTP 403 and returned a full results page anyway, with 24 listing links and 33 prices. Rejecting on status alone discards it, so the tool reports the status and returns the content.
- Returning raw HTML. A 652,316-character page like those Zillow results arrives as roughly 240K tokens of mostly markup and hides the listings you wanted. markdown: true is the fix.
- Treating scraped text as trusted input. A page your agent fetched is attacker-controlled text arriving in a model's context. Indirect prompt injection through tool results is a documented attack on MCP-connected systems. Never let scraped content reach a tool that writes or spends, and have a person approve anything irreversible.
- Letting the model choose URLs with no limits. ALLOWED_HOSTS starts empty so the examples run anywhere. Set it to a comma-separated list like example.com,docs.example.com before the server runs unattended. Add a rate limit too, and remember that scraping a site is a decision about its terms, not only a technical one.
Final thoughts
MCP makes the tool portable, and the Web Scraping API handles the proxy and browser work, so neither part does the other's job. Keep the server small, make render_js a parameter, and let the API return Markdown so your context window fills with content instead of markup. Report ranges rather than single runs when you benchmark a site, because the same URL that gave us job listings in the morning needed rendering by evening. A second tool that calls our fast search endpoint lets the same agent search before it scrapes. Get a token from the Decodo's web scraping solutions free plan, run scrape_page on a URL your current tool can't handle, and compare what you get.


Top comments (0)