Every RAG project starts the same way: "first, get the docs into clean text." And every time it takes longer than it should: menus and cookie banners leak into your chunks, JavaScript-rendered pages come back empty, PDFs need a separate pipeline, and crawling the same page under five URLs quietly inflates your embedding bill.
I built a small crawler that handles all of that and runs on Apify, so you can call it from code, a schedule, or an AI agent.
What it does
Give it a start URL and a page limit:
{
"startUrls": [{ "url": "https://docs.example.com/" }],
"maxPages": 200,
"cssSelector": "main",
"chunkSize": 2000
}
You get one item per page:
{
"url": "https://docs.example.com/getting-started",
"title": "Getting started",
"markdown": "# Getting started\n\nInstall the package...",
"chunks": [{ "index": 0, "text": "# Getting started..." }],
"mode": "fast"
}
The details that matter for RAG
-
Stays in scope. By default it only follows links under the start URL's folder, so
/docs/doesn't wander into/blog/and/careers/. -
Clean content. A CSS selector like
maindrops navigation and footers before conversion. - Documents included. Linked PDFs, Word and Excel files are converted to Markdown in the same run.
- Chunking built in. Paragraph-aware chunks with overlap, ready for your embedding model.
- Duplicates removed. Pages with identical content under different URLs (tracking params, session ids) are stored once and not charged twice.
- Fast first, browser when needed. Most pages are fetched over plain HTTP (fast, $1 / 1,000 pages). If a site blocks that or is a JavaScript shell, it switches to a real headless browser automatically ($2.50 / 1,000 pages).
- Pay only for success. Blocked pages, errors and duplicates are free.
Calling it from Python
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("tidytools/website-markdown-crawler").call(run_input={
"startUrls": [{"url": "https://docs.example.com/"}],
"maxPages": 100,
"chunkSize": 2000,
})
for page in client.dataset(run["defaultDatasetId"]).iterate_items():
for chunk in page.get("chunks", []):
index.add(text=chunk["text"], metadata={"url": page["url"], "title": page["title"]})
From an AI agent
Because it's an Apify Actor, agents can use it as a tool through the Apify MCP server: "read the docs at X and answer Y" becomes one tool call.
Try it
Website to Markdown Crawler on Apify. New Apify accounts get free monthly credit, so a few hundred pages cost nothing to try.
If you need structured data instead of Markdown (prices, job details, company info), there's a sister tool, AI Web Data Extractor: list the fields you want and get JSON back.
Feedback and feature requests are very welcome.
Top comments (0)