DEV Community

Tidy Tools
Tidy Tools

Posted on

Turn any docs site into RAG-ready Markdown for $1 per 1,000 pages

Every RAG project starts the same way: "first, get the docs into clean text." And every time it takes longer than it should: menus and cookie banners leak into your chunks, JavaScript-rendered pages come back empty, PDFs need a separate pipeline, and crawling the same page under five URLs quietly inflates your embedding bill.

I built a small crawler that handles all of that and runs on Apify, so you can call it from code, a schedule, or an AI agent.

What it does

Give it a start URL and a page limit:

{
  "startUrls": [{ "url": "https://docs.example.com/" }],
  "maxPages": 200,
  "cssSelector": "main",
  "chunkSize": 2000
}
Enter fullscreen mode Exit fullscreen mode

You get one item per page:

{
  "url": "https://docs.example.com/getting-started",
  "title": "Getting started",
  "markdown": "# Getting started\n\nInstall the package...",
  "chunks": [{ "index": 0, "text": "# Getting started..." }],
  "mode": "fast"
}
Enter fullscreen mode Exit fullscreen mode

The details that matter for RAG

  • Stays in scope. By default it only follows links under the start URL's folder, so /docs/ doesn't wander into /blog/ and /careers/.
  • Clean content. A CSS selector like main drops navigation and footers before conversion.
  • Documents included. Linked PDFs, Word and Excel files are converted to Markdown in the same run.
  • Chunking built in. Paragraph-aware chunks with overlap, ready for your embedding model.
  • Duplicates removed. Pages with identical content under different URLs (tracking params, session ids) are stored once and not charged twice.
  • Fast first, browser when needed. Most pages are fetched over plain HTTP (fast, $1 / 1,000 pages). If a site blocks that or is a JavaScript shell, it switches to a real headless browser automatically ($2.50 / 1,000 pages).
  • Pay only for success. Blocked pages, errors and duplicates are free.

Calling it from Python

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("tidytools/website-markdown-crawler").call(run_input={
    "startUrls": [{"url": "https://docs.example.com/"}],
    "maxPages": 100,
    "chunkSize": 2000,
})
for page in client.dataset(run["defaultDatasetId"]).iterate_items():
    for chunk in page.get("chunks", []):
        index.add(text=chunk["text"], metadata={"url": page["url"], "title": page["title"]})
Enter fullscreen mode Exit fullscreen mode

From an AI agent

Because it's an Apify Actor, agents can use it as a tool through the Apify MCP server: "read the docs at X and answer Y" becomes one tool call.

Try it

Website to Markdown Crawler on Apify. New Apify accounts get free monthly credit, so a few hundred pages cost nothing to try.

If you need structured data instead of Markdown (prices, job details, company info), there's a sister tool, AI Web Data Extractor: list the fields you want and get JSON back.

Feedback and feature requests are very welcome.

Top comments (0)