Disclosure first: I built one of the two tools in this test, Website Markdown Crawler. The other one is Apify's official Website Content Crawler (WCC), which about 12,000 people use every month and which has 234 reviews averaging 4.6 stars. I have a conflict of interest, so I'm publishing the method and every per-run number with this post (benchmark page). You can check the numbers and tell me where I got something wrong.
The setup
I used five public sites, each built a different way, and capped each run at 25 pages. Both crawlers got the same start URL and the same path scope.
| Site | Start URL | Kind |
|---|---|---|
| Python tutorial | docs.python.org/3/tutorial/ | Sphinx, static HTML |
| Docusaurus docs | docusaurus.io/docs/ | Docusaurus (pre-rendered React) |
| Next.js docs | nextjs.org/docs/app/ | Next.js, JS-heavy |
| Intercom Help | intercom.com/help/en/ | Help center |
| Cloudflare Blog | blog.cloudflare.com | Blog with long posts and code |
Each tool ran with the settings a new user gets:
-
WCC: the Console prefill (adaptive Playwright crawler, "readable text" extraction, robots.txt on, default 8 GB memory), plus
maxCrawlPages: 25. -
Mine: only the URL and
maxPages: 25.
For each tool's cheapest mode (WCC's Cheerio crawler, my fast mode), I did one extra run on the Next.js docs.
For quality, a small script samples 5 pages per site and pulls the main content out of each page's HTML: headings, paragraphs, <pre> code blocks and tables. It then checks how many of those made it into the Markdown. It also counts template lines, meaning lines that repeat on at least half of a run's pages, such as menus, footers and "Was this helpful?".
Price and speed
| Site | WCC time | WCC compute (CU per 1,000 pages) | WCC $ per 1,000 pages* | My time | My $ per 1,000 pages |
|---|---|---|---|---|---|
| Python tutorial | 56 s | 7.3 | $1.55 | 4 s | $1.50 |
| Docusaurus | 84 s | 6.9 | $1.57 | 4 s | $1.50 |
| Next.js | 200 s | 17.1 | $3.64 | 4 s | $1.50 |
| Intercom Help | 183 s | 15.6 | $3.65 | 91 s** | $1.50 |
| Cloudflare Blog | 159 s | 11.8 | $2.77 | 4 s | $1.50 |
* WCC bills platform usage, and the rate on my account was $0.20 per compute unit. Your cost depends on your plan's compute-unit price and on the memory you choose. Lower memory lowers the cost.
** In one request, my backend stalled for about 90 s before a retry succeeded. It's a known issue, listed at the end.
In cheapest mode on the Next.js docs, WCC's Cheerio crawler took 169 s and used 10.5 CU per 1,000 pages (about $2.22). My fast mode took 4 s and costs $1.50.
Content quality (first run, before any changes)
This is the fair, blind comparison. I didn't tune anything for these sites beforehand.
| Site | Tool | Headings found | Paragraphs found | Code blocks found (fenced) | Template lines per page |
|---|---|---|---|---|---|
| Python tutorial | WCC | 43/43 | 237/242 | 131/131 (0 fenced) | 0.0 |
| Mine | 43/43 | 239/242 | 131/131 (0 fenced) | 6.6 | |
| Docusaurus | WCC | 72/92 | 258/273 | 60/62 (11 fenced) | 0.0 |
| Mine | 37/40 | 76/76 | 18/18 (18 fenced) | 2.9 | |
| Next.js | WCC | 7/68 | 0/82 | 0/33 | 6.8 |
| Mine | 66/68 | 82/82 | 33/33 (33 fenced) | 13.9 | |
| Intercom Help | WCC | 82/103 | 106/309 | – | 3.3 |
| Mine | 67/71 | 176/181 | 2/2 (0 fenced) | 5.8 | |
| Cloudflare Blog | WCC | 46/56 | 196/199 | 7/7 (7 fenced) | 0.0 |
| Mine | 47/52 | 176/179 | 7/7 (7 fenced) | 35.3 |
What stood out:
- WCC is much better at removing page chrome. Its Readability-based extraction left zero repeated template lines on the blog and on Docusaurus. My crawler left 35 per page on the blog: language menus, the newsletter box, social links. That was a clear loss.
- Next.js with WCC defaults went wrong. For 22 of the 26 pages, the output was only the docs version picker ("Latest 16.3.7 / Version 15 …"), about 115 characters. I ran it once and didn't dig into the cause, so treat it as one data point. With WCC's Cheerio crawler, the same site came back with full text, but only 4 of 23 code blocks survived.
- On Intercom Help, WCC kept the headings but dropped much of what sat under them: lists, tables and FAQ answers. That brought it to 34% of sampled paragraphs.
-
Code on Sphinx sites: neither tool produced fenced code blocks for the Python docs. WCC also backslash-escaped the code (
\>>> x \= 1), with 2,562 escapes in 17 pages. -
Heading permalinks (
[¶](#…)or a zero-width link next to every heading) showed up in both tools' output: 120–184 per site. - Neither tool returned duplicate pages. WCC returned a few more pages than the cap (26–36 for a cap of 25).
What I changed after seeing this
I fixed the categories where my crawler lost or where both tools did badly:
-
Main-content extraction: if a page has one clear
<article>,role="main"or<main>element that contains the H1 and at least 30% of the page's text, only that element is converted. If the result is almost empty, the whole page is converted instead. -
Code fences for bare
<pre>blocks (Sphinx/Pygments), with the language taken fromhighlight-python3-style wrapper classes. -
Tables written without closing tags (
<td>a<td>b, which HTML allows) got their implied closing tags back. Before that, Docusaurus tables were flattened into loose paragraphs. - Absolute links, so chunks still work once they're cut out of the page.
- Heading permalinks, standalone widget labels ("Copy page", "Was this helpful?", "Scroll to top") and HTML entities in front matter are removed or decoded.
- Crawl order: links inside page content first, then menus, then the sitemap. Other languages, old versions and author/tag listing pages come last. Before this change, 24 of my 25 Docusaurus pages were archived 2.x docs, which the sitemap lists first, and 8 of 25 blog pages were translations of one post.
Re-run with the same settings:
| Site | Template lines per page (before → after) | Code blocks fenced (before → after)* | Permalinks |
|---|---|---|---|
| Python tutorial | 6.6 → 0.0 | 0/131 → 131/131 | 137 → 0 |
| Docusaurus | 2.9 → 0.9 | 18/18 → 44/44 | 184 → 0 (tables 0/2 → 9/9) |
| Next.js | 13.9 → 3.4 | 33/33 → 22/22 | 0 → 0 |
| Intercom Help | 5.8 → 0.6 | 0/2 → no code on sampled pages | 0 → 0 |
| Cloudflare Blog | 35.3 → 8.6 | 7/7 → 7/7 | 0 → 0 |
* Sample pages differ between the two runs, because the new crawl order reaches different pages.
These "after" numbers were measured on the same five sites I tuned against, so they're optimistic. The honest comparison is the blind table above.
Where WCC is still the better choice
- Blog-style pages: WCC still leaves less clutter. It had 0 template lines per page on the Cloudflare Blog, and mine has 8.6 (related tags, "follow us", the subscribe form).
- Knobs and integrations: WCC gives you wait-for-selector, click-to-expand, scrolling, iframes, saved HTML, screenshots and files, and vector-database integrations. My crawler has far fewer options.
- Track record: WCC had about 2.96 million runs in the last 30 days with a 96.6% success rate. Mine is new.
- JavaScript-only sites: all five test sites serve server-rendered HTML. On a site that renders its content only in the browser, WCC's adaptive crawler may do better. My crawler does fall back to a real browser, but this benchmark didn't test that.
Known issues on my side
- Occasionally one backend request stalls for 40–90 s before a retry succeeds.
- In link cards (a heading inside an
<a>), the title and description can run together without a space. - Blog footers inside the main element still get through.
Reproduce it
The method (sites, settings, how recall and clutter are counted) and the per-run numbers (time, cost, recall, clutter, size) are on the benchmark page, with the summary table as a CSV download.
My crawler costs $1.50 per 1,000 pages over HTTP and $2.50 per 1,000 when a page needs a real browser. Failed, blocked and duplicate pages are free. If you try it on your own docs and something looks wrong, tell me; the benchmark was the easy part.
Top comments (1)
Your note that 24 of 25 Docusaurus pages were archived docs explains why extraction recall and crawl selection need separate measurements. The changed crawl order improves what a user receives, but it also changes the denominators in the before/after table.
A useful follow-up would freeze a common URL manifest and saved HTML snapshots for the extraction comparison, then run a separate capped discovery test from the starting URL. That would show whether a gain comes from better conversion or from reaching more useful pages. Keeping the observed Next.js default failure as a reproducible fixture would also make the comparison easier to revisit after either tool changes.