DEV Community

Tidy Tools
Tidy Tools

Posted on

I benchmarked my website-to-Markdown crawler against Apify's Website Content Crawler on 5 real sites

Disclosure first: I built one of the two tools in this test, Website Markdown Crawler. The other one is Apify's official Website Content Crawler (WCC), which about 12,000 people use every month and which has 234 reviews averaging 4.6 stars. I have a conflict of interest, so I'm publishing the method and every per-run number with this post (benchmark page). You can check the numbers and tell me where I got something wrong.

The setup

I used five public sites, each built a different way, and capped each run at 25 pages. Both crawlers got the same start URL and the same path scope.

Site Start URL Kind
Python tutorial docs.python.org/3/tutorial/ Sphinx, static HTML
Docusaurus docs docusaurus.io/docs/ Docusaurus (pre-rendered React)
Next.js docs nextjs.org/docs/app/ Next.js, JS-heavy
Intercom Help intercom.com/help/en/ Help center
Cloudflare Blog blog.cloudflare.com Blog with long posts and code

Each tool ran with the settings a new user gets:

  • WCC: the Console prefill (adaptive Playwright crawler, "readable text" extraction, robots.txt on, default 8 GB memory), plus maxCrawlPages: 25.
  • Mine: only the URL and maxPages: 25.

For each tool's cheapest mode (WCC's Cheerio crawler, my fast mode), I did one extra run on the Next.js docs.

For quality, a small script samples 5 pages per site and pulls the main content out of each page's HTML: headings, paragraphs, <pre> code blocks and tables. It then checks how many of those made it into the Markdown. It also counts template lines, meaning lines that repeat on at least half of a run's pages, such as menus, footers and "Was this helpful?".

Price and speed

Site WCC time WCC compute (CU per 1,000 pages) WCC $ per 1,000 pages* My time My $ per 1,000 pages
Python tutorial 56 s 7.3 $1.55 4 s $1.50
Docusaurus 84 s 6.9 $1.57 4 s $1.50
Next.js 200 s 17.1 $3.64 4 s $1.50
Intercom Help 183 s 15.6 $3.65 91 s** $1.50
Cloudflare Blog 159 s 11.8 $2.77 4 s $1.50

* WCC bills platform usage, and the rate on my account was $0.20 per compute unit. Your cost depends on your plan's compute-unit price and on the memory you choose. Lower memory lowers the cost.
** In one request, my backend stalled for about 90 s before a retry succeeded. It's a known issue, listed at the end.

In cheapest mode on the Next.js docs, WCC's Cheerio crawler took 169 s and used 10.5 CU per 1,000 pages (about $2.22). My fast mode took 4 s and costs $1.50.

Content quality (first run, before any changes)

This is the fair, blind comparison. I didn't tune anything for these sites beforehand.

Site Tool Headings found Paragraphs found Code blocks found (fenced) Template lines per page
Python tutorial WCC 43/43 237/242 131/131 (0 fenced) 0.0
Mine 43/43 239/242 131/131 (0 fenced) 6.6
Docusaurus WCC 72/92 258/273 60/62 (11 fenced) 0.0
Mine 37/40 76/76 18/18 (18 fenced) 2.9
Next.js WCC 7/68 0/82 0/33 6.8
Mine 66/68 82/82 33/33 (33 fenced) 13.9
Intercom Help WCC 82/103 106/309 – 3.3
Mine 67/71 176/181 2/2 (0 fenced) 5.8
Cloudflare Blog WCC 46/56 196/199 7/7 (7 fenced) 0.0
Mine 47/52 176/179 7/7 (7 fenced) 35.3

What stood out:

  • WCC is much better at removing page chrome. Its Readability-based extraction left zero repeated template lines on the blog and on Docusaurus. My crawler left 35 per page on the blog: language menus, the newsletter box, social links. That was a clear loss.
  • Next.js with WCC defaults went wrong. For 22 of the 26 pages, the output was only the docs version picker ("Latest 16.3.7 / Version 15 …"), about 115 characters. I ran it once and didn't dig into the cause, so treat it as one data point. With WCC's Cheerio crawler, the same site came back with full text, but only 4 of 23 code blocks survived.
  • On Intercom Help, WCC kept the headings but dropped much of what sat under them: lists, tables and FAQ answers. That brought it to 34% of sampled paragraphs.
  • Code on Sphinx sites: neither tool produced fenced code blocks for the Python docs. WCC also backslash-escaped the code (\>>> x \= 1), with 2,562 escapes in 17 pages.
  • Heading permalinks ([¶](#…) or a zero-width link next to every heading) showed up in both tools' output: 120–184 per site.
  • Neither tool returned duplicate pages. WCC returned a few more pages than the cap (26–36 for a cap of 25).

What I changed after seeing this

I fixed the categories where my crawler lost or where both tools did badly:

  1. Main-content extraction: if a page has one clear <article>, role="main" or <main> element that contains the H1 and at least 30% of the page's text, only that element is converted. If the result is almost empty, the whole page is converted instead.
  2. Code fences for bare <pre> blocks (Sphinx/Pygments), with the language taken from highlight-python3-style wrapper classes.
  3. Tables written without closing tags (<td>a<td>b, which HTML allows) got their implied closing tags back. Before that, Docusaurus tables were flattened into loose paragraphs.
  4. Absolute links, so chunks still work once they're cut out of the page.
  5. Heading permalinks, standalone widget labels ("Copy page", "Was this helpful?", "Scroll to top") and HTML entities in front matter are removed or decoded.
  6. Crawl order: links inside page content first, then menus, then the sitemap. Other languages, old versions and author/tag listing pages come last. Before this change, 24 of my 25 Docusaurus pages were archived 2.x docs, which the sitemap lists first, and 8 of 25 blog pages were translations of one post.

Re-run with the same settings:

Site Template lines per page (before → after) Code blocks fenced (before → after)* Permalinks
Python tutorial 6.6 → 0.0 0/131 → 131/131 137 → 0
Docusaurus 2.9 → 0.9 18/18 → 44/44 184 → 0 (tables 0/2 → 9/9)
Next.js 13.9 → 3.4 33/33 → 22/22 0 → 0
Intercom Help 5.8 → 0.6 0/2 → no code on sampled pages 0 → 0
Cloudflare Blog 35.3 → 8.6 7/7 → 7/7 0 → 0

* Sample pages differ between the two runs, because the new crawl order reaches different pages.

These "after" numbers were measured on the same five sites I tuned against, so they're optimistic. The honest comparison is the blind table above.

Where WCC is still the better choice

  • Blog-style pages: WCC still leaves less clutter. It had 0 template lines per page on the Cloudflare Blog, and mine has 8.6 (related tags, "follow us", the subscribe form).
  • Knobs and integrations: WCC gives you wait-for-selector, click-to-expand, scrolling, iframes, saved HTML, screenshots and files, and vector-database integrations. My crawler has far fewer options.
  • Track record: WCC had about 2.96 million runs in the last 30 days with a 96.6% success rate. Mine is new.
  • JavaScript-only sites: all five test sites serve server-rendered HTML. On a site that renders its content only in the browser, WCC's adaptive crawler may do better. My crawler does fall back to a real browser, but this benchmark didn't test that.

Known issues on my side

  • Occasionally one backend request stalls for 40–90 s before a retry succeeds.
  • In link cards (a heading inside an <a>), the title and description can run together without a space.
  • Blog footers inside the main element still get through.

Reproduce it

The method (sites, settings, how recall and clutter are counted) and the per-run numbers (time, cost, recall, clutter, size) are on the benchmark page, with the summary table as a CSV download.

My crawler costs $1.50 per 1,000 pages over HTTP and $2.50 per 1,000 when a page needs a real browser. Failed, blocked and duplicate pages are free. If you try it on your own docs and something looks wrong, tell me; the benchmark was the easy part.

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

Your note that 24 of 25 Docusaurus pages were archived docs explains why extraction recall and crawl selection need separate measurements. The changed crawl order improves what a user receives, but it also changes the denominators in the before/after table.

A useful follow-up would freeze a common URL manifest and saved HTML snapshots for the extraction comparison, then run a separate capped discovery test from the starting URL. That would show whether a gain comes from better conversion or from reaching more useful pages. Keeping the observed Next.js default failure as a reproducible fixture would also make the comparison easier to revisit after either tool changes.