The first version of my URL-to-Markdown converter used a simple heuristic: walk the DOM, score every text block by word count, penalize link density, pick the winner. It worked beautifully on the ten blogs I tested it on — which should have been a warning.
Then I pointed it at a long news article with an active comments section. The output started with "Sort by: newest" and ran for 400 replies. The article's 700 words sat in short paragraphs interleaved with ads; the comment thread was one contiguous column of 3,100 words in uniform blocks. My scorer did exactly what I asked: it found the densest low-link text on the page. It just wasn't the text anyone wanted.
Three fixes, in order of impact:
Score containers, not blocks. Instead of judging each
<p>alone, I group a node with all its siblings under the same parent and score the parent. Articles live in paragraphs; a single huge text block is usually a banner or a legal notice.Hard link-density ceiling. Nav, footers, and "related posts" widgets are mostly links. Anything above ~0.35 link density gets multiplied down hard — this alone killed most of the false positives.
Semantic tie-breakers. When two candidates finish within 10% of each other, prefer the one under
<article>,<main>, or[role=main], and prefer the one higher in the DOM.
I re-ran it against a hand-checked set of 200 pages — blogs, docs, news, forum threads. Main-content hit rate went from 74% to 93%. Forum threads are still the worst case: the replies genuinely are the main content there, and nothing cleanly separates a good thread from an article.
The lesson that stuck: the main content isn't the biggest text on the page — it's the text whose siblings look like it. Structure beats count.
I ended up packaging that scorer as my URL-to-Markdown API (https://x402.freeq.one/tools/markdown.html), and it's been most valuable for RAG ingestion, where one wrong sidebar silently poisons every chunk downstream of it.
Top comments (1)
The container-level scoring point is a strong correction to paragraph-only heuristics. I would also treat repeated UI labels such as sort controls, reaction counts, and reply affordances as negative structural signals before the density score runs; they are often cheap to identify and help distinguish an article body from an otherwise well-formed discussion column.