DEV Community

Cover image for My URL extractor scored the comments section as the article — text density alone can't find the main content
InApp
InApp

Posted on Originally published at imapp.blogspot.com

My URL extractor scored the comments section as the article — text density alone can't find the main content

The first version of my URL-to-Markdown converter used a simple heuristic: walk the DOM, score every text block by word count, penalize link density, pick the winner. It worked beautifully on the ten blogs I tested it on — which should have been a warning.

Then I pointed it at a long news article with an active comments section. The output started with "Sort by: newest" and ran for 400 replies. The article's 700 words sat in short paragraphs interleaved with ads; the comment thread was one contiguous column of 3,100 words in uniform blocks. My scorer did exactly what I asked: it found the densest low-link text on the page. It just wasn't the text anyone wanted.

Three fixes, in order of impact:

  1. Score containers, not blocks. Instead of judging each <p> alone, I group a node with all its siblings under the same parent and score the parent. Articles live in paragraphs; a single huge text block is usually a banner or a legal notice.

  2. Hard link-density ceiling. Nav, footers, and "related posts" widgets are mostly links. Anything above ~0.35 link density gets multiplied down hard — this alone killed most of the false positives.

  3. Semantic tie-breakers. When two candidates finish within 10% of each other, prefer the one under <article>, <main>, or [role=main], and prefer the one higher in the DOM.

I re-ran it against a hand-checked set of 200 pages — blogs, docs, news, forum threads. Main-content hit rate went from 74% to 93%. Forum threads are still the worst case: the replies genuinely are the main content there, and nothing cleanly separates a good thread from an article.

The lesson that stuck: the main content isn't the biggest text on the page — it's the text whose siblings look like it. Structure beats count.

I ended up packaging that scorer as my URL-to-Markdown API (https://x402.freeq.one/tools/markdown.html), and it's been most valuable for RAG ingestion, where one wrong sidebar silently poisons every chunk downstream of it.

Top comments (1)

Collapse
 
alexshev profile image
Alex Shev •

The container-level scoring point is a strong correction to paragraph-only heuristics. I would also treat repeated UI labels such as sort controls, reaction counts, and reply affordances as negative structural signals before the density score runs; they are often cheap to identify and help distinguish an article body from an otherwise well-formed discussion column.