DEV Community

neuralbyte
neuralbyte

Posted on

Extracting Text from HTML with Python: 4 Methods I Actually Use

TL;DR

  • Beautiful Soup is the clearest general-purpose method for extracting text from imperfect HTML.
  • lxml is a strong choice when XPath and high-throughput parsing matter.
  • Trafilatura is better when the goal is main article text rather than every visible navigation label.
  • Inscriptis is useful when text layout, tables, and lists need better preservation.
  • Remove scripts and style content, preserve block boundaries, normalize Unicode, and validate against expected text before storage.

Why I approached it this way

I stopped asking for the single best HTML-to-text library because the output contract changes the answer. DOM text, article text, XPath-selected text, and layout-aware text are different products. The four methods below are the ones I reach for in those four situations.

What is the best way to extract text from HTML with Python?

The best method depends on whether the input is a snippet, a complete document, or a rendered page. Regex alone is not a reliable HTML parser.

Text extraction should preserve enough structure for its destination: search indexing may need headings and paragraphs, while a compact classifier may need normalized plain text.

What do you need before you start?

Install the four libraries in an isolated environment. Use one shared fixture so the outputs can be compared fairly.

python -m venv .venv
source .venv/bin/activate
python -m pip install beautifulsoup4 lxml trafilatura inscriptis
Enter fullscreen mode Exit fullscreen mode
HTML = """<div><style>.x{display:none}</style><nav>Home</nav><main>
<h1>Parsing HTML</h1><p>Keep <strong>meaningful</strong> text.</p>
<script>ignore()</script></main></div>"""
Enter fullscreen mode Exit fullscreen mode

The practical workflow

Method 1: Use Beautiful Soup for readable general-purpose parsing

Step 1: Build the parse tree

from bs4 import BeautifulSoup
soup = BeautifulSoup(HTML, "html.parser")
Enter fullscreen mode Exit fullscreen mode

Step 2: Remove non-content nodes

for node in soup.select("script, style, noscript, template"):
    node.decompose()
Enter fullscreen mode Exit fullscreen mode

Step 3: Preserve boundaries and normalize whitespace

lines = [line.strip() for line in soup.get_text("\n").splitlines()]
text = "\n".join(line for line in lines if line)
print(text)
Enter fullscreen mode Exit fullscreen mode

Beautiful Soup tolerates malformed markup and has an approachable selector API. Its limitation is that get_text() does not know which visible blocks are primary content. Use semantic selectors such as main, article, or a verified content container when available. Check the official Beautiful Soup documentation for parser behavior.

Method 2: Use lxml for XPath and throughput

Step 1: Parse the document

from lxml import html
tree = html.fromstring(HTML)
Enter fullscreen mode Exit fullscreen mode

Step 2: Remove unwanted elements

for node in tree.xpath("//script|//style|//noscript|//template"):
    node.drop_tree()
Enter fullscreen mode Exit fullscreen mode

Step 3: Extract selected text nodes

parts = tree.xpath("//main//text()[normalize-space()]")
text = "\n".join(part.strip() for part in parts if part.strip())
Enter fullscreen mode Exit fullscreen mode

lxml is fast and gives precise XPath control, but careless XPath can flatten meaningful structure or collect hidden content. The official lxml.html documentation describes its HTML helpers.

Method 3: Use Trafilatura for main-content extraction

Step 1: Pass the complete document

import trafilatura
text = trafilatura.extract(HTML, include_links=False, include_images=False)
Enter fullscreen mode Exit fullscreen mode

Step 2: Handle a missing result

Trafilatura can return no main content for short, unusual, or navigation-heavy pages. Treat that as a validation failure and fall back to a scoped DOM method rather than storing an empty document.

Step 3: Keep metadata separately

Main-content extraction deliberately removes boilerplate. If title, canonical URL, author, or publication date matters, extract those fields into metadata instead of expecting them to survive plain-text conversion. Verify current options in the Trafilatura documentation.

Method 4: Use Inscriptis when layout carries meaning

Step 1: Convert HTML to layout-aware text

from inscriptis import get_text
text = get_text(HTML)
Enter fullscreen mode Exit fullscreen mode

Step 2: Review lists and tables

Inscriptis aims to preserve visible layout better than a simple descendant-text join. This can help with tables, lists, and documents where line breaks communicate structure.

Step 3: Normalize for the destination

Do not apply aggressive whitespace collapsing if table columns or list indentation matter. Create separate normalization profiles for search, LLM ingestion, and archival review.

How do you fetch HTML before extracting text?

Use a bounded HTTP client for static pages and validate status, content type, encoding, final URL, and expected page identity. If the page requires JavaScript, use a browser or managed rendering layer before parsing.

A parser cannot recover data that never arrived in its input.

How do you clean extracted text without damaging it?

Normalize Unicode, convert non-breaking spaces, remove repeated blank lines, and retain block boundaries. Avoid deleting all short lines because headings and labels can be short. Avoid lowercasing content intended for display or entity extraction. Keep the raw source or a permitted hash so a cleaning regression can be investigated.

Build golden fixtures containing headings, paragraphs, entities, lists, tables, hidden nodes, malformed tags, and non-ASCII text. Compare exact output or a structured block representation in automated tests. The WHATWG HTML standard is the primary reference when parser behavior depends on document structure.

What I would keep in production

I use Beautiful Soup for readable general parsing, lxml when XPath and throughput matter, Trafilatura for article bodies, and Inscriptis when layout carries meaning. I keep several fixtures and compare missing content, boilerplate, and structure instead of judging one attractive output by eye.

FAQ

Q: Can Python extract text from malformed HTML?

Yes. Beautiful Soup and lxml both recover many malformed documents, but their repairs may differ, so test the selected parser against representative fixtures.

Q: Why does get_text include JavaScript or CSS?

The parser treats script and style contents as text nodes unless those elements are removed. Decompose them before extracting text.

Q: Is regex suitable for removing HTML tags?

Regex can clean a known fragment, but it is not a reliable general HTML parser because nesting, malformed markup, entities, and embedded languages complicate the grammar.

Q: How do you extract only article text?

Use a verified article or content selector when the template is known, or a main-content library such as Trafilatura when templates vary.

Q: Which method is fastest?

lxml and other native-backed parsers are often faster than Beautiful Soup, but throughput should be measured with the real documents and required cleanup.

Top comments (0)