TL;DR
- Beautiful Soup is the clearest general-purpose method for extracting text from imperfect HTML.
- lxml is a strong choice when XPath and high-throughput parsing matter.
- Trafilatura is better when the goal is main article text rather than every visible navigation label.
- Inscriptis is useful when text layout, tables, and lists need better preservation.
- Remove scripts and style content, preserve block boundaries, normalize Unicode, and validate against expected text before storage.
Why I approached it this way
I stopped asking for the single best HTML-to-text library because the output contract changes the answer. DOM text, article text, XPath-selected text, and layout-aware text are different products. The four methods below are the ones I reach for in those four situations.
What is the best way to extract text from HTML with Python?
The best method depends on whether the input is a snippet, a complete document, or a rendered page. Regex alone is not a reliable HTML parser.
Text extraction should preserve enough structure for its destination: search indexing may need headings and paragraphs, while a compact classifier may need normalized plain text.
What do you need before you start?
Install the four libraries in an isolated environment. Use one shared fixture so the outputs can be compared fairly.
python -m venv .venv
source .venv/bin/activate
python -m pip install beautifulsoup4 lxml trafilatura inscriptis
HTML = """<div><style>.x{display:none}</style><nav>Home</nav><main>
<h1>Parsing HTML</h1><p>Keep <strong>meaningful</strong> text.</p>
<script>ignore()</script></main></div>"""
The practical workflow
Method 1: Use Beautiful Soup for readable general-purpose parsing
Step 1: Build the parse tree
from bs4 import BeautifulSoup
soup = BeautifulSoup(HTML, "html.parser")
Step 2: Remove non-content nodes
for node in soup.select("script, style, noscript, template"):
node.decompose()
Step 3: Preserve boundaries and normalize whitespace
lines = [line.strip() for line in soup.get_text("\n").splitlines()]
text = "\n".join(line for line in lines if line)
print(text)
Beautiful Soup tolerates malformed markup and has an approachable selector API. Its limitation is that get_text() does not know which visible blocks are primary content. Use semantic selectors such as main, article, or a verified content container when available. Check the official Beautiful Soup documentation for parser behavior.
Method 2: Use lxml for XPath and throughput
Step 1: Parse the document
from lxml import html
tree = html.fromstring(HTML)
Step 2: Remove unwanted elements
for node in tree.xpath("//script|//style|//noscript|//template"):
node.drop_tree()
Step 3: Extract selected text nodes
parts = tree.xpath("//main//text()[normalize-space()]")
text = "\n".join(part.strip() for part in parts if part.strip())
lxml is fast and gives precise XPath control, but careless XPath can flatten meaningful structure or collect hidden content. The official lxml.html documentation describes its HTML helpers.
Method 3: Use Trafilatura for main-content extraction
Step 1: Pass the complete document
import trafilatura
text = trafilatura.extract(HTML, include_links=False, include_images=False)
Step 2: Handle a missing result
Trafilatura can return no main content for short, unusual, or navigation-heavy pages. Treat that as a validation failure and fall back to a scoped DOM method rather than storing an empty document.
Step 3: Keep metadata separately
Main-content extraction deliberately removes boilerplate. If title, canonical URL, author, or publication date matters, extract those fields into metadata instead of expecting them to survive plain-text conversion. Verify current options in the Trafilatura documentation.
Method 4: Use Inscriptis when layout carries meaning
Step 1: Convert HTML to layout-aware text
from inscriptis import get_text
text = get_text(HTML)
Step 2: Review lists and tables
Inscriptis aims to preserve visible layout better than a simple descendant-text join. This can help with tables, lists, and documents where line breaks communicate structure.
Step 3: Normalize for the destination
Do not apply aggressive whitespace collapsing if table columns or list indentation matter. Create separate normalization profiles for search, LLM ingestion, and archival review.
How do you fetch HTML before extracting text?
Use a bounded HTTP client for static pages and validate status, content type, encoding, final URL, and expected page identity. If the page requires JavaScript, use a browser or managed rendering layer before parsing.
A parser cannot recover data that never arrived in its input.
How do you clean extracted text without damaging it?
Normalize Unicode, convert non-breaking spaces, remove repeated blank lines, and retain block boundaries. Avoid deleting all short lines because headings and labels can be short. Avoid lowercasing content intended for display or entity extraction. Keep the raw source or a permitted hash so a cleaning regression can be investigated.
Build golden fixtures containing headings, paragraphs, entities, lists, tables, hidden nodes, malformed tags, and non-ASCII text. Compare exact output or a structured block representation in automated tests. The WHATWG HTML standard is the primary reference when parser behavior depends on document structure.
What I would keep in production
I use Beautiful Soup for readable general parsing, lxml when XPath and throughput matter, Trafilatura for article bodies, and Inscriptis when layout carries meaning. I keep several fixtures and compare missing content, boilerplate, and structure instead of judging one attractive output by eye.
FAQ
Q: Can Python extract text from malformed HTML?
Yes. Beautiful Soup and lxml both recover many malformed documents, but their repairs may differ, so test the selected parser against representative fixtures.
Q: Why does get_text include JavaScript or CSS?
The parser treats script and style contents as text nodes unless those elements are removed. Decompose them before extracting text.
Q: Is regex suitable for removing HTML tags?
Regex can clean a known fragment, but it is not a reliable general HTML parser because nesting, malformed markup, entities, and embedded languages complicate the grammar.
Q: How do you extract only article text?
Use a verified article or content selector when the template is known, or a main-content library such as Trafilatura when templates vary.
Q: Which method is fastest?
lxml and other native-backed parsers are often faster than Beautiful Soup, but throughput should be measured with the real documents and required cleanup.
Top comments (0)