Feeding raw HTML into LLM context windows is one of the most expensive and inefficient mistakes in modern AI engineering.
A standard modern news or blog page easily spans 1.5MB to 4MB of raw DOM payload. When passed straight into an LLM or vector database, 90% of those tokens are spent on tracking scripts, serialized JSON-LD blobs, cookie banners, navigation menus, and inline CSS styles. This not only causes severe context bloat and escalates inference bills, but it also degrades retrieval-augmented generation (RAG) semantic search precision by polluting your vector space with boilerplate noise.
Here is how to design and build an enterprise-grade, self-hosted extraction microservice using FastAPI, Trafilatura, Readability, and an asynchronous Playwright fallback for SPA rendering.
1. The Bottleneck: Commercial APIs vs. Fragile Scrapers
Most teams start with simple libraries like BeautifulSoup or newspaper3k. These quickly break down:
- Layout Brittleness: Custom CSS selectors degrade when publishers redesign their DOM or randomize class names (e.g., CSS Modules/Tailwind compiles).
-
Client-Side Rendering (CSR): Single-page applications built on React, Next.js, or Vue return empty
<div id="root"></div>shells to standard HTTP clients. - Commercial SaaS Overkill: Hosted extraction APIs charge upwards of $0.002 to $0.01 per page. If your pipeline indexes 200,000 URLs monthly, you are paying hundreds of dollars for what essentially amounts to an HTTP request and an AST traversal.
To balance latency, compute cost, and reliability, the optimal architecture uses a tiered extraction waterfall:
-
Tier 1 (Fast Path): High-speed async HTTP fetch +
trafilatura. Latency: ~150-300ms. -
Tier 2 (Heuristic Fallback): If Trafilatura fails or yields empty strings, fall back to Mozilla's
readability-lxmlalgorithm. - Tier 3 (Headless Browser Fallback): If client-side hydration or dynamic script loading is detected, trigger an isolated headless Playwright worker to render the DOM before extraction.
2. System Architecture
[Incoming Request: URL]
│
▼
┌───────────────────┐
│ Redis Cache Check │ ──(Hit)──► Return Markdown & Metadata
└───────────────────┘
│ (Miss)
▼
┌───────────────────┐
│ HTTPX Async Fetch │
└───────────────────┘
│
[Static HTML]
│
▼
┌───────────────────┐
│ Trafilatura │ ──(Success: Length > Threshold)──► Parse Meta & Return
└───────────────────┘
│ (Failed / Empty)
▼
┌───────────────────┐
│ Readability-lxml │ ──(Success: Length > Threshold)──► Parse Meta & Return
└───────────────────┘
│ (Failed / SPA Detected)
▼
┌───────────────────┐
│ Playwright Worker │ ──(Render DOM)──► Re-extract via Trafilatura
└───────────────────┘
│
▼
[Store in Cache & Return Markdown]
This setup allows 85% of standard web content to pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero.
3. Implementation: Code & Core Logic
Below is the complete implementation of the dual-engine pipeline using FastAPI, Pydantic v2, and async execution.
3.1 Pydantic Validation & Schemas
from pydantic import BaseModel, HttpUrl, Field
from typing import Optional, Dict, Any
from datetime import datetime
class ExtractionRequest(BaseModel):
url: HttpUrl
force_playwright: bool = Field(default=False, description="Bypass fast path and force browser rendering")
max_chars: Optional[int] = Field(default=None, description="Truncate body content for token limits")
class PageMetadata(BaseModel):
title: Optional[str] = None
author: Optional[str] = None
published_date: Optional[str] = None
site_name: Optional[str] = None
reading_time_minutes: int = 0
opengraph: Dict[str, Any] = {}
class ExtractionResponse(BaseModel):
url: str
markdown: str
metadata: PageMetadata
engine_used: str
execution_time_ms: float
3.2 The Dual-Engine Core Service
import time
import httpx
import trafilatura
from readability import Document
from bs4 import BeautifulSoup
from markdownify import markdownify as md
from playwright.async_api import async_playwright
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"
async def fetch_html_fast(url: str) -> str:
async with httpx.AsyncClient(timeout=10.0, follow_redirects=True) as client:
response = await client.get(url, headers={"User-Agent": USER_AGENT})
response.raise_for_status()
return response.text
async def fetch_html_playwright(url: str) -> str:
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True, args=["--no-sandbox", "--disable-dev-shm-usage"])
context = await browser.new_context(user_agent=USER_AGENT)
page = await context.new_page()
await page.goto(url, wait_until="networkidle", timeout=20000)
content = await page.content()
await browser.close()
return content
def extract_metadata(html: str) -> PageMetadata:
soup = BeautifulSoup(html, "lxml")
og_data = {}
for tag in soup.find_all("meta"):
prop = tag.get("property", tag.get("name", ""))
if prop.startswith("og:") or prop.startswith("twitter:"):
og_data[prop] = tag.get("content", "")
title = og_data.get("og:title") or (soup.title.string if soup.title else None)
author = og_data.get("article:author") or og_data.get("twitter:creator")
pub_date = og_data.get("article:published_time")
return PageMetadata(
title=title,
author=author,
published_date=pub_date,
site_name=og_data.get("og:site_name"),
opengraph=og_data
)
def extract_content(html: str, url: str) -> tuple[str, str]:
# Primary Engine: Trafilatura
extracted = trafilatura.extract(
html,
url=url,
output_format="markdown",
include_links=True,
include_images=False,
favor_recall=False
)
if extracted and len(extracted.strip()) > 200:
return extracted, "trafilatura"
# Fallback Engine: Readability + Markdownify
doc = Document(html)
summary_html = doc.summary()
markdown_output = md(summary_html, heading_style="ATX").strip()
if len(markdown_output) > 100:
return markdown_output, "readability"
return "", "none"
3.3 The FastAPI Microservice Route
from fastapi import FastAPI, HTTPException, status
app = FastAPI(title="Article-to-Markdown Extraction API", version="1.0.0")
@app.post("/api/v1/extract", response_model=ExtractionResponse)
async def extract_article(payload: ExtractionRequest):
start_time = time.perf_counter()
url_str = str(payload.url)
engine_used = "trafilatura"
try:
if payload.force_playwright:
html = await fetch_html_playwright(url_str)
markdown, engine_used = extract_content(html, url_str)
engine_used = f"playwright+{engine_used}"
else:
# Tier 1 & 2: Fast HTTP path
try:
html = await fetch_html_fast(url_str)
markdown, engine_used = extract_content(html, url_str)
except Exception:
markdown = ""
# Tier 3: SPA / Fallback if empty
if not markdown or len(markdown.strip()) < 150:
html = await fetch_html_playwright(url_str)
markdown, engine_used = extract_content(html, url_str)
engine_used = f"playwright_fallback+{engine_used}"
if not markdown:
raise HTTPException(
status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,
detail="Failed to extract meaningful content from the target URL."
)
metadata = extract_metadata(html)
words = len(markdown.split())
metadata.reading_time_minutes = max(1, round(words / 200))
if payload.max_chars and len(markdown) > payload.max_chars:
markdown = markdown[:payload.max_chars] + "\n\n[Content Truncated]"
exec_duration = (time.perf_counter() - start_time) * 1000
return ExtractionResponse(
url=url_str,
markdown=markdown,
metadata=metadata,
engine_used=engine_used,
execution_time_ms=round(exec_duration, 2)
)
except HTTPException:
raise
except Exception as e:
raise HTTPException(status_code=status.HTTP_500_INTERNAL_SERVER_ERROR, detail=str(e))
4. Deployment, Caching & Concurrency Hardening
When exposing this service in production pipelines (e.g., connecting n8n webhooks or feeding continuous Celery queues), consider these three safeguards:
Browser Resource Constraints
Launching a new Chromium instance via async_playwright on every request will instantly lead to Linux OOM crashes. In a production build:
- Maintain a single long-lived Browser instance.
- Create and destroy lightweight
BrowserContextsessions per request. - Enforce
--disable-dev-shm-usageand--no-sandboxflags inside your Docker container.
Redis Caching Layer
Articles rarely mutate hourly. Adding an upstream Redis cache using a SHA-256 hash of the normalized canonical URL drastically drops response times to sub-10ms for repeated requests:
import hashlib
def generate_cache_key(url: str) -> str:
normalized = url.split("?")[0].rstrip("/").lower()
return f"extract:{hashlib.sha256(normalized.encode()).hexdigest()}"
Docker Container Configuration
FROM python:3.11-slim
ENV PYTHONUNBUFFERED=1 \
PIP_NO_CACHE_DIR=1 \
PLAYWRIGHT_BROWSERS_PATH=/ms-playwright
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
libxml2-dev libxslt-dev gcc libc-dev curl \
&& rm -rf /var/lib/apt/lists/*
COPY requirements.txt .
RUN pip install -r requirements.txt
RUN playwright install --with-deps chromium
COPY . .
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]
5. Conclusion & Ready-to-Use Workflow
You can manually wire this FastAPI microservice into your infrastructure following the architectural design and code above.
If you want the complete, production-ready implementation out of the box—including:
- Pre-configured Docker Compose with Redis cache and persistent volume configurations
- Playwright connection-pooling to avoid concurrency memory leaks
- Rate-limiting middleware (slowapi) and health check probes
- Complete test suites with synthetic fixtures and ready-to-import n8n orchestration workflows
You can grab the complete source repository directly:
- Instant Access on Whop: Get the Microservice on Whop
-
Direct Download on Gumroad: Download on Gumroad — Use promo code
EARLYBIRDat checkout for 20% off.
Top comments (0)