DEV Community

Cover image for Building a High-Throughput Article-to-Markdown API for LLM Ingestion with FastAPI and Playwright
Ruesch Manny
Ruesch Manny

Posted on Originally published at mannyverse767.gumroad.com

Building a High-Throughput Article-to-Markdown API for LLM Ingestion with FastAPI and Playwright

Feeding raw HTML into LLM context windows is one of the most expensive and inefficient mistakes in modern AI engineering.

A standard modern news or blog page easily spans 1.5MB to 4MB of raw DOM payload. When passed straight into an LLM or vector database, 90% of those tokens are spent on tracking scripts, serialized JSON-LD blobs, cookie banners, navigation menus, and inline CSS styles. This not only causes severe context bloat and escalates inference bills, but it also degrades retrieval-augmented generation (RAG) semantic search precision by polluting your vector space with boilerplate noise.

Here is how to design and build an enterprise-grade, self-hosted extraction microservice using FastAPI, Trafilatura, Readability, and an asynchronous Playwright fallback for SPA rendering.


1. The Bottleneck: Commercial APIs vs. Fragile Scrapers

Most teams start with simple libraries like BeautifulSoup or newspaper3k. These quickly break down:

  • Layout Brittleness: Custom CSS selectors degrade when publishers redesign their DOM or randomize class names (e.g., CSS Modules/Tailwind compiles).
  • Client-Side Rendering (CSR): Single-page applications built on React, Next.js, or Vue return empty <div id="root"></div> shells to standard HTTP clients.
  • Commercial SaaS Overkill: Hosted extraction APIs charge upwards of $0.002 to $0.01 per page. If your pipeline indexes 200,000 URLs monthly, you are paying hundreds of dollars for what essentially amounts to an HTTP request and an AST traversal.

To balance latency, compute cost, and reliability, the optimal architecture uses a tiered extraction waterfall:

  1. Tier 1 (Fast Path): High-speed async HTTP fetch + trafilatura. Latency: ~150-300ms.
  2. Tier 2 (Heuristic Fallback): If Trafilatura fails or yields empty strings, fall back to Mozilla's readability-lxml algorithm.
  3. Tier 3 (Headless Browser Fallback): If client-side hydration or dynamic script loading is detected, trigger an isolated headless Playwright worker to render the DOM before extraction.

2. System Architecture

[Incoming Request: URL]
          │
          ▼
┌───────────────────┐
│ Redis Cache Check │ ──(Hit)──► Return Markdown & Metadata
└───────────────────┘
          │ (Miss)
          ▼
┌───────────────────┐
│ HTTPX Async Fetch │
└───────────────────┘
          │
    [Static HTML]
          │
          ▼
┌───────────────────┐
│    Trafilatura    │ ──(Success: Length > Threshold)──► Parse Meta & Return
└───────────────────┘
          │ (Failed / Empty)
          ▼
┌───────────────────┐
│ Readability-lxml  │ ──(Success: Length > Threshold)──► Parse Meta & Return
└───────────────────┘
          │ (Failed / SPA Detected)
          ▼
┌───────────────────┐
│ Playwright Worker │ ──(Render DOM)──► Re-extract via Trafilatura
└───────────────────┘
          │
          ▼
[Store in Cache & Return Markdown]
Enter fullscreen mode Exit fullscreen mode

This setup allows 85% of standard web content to pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero.


3. Implementation: Code & Core Logic

Below is the complete implementation of the dual-engine pipeline using FastAPI, Pydantic v2, and async execution.

3.1 Pydantic Validation & Schemas

from pydantic import BaseModel, HttpUrl, Field
from typing import Optional, Dict, Any
from datetime import datetime

class ExtractionRequest(BaseModel):
    url: HttpUrl
    force_playwright: bool = Field(default=False, description="Bypass fast path and force browser rendering")
    max_chars: Optional[int] = Field(default=None, description="Truncate body content for token limits")

class PageMetadata(BaseModel):
    title: Optional[str] = None
    author: Optional[str] = None
    published_date: Optional[str] = None
    site_name: Optional[str] = None
    reading_time_minutes: int = 0
    opengraph: Dict[str, Any] = {}

class ExtractionResponse(BaseModel):
    url: str
    markdown: str
    metadata: PageMetadata
    engine_used: str
    execution_time_ms: float
Enter fullscreen mode Exit fullscreen mode

3.2 The Dual-Engine Core Service

import time
import httpx
import trafilatura
from readability import Document
from bs4 import BeautifulSoup
from markdownify import markdownify as md
from playwright.async_api import async_playwright

USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"

async def fetch_html_fast(url: str) -> str:
    async with httpx.AsyncClient(timeout=10.0, follow_redirects=True) as client:
        response = await client.get(url, headers={"User-Agent": USER_AGENT})
        response.raise_for_status()
        return response.text

async def fetch_html_playwright(url: str) -> str:
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True, args=["--no-sandbox", "--disable-dev-shm-usage"])
        context = await browser.new_context(user_agent=USER_AGENT)
        page = await context.new_page()
        await page.goto(url, wait_until="networkidle", timeout=20000)
        content = await page.content()
        await browser.close()
        return content

def extract_metadata(html: str) -> PageMetadata:
    soup = BeautifulSoup(html, "lxml")
    og_data = {}
    for tag in soup.find_all("meta"):
        prop = tag.get("property", tag.get("name", ""))
        if prop.startswith("og:") or prop.startswith("twitter:"):
            og_data[prop] = tag.get("content", "")

    title = og_data.get("og:title") or (soup.title.string if soup.title else None)
    author = og_data.get("article:author") or og_data.get("twitter:creator")
    pub_date = og_data.get("article:published_time")

    return PageMetadata(
        title=title,
        author=author,
        published_date=pub_date,
        site_name=og_data.get("og:site_name"),
        opengraph=og_data
    )

def extract_content(html: str, url: str) -> tuple[str, str]:
    # Primary Engine: Trafilatura
    extracted = trafilatura.extract(
        html,
        url=url,
        output_format="markdown",
        include_links=True,
        include_images=False,
        favor_recall=False
    )
    if extracted and len(extracted.strip()) > 200:
        return extracted, "trafilatura"

    # Fallback Engine: Readability + Markdownify
    doc = Document(html)
    summary_html = doc.summary()
    markdown_output = md(summary_html, heading_style="ATX").strip()

    if len(markdown_output) > 100:
        return markdown_output, "readability"

    return "", "none"
Enter fullscreen mode Exit fullscreen mode

3.3 The FastAPI Microservice Route

from fastapi import FastAPI, HTTPException, status

app = FastAPI(title="Article-to-Markdown Extraction API", version="1.0.0")

@app.post("/api/v1/extract", response_model=ExtractionResponse)
async def extract_article(payload: ExtractionRequest):
    start_time = time.perf_counter()
    url_str = str(payload.url)
    engine_used = "trafilatura"

    try:
        if payload.force_playwright:
            html = await fetch_html_playwright(url_str)
            markdown, engine_used = extract_content(html, url_str)
            engine_used = f"playwright+{engine_used}"
        else:
            # Tier 1 & 2: Fast HTTP path
            try:
                html = await fetch_html_fast(url_str)
                markdown, engine_used = extract_content(html, url_str)
            except Exception:
                markdown = ""

            # Tier 3: SPA / Fallback if empty
            if not markdown or len(markdown.strip()) < 150:
                html = await fetch_html_playwright(url_str)
                markdown, engine_used = extract_content(html, url_str)
                engine_used = f"playwright_fallback+{engine_used}"

        if not markdown:
            raise HTTPException(
                status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,
                detail="Failed to extract meaningful content from the target URL."
            )

        metadata = extract_metadata(html)
        words = len(markdown.split())
        metadata.reading_time_minutes = max(1, round(words / 200))

        if payload.max_chars and len(markdown) > payload.max_chars:
            markdown = markdown[:payload.max_chars] + "\n\n[Content Truncated]"

        exec_duration = (time.perf_counter() - start_time) * 1000

        return ExtractionResponse(
            url=url_str,
            markdown=markdown,
            metadata=metadata,
            engine_used=engine_used,
            execution_time_ms=round(exec_duration, 2)
        )

    except HTTPException:
        raise
    except Exception as e:
        raise HTTPException(status_code=status.HTTP_500_INTERNAL_SERVER_ERROR, detail=str(e))
Enter fullscreen mode Exit fullscreen mode

4. Deployment, Caching & Concurrency Hardening

When exposing this service in production pipelines (e.g., connecting n8n webhooks or feeding continuous Celery queues), consider these three safeguards:

Browser Resource Constraints

Launching a new Chromium instance via async_playwright on every request will instantly lead to Linux OOM crashes. In a production build:

  1. Maintain a single long-lived Browser instance.
  2. Create and destroy lightweight BrowserContext sessions per request.
  3. Enforce --disable-dev-shm-usage and --no-sandbox flags inside your Docker container.

Redis Caching Layer

Articles rarely mutate hourly. Adding an upstream Redis cache using a SHA-256 hash of the normalized canonical URL drastically drops response times to sub-10ms for repeated requests:

import hashlib

def generate_cache_key(url: str) -> str:
    normalized = url.split("?")[0].rstrip("/").lower()
    return f"extract:{hashlib.sha256(normalized.encode()).hexdigest()}"
Enter fullscreen mode Exit fullscreen mode

Docker Container Configuration

FROM python:3.11-slim

ENV PYTHONUNBUFFERED=1 \
    PIP_NO_CACHE_DIR=1 \
    PLAYWRIGHT_BROWSERS_PATH=/ms-playwright

WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
    libxml2-dev libxslt-dev gcc libc-dev curl \
    && rm -rf /var/lib/apt/lists/*

COPY requirements.txt .
RUN pip install -r requirements.txt
RUN playwright install --with-deps chromium

COPY . .
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]
Enter fullscreen mode Exit fullscreen mode

5. Conclusion & Ready-to-Use Workflow

You can manually wire this FastAPI microservice into your infrastructure following the architectural design and code above.

If you want the complete, production-ready implementation out of the box—including:

  • Pre-configured Docker Compose with Redis cache and persistent volume configurations
  • Playwright connection-pooling to avoid concurrency memory leaks
  • Rate-limiting middleware (slowapi) and health check probes
  • Complete test suites with synthetic fixtures and ready-to-import n8n orchestration workflows

You can grab the complete source repository directly:

Top comments (0)