DEV Community

neuralbyte
neuralbyte

Posted on

Scraping JavaScript-Rendered Pages with Python: What I Try First

TL;DR

  • Use an authorized JSON endpoint when possible; it is usually more stable and cheaper than rendering a browser.
  • Playwright is the primary Python method in this guide because its locators, contexts, and network controls fit modern client-rendered pages.
  • Selenium remains useful in organizations with an established WebDriver stack.
  • Scrapy plus scrapy-playwright fits queued crawls where only selected requests need rendering.
  • A managed rendering API is appropriate when browser operations, retries, and artifacts should be service responsibilities.

Why I approached it this way

When requests.get() returns an almost empty document, I compare three views: the original response, the rendered DOM, and Fetch/XHR traffic. That comparison usually tells me whether I need a browser at all.

Why do JavaScript-rendered pages return empty content in Python?

JavaScript-rendered pages return empty or partial content because requests and similar clients download the initial response but do not execute the application. The first diagnostic step is comparing View Source, the rendered DOM, and Fetch/XHR responses.

A skeleton page can return status 200 while the data request fails later.

What do you need before scraping a rendered page?

You need permission, a precise schema, a small target corpus, explicit time and page budgets, and one or more meaningful readiness signals. Install the selected tools in an isolated environment.

python -m pip install playwright selenium scrapy scrapy-playwright beautifulsoup4
python -m playwright install chromium
Enter fullscreen mode Exit fullscreen mode

Use the official Playwright Python guide, Selenium WebDriver documentation, and scrapy-playwright repository for current APIs.

The practical workflow

Method 1: Capture the underlying JSON request

Step 1: Inspect Fetch/XHR traffic

On a site you own or are authorized to inspect, open DevTools → Network → Fetch/XHR, reload, and trigger the content. Identify a public request that returns the required records.

Step 2: Reproduce one request

import httpx
r = httpx.get("https://example.com/api/items", params={"page": 1}, timeout=20)
r.raise_for_status()
items = r.json().get("items", [])
Enter fullscreen mode Exit fullscreen mode

Step 3: Bound pagination

Stop on a documented cursor, empty page, or configured maximum. Do not replay private application calls or copy session credentials without approval.

Method 2: Render with Playwright

Step 1: Start a browser context

from playwright.sync_api import sync_playwright
with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1365, "height": 900})
    page.goto("https://example.com/app", wait_until="domcontentloaded")
    page.locator("[data-testid='item']").first.wait_for(timeout=10000)
    rows = page.locator("[data-testid='item']").all_inner_texts()
    browser.close()
Enter fullscreen mode Exit fullscreen mode

Step 2: Wait for evidence, not time

Prefer a stable selector or the specific response that delivers data. networkidle can be unsuitable for applications with analytics or long-lived connections.

Step 3: Validate the page state

Check final URL, title, expected controls, row count, and a screenshot on failure. Reject generic errors and consent-only states.

Method 3: Use Selenium with explicit waits

Step 1: Navigate with WebDriver

Create the driver, set a page-load timeout, and open one approved page. Avoid implicit waits mixed with explicit waits because combined timing is difficult to reason about.

Step 2: Wait for the result container

Use WebDriverWait and an expected condition for the actual content. Avoid time.sleep() as the primary readiness mechanism.

Step 3: Always close the driver

Use try/finally or a context wrapper so failed jobs do not leak browser processes.

Method 4: Add rendering to Scrapy selectively

Step 1: Enable scrapy-playwright

Configure the download handler and reactor exactly as documented for the installed version. Pin Scrapy, Playwright, and plugin versions together.

Step 2: Mark only rendered requests

Send static pages through Scrapy's normal downloader and add meta={"playwright": True} only where rendering is required. This preserves concurrency and reduces browser cost.

Step 3: Close page objects

If a callback receives a Playwright page, close it on both success and exception paths. Monitor open pages and browser memory.

Method 5: Use a managed rendering service

Step 1: Define a minimal request

Request only necessary formats and interactions.

Step 2: Poll bounded asynchronous work

Use terminal states, a maximum duration, and exponential backoff with jitter. Inspect body-level success, status, and error information rather than the transport code alone.

Step 3: Validate cost per accepted page

Include rejected and retried pages when comparing cost with self-hosted browsers.

How do you scrape infinite scroll and lazy content?

Prefer the underlying cursor or page endpoint when permitted. If scrolling is required, record the current item count, scroll once, wait for an increase, and stop when the count is unchanged or a maximum loop count is reached. Lazy images may require viewport intersection, but scrolling the entire page can accidentally trigger unbounded feeds.

Keep URL discovery and item acceptance separate so repeated cards do not become duplicate records.

What errors should a rendered-page scraper handle?

Handle navigation timeout, selector timeout, browser crash, access denied, wrong locale, empty result, parse failure, and storage failure as distinct states. Retry only transient classes. Capture sanitized diagnostics, not cookies or tokens. The Robots Exclusion Protocol is one policy input; it does not replace legal and contractual review.

What I would keep in production

I prefer the underlying authorized request when it is stable, Playwright for modern interaction-heavy pages, Selenium for existing WebDriver environments, and scrapy-playwright when only selected crawl requests need rendering. Fixed sleeps are the one method I avoid in every version.

FAQ

Q: Can Beautiful Soup scrape JavaScript-rendered content?

Beautiful Soup can parse rendered HTML after another tool executes JavaScript, but it does not run JavaScript itself.

Q: Should Playwright wait for networkidle?

Only when network quiet is a meaningful condition for the target. A stable selector or known data response is often more precise.

Q: Can Scrapy render JavaScript?

Scrapy does not render JavaScript by itself, but integrations such as scrapy-playwright can route selected requests through a browser.

Q: Why does headless mode return different content?

Viewport, locale, browser version, timing, permissions, and target behavior can differ. Compare configurations and retain a screenshot before assuming the selector is wrong.

Q: How many browser pages should run concurrently?

Concurrency depends on memory, CPU, target policy, and page weight. Start low, measure browser resource use and accepted-page rate, and set a hard cap.

Top comments (0)