DEV Community

neuralbyte
neuralbyte

Posted on

Scraping Dynamic Websites with Python: My Practical Decision Tree

TL;DR

  • The most efficient way to scrape a dynamic website with Python is to call its authorized JSON/XHR endpoint when one exists.
  • Use Playwright when the data depends on browser rendering or interaction; use Selenium when its ecosystem or an existing test stack makes it the better fit.
  • Wait for a meaningful selector or response, not a fixed sleep.
  • Separate retrieval, validation, parsing, and storage so a rendered error page cannot become accepted data.
  • Bound pagination, scrolling, retries, and concurrency before running beyond a small test set.

Why I approached it this way

My first question is not “Which browser library should I install?” It is “Where does the page get its data?” If an authorized JSON response already contains the records, replaying that bounded request is usually simpler than running Chromium.

What does it mean to scrape dynamic websites with Python?

Scraping dynamic websites with Python means collecting content that appears after client-side JavaScript runs or after the page makes additional network requests. A page is dynamic when the initial HTML does not contain the data the user sees after loading, scrolling, selecting a filter, or clicking a control.

Before opening a browser, compare View Source with the rendered DOM and inspect the browser Network panel.

What do you need before you start?

Install Python, httpx, Playwright, Beautiful Soup, and optionally Selenium. Define the fields, permitted domains, pagination boundary, request budget, timeout, and terminal failure states. Use a local fixture or a site you own for development.

python -m venv .venv
source .venv/bin/activate
python -m pip install httpx beautifulsoup4 playwright selenium
python -m playwright install chromium
Enter fullscreen mode Exit fullscreen mode

The Playwright Python documentation and Selenium WebDriver documentation should be checked for current installation and wait APIs.

How do dynamic pages load data?

Dynamic pages commonly fetch JSON through XHR or Fetch, hydrate server-rendered HTML, or build components entirely in the browser. Infinite scroll is usually paginated data hidden behind an event. The correct method depends on where the authoritative content exists, not on which library is most popular.

The practical workflow

Method 1: Reuse an authorized JSON endpoint

Step 1: Identify the request

Open DevTools → Network → Fetch/XHR on a site you own or are authorized to inspect. Trigger the action that loads the data, then record the request URL, method, public parameters, response shape, and pagination token. Do not copy private tokens or replay undocumented account endpoints without permission.

Step 2: Reproduce one bounded request

import httpx

url = "https://example.com/api/products"
params = {"page": 1, "limit": 20}
response = httpx.get(url, params=params, timeout=20)
response.raise_for_status()
payload = response.json()
items = payload.get("items", [])
Enter fullscreen mode Exit fullscreen mode

Step 3: Validate and paginate

Validate types and required identifiers before storage. Stop at an explicit next token or a configured page limit; never continue until an endpoint returns an error.

Method 2: Render with Playwright

Step 1: Open a controlled browser

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1440, "height": 900})
    page.goto("https://example.com/dynamic", wait_until="domcontentloaded")
    page.locator("[data-testid='result-card']").first.wait_for(timeout=10000)
    cards = page.locator("[data-testid='result-card']").all_inner_texts()
    browser.close()
Enter fullscreen mode Exit fullscreen mode

Step 2: Replace sleeps with evidence

Wait for an element, response, or application state that proves the data loaded. A fixed delay is simultaneously too slow on fast pages and too short on slow pages.

Step 3: Handle scroll or clicks with a cap

Record the item count before and after each action. Stop when the count no longer changes, the next control disappears, or a maximum iteration count is reached.

Method 3: Use Selenium for an existing WebDriver stack

Step 1: Create the driver

Use Selenium when an organization already maintains WebDriver tests, browser profiles, or Grid infrastructure. Keep dependencies pinned and use headless mode only after the flow works visibly in development.

Step 2: Use explicit waits

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("https://example.com/dynamic")
cards = WebDriverWait(driver, 10).until(
    EC.presence_of_all_elements_located((By.CSS_SELECTOR, "[data-testid='result-card']"))
)
texts = [card.text for card in cards]
driver.quit()
Enter fullscreen mode Exit fullscreen mode

Step 3: Capture failure evidence

On failure, save the final URL, title, screenshot, and a sanitized HTML excerpt. Never log cookies, authorization headers, or personal data.

Method 4: Use a managed rendering API

Step 1: Define the output contract

Request only required artifacts and verify current browser actions in the official documentation.

Step 2: Inspect body-level task status

A successful submission is not proof that the target page loaded. Check task state, page status, expected title, content signals, and error fields.

Step 3: Parse through the same adapter

Normalize managed output and self-hosted browser output into the same source schema. This keeps downstream validation independent of the retrieval provider.

Why is the dynamic scraper returning empty HTML?

The initial response may contain only an application shell, the selector may run before hydration, the request may have reached a consent or error page, or the content may be inside an iframe or shadow DOM. Inspect the Network panel, final URL, title, response size, and a screenshot before changing selectors. Avoid retrying persistent authorization failures.

How do you run dynamic scraping in production?

Use a queue with bounded concurrency, per-domain limits, idempotent document IDs, structured errors, and separate retry policies for timeouts, rate limits, and parse failures. Monitor accepted-page rate, render duration, browser memory, selector-missing rate, and duplicate content.

The Robots Exclusion Protocol standard is one technical signal, not a complete permission decision. Keep the collection public or otherwise authorized, and do not bypass authentication, paywalls, or access controls.

What I would keep in production

My order of preference is an authorized data endpoint, then Playwright, then Selenium when an existing WebDriver stack makes it practical. Whichever route I take, I wait for evidence, cap every loop, and reject pages that do not match the expected identity.

FAQ

Q: Can Requests scrape a dynamic website?

Requests can scrape a dynamic website when the required data is present in the initial HTML or an authorized JSON endpoint. It cannot execute the page's JavaScript.

Q: Is Playwright better than Selenium for scraping?

Playwright often provides convenient modern browser waits and contexts, while Selenium fits mature WebDriver ecosystems. The better choice depends on team tooling and target behavior.

Q: How do you scrape infinite scroll safely?

Prefer the underlying paginated endpoint when permitted; otherwise scroll with item-count checks, a maximum iteration limit, and a clear stop condition.

Q: Why should fixed sleeps be avoided?

Fixed sleeps do not prove content readiness and make jobs slow or flaky. Wait for a specific element, response, or state instead.

Q: How do you prevent duplicate dynamic-page records?

Use canonical URLs or stable entity IDs, content hashes, pagination cursors, and idempotent writes.

Top comments (0)