DEV Community

neuralbyte
neuralbyte

Posted on

Building a Zalando Scraper with Python: The Approach I’d Use

TL;DR

  • Build a Zalando scraper only for public pages you are permitted to collect, and check the site's current terms and robots policy before running it.
  • Start with one product URL and extract embedded structured data before adding browser automation.
  • Use Playwright only when the required fields appear after JavaScript execution; plain HTTP is faster when the initial response already contains them.
  • Store stable product identifiers, canonical URLs, currency, availability, retrieval time, and a content hash so updates are idempotent.
  • Treat consent pages, locale redirects, missing variants, and access-denied responses as explicit terminal states.

Why I approached it this way

The tempting version of this project starts with CSS selectors and a large product list. I prefer to start with one permitted product URL, look for JSON-LD, and write an acceptance test before adding pagination. That order makes markup changes much easier to diagnose.

What is a Zalando scraper and why would you need one?

A Zalando scraper is a controlled program that reads permitted public product pages and converts selected fields into a consistent record. Appropriate uses include internal QA, approved catalog research, and monitoring products your organization is authorized to observe.

The goal is not to copy the entire storefront. A reliable job starts from an approved URL inventory and a narrow schema.

What do you need before you start?

You need Python 3.11 or newer, an isolated environment, httpx, beautifulsoup4, and Playwright for pages that truly require a browser. You also need written permission or a documented lawful basis, a small test URL set, a locale and currency policy, and a destination schema.

python -m venv .venv
source .venv/bin/activate
python -m pip install httpx beautifulsoup4 playwright
python -m playwright install chromium
Enter fullscreen mode Exit fullscreen mode

Define these fields before writing selectors: product_id, canonical_url, name, brand, currency, price_text, availability, color, size_options, retrieved_at, and source_hash. Do not collect customer, account, or checkout data. Read the current Zalando terms and the site's robots file for the exact host and locale you plan to access.

How does a Zalando product page actually work?

A product page can combine server-rendered HTML, embedded JSON, and client-side requests. The visible DOM is not always the best source. Inspect the initial response and <script type="application/ld+json"> blocks first because structured product data is usually more stable than presentation classes. If required data appears only after rendering, use a browser and wait for a meaningful product element rather than sleeping for an arbitrary duration.

The practical workflow

Method 1: Extract JSON-LD from a permitted product page

Step 1: Fetch one URL with a bounded client

Use a descriptive user agent, a timeout, redirect limits, and a low request rate. The target URL must come from your approved inventory.

import os
import httpx

url = os.environ["TARGET_URL"]
with httpx.Client(timeout=20, follow_redirects=True) as client:
    response = client.get(url, headers={"User-Agent": "CatalogQA/1.0 contact@example.com"})
    response.raise_for_status()
    html = response.text
Enter fullscreen mode Exit fullscreen mode

Step 2: Parse product JSON-LD

import json
from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
products = []
for node in soup.select('script[type="application/ld+json"]'):
    try:
        value = json.loads(node.string or "")
    except json.JSONDecodeError:
        continue
    values = value if isinstance(value, list) else [value]
    products.extend(x for x in values if isinstance(x, dict) and x.get("@type") == "Product")

if not products:
    raise RuntimeError("No Product JSON-LD found; inspect the response before changing selectors")
Enter fullscreen mode Exit fullscreen mode

Step 3: Normalize without inventing values

Map only present fields. Preserve the original price string and currency, and use None for unavailable data. Never infer a size or stock state from an absent button.

Method 2: Render the page with Playwright

Step 1: Launch a bounded browser context

import os
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(locale="en-GB", viewport={"width": 1440, "height": 1000})
    page.goto(os.environ["TARGET_URL"], wait_until="domcontentloaded", timeout=30000)
    page.locator('script[type="application/ld+json"]').first.wait_for(timeout=10000)
    rendered_html = page.content()
    browser.close()
Enter fullscreen mode Exit fullscreen mode

Step 2: Detect the wrong page state

Check the final URL, page title, expected product identifier, locale, and presence of product evidence. Reject consent-only, login, challenge, and generic error pages. The official Playwright Python documentation is the source for current installation and waiting APIs.

Step 3: Save a review artifact

Save a screenshot or HTML hash for failures, but apply retention limits. Review artifacts should help diagnose a selector or rendering change without becoming an uncontrolled copy of the page.

Method 3: Use a managed crawl API for acquisition

Step 1: Keep acquisition separate from parsing

Submit only approved public URLs, request the smallest set of formats, and inspect body-level success rather than assuming a successful HTTP response means the product loaded.

Step 2: Validate the returned artifact

Confirm canonical URL, product evidence, locale, currency, and content completeness before parsing.

Step 3: Feed the same normalization layer

The HTTP, Playwright, and managed methods should all produce the same internal source record. Provider-specific fields belong in the acquisition adapter, not in downstream catalog tables.

Why is the Zalando scraper returning empty or inconsistent data?

Empty data usually means the response is a different page state, the field moved into embedded JSON, the locale redirected, or the content requires rendering. Compare response.url, status, title, body length, and a saved hash. If Playwright finds the field but HTTP does not, inspect authorized network requests before adding more browser interactions.

Inconsistent prices often come from locale, currency, promotion, membership, or variant context. Store those dimensions with every observation.

How do you make the scraper safe and maintainable?

Use a bounded queue, low concurrency, retry only transient failures, and stop on persistent denial. Version the parser, keep fixtures for each supported page state, and alert on missing-field rate rather than silently writing nulls. The Robots Exclusion Protocol describes standardized robots rules, but compliance also requires terms, privacy, copyright, and internal review.

What I would keep in production

I would ship the smallest version first: one permitted URL, one stable product identifier, a validated schema, and an idempotent write. Browser rendering should be an escalation path rather than the default, and every expansion in scope should have a clear stop condition.

FAQ

Q: Is it legal to scrape Zalando?

Legality depends on the target, method, jurisdiction, terms, and data use. Collect only public or otherwise authorized pages and obtain legal guidance for the specific project.

Q: Should a Zalando scraper use Beautiful Soup or Playwright?

Use Beautiful Soup when the initial response or embedded JSON contains the required data; use Playwright only when JavaScript rendering is necessary. This reduces cost and failure surface.

Q: Why does the scraper see a different language or price?

Locale redirects, currency settings, geography, promotions, and variant context can change the page. Store those dimensions and reject records that do not match the requested market.

Q: How often should product pages be refreshed?

Refresh frequency should follow business need, page volatility, permission, and site load. Use change history to slow stable products and prioritize volatile ones.

Q: How do you prevent duplicate products?

Use a stable product identifier plus market and variant dimensions, canonicalize URLs conservatively, and make writes idempotent with content hashes.

Top comments (0)