DEV Community

Žygimantas
Žygimantas

Posted on Originally published at stealthasf.com

How to scrape DataDome-protected sites with Python (without running a browser farm)

Point requests at Tripadvisor or Allegro and you will usually get a 403 and a page asking you to prove you are human. Both sites sit behind DataDome. This tutorial covers what DataDome looks at, how to detect its block page in code, and how to get the real HTML and structured data from Python with one HTTP call to StealthASF, the scraping API we build. You don't host any browsers and you don't manage a proxy pool.

What DataDome checks

DataDome scores each request before the site itself answers. At a high level it looks at:

  • The IP address. Datacenter ranges and addresses with a bad history start with a poor score.
  • The connection. An HTTP library negotiates TLS and HTTP/2 differently from Chrome or Firefox, and that fingerprint does not change when you change the User-Agent header.
  • Header consistency. The headers you send, and their order, have to match the browser you claim to be.
  • A JavaScript device check. The first page runs a script that collects browser signals and reports them back. A client that cannot run JavaScript never passes it, and a headless browser with visible automation traces fails it.
  • Behaviour over time. Request rate and navigation patterns from the same client.

If the score is bad, DataDome serves its own page in place of the site's.

Why requests with browser headers still fails

This is the usual first attempt:

import requests

r = requests.get(
    "https://www.tripadvisor.com/",
    headers={
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                      "(KHTML, like Gecko) Chrome/126.0 Safari/537.36",
        "Accept-Language": "en-US,en;q=0.9",
    },
    timeout=30,
)
print(r.status_code)
print("captcha-delivery.com" in r.text, r.headers.get("x-datadome"))
Enter fullscreen mode Exit fullscreen mode

On a DataDome site the typical result is a 403 and a short HTML document that loads its challenge from captcha-delivery.com, often with an x-datadome response header. The copied User-Agent fixes one signal out of many. The TLS handshake still identifies a Python client, and nothing on your side runs the device check.

Doing it yourself means running real browsers on residential IPs, keeping them current as DataDome updates its checks, and paying for that capacity whether a page comes back or not.

Getting the real page with one API call

Create an account at stealthasf.com (200 free credits, no card), verify your email, create an API key in the dashboard and export it:

export STEALTHASF_API_KEY="sasf_live_..."
Enter fullscreen mode Exit fullscreen mode

A first test with curl:

curl --max-time 600 "https://stealthasf.com/v1/scrape" \
  -H "x-api-key: $STEALTHASF_API_KEY" \
  -H "content-type: application/json" \
  --data-raw '{"url":"https://www.tripadvisor.com/","engine":"ultra","extract":"links"}'
Enter fullscreen mode Exit fullscreen mode

The same request in Python, standard library only:

import json
import os
from urllib.error import HTTPError
from urllib.request import Request, urlopen

payload = {
    "url": "https://www.tripadvisor.com/",
    "engine": "ultra",
    "extract": "links"
}
request = Request(
    "https://stealthasf.com/v1/scrape",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "x-api-key": os.environ["STEALTHASF_API_KEY"],
        "content-type": "application/json",
    },
    method="POST",
)
try:
    with urlopen(request, timeout=600) as response:
        result = json.load(response)
except HTTPError as error:
    detail = error.read().decode("utf-8")
    raise SystemExit(f"API error {error.code}: {detail}")

print("Target status:", result["status"])
print("Credits:", result["credits_charged"])
print(result.get("data"))
Enter fullscreen mode Exit fullscreen mode

About engine: ultra is our strongest mode and the one that passed Tripadvisor in our tests. stealth is the 25-credit protected mode. auto (the default) starts with plain HTTP and escalates only when the site blocks, and it remembers sites that needed protected mode. Keep the 600-second timeout: rendering a protected page takes longer than a plain fetch.

A successful response looks like this (trimmed):

{
  "ok": true,
  "job_id": "cm...",
  "engine": "ultra",
  "status": 200,
  "html": "<!doctype html>...",
  "data": { "kind": "table", "columns": ["text", "url"], "rows": [["...", "https://..."]] },
  "usage": { "engine": "ultra", "bytes": 3145728, "proxy_mb": 3.1, "ms": 53210 },
  "credits_charged": 70,
  "balance": 69930
}
Enter fullscreen mode Exit fullscreen mode

Telling a block from a real page

Check three things, in this order.

  1. The API's HTTP status. When StealthASF detects an anti-bot block it returns HTTP 422 with a job_id, and the request is not charged. The example above stops on any API error, so a block never reaches your data.
  2. The target status in the JSON. This is what the site returned. A real page is normally 200.
  3. The content. Look for something only the real page has: a listing ID, a price, a URL pattern. A title alone is not enough.
def is_real_page(result, must_contain):
    html = result["html"]
    return (
        result["status"] == 200
        and "captcha-delivery.com" not in html
        and must_contain in html
    )

if not is_real_page(result, "_Review-"):
    raise SystemExit(f"Unexpected page, job {result['job_id']}")
Enter fullscreen mode Exit fullscreen mode

Save engine, credits_charged and job_id with each record. They tell you which mode produced the page, what it cost, and which request to quote to support.

Extracting data

You don't have to parse HTML for common cases. The extract field returns structured data:

  • links: a table with text and url columns, one row per unique link.
  • text: the visible page text, scripts and styles removed.
  • meta: title, description and Open Graph tags.
  • images, headings and table cover the rest.

The full rendered HTML is always in html for your own parser.

On Tripadvisor, detail pages have _Review- in their path (Hotel_Review-, Restaurant_Review-, Attraction_Review-). Keep those links, request each one with the same payload, and read the JSON-LD the page embeds:

import re

# result is the listing-page response from the example above.
detail_urls = sorted({url for _text, url in result["data"]["rows"] if "_Review-" in url})
print(len(detail_urls), "detail pages")

def json_ld(html):
    """Every JSON-LD block on the page that parses as JSON."""
    blocks = re.findall(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", html, re.S | re.I)
    found = []
    for block in blocks:
        try:
            found.append(json.loads(block))
        except json.JSONDecodeError:
            pass
    return found

# Send each detail URL with the same payload, then read its fields:
# for item in json_ld(detail["html"]): print(item.get("@type"), item.get("name"))
Enter fullscreen mode Exit fullscreen mode

Pages that embed schema.org data give you name, address and rating as JSON with no selectors to maintain. Deduplicate URLs before the detail pass, since popular places appear on several listing pages.

Cost math

Credits per request:

Engine Base credits
http 1
browser 5
stealth 25
ultra 50

Browser-based tiers include the first MB of transfer and add 10 credits per MB beyond it, rounded up. Blocked requests cost 0. With auto, the plain HTTP attempt that DataDome blocks is free and you pay only for the tier that returned the page.

Worked examples:

  • One Tripadvisor city page plus 30 place pages on Ultra at the base rate: 31 × 50 = 1,550 credits.
  • A 3 MB page on Ultra: 50 + 20 = 70 credits.
  • Hobby ($17/month, 70,000 credits) covers 1,400 Ultra pages or 2,800 stealth pages a month. Pro ($49, 250,000) covers 5,000 Ultra pages, Scale ($169, 1,000,000) covers 20,000.
  • One-time packs start at $12 for 25,000 credits, valid for 12 months.

Read credits_charged on each response for the exact figure.

What we tested

On 8 October 2026 we sent one request to each of these sites from our production worker through a residential IP, and each returned the real page:

  • DataDome: Tripadvisor, Allegro
  • PerimeterX: Zillow, Walmart, StockX
  • Akamai: Nike, Lowe's, Home Depot
  • Cloudflare: Glassdoor, WhoScored
  • Kasada: Hyatt, realestate.com.au

No single browser passed all of them. Kasada rejected our hardened Firefox-based browser regardless of IP, while real Chrome passed. One Cloudflare site with an interactive Turnstile check needed the stealth tier rather than ultra. That is why the API has several engines and auto picks the tier per site.

Try it

The API is at stealthasf.com, and new accounts get 200 free credits with no card. That covers four Ultra requests, eight stealth requests or up to 200 plain HTTP requests, enough to point it at the DataDome site you care about and see what comes back. It is also on RapidAPI.

Top comments (0)