Cross-border wholesale platforms like DHgate hide structural data behind complex React Server Component (RSC) streams. When analyzing supplier pricing or automating catalog ingestion, fetching plain HTML leaves you missing critical parameters—such as quantity-based price tiers, per-country shipping rates, and variant image maps. Furthermore, DHgate relies on Akamai Web Application Firewall (WAF) rules that challenge standard HTTP clients, frequently returning HTTP 403 status codes to simple curl or requests calls.
Extracting structured data from DHgate requires parsing these embedded RSC payloads directly or escalating to browser impersonation when challenged. The dhgate-scraper handles this orchestration, surfacing structured fields directly from storefront pages, search runs, and shop catalogs.
How RSC Stream Extraction Solves Bulk Sourcing
Traditional HTML scrapers rely on static DOM nodes, which break whenever a marketplace updates its front-end CSS framework. DHgate product pages render initial data through an embedded RSC payload. This stream contains the exact JSON data structures used to populate the user interface, including:
- Tiered pricing arrays (
priceTiers[]) specifying wholesale quantity breaks (startQty,endQty,price). - Logistics properties like
stockCountries[],shippingCostUsd, anddeliveryEstimate. - Seller intelligence metrics, including
positiveFeedbackPercent,sellerLevel,sellerTier, andsellerRatingPercent.
By targeting these underlying streams, data pipeline workflows avoid flaky DOM selectors and gain immediate access to structured numbers without risking missing hidden variants or shipping tables.
When Akamai blocks direct client fetches, the scraper defaults to browser-impersonated TLS via curl_cffi before escalating automatically to Playwright rendering if persistent challenges occur.
Choosing Execution Modes for Specific Sourcing Pipelines
The scraper exposes four operational modes in its input configuration to target different endpoints of the marketplace.
{
"mode": "search",
"searchQuery": "mechanical keyboard",
"minPrice": 15,
"maxPrice": 80,
"country": "US",
"maxItems": 100
}
1. Market Research via search and browseByCategory
To map market availability or monitor dynamic pricing across an entire niche, set mode to search alongside a searchQuery. Alternatively, pass one of the 21 pre-defined category strings (such as electronics or cell-phones-accessories) into the category field using mode: "browseByCategory".
2. Supplier Diligence via shopProducts
Evaluating single-seller risk requires extracting an entire merchant catalog. By passing the numerical shop identifier (extracted from standard merchant URLs like dhgate.com/store/22268785) into shopId, the shopProducts mode steps through all items published by that store. This outputs records with full sellerRatingPercent and transactionCount values for catalog-wide analysis.
3. Target Ingestion via byUrl
If your upstream system already tracks product URLs, pass an array of links directly to startUrls with mode: "byUrl". The pipeline resolves canonical item details, outputting precise variant matrices (variants[]) alongside specification pairs (specifications[]).
{
"mode": "byUrl",
"startUrls": [
"https://www.dhgate.com/product/10a-premium-original-luxury-handbags-designer/1086662889.html"
]
}
Running the Scraper Programmatically
Integrating DHgate extraction into an existing ETL pipeline takes a few API calls using the Apify Python client or REST endpoint.
- Construct the payload: Define the targeted search query, price bounds, or supplier IDs in a JSON object matching the input schema.
- Execute the run: Post the input parameters to the API endpoint to initiate the Actor task.
-
Handle optional proxies: If running from heavily flagged datacenter blocks, configure the
proxyConfigurationparameter to enable automatic proxy rotation across blocked endpoints. - Fetch the dataset: Pull records directly from the default dataset associated with the completed run ID.
The code below shows how to trigger an extraction for wholesale electronics using Python:
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run_input = {
"mode": "search",
"searchQuery": "usb c hub",
"minPrice": 5.0,
"maxPrice": 25.0,
"minRating": 4.5,
"country": "US",
"maxItems": 50
}
run = client.actor("crawlerbros/dhgate-scraper").call(run_input=run_input)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(f"Product: {item.get('title')}")
print(f"Base Price: ${item.get('priceMin')} - ${item.get('priceMax')}")
print(f"MOQ: {item.get('minOrderQty')} {item.get('measureUnit')}")
print("---")
Pricing and Consumption Mechanics
Runs on Apify for this Actor are billed using a pay-per-event pricing structure alongside platform usage.
-
Dataset Item Event (
result): Each result emitted to the default dataset costs $0.005 on the FREE tier ($0.00433 on BRONZE, $0.00367 on SILVER, and $0.003 on GOLD, PLATINUM, and DIAMOND). -
Start Event (
Actor Start): Charged at $0.005 per GB of memory allocated to the run upon execution. - Platform Usage: Platform usage for the run is billed separately at your Apify plan's rates.
Because costs scale directly with records extracted and allocated execution RAM, set the maxItems integer parameter on large category sweeps to prevent your dataset size from running beyond required limits.
Limitations and Edge Cases
DHgate's search architecture uses a fuzzy fallback mechanism. When a specific searchQuery returns zero direct matches, DHgate returns popular recommended products instead of an empty set. The scraper emits these fallback records as valid product items containing their actual product URLs. Pipeline downstream filters should validate that the returned item's title contains critical strict keywords (or supply the containsKeyword input parameter) if exact keyword matches are required for your data schema.
DHgate Scraper is what these steps drive. The README covers the inputs this article skipped, including the ones that change how much a run costs.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-03. Check the Actor page for the current rates.
Top comments (0)