Data engineering for the Russian DIY market requires navigating the recent rebranding of major players. Lemanapro, formerly Leroy Merlin Russia, maintains a massive catalog of home improvement goods, but its anti-scraping measures often hinder direct data extraction. For developers building price monitoring tools or supply chain dashboards, the challenge lies in obtaining more than just a surface-level price. Real utility comes from capturing regional stock levels and specific courier delivery windows across different Russian territories.
The lemanapro-scraper handles the heavy lifting of bypassing QRATOR protection and navigating the complex structure of Lemanapro.ru. It provides a structured path to extracting granular product details including technical specifications, inventory status, and historical price data.
Accessing Regional Inventory and Delivery Logic
One of the most difficult datasets to capture from a large-scale retailer is the localized availability of goods. Pricing and stock for a heavy item like ceramic tile or a miter saw in Moscow often differ significantly from the availability in Novosibirsk.
By setting the mode to productDetails, the scraper targets specific product pages to extract the stockAvailabilityText. This field returns the live stock-quantity sentence as it appears on the site, such as "Доступно для заказа 156 кор." (156 boxes available for order). Furthermore, the scraper identifies the deliveryRegion, ensuring that the extracted data is tied to a specific geographical context.
For logistics analysts, the tool extracts three key logistical strings:
-
pickupText: Details regarding in-store pickup availability (e.g., "Same day pickup for free"). -
deliveryEstimateText: The estimated arrival window for courier delivery and the associated cost. -
specifications: An array of name-value pairs that define the technical attributes of the product, which is vital for comparing similar items across different brands.
The productDetails mode is specifically designed for this level of depth. While search and category modes are useful for discovery, they often provide only the "snippet" data. To build a reliable inventory tracker, a developer would typically run a search first to gather URLs, then feed those URLs into a second run using the productUrls array to pull the deep technical and logistical data.
Optimizing Scraper Runs with Mode Selection
The scraper operates in four distinct modes, and choosing the right one directly impacts the efficiency of the data pipeline.
-
search: Best for broad market research. By passing a
searchQuerylike "ламинат" (laminate), you can see what the retailer prioritizes in its internal search rankings. -
byCategory: Useful for catalog mirroring. You can use a curated list via the
categoryfield or provide a specificcustomCategorySlugfrom the website's URL (e.g.,stroitelnye-materialy). -
productDetails: The high-fidelity option. It requires a
productUrlorproductUrlsand returns the full specification table and image gallery. -
bySku: The most precise mode for price monitoring. If you already have a list of internal Lemanapro article numbers, providing them in the
productSkusarray ensures you are tracking the exact inventory items without the noise of search results.
A common workflow involves using the minPrice and maxPrice filters at the input level. This allows the scraper to discard items outside of a specific budget or quality tier before they are even processed into the dataset, which saves post-processing time.
Setting Up a Data Extraction Task
To extract data from Lemanapro, a developer follows a standard configuration pattern. No custom browser logic or header rotation needs to be written manually, as the actor manages these via its internal fallback mechanisms.
- Select the operation mode by setting the
modeproperty (e.g.,bySku). - Define the targets, such as entering a list of article numbers into the
productSkusarray. - Set a
maxItemslimit to control the volume of the output; this acts as a hard cap on the number of product records emitted. - Configure optional filters like
sortByto organize the results by price or relevance before extraction. - If the target pages are heavily protected, configure the
proxyConfigurationto enable browser-based enrichment, though the actor's default behavior uses a Google-indexed fallback to retrieve data even without proxies.
The resulting dataset is structured as a collection of objects where each record represents a single product. Fields like oldPrice and discountPercent are included automatically when a sale is active, making it straightforward to calculate the depth of retail promotions.
Understanding the Cost Structure
The cost of running the lemanapro-scraper is determined by specific events and the underlying infrastructure usage. This model allows for predictable scaling based on the number of items retrieved.
- Actor Start: There is a flat charge of $0.005 per GB of memory allocated to the run. This is charged once when the execution begins.
- Result: Each single product record successfully added to the default dataset is charged as a "result" event. The price for this event depends on the user's tier: $0.005 for the FREE tier, $0.00433 for BRONZE, $0.00367 for SILVER, and $0.003 for GOLD, PLATINUM, and DIAMOND tiers.
Platform usage for the run, which includes the resources consumed during the scraping process, is billed separately at the rates defined by your specific Apify plan. For example, if you are on the FREE tier and retrieve 1,000 product results, the event-based cost would be $5.00 for the results plus the $0.005 per GB start charge, in addition to the platform usage.
Data Schema and Limitations
The output schema is designed to be sparse; if a piece of information is not present on the Lemanapro website, the corresponding field is omitted from the JSON record. This prevents issues with null-pointer exceptions in downstream applications. A typical product record includes the sku, brand, price, availability, and productUrl.
When using the productDetails mode, you also receive the categoryPath and breadcrumbs, which are essential for maintaining a hierarchical database. The scrapedAt timestamp is included in every record to help analysts determine the freshness of the price data.
One limitation of this approach is that it is not intended for real-time transactional use. Because the scraper may rely on indexed fallbacks or browser-based enrichment to bypass bot detection, there is an inherent latency. It is an analytical tool for market research and price indexing, rather than a sub-second API for live checkout flows. Furthermore, the maxItems property is capped at 500 per run, meaning very large catalog extractions must be split across multiple tasks or different category slugs.
Lemanapro Scraper is the Actor behind these examples. If a selector in your own version breaks, compare your output against the fields listed in its README first.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-23. Check the Actor page for the current rates.
Top comments (0)