Extracting ecommerce data from Russian marketplaces like Megamarket.ru often results in bloated datasets. When pulling product listings for competitive price intelligence or merchant auditing, receiving thousands of low-value accessories or out-of-stock items inflates your dataset size and incurs unnecessary per-event costs.
Using the Megamarket.ru Scraper actor allows you to push filtering logic directly to the source fetcher rather than filtering records in post-processing. Restricting the ingestion pipeline at the request level guarantees that every record emitted to your default dataset represents a usable product.
Server-side filters prevent paying for unwanted records
This actor uses a PAY_PER_EVENT pricing structure. Every output row written to the default dataset counts as a result event priced at $0.005 per event. On top of event charges, users pay for platform usage consumed by the run at their Apify plan's standard rates.
When running broad catalog searches, up to 70% of the returned inventory may fall outside your target price range or brand requirements. If you ingest 1,000 unformatted search results and drop 700 of them in your local ETL script, you still pay for all 1,000 result events ($5.00 at the FREE tier, or $3.00 at the GOLD tier) plus the platform usage required to parse those records.
Applying filters within the actor's input schema drops unmatching inventory before records are generated. The primary fields for controlling payload boundaries include:
-
minPrice(integer): Ignores products priced below this threshold in Russian rubles (RUB). -
maxPrice(integer): Ignores products priced above this threshold in RUB. -
minRating(integer): Filters out items rated below a 1–5 scale. -
onSaleOnly(boolean): Drops items that do not currently have an active discount badge or crossed-out original price. -
maxItems(integer): Sets a hard cap on total records emitted (from 1 to 1,000).
Passing minPrice: 10000 on a query for consumer electronics prevents cables, cases, and minor accessories from populating the dataset. This keeps event counts aligned directly with core inventory.
Execution modes and schema variations
The actor operates across four distinct mode settings. Selecting the correct mode prevents unnecessary requests and keeps record types consistent:
-
search: Accepts a Latin or Cyrillic keyword viasearchQueryand returns lightweight product records. -
byCategory: Accepts full catalog URLs viacategoryUrlsto fetch category-wide listings. -
productDetails: Takes specific URLs or numeric product IDs viaproductUrlsand yields full product specifications, seller legal details, and variant matrices. -
reviews: Takes product IDs viaproductUrlsto extract star breakdowns and individual customer comments.
Search and category modes yield basic listing data such as title, price, rating, sellerName, and images. If your goal is to evaluate legal seller entities or extract exact stock quantities, skip search modes and pass target IDs directly to productDetails.
Full product detail runs parse deep merchant metadata, returning fields like sellerLegalName, sellerInn (taxpayer ID), sellerOrgn (registration number), stockQuantity, and Sber cashback metrics (cashbackAmount, cashbackPercent).
Configuring a filtered category extraction
To extract discounted smartphones within a designated price range, configure the JSON input to combine catalog targeted filtering with a strict item cap.
{
"mode": "byCategory",
"categoryUrls": [
"https://megamarket.ru/catalog/smartfony/"
],
"minPrice": 20000,
"maxPrice": 120000,
"onSaleOnly": true,
"sortBy": "priceAsc",
"maxItems": 100
}
This configuration navigates the target category, ignores listings below 20,000 RUB or above 120,000 RUB, drops items without discounts, sorts the valid items by ascending price, and stops immediately after emitting 100 records.
Step-by-step setup in Python
You can programmatically execute the run and pull the filtered dataset directly into a Python pipeline using the official SDK.
Step 1: Install the Apify Client
pip install apify-client
Step 2: Write the execution script
Initialize the client, pass the input parameters with strict price bounds, and stream the dataset items as they complete.
import os
from apify_client import ApifyClient
client = ApifyClient(os.getenv("APIFY_TOKEN"))
run_input = {
"mode": "search",
"searchQuery": "холодильник",
"minPrice": 30000,
"minRating": 4,
"sortBy": "popular",
"maxItems": 50
}
# Run the actor and wait for completion
run = client.actor("crawlerbros/megamarket-scraper").call(run_input=run_input)
# Fetch results from the run's default dataset
dataset_items = client.dataset(run["defaultDatasetId"]).list_items().items
for item in dataset_items:
print(f"Product: {item.get('title')} | Price: {item.get('price')} RUB | Rating: {item.get('rating')}")
Every item returned by list_items() corresponds to a single result event.
Cost calculations across billing tiers
Billing for the actor depends on the total count of charged events plus platform usage.
Every execution incurs an initial Actor Start event, charged at $0.005 per GB of memory allocated to the run. Each emitted dataset row incurs a result event charge. The unit price for result events scales down based on your account's discount tier:
- FREE: $0.005 per result
- BRONZE: $0.00433 per result
- SILVER: $0.00367 per result
- GOLD: $0.003 per result
- PLATINUM: $0.003 per result
- DIAMOND: $0.003 per result
Extracting 500 product records on a standard FREE tier run consumes:
- 1 x
Actor Startevent (at 1 GB allocation) = $0.005 - 500 x
resultevents ($0.005 each) = $2.50 - Platform usage: billed separately according to your plan's rates.
If you omit the minPrice or onSaleOnly parameters, pulling 500 items might yield 350 irrelevantly cheap or non-discounted products. Setting parameters upfront delivers 500 targeted records for the same $2.50 event total, avoiding the need to run 1,500 total extractions to gather 500 valid data points.
Scraper scope and execution limits
This scraper extracts publicly accessible storefront data available on Megamarket.ru. It does not access private user accounts, order histories, internal admin dashboards, or checkout flows requiring user authentication.
When planning data pipelines, verify whether your target analytics require historical trend data or point-in-time snapshots, as the scraper reflects current live catalog state upon execution.
Source for the runs in this article: Megamarket.ru Scraper. The input schema there is authoritative; treat anything in this post that contradicts it as out of date.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-02. Check the Actor page for the current rates.
Top comments (0)