Extracting unstructured data from classifieds platforms often leads to inflated execution costs when developers pull broad datasets and apply filters locally. On Haraj (haraj.com.sa), Saudi Arabia's primary consumer marketplace, a generic query for vehicles or electronics returns thousands of listings. If your data pipeline only requires a specific sub-category or model, running client-side cleanups means paying for records you instantly throw away.
The Haraj Scraper interface provides targeted parameters that pass directly to Haraj's underlying GraphQL endpoint. Structuring requests with exact tags like subTag, carBodyType, and regional bounds reduces payload sizes before data hits your dataset.
Upstream tag resolution saves event charges
When scraping classifieds, specifying a generic category like حراج السيارات (Car Haraj) yields an unfiltered stream of cars, spare parts, and heavy equipment. The actor's input schema accepts granular parameters that take priority during query construction.
Setting subTag applies a specific brand or model tag from Haraj's internal tag tree (for instance, a specific vehicle brand or phone model). When subTag is provided, it overrides category for upstream API filtering. Similarly, setting carBodyType (such as جيوب for SUVs or سيدان for Sedans) takes priority over both subTag and category.
By enforcing strict criteria at the platform API level, you avoid collecting irrelevant records. This direct filtering directly impacts platform execution efficiency and dataset event fees.
Structured payload comparison
Consider a scenario targeting mid-sized SUVs in Riyadh priced between 50,000 SAR and 120,000 SAR. An unoptimized request pulls every automotive post and requires local filtering:
{
"mode": "browse",
"category": "حراج السيارات",
"maxItems": 500
}
The optimized alternative passes exact requirements directly to the GraphQL request layer:
{
"mode": "search",
"searchQuery": "تويوتا",
"category": "حراج السيارات",
"carBodyType": "جيوب",
"cities": ["الرياض"],
"priceMin": 50000,
"priceMax": 120000,
"onlyWithImage": true,
"maxItems": 100
}
The second payload forces Haraj's backend to execute the logic, returning only valid listings that contain structured domain models.
Structured schema mapping for vertical data
The actor returns cleaned, typed JSON objects containing domain-specific fields based on the item type. Instead of parsing raw HTML descriptions, the backend extracts metadata into top-level keys like car, realEstate, and job.
Vehicle records include standardized specifications:
{
"id": "186610241",
"title": "تويوتا لاندكروزر 2022 فلكامل",
"priceSAR": 210000,
"category": "حراج السيارات",
"city": "الرياض",
"geoCity": "الرياض",
"geoNeighborhood": "حطين",
"authorUsername": "معرض الرياض",
"authorId": "987654",
"postedDate": "2023-10-15T10:30:00.000Z",
"upvotes": 12,
"downvotes": 0,
"car": {
"modelYear": "2022",
"mileageKm": "45000",
"fuelType": "بنزين",
"transmission": "أوتوماتيك",
"condition": "مستعمل",
"is4WheelDrive": true
},
"recordType": "carListing",
"scrapedAt": "2023-10-20T08:00:00.000Z"
}
For real estate entries, the output surfaces regulatory compliance numbers inside the realEstate object, such as regaAdvertiserRegistrationNumber and regaAuthorizationNumber. For general items, extracted key-value pairs sit inside a specs object. Listings where the seller did not state an explicit price omit the priceSAR key entirely, avoiding misleading zero-value entries in numeric analysis pipelines.
How to execute a target extraction pipeline
Setting up an automated run requires defining execution modes and targeted parameters inside your orchestration script or HTTP client.
Step 1: Select the run mode
Choose one of the five supported execution modes based on your ingestion logic:
-
search: Keyword queries paired with filtering parameters. -
browse: Category or city feeds without keyword dependencies. -
byId: Lookups for specific numeric IDs or canonical listing URLs. -
byAuthor: Complete listing dumps for a specific seller account. -
trending: Keyword popularity metrics spanning a configurable day range.
Step 2: Configure geographic and price limits
Define target markets using the array-based cities field to query multiple regions (e.g., Riyadh and Jeddah) in a single run. To filter down to specific municipalities, populate neighborhood with the exact text string matching Haraj's geoCity value.
{
"mode": "search",
"searchQuery": "شقة للاراض",
"category": "حراج العقار",
"cities": ["الرياض", "جده"],
"neighborhood": "الدرعية",
"priceMin": 30000,
"priceMax": 80000,
"maxItems": 200
}
Step 3: Run the Actor via Client SDK
Trigger the job using the Apify Python SDK, passing input arguments directly:
from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run_input = {
"mode": "search",
"searchQuery": "ايفون 15 برو",
"category": "حراج الأجهزة",
"city": "الرياض",
"priceMin": 3000,
"priceMax": 4500,
"sortBy": "recentlyActive",
"maxItems": 50
}
run = client.actor("crawlerbros/haraj-scraper").call(run_input=run_input)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(f"{item.get('title')} - {item.get('priceSAR')} SAR")
Calculating operational costs on Apify
The actor uses a pay-per-event pricing structure on the Apify platform. Each returned record emitted to the default dataset charges a fixed "result" event fee, alongside a static initialization fee for starting the execution container.
-
Actor Start (
apify-actor-start): $0.005 per GB of memory allocated to the run. -
Dataset Result (
apify-default-dataset-item): $0.005 per event on the FREE tier ($0.00433 on BRONZE, $0.00367 on SILVER, and $0.003 on GOLD, PLATINUM, and DIAMOND tiers).
Each result costs $0.005 on the free tier, plus a one-time start charge of $0.005 per GB allocated. Platform usage for the run is billed separately at your Apify plan's rates.
Passing tight criteria via subTag, priceMin, and cities limits emitted records strictly to target items. Collecting 100 relevant items costs $0.50 in dataset result event fees on the free tier, whereas pulling 1,000 broad category items to find those same 100 listings locally increases dataset fees to $5.00 for the same yield.
Scope limitations
This actor does not traverse or scrape Haraj's standalone Jobs section (وظائف) through category browse mode, as that vertical relies on a separate upstream architecture not exposed by the general marketplace search endpoints. Individual job postings appear only when directly discovered through explicit keyword searches or author lookup runs.
Haraj Scraper is the Actor behind these examples. If a selector in your own version breaks, compare your output against the fields listed in its README first.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-02. Check the Actor page for the current rates.
Top comments (0)