Data engineers often face the challenge of sourcing structured, up-to-date information from diverse web platforms for business intelligence or analytical projects. Consider the task of constructing a daily salary index for European startup jobs. This requires regularly collecting job listings, extracting key compensation data, and ensuring the data is clean and consistent for downstream analysis. Manually browsing hundreds or thousands of listings on job boards like TheHub.io is not scalable. Automating this process while accounting for specific search criteria and data cleanliness is crucial.
The Hub Startup Jobs Scraper (available at https://apify.com/crawlerbros/thehub-jobs-scraper) addresses this by providing a programmatic interface to TheHub.io's job listings. This actor uses the public TheHub JSON API, eliminating the need for credential management or complex browser automation setups. Its output includes structured fields like salaryMin, salaryMax, and equity, which are directly relevant for building a compensation database.
Tailoring Data Collection for Salary Benchmarking
To build a salary index, the data engineer needs to specify which jobs to scrape and what information to prioritize. The actor offers two primary modes: searchJobs for keyword-based queries and browseByRole for filtering by predefined job categories. For salary benchmarking, a browseByRole approach combined with a focused jobRole can be more effective than a broad keyword search.
For instance, to track salaries for data-scientist roles, configuring the actor with "mode": "browseByRole" and "jobRole": "data-scientist" ensures that only relevant job postings are collected. The remoteOnly boolean field allows further refinement, enabling the collection of data specifically for remote positions if the salary index needs to distinguish between on-site and remote compensation.
A critical aspect of collecting salary data is ensuring a sufficient volume of records. The maxItems input field directly controls the hard cap on emitted records per run, allowing up to 500 items. To build a daily salary index, running the actor daily with maxItems set to its limit can provide a consistent stream of data.
{
"mode": "browseByRole",
"jobRole": "data-scientist",
"remoteOnly": false,
"maxItems": 500
}
This configuration would yield up to 500 job records for data scientist roles, including salaryRange, salaryMin, salaryMax, and equity where available. The actor ensures that empty fields are omitted, meaning that if a job listing does not specify a salary, the salaryMin and salaryMax fields will not be present in the output, simplifying downstream data cleaning. The jobDescription field is also provided in plain text, with HTML stripped, which can be useful for keyword analysis related to compensation if needed.
Understanding the Output Structure for Analysis
Each record returned by the actor represents a single job listing and includes specific fields crucial for building a salary index.
Key output fields for salary analysis:
-
jobTitle: The position's title. -
companyName: The hiring company. -
location: Geographic location (e.g., "Copenhagen, Denmark" or "Remote"). -
isRemote: A boolean indicating remote eligibility. -
jobRole: The categorized role slug (e.g.,data-scientist). -
salaryRange: A formatted string representing the salary range (e.g., "€50,000 - €70,000"). -
salaryMin: The raw minimum salary as a number. -
salaryMax: The raw maximum salary as a number. -
equity: The equity percentage if offered. -
expirationDate: When the listing is expected to expire. -
createdAt: The date the listing was created. -
scrapedAt: The UTC timestamp when the record was scraped, useful for tracking data freshness.
The presence of salaryMin and salaryMax as distinct numeric fields is particularly valuable for quantitative analysis. These can be directly used for calculating averages, medians, and percentile ranges within a salary index. The equity field adds another dimension to total compensation analysis, allowing for a more holistic view beyond base salary.
The scrapedAt timestamp is crucial for time-series analysis. By logging the scrapedAt alongside other job data, a data engineer can track changes in salary ranges, company hiring patterns, and the prevalence of remote roles over time. This historical data is fundamental for building a dynamic, accurate salary index that reflects market fluctuations.
How to Use The Hub Startup Jobs Scraper
To integrate this actor into a data pipeline for salary benchmarking:
- Navigate to the Actor page: Go to https://apify.com/crawlerbros/thehub-jobs-scraper.
-
Configure Input: Define the
mode,jobRole(ifbrowseByRole),remoteOnlystatus, andmaxItemsbased on your data collection requirements. For a daily salary index, settingmaxItemsto 500 is common.
{ "mode": "browseByRole", "jobRole": "software-engineer", "remoteOnly": true, "maxItems": 500 } Start the Run: Initiate the actor run. This will trigger the data extraction process.
Retrieve Results: Once the run completes, download the collected data from the default dataset. The output will be a collection of JSON objects, each representing a job listing, ready for ingestion into a database or data warehouse.
Schedule Daily Runs: For a continuous salary index, set up a daily schedule for the actor run using Apify's scheduling features or external orchestration.
Cost Considerations for Ongoing Data Collection
Using The Hub Startup Jobs Scraper incurs costs based on events charged during its operation, in addition to platform usage. Each job record successfully extracted and stored in the default dataset is billed as a "result" event at $0.005. Discount tiers apply to this event: BRONZE users pay $0.00433, SILVER $0.00367, and GOLD, PLATINUM, DIAMOND users pay $0.003 per event.
An "Actor Start" event is also charged at $0.005 per GB of memory allocated to the run, with a minimum of one event. For example, a run configured to fetch 500 items daily would incur a minimum of 500 "result" events per run (assuming 500 distinct items are found) plus the "Actor Start" event. Over a month, this would be 30 runs, each potentially producing up to 500 results, totaling 15,000 "result" events, plus 30 "Actor Start" events, on top of platform usage. These costs are predictable based on the maxItems setting and run frequency.
One limitation of this approach is that while maxItems can be set to 500, TheHub.io might not always have 500 active listings matching a specific, narrow jobRole and remoteOnly filter. In such cases, the run would complete with fewer than 500 results, yielding less data for the salary index for that particular day. Data engineers building robust salary indexes must account for these fluctuations in job board activity.
The examples here were produced with The Hub Startup Jobs Scraper. Its README lists the output fields, so you can check a response against the schema before you build on it.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-24. Check the Actor page for the current rates.
Top comments (0)