Building a unified job feed typically requires navigating a fragmented landscape of proprietary APIs and anti-bot measures. Developers attempting to aggregate listings from the UK and EU often find that high-volume sources like Indeed or TotalJobs require expensive residential proxies and complex browser automation to bypass Web Application Firewalls (WAF). This technical overhead increases the cost per record and introduces significant latency into the data pipeline.
The UK Jobs Aggregator addresses this by focusing on sources that remain accessible via standard datacenter IPs. By querying RemoteOK, Arbeitnow, and Reed.co.uk through an HTTP-only architecture, the actor provides a structured, multi-source dataset without the overhead of headless browsers or specialized proxy rotations.
Resolving Schema Fragmentation Across Job Boards
The primary challenge in job aggregation is not just the retrieval of data, but the normalization of disparate schemas. RemoteOK and Arbeitnow provide JSON APIs, while Reed.co.uk requires a DOM-based scrape of search result cards. These sources describe salaries, locations, and employment types using different terminologies and structures.
This actor normalizes these inputs into a consistent output object. For instance, while one board might provide a raw string like "£45k - £55k per annum," the actor parses this into structured numeric fields. The resulting dataset includes salary_min, salary_max, and a derived salary_average, alongside the salary_currency (e.g., GBP, EUR) and salary_period. This allows for programmatic filtering using the minSalary input field, which drops any posting where the parsed salary_max falls below your threshold.
By centralizing these sources into one stream, you avoid writing separate parsers for each site's specific HTML structure or API response. The logic also handles "isRemote" flags by searching for specific string patterns within the location data or source-provided metadata, providing a boolean field that is more reliable for filtering than raw text search.
Implementing Cross-Board Deduplication
Duplicate listings are a recurring problem when aggregating from multiple boards. A single role is often posted on a niche remote board, a general UK board like Reed, and an aggregator like Adzuna simultaneously. Without a deduplication strategy, a search for "Senior DevOps Engineer" will return redundant records, skewing analytics or cluttering a user-facing job feed.
The actor implements a deduplication logic controlled by the removeDuplicates boolean. When enabled, it groups records by a composite key consisting of the title, company, and the first segment of the location string (split by commas). These values are lowercased and stripped of whitespace to ensure that "London, UK" and "London, England" are treated as the same locality. In this "first-occurrence-wins" model, the aggregator retains the first instance of a job it encounters and discards subsequent matches from other sources.
Extending Reach with Adzuna Integration
While the actor works out of the box for UK and EU remote roles, it includes optional support for Adzuna via the adzunaAppId and adzunaAppKey input fields. Adzuna acts as a meta-search engine, and by providing a free API key from their developer portal, you can expand the geographic scope of the run.
This is particularly useful when the country parameter is set to "UK" or "DE," as it allows the actor to supplement HTML-scraped data from Reed with Adzuna's API results. Because Adzuna supports multiple countries, this is also the primary mechanism for retrieving listings outside of the UK and Germany. This hybrid approach—combining direct HTML scraping of local boards with API-based aggregation—provides a broader dataset than a single-source scraper could offer.
Usage and Implementation
To begin collecting aggregated job data, follow these steps:
- Navigate to the UK Jobs Aggregator on the Apify platform.
- Configure the
jobRole(required) with your search term, such as "Product Manager" or "Data Scientist." - Set the
countryto "UK", "DE", or "Remote". Note that RemoteOK and Arbeitnow are queried regardless of this setting to maximize coverage. - Optionally set
minSalaryto exclude lower-paying roles anddatePosted(e.g., "LastWeek") to ensure recency. - Run the actor and monitor the default dataset for the standardized JSON output.
The maxResults field acts as a hard cap on the total records emitted. If you set this to 100, the actor will stop processing once it has successfully normalized and deduplicated 100 unique items across all active sources.
Structured Data and Company Enrichment
One of the more difficult aspects of job scraping is gathering metadata about the hiring organization. Many boards only provide a company name string. However, when a source includes JSON-LD or structured hiringOrganization data, this actor extracts enriched fields. This includes the companyLogo URL, companyWebsite, and companySocialLinks.
The social links are returned as a flat dictionary, such as {"linkedin": "company-handle", "twitter": "company-handle"}. This is highly effective for lead generation or automated outreach workflows where you need to verify the legitimacy of a posting or find the company’s digital footprint without performing secondary searches.
Cost Structure and Platform Usage
Billed on a per-event basis, the actor's costs are transparent and tied directly to the volume of data retrieved. The charges are as follows:
- Actor Start: $0.005 per GB of memory allocated to the run. This is a one-time charge at the beginning of the execution.
- Result: $0.002 per result generated in the default dataset. This price varies based on the user's discount tier.
- FREE: $0.002
- BRONZE: $0.00167
- SILVER: $0.00133
- GOLD: $0.001
- PLATINUM: $0.001
- DIAMOND: $0.001
On top of these event charges, users also pay the Apify platform usage the run consumes, which is billed separately at the rates defined by their specific plan. Using a "result" based model allows you to estimate the cost of a run by multiplying your maxResults input by the per-item rate. For example, a run yielding 1,000 results at the GOLD tier would incur $1.00 in result charges plus the $0.005 per GB start fee and the platform usage.
Data Limitations and Scope
It is important to recognize what this tool is not designed to do. Because it prioritizes an HTTP-only, no-proxy architecture to keep costs low, it does not support sites with aggressive anti-bot protections like Indeed, TotalJobs, or GOV.UK's Find a Job service. Attempting to scrape these sites with datacenter IPs generally results in generic 403 errors or WAF challenges. If your project requires data specifically from the Department for Work and Pensions (DWP), this aggregator is not the appropriate tool, as those sources are currently disabled to maintain the reliability of the actor's datacenter-friendly design. Furthermore, the description field is capped at 2000 characters; if you require the full text of a 5000-word job posting, you must use the provided url field to perform a targeted follow-up scrape of the specific job page.
Everything above runs on UK Jobs Aggregator. Start with a small input and a low result limit before you widen the run -- the output shape is easier to check that way.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-24. Check the Actor page for the current rates.
Top comments (0)