If you build a job board, a newsletter or a salary dataset, you quickly find out that "remote jobs" is not one source. Each board has its own shape, and several of them have free public endpoints you can use without scraping HTML.
Free endpoints worth knowing
-
RemoteOK:
https://remoteok.com/apireturns JSON. The first element is a legal notice, skip it. Link back to the job URL, their terms ask for attribution. -
Remotive:
https://remotive.com/api/remote-jobs?category=software-dev. Documented, with a fair-use note about polling rarely. -
Arbeitnow:
https://www.arbeitnow.com/api/job-board-apipaginated JSON, mostly Europe. - We Work Remotely: RSS feeds per category.
- Hacker News "Who is hiring?": the monthly thread is on the HN Algolia API; each top-level comment is one posting.
The normalization problem
Every source names things differently: position vs title, tags vs category, salary as a string, a range or missing. Pick one schema early:
{ "source": "remoteok", "title": "...", "company": "...", "location": "Worldwide",
"tags": ["python"], "salaryMin": null, "salaryMax": null, "postedAt": "2026-10-01T00:00:00Z", "url": "..." }
Deduplicate
The same job often appears on 2-3 boards. A key of lowercase company + title + normalized location catches most of them. Keep the earliest postedAt and the list of sources.
If you do not want to maintain it
I run this pipeline as an Apify actor that merges seven sources into one deduplicated feed, with keyword, tag and date filters. You pay per job delivered, and it uses official APIs and feeds only, so it works without proxies.
Remote Jobs Aggregator on Apify
Code and notes: https://github.com/quiethand098/remote-jobs-aggregator
Top comments (0)