If you've ever built a niche job board, a hiring-signal feed for sales, or just wanted to track who's hiring data engineers, you've probably hit the same wall: every company's careers page looks different.
The good news is that under the hood, most of them run on a small number of applicant tracking systems (ATS), and many of those publish open jobs through public, unauthenticated endpoints. They exist precisely so companies can show their openings anywhere. You don't need to scrape HTML.
Here's how the big ones work.
Greenhouse
curl "https://boards-api.greenhouse.io/v1/boards/airbnb/jobs?content=true"
The airbnb part is the company's board token, the same slug you see in boards.greenhouse.io/airbnb. You get title, location, departments, first-published date and (with content=true) the full description as HTML-escaped text.
Lever
curl "https://api.lever.co/v0/postings/palantir?mode=json"
Lever gives you categories.commitment (full-time, contract…), workplaceType (remote/hybrid/onsite), and a salaryRange object when the company publishes pay.
Ashby
curl "https://api.ashbyhq.com/posting-api/job-board/openai?includeCompensation=true"
Ashby is the friendliest for salary data: includeCompensation=true returns structured min/max, currency and interval. In a quick test, OpenAI's "Data Engineer" role came back as $235K – $385K.
Workday
Workday is the one everyone asks about, and the trickiest. Each company's career site exposes a JSON endpoint behind its public page:
curl -X POST "https://nvidia.wd5.myworkdayjobs.com/wday/cxs/nvidia/NVIDIAExternalCareerSite/jobs" \
-H "Content-Type: application/json" \
-d '{"appliedFacets":{},"limit":20,"offset":0,"searchText":""}'
Things to know:
- You can't guess the URL from a company name. Tenant (
nvidia), data center (wd5) and site name (NVIDIAExternalCareerSite) all vary, so start from the company's careers link. - Pages are capped at 20 jobs, and the
totalfield is only filled in on the first page. If you read it on page 2, you'll stop early. (Ask me how I know.) - Dates come as text like "Posted 3 Days Ago", so you only get day precision.
The annoying part: normalizing everything
Each system has different field names, date formats, location shapes and salary structures. If you want one clean table, you end up writing a small adapter per ATS, plus:
- detecting which ATS a company uses,
- handling pagination and rate limits politely,
- deduplicating, and
- remembering which jobs you've already seen if you run it daily.
A shortcut
I packaged all of this as an Apify Actor: Company Jobs. You give it a list of companies (names, websites or careers URLs) and it returns one normalized row per open job from Workday, Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Recruitee, Personio, Teamtailor and Rippling:
{
"company": "openai",
"title": "Data Engineer",
"location": "San Francisco; Mountain View",
"salaryMin": 235000,
"salaryMax": 385000,
"salaryCurrency": "USD",
"platform": "ashby",
"url": "https://jobs.ashbyhq.com/openai/…"
}
Two features I use most:
- "Only new jobs since my last run": put it on a daily schedule and get just the new postings. Great for hiring alerts or "which of my target accounts just opened a data team?"
- Title, location and remote filters, so you only pay for jobs you actually want.
It's pay-per-result (about $1 per 1,000 jobs), and it only reads the public job boards that companies publish themselves.
Be a good citizen
Whatever route you take:
- Stick to the official public endpoints above, not HTML scraping.
- Keep concurrency low and back off on 429s.
- Don't collect recruiter names or emails that some boards include. You don't need them for job data.
If there's an ATS you'd like covered next, tell me in the comments.
Written with AI assistance; the endpoints, numbers and examples were tested before publishing.
Top comments (0)