DEV Community

Cover image for What I learned counting 600,000 live job adverts every night
Praveen Kumar
Praveen Kumar

Posted on

What I learned counting 600,000 live job adverts every night

I'm building TUNAI, a job-matching assistant. Under it sits a crawler that
reads live job adverts straight from employers' own applicant tracking
systems (Greenhouse, Workday, Lever, Ashby, Oracle HCM and about fifteen
others) across the UK, the US, Canada and Australia. Today that is about
600,000 live adverts, recounted every night.

Once you hold that many adverts, you start noticing things no single job
seeker could. I put the odd ones on a page that updates itself:
jobs.tun-ai.com/insights. Here are a
few, and then the three bugs that nearly made those numbers lie.

The numbers

  • 34,000 live adverts say they were posted more than a year ago, 6% of the adverts that give a date. The oldest claims 2010. Another 66,000 are more than six months old.
  • 3,100 adverts have been reposted five or more times. One has been reposted 31 times.
  • Friday is the busiest day: 20% of UK and US adverts go up on a Friday, and only 4% at the weekend. In the UK, 31% of timed adverts appear between 2pm and 5pm.
  • 2,500 US job titles contain an exclamation mark. The UK has 79.
  • Three live titles still ask for a "rockstar", and five for a "ninja". The hype word that won is "champion", with 140.

Every figure on the page carries the number of adverts it rests on and the
caveat that must travel with it. The posting dates, for example, are the
employers' own, and plenty of the old ones are evergreen adverts kept open
to collect CVs rather than to fill one job. The whole table is a CSV under
CC BY, also on Kaggle,
Hugging Face
and Zenodo.

Three bugs that nearly lied to me

1. The same job, listed twice by the same employer. Oracle HCM lets one
company run several career sites, and each site serves the same
requisitions. Stored naively, 43,704 of 61,111 eligible Oracle adverts sat
in 15,035 groups of byte-identical adverts. Merging on "similar text" had
once destroyed hundreds of real vacancies, so the fix is deliberately
narrow: a job points at the oldest of its group only when the source,
title, place and the advert text are all identical, and both rows stay in
the corpus. Only one gets a public page.

2. An index that was never used. Postgres partial indexes only apply
when the query's predicate matches the index's. Our indexes said
WHERE live; our queries said live IS true. To a human those are the
same; to the planner they are not, so the vector (HNSW) and GIN indexes
sat unused and a match took 50 to 130 seconds. If you use partial indexes,
check idx_scan on each of them. Ours read zero.

3. A reader that stopped too early. The "how many adverts state pay"
figure said 94% of US adverts give no pay. Industry research says roughly
half do, so I sampled 200 of our "no pay" US adverts by hand: 43 stated
pay in text we already held, and 42 of those put it after the 600th
character, which is exactly where our reader stopped. US adverts tend to
put the salary range near the end. The figure is
withdrawn until the fix is certified. The lesson I keep relearning: when
your number disagrees with everyone else's, check your reader before you
publish the surprise.

What it's for

The counting exists because the product matches people to jobs: it ranks
live adverts against your CV and keeps a shortlist current. There is also
a free ATS CV checker that
shows how an applicant tracking system reads your CV, no sign-in needed.

If a number on the insights page looks wrong to you, I would genuinely like
to hear it.

Top comments (0)