I run Nokku.ai, a free site for job-market trends, courses and remote jobs. Its remote jobs page had a problem I didn't notice for weeks: it listed 30 jobs from the last 30 days.
That looked like a quiet market. It wasn't. We collect from three job boards (RemoteOK, JSearch and Arbeitnow), but the page only read from one of them. The other remote listings were sitting in the database, unused.
Here's what it took to turn that into a page with ~200 live remote jobs from 150 companies, and to make it something Google can actually index.
1. Merge the sources, then deal with the mess
Reading all three boards is one query each. The work is in what comes back:
- The same job shows up twice. A company posts on LinkedIn, and two boards pick it up. I remove duplicates by company and title, and keep the copy that lists pay.
- "Remote" isn't always remote. Boards marked as remote-only still return titles like "Senior Developer – DC + Hybrid". If the title says hybrid or on-site, it's out.
-
Some links aren't links. Anything that isn't
http(s)is dropped before it reaches the page.
const NOT_REMOTE = /\b(hybrid|on-?site|in[- ]office)\b/i
const key = (job: RemoteJob) =>
`${norm(job.company)}|${norm(job.position)}`
// Keep one copy per job, preferring the one that lists pay
if (!existing || (!existing.salary && job.salary)) byKey.set(key(job), job)
2. The bug that made "Go" the hottest skill
Once everything was merged, Go was tagged on 102 of 201 jobs. That's not a market signal, it's a bug.
Older rows had been tagged with substring matching, so "Go" matched "good", "AI" matched "email" and "Java" matched "JavaScript". I'd already switched to word-boundary patterns, but rows scraped before the fix still carried the old tags.
Rather than rewrite the database first, the pagfix rows when it reads them, from the title and
the board's own tags:
if (row.scraped_at >= SKILL_FIX_DEPLOYED) returreturn extractSkillsFromText(`${title} ${boardTags}`)
"Go" dropped from 102 jobs to 11. That's why I til I've looked at the rows behind a surprisingnumber.
3. Google was seeing an empty page
The old page was a client component that fetched jobs after the page loaded. Google got a spinner.
Now /jobs is rendered on the server and refreshed every 30 minutes, so the first 25 listings are in the HTML. Filters
live in the URL (/jobs?skill=Python), so filtand the canonical still points at /jobs.
4. One page per role and skill, but only whe
People don't search for "remote jobs". They seas" or "remote data analyst jobs". So every roleand skill gets its own page:
/jobs/remote-python-jobs/jobs/remote-software-engineer-jobs/jobs/remote-jobs-with-salary
The risk with generated pages is thin content. Two rules keep it honest:
-
A page with fewer than 5 listings is
noindexand left out of the sitemap. It still works for visitors, and it goes back into search on its own when jobs arri. - Every intro is written from that page's data, not a template with a keyword swapped in:
18 remote jobs from the last 30 days ask for Python, at 17 companies. The skills most often listed alongside it are AI,
AWS and Java.
5. Only show numbers that can bear the weigh
My first header showed "Median listed pay: $1t came from **3 listings.
A median of three is an anecdote, so the page nngs with pay before it shows one. Until then it
just says "With pay listed: 3". It's less impressive, and it's true.
The same goes for the "New" badge. With a two-day window, 13 of the first rows had it, so it meant nothing. Now it marks
only the newest day's listings.
6. Small things that made the list easier to
- Company first, in colour. Each row leads the same colour as its initials tile, then the title, location and up to four skills.
-
Where you'll apply. "via LinkedIn", "via te", with the site's favicon. The icons are
served from our own
/publicfolder instead of a favicon service, which would send every visitor's IP address to a third party. -
Links out are
nofollow, and each board is credited on the page.
What I'd tell anyone building a job aggregator
- Count what you show against what you store. My problem wasn't a lack of data.
- **Treat surprising numbers as bugs until prov
- Generated SEO pages need a quality floor, and a way to get back out of the index.
- **Don't headline a statistic you wouldn't def
It's free, with no sign-up: [nokku.payanai.conai.com/jobs)
Next I'm adding more free job-board feeds, and . If you've built something similar, I'd like to hear how you handle duplicates across boards. Mine is still fairly simple.
Top comments (0)