A user's job-matching agent kept ranking near-identical listings against each other, burning tokens and crowding out real candidates. I pulled a sample to investigate: one search for "data engineer" in Berlin returned 1,000 results — and 359 of them were duplicates of other entries in the same result set.
My dedup key was the URL. That was the bug.
Job boards rewrite URLs for tracking. One board appended utm_source, another used gh_src, a third proxied every apply link through a /jobclick?id=... redirector before it ever reached the company page. Three boards, one job, three URLs guaranteed distinct. Exact-match URL dedup caught none of it.
The fix ended up being three layers, cheapest first:
- Canonicalize URLs. Strip query strings, lowercase the host, drop trailing slashes and fragments. This alone removed about half the duplicates.
-
Normalize title + company as the real fingerprint. Lowercase, strip punctuation and bracketed tags like
(Remote), collapse whitespace. "Senior Software Engineer — Remote" and "Senior Software Engineer (Remote)" at the same company now hash identical. - Canonical copy selection. When several copies survive, keep the one with the richest structured fields — salary range, seniority, employment type — rather than whichever landed first.
Result: 1,000 results → 641 unique postings, with the best-filled record kept as the survivor.
Two surprises worth flagging. Some boards mutate the title itself — "Hiring urgently!" appended, company rendered as "Acme GmbH" vs "Acme Inc." — so even a pure text match fails; hashing normalized fields is more forgiving than substring comparison. And duplicate copies aren't always identical: the board's copy often lacks salary data the original company page has, which is a second reason to pick the canonical record by field completeness.
The dedup layer now runs before anything else in that pipeline. The search itself stayed simple — one call per keyword + location pair, structured results back. That's what I ended up packaging it as: my Job Postings Search API. Honestly, though, the dedup layer is what made those results usable.
Top comments (0)