I run a Job Postings Search API (https://x402.freeq.one/tools/jobs.html) that queries live postings across boards and returns structured results — title, company, location, apply URL, salary, description. Building it taught me my favorite kind of lesson: the data lies in boring, predictable ways.
First real test: "backend engineer, Berlin". 240 hits. 178 unique jobs.
The same posting appeared five times. Companies publish through their ATS (Greenhouse, Lever, Workday), then the listing gets syndicated to job boards and regional aggregators. Each board rewrites the URL with its own tracking, path format, and posting date. One listing, five hosts, five "posted 2 days ago" claims.
My first dedup was apply-URL based. Dead on arrival — no two copies share a URL.
Then (company, title) matching. Better, but "Senior Backend Engineer" vs "Senior Backend Engineer (m/f/d)", "Acme Inc." vs "Acme", and "Engineer, Backend" vs "Backend Engineer" all slipped through.
What finally worked was three-layer normalization:
- URL: strip query params (pure tracking), fix scheme and trailing slash, lowercase host. Not for matching — kept as provenance evidence.
- Title: lowercase, strip parenthetical suffixes like (m/f/d) and (remote), drop location prefixes, collapse whitespace.
- Company: drop legal suffixes (Inc, GmbH, Ltd), collapse whitespace.
Then group on normalized company + title, take the earliest posting date as creation date, and keep every apply URL ordered with the canonical one first.
It's not perfect. Genuine reposts (same role re-opened months later) merge into one entry, and some chain-company duplicates still escape. But removing ~25% noise beats shipping it raw.
The wider lesson for agents: whenever a dataset is syndicated — jobs, news, products, any board-fed feed — the URL is an attribute of the copy, not the entity. Dedupe on normalized content fields; treat URLs as provenance, not identity.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.