A job listing can be new to an aggregator without being new to the job market. A recruiter can also refresh an old listing, and a company can publish the same role through several channels. A feed that treats all three events as “new vacancy” gives a developer a misleading queue.
That makes job search an interesting application of a familiar data-engineering problem: keeping event time separate from processing time.
HireSeeker provides a useful example of the product side. Its public interface brings IT and digital vacancies from Telegram, company career pages, LinkedIn, and other career platforms into one place. Profession filters, source links, favorites, and daily selections turn that collection into something a person can use without reading every source feed separately.
Three clocks worth separating
For a vacancy record, these timestamps answer different questions:
| Timestamp | Question it answers |
|---|---|
| Source publication time | When did this version of the listing become public? |
| First observation time | When did this collector first encounter it? |
| Last successful check | When was its source last checked? |
The first timestamp may be missing or ambiguous. The second belongs to the collector. The third describes a check, not necessarily a change in the job itself.
Suppose a crawler discovers a listing today whose source says it was published ten days ago. Sorting by first observation is reasonable for an “unseen by me” feed. Calling the same item a newly opened role would answer a different question with the wrong clock.
A refresh introduces another distinction: the source can update a timestamp while keeping the responsibilities and requirements intact. A useful display can expose that event without pretending it is an entirely different job.
Deduplication should preserve provenance
Two identical titles are not enough to identify the same opening. A company can hire several backend engineers at once. Conversely, one opening can appear as “Python Developer” on a job board and “Backend Engineer” in a recruiter’s Telegram message.
A deduplication decision therefore needs more context: employer, role, location, responsibilities, source identifiers, and the relationship between versions. Even a good decision can lose useful information if it throws away every source except one.
HireSeeker’s published collection and filtering methodology describes content-based duplicate detection and combining matching listings into a card with information about the sources. The useful part of that design is the retained route back to the original: the reader can check details where the vacancy was published.
A small model for the UI
An aggregator can expose freshness without collapsing its clocks. Here is an illustrative Python pattern; it is a design example, not HireSeeker’s implementation:
def freshness_labels(vacancy, now, recent_window):
labels = []
if now - vacancy.first_seen_at <= recent_window:
labels.append("Recently added to this feed")
published_at = vacancy.source_published_at
if published_at is not None and now - published_at <= recent_window:
labels.append("Recently published at the source")
if vacancy.source_updated_at is not None:
labels.append("Source reports an update")
return labels
This example assumes valid timezone-aware timestamps. A production collector also needs to handle parsing failures, future dates, unavailable pages, and source-specific refresh semantics. Recording missing information explicitly is more useful than substituting the ingestion timestamp into every empty field.
What this changes for a developer looking for work
The practical advantage of one window is a shorter route from collection to a shortlist. HireSeeker offers profession selection for backend, frontend, mobile and game development, systems engineering, QA, analytics, ML and AI, alongside design, product and project management, marketing, and HR.
After selecting a profession, the public interface exposes salary, source, language, and period filters. Favorites support a shortlist; a daily selection supports a regular review instead of continuous monitoring of job boards and Telegram feeds. Vacancy cards also offer resume adaptation and cover-letter generation, alongside links to the original listing.
For a backend or QA search, a workable sequence is to select the profession, narrow the feed, save suitable roles, and open the originals before applying. The aggregation and filtering do the repetitive collection work. The original listing supplies the details needed for the application.
The strongest design choice here is making a scattered search manageable in one place. Keeping provenance and freshness visible makes that convenience useful: readers can distinguish a newly discovered vacancy, an updated listing, and another copy of a role they have already reviewed.
Top comments (0)