Notifio is a desktop app that watches rental search result pages and tells you the moment a new listing appears on one. The whole product is a diff. It reads a search page, compares what is on it with what was on it last time, and the difference is the alert.
Which means the only interesting question in the codebase is what counts as "the same listing". Get that wrong in one direction and the user gets an email about a room they already replied to. Get it wrong in the other and a real room never reaches them.
A scraped page does not hand you a primary key. It hands you anchors. So the identity has to be derived from the URL, and the obvious choice, the href as written, is the one that does not work.
Three anchors, one room
Start with the smaller problem, because it is the one that shows up first. A listing card on a rental portal is usually wrapped in more than one link: the photo is a link, the heading is a link, and there is often a "view" or "respond" button that is also a link. Same room, three anchors, sometimes with slightly different URLs on the end of them.
Then across cycles the same room comes back with a different query string. Portals append things to their own internal links: a position index so they can tell which card in the grid you clicked, a tracking parameter, a session reference, a from=search flag. None of it identifies the listing. All of it changes.
And then there is the trailing slash, which is the stupidest version of the same bug and the easiest to ship. /for-rent/room-amsterdam-1234 and /for-rent/room-amsterdam-1234/ are one room and two strings. A portal that changes which one it emits, or that emits both depending on where the link sits in the page, turns an entire search into new listings at once.
The id is the path, the link is not
So the id is normalised and the link is not:
const urlObj = new URL(fullUrl);
// Strip the query string and the trailing slash for deduplication.
const normalised = urlObj.origin + urlObj.pathname.replace(/\/$/, '');
if (seen.has(normalised)) continue;
// ... filters ...
seen.add(normalised);
results.push({
id: normalised, // stable, path-only identifier
title: label.substring(0, 120),
url: fullUrl, // full URL, query params included, for the email link
});
Two fields that both look like the URL, and they are deliberately different values.
id is for machinery. It is the key in the previous cycle's snapshot, the dedupe key for the anchors on one page, and the key in the on-disk history of recent finds. It needs to be the narrowest string that still names the room.
url is for a human. It goes in the alert email and in the desktop notification, and clicking it should land you exactly where the portal expects you to land. Some portals genuinely use their own parameters on a detail page, for a back link to your search, or to keep a filter applied. Stripping them to make a tidy email is optimising the wrong thing: the point of the alert is the click.
The same split shows up in the seen set, which is per page rather than per session. Three anchors pointing at one room collapse to one entry before any comparison happens, so the diff never sees the duplicate and the user never gets three notifications for one room.
What this deliberately cannot handle
A path-only id assumes the portal puts the listing identity in the path. That is true for every site in the catalogue of portals we write setup pages for: Kamernet and Pararius in the Netherlands, Rightmove and SpareRoom in the UK, Idealista and WG-Gesucht further out. Listing URLs on all of them carry the id as a path segment, because they want those pages indexed and a crawlable URL is a path.
A site that served listings as /detail?id=1234 would break this, and it would break it in the worst available way: every listing on the page would normalise to the same /detail string, the dedupe would collapse them into one, and the app would quietly report a single listing forever. Not an error, just a wrong answer.
That is a real constraint rather than an oversight, and the honest version of it is that the guard which catches it is somewhere else: a path has to be at least two segments deep to be treated as a listing at all, so a bare /detail never enters the set. The app would find nothing on such a site rather than finding one thing. Nothing is a visible failure. One is not.
The adjacent rule, in a different file
Identity is half the question. The other half is whether a difference should produce an alert at all, and that answer lives in a separate module, because the first check of a search after the app starts is compared against nothing: it re-bases instead. I wrote about that separately in most of our diff code exists to not send an alert.
What this post adds to it is the layer underneath. A re-basing rule and a too-much-changed guard both assume the set they are comparing is keyed by something stable. If the key moves, every guard above it is reasoning about noise.
You can see the output of all of this on the download page, and the setup side of it, including what a search URL has to look like before the app will take it, on the help page.
Top comments (0)