Notifio decides what counts as a new rental listing with the simplest mechanism available: it keeps a snapshot of what each saved search's results page held last time, scrapes it again, and treats anything not in the snapshot as new.
That works until you remember what a results page actually is. A lazily loaded grid, on somebody else's server, over a domestic connection, read every half minute for hours. It does not return the same thing every time, and not because anything changed.
Two guards exist for the two ways that goes wrong, and the interesting thing about both is that they do not end in a refusal. They end in the app admitting the page is allowed to change.
Too few results: do not diff, and do not write
First failure mode: the page came back with a fraction of its listings. The grid had not finished, a request was dropped, the network stalled at the wrong moment.
// Guard: a sudden large drop in the listing count almost always means the
// page only partially loaded this cycle (lazy grids, slow network). Don't
// diff or overwrite the baseline, because otherwise the next full load looks like
// a flood of "new" listings and gets suppressed by the ratio guard below.
if (previous.length >= 4 && current.length < previous.length * 0.5) {
Both halves of that comment matter, and the second one is the one that bites.
Not diffing is obvious: a half loaded page has no new listings in it, it has missing old ones, so diffing it produces nothing useful. Not writing is the subtle part. If the half loaded page became the new baseline, then the next complete load would be mostly listings that are not in the baseline, which is to say an email about 18 rooms the user has already seen. One dropped request would become one wrong email, reliably.
The two conditions are both deliberately blunt. previous.length >= 4 keeps the guard away from genuinely small searches, where going from 3 results to 1 is ordinary. The 50% threshold is nowhere near either case it has to separate: a partial load is usually a small fraction of the page, and a real page does not quietly lose half its content between two checks half a minute apart.
Except when it does, three times in a row
A guard that only ever refuses is a guard that breaks a search permanently the first time it is wrong.
Pages do shrink for real reasons. A landlord removes six listings. The site changes what its default filter includes. A search that used to match 20 rooms now matches 8, forever. If "lost half its listings" always meant "ignore this", that search would never be looked at again, and the app would go quiet for the worst possible reason: it was being careful.
const skips = (_partialSkips[site.url] ?? 0) + 1;
_partialSkips[site.url] = skips;
if (skips < PARTIAL_SKIP_LIMIT) {
log(`Skipping ${site.name}: only ${current.length} listing(s) vs ${previous.length} in baseline (likely partial load)`);
siteState.markFailure(site.url, 'error', {
message: 'Page looked half-loaded',
durationMs: outcome.durationMs,
});
return;
}
// Three checks in a row have agreed, so this is the page now. Accept it
// as the baseline rather than refusing to look at this search forever.
saveSnapshot(site.url, current);
delete _partialSkips[site.url];
PARTIAL_SKIP_LIMIT is 3. Three consecutive checks, roughly a minute and a half of the page insisting, and the app changes its mind. The counter is per search and is cleared on any normal check, so three agreeing checks is a real condition and not three unlucky ones spread over an afternoon.
The log line it writes when it concedes is deliberately specific about what it has concluded:
Funda has settled at 8 listing(s), down from 20. Using that as the new baseline
Too much new: complete, but not the same page
The second failure mode looks like the opposite and comes from the same cause. The page is full, nothing is missing, and almost everything on it is unrecognised.
// Guard: if most of the results are "new", it's likely a page-load anomaly
// (e.g. session dropped, site showed different view) rather than real new listings.
// Only alert if fewer than 80% of current results appear new.
const newRatio = newListings.length / current.length;
if (newRatio >= 0.8 && current.length > 3) {
In practice this is a session that quietly lapsed, so the site served a logged out or regional variant of the search, or a layout change that moved the results into a different shape. Either way the diff is technically correct and completely useless: 22 of 24 listings are "new" because we are looking at a different page, not a busier one.
The threshold is 80% rather than 100% because a real burst exists. Student sites in late August genuinely turn over most of a page in an hour, and a search narrow enough to hold four rooms can legitimately replace three of them. The current.length > 3 condition is what protects those, since on a small page a high ratio means very little.
This guard re-bases immediately rather than counting to three, and the asymmetry is the point. A page with too few results is probably mid load, so waiting is cheap and likely to resolve it. A page that is complete but unfamiliar is not going to become familiar by looking again, so the useful move is to adopt it and have the next cycle be correct.
Neither guard is allowed to be silent
A skipped cycle that says nothing is indistinguishable from a working cycle that found nothing, which is the failure this product can least afford. So each outcome writes a sentence into the per search state the window reads:
Page looked half-loadedRe-based after a smaller pagePage changed shape, re-basedBaseline saved
Those strings live in the runtime state the monitor owns rather than in the renderer, so a user who opens the window after lunch sees why a search has been quiet instead of inferring it.
This is also the point in the cycle where the guards hand over to a different set of rules. Nothing a guard re-bases produces an email, and when there is an email to send, the snapshot is not written until the email has actually been delivered. Both guards exist so that the thing being delivered is worth delivering, and the desktop banner goes first because it costs no network round trip at all.
The underlying principle is one I keep relearning in diff-based products: a comparison against stored state has three outcomes, not two. There is "nothing new", there is "here is what is new", and there is "this does not look like the thing I stored", and the third one needs its own branch or it will be reported as one of the first two.
What this looks like from outside
The honest summary of this post is that a user never sees any of it working, and would very much notice it missing, either as an email about rooms they already know about or as a search that quietly stopped reporting.
Portals differ a lot in how their results pages load, which is why each has its own page rather than one template with the name swapped in: notifio.app/alerts/funda, notifio.app/alerts/spareroom and notifio.app/alerts/huurwoningen each describe their own rendering behaviour.
If you want the product reasoning rather than the implementation, notifio.app/guides/how-fast-do-rental-listings-go is why a minute and a half of caution is a real cost, notifio.app/help covers setup, and the app is at notifio.app/download.
Top comments (0)