Notifio watches rental search pages and alerts you when a new listing appears. One function turns a single page read into everything downstream: the alert, the desktop banner, the row in the window, and the stored snapshot that the next read will be compared against.
That last one is the dangerous one. The snapshot is the app's memory of what it has already seen, so writing it is a commit. Anything recorded in it will never be alerted on again, for as long as the search exists.
So the interesting shape of the function is not the happy path. It is eleven exits, of which five must not write anything, five write immediately, and one writes only after something else has succeeded.
The five that must not write
if (outcome.blocked) { /* backoff, no write */ return; }
if (outcome.listings === null) { /* error ladder, no write */ return; }
if (outcome.authFailure) { /* login needed, no write */ return; }
if (current.length === 0) { /* no write */ return; }
if (partialLoad && skips < PARTIAL_SKIP_LIMIT) { /* no write */ return; }
Five different ways of reading nothing useful, and they are deliberately not collapsed into one branch, because the user-visible consequence of each is different: a wall needs a backoff ladder keyed to what kind of wall it was, a network error needs a grace count before it earns a wait at all, a login wall needs the user to go and log in, and a half-loaded page needs nothing except to be ignored.
What they share is the thing that would be a bug in every one of them. Suppose the login wall path wrote its snapshot. The page behind a login wall has no listings on it, so the stored memory of that search becomes empty. Then the user logs in, the next check reads forty rooms, and every one of them is new.
That does not produce forty alerts. It produces zero, because a result that is almost entirely new trips a separate guard that assumes the page changed shape rather than the market exploding. So the user's search quietly stops working: no error, no alert, nothing to notice. A wrong write turns into silence, which is the worst failure mode a monitoring tool has.
That is why the five exits are not an optimisation. They are the correctness of the feature.
The one that reverses itself
The partial-load guard is the only one that gives up:
if (previous.length >= 4 && current.length < previous.length * 0.5) {
const skips = (_partialSkips[site.url] ?? 0) + 1;
_partialSkips[site.url] = skips;
if (skips < PARTIAL_SKIP_LIMIT) {
// Likely a lazy grid that did not finish. Ignore and look again.
return;
}
// Three checks in a row have agreed, so this is the page now. Accept it
// as the baseline rather than refusing to look at this search forever.
saveSnapshot(site.url, current);
...
}
A results page that suddenly holds half as much as it did is almost always a lazy-loading grid that did not finish before we read it. Refusing to diff that is right. Refusing it forever is not, because sometimes a page really does shrink: the user narrowed their filters, or the portal changed its page size from 30 to 12.
Without the counter, that search is dead. Every check looks like a partial load, every check is skipped, and the row says "page looked half-loaded" for the rest of the licence. With it, three consecutive checks that agree are treated as the truth, and the baseline moves.
A guard that can never be convinced is not a guard, it is a stuck state. Any heuristic that refuses to act needs a path by which reality can override it.
The one that writes last
Then there is the exit that actually matters to the user, the one with new listings on it. It writes too, and it writes somewhere else first:
// Snapshots for sites that produced alerts are persisted ONLY after the
// notification is delivered. Otherwise a failed send would record those
// listings as "seen" and they'd never be alerted again.
const pendingSnapshots: Array<{ url: string; listings: Listing[] }> = [];
The write is queued, the alert is sent, and only a delivered alert releases the snapshot.
Reverse that order and consider the failure. The email provider is down for ninety seconds. The snapshot is already saved, with the three new rooms in it. The send fails, we log it, and the next cycle compares the page against a baseline that already contains those rooms. They are not new any more. They were never announced, and they never will be.
Nobody sees that bug in testing, because the write and the send both succeed on a laptop with a good connection. It only exists on the days it matters.
Ordering a durable write after the side effect it depends on is the whole trick, and the local channels are ordered for the same reason in the other direction:
// The local channel goes first. It costs no network round trip, so the
// banner is on screen before the email has left the building, and on a
// rental site the first few minutes are the whole difference.
Fastest channel first, durable commit last, and the network in between.
Why the empty page is treated as a failure
One of the five deserves its argument made out loud, because it is not obviously right.
if (current.length === 0) {
log(`Skipping diff for ${site.name}, the check returned 0 listings`);
siteState.markFailure(site.url, 'error', { message: 'No listings on the page' });
return;
}
An empty page is ambiguous. It might be a block we failed to classify, a page that did not render, or a search that genuinely has no results right now: filters set tight, a small city, a budget under the market.
We treat it as a failure in all three cases. For the genuine case that is a false alarm, and the user sees a row saying "no listings on the page" for a search that is simply quiet.
The trade is not close. A false alarm is visible and self-correcting: the user looks at the row, opens the URL, sees an empty results page, and either widens the search or ignores it. The other error writes an empty baseline for a working search and loses rooms silently. Visible wrong beats invisible wrong, every time, in a tool whose only job is to tell you something happened.
Where this sits
The sibling post to this one is most of our diff code exists to not send an alert, which covers the same function from the alerting side: the ratio guard, and the rule that re-bases a search on a fresh app start. What this one adds is the other output. An alert is a message, and a missed one costs you a room. A snapshot is memory, and a wrong one costs you every room on that search from then on, which is why eleven exits is the right number and not a smell.
The identity used in that comparison, which is narrower than the URL it came from, is in a listing id that ignores the query string. The wait after a refusal is in our retry ladder was keyed by the wrong thing.
If you want to watch the states this produces, every one of them is a row in the app window: notifio.app/download, with what a search URL needs to look like on notifio.app/help, and a per-portal note on login requirements and rendering quirks on pages like Funda and Rightmove.
Top comments (0)