DEV Community

Daniel Pertu
Daniel Pertu

Posted on

A host coming back from a block gets one search, not five, and the other four follow in the same cycle

Notifio checks a user's saved rental searches on a loop and emails them about new listings. Several of those searches are usually on the same portal: three rooms in different Amsterdam districts, two price bands on the same site.

That detail is the reason this post exists. A poll cycle is not a list of independent jobs. It is a list of requests to a handful of hosts, and the hosts notice.

Being a polite poller covers the steady state. This is about the awkward moment: a site has refused us, the host-keyed ladder has stood it down, and the wait is now over. What do you send?

One request is enough to ask the question

The naive answer is "everything that was waiting", which sends five requests in a few seconds to a host that refused us a minute ago. That is indistinguishable from the behaviour that earned the block, so the first thing it probably does is earn another one, with a longer wait on the end of it.

So the cycle groups the due searches by host and sends exactly one for any host that is on its way back:

const due: Site[] = [];
const probing = new Set<string>();

for (const [host, list] of byHost) {
  const backoff = _backoff.get(host);
  if (!backoff) {
    due.push(...list);
    continue;
  }
  if (backoff.nextRetryAt <= now) {
    probing.add(host);
    due.push(list[0]);
    for (const site of list.slice(1)) siteState.setNextRetry(site.url, backoff.nextRetryAt);
    continue;
  }
  waitingCount += list.length;
  for (const site of list) siteState.setNextRetry(site.url, backoff.nextRetryAt);
}
Enter fullscreen mode Exit fullscreen mode

list[0] is the probe. If the site is still unhappy, one request finds that out, and the cost of being wrong is one request instead of five.

A clean probe does not have to wait for the next cycle

The obvious objection is latency. If the probe works and the other four searches wait for the next cycle, the product has made its own recovery slower in exchange for being careful, and this is a product where a few minutes is the difference between a viewing and a sorry-it-is-gone email.

So they do not wait. A probe that came back clean clears the host's backoff entry, and that is a fact the same cycle can read:

// A probe that came back clean means the host is serving us again, so the
// rest of its searches are checked now instead of waiting for the next cycle.
const recovered = new Set(
  Array.from(probing).filter((host) => _backoff.get(host) === undefined)
);
if (recovered.size > 0) {
  const held = enabledSites.filter(
    (site) => recovered.has(hostOf(site.url)) && !due.includes(site)
  );
  if (held.length > 0) {
    log(`Recovered ${recovered.size} site(s), checking ${held.length} held search(es)`);
    const secondPass = await scrapeAll(shuffled(held), log, hooks);
    // ...
  }
}
Enter fullscreen mode Exit fullscreen mode

Two passes in one cycle, and the second one only exists when the first one earned it. The sequence is "ask once, then proceed", which is the same shape as a connection pool's health check and for the same reason: the expensive thing is not the request, it is being wrong about whether you are welcome.

The order of the searches is itself a pattern

A detail that has nothing to do with recovery and everything to do with not needing it:

// Vary the order. Checking the same searches in the same sequence every 30
// seconds for hours is a pattern in its own right, and it also means the
// search at the bottom of the list is always the last to hear any news.
const results = await scrapeAll(shuffled(due), log, hooks);
Enter fullscreen mode Exit fullscreen mode

Two unrelated benefits from one line. Anything that repeats identically for hours is a signature, and a stable order also quietly makes the last search in the list permanently worse than the first one. Shuffling fixes a fairness bug and a fingerprinting problem at the same time.

The snapshot waits for the email

This is the ordering decision I would defend hardest, and it is three lines.

Each search is compared against a stored snapshot of what its page held last time, and anything not in that snapshot is new. When the snapshot is allowed to be thrown away is its own question. This is about when it is allowed to be written.

The tempting place to write it is immediately after the scrape, while the data is in hand. That is also the place that loses listings:

// Snapshots for sites that produced alerts are persisted ONLY after the
// notification is delivered. Otherwise a failed send would record those
// listings as "seen" and they'd never be alerted again.
const pendingSnapshots: Array<{ url: string; listings: Listing[] }> = [];
Enter fullscreen mode Exit fullscreen mode
if (alerts.length > 0) {
  try {
    await sendListingAlert(alerts, config);
    // Delivery confirmed, so now it is safe to record these listings as seen.
    for (const snap of pendingSnapshots) saveSnapshot(snap.url, snap.listings);
  } catch (err) {
    const message = err instanceof Error ? err.message : String(err);
    log(`[notify] Failed to send: ${message}. Keeping listings unseen, will retry next poll`);
  }
}
Enter fullscreen mode Exit fullscreen mode

Write the snapshot first and a failed send is permanent: the listings are marked seen, the next cycle finds nothing new, and the user never hears about rooms the app definitely saw. Write it after and a failed send costs nothing, because the next cycle finds the same listings new again and tries again.

The failure this ordering chooses instead is a duplicate email, which happens if the send succeeds and the snapshot write then fails. Comparing the two is not close. A user who gets the same room twice is mildly annoyed; a user who never hears about it has been let down in the only way this product can let someone down.

The general rule, for anything that notifies: do not record that you told someone until you have told them.

Replying comes last, and that is also deliberate

The optional auto-reply step runs after the alert has gone out, never before:

// Auto-replies run after the alert. The user is told about the listing either
// way, so a broken recipe or a blocking site still leaves them able to reply
// by hand. Failing to reply must never mean failing to notify.
Enter fullscreen mode Exit fullscreen mode

The reply phase can take a while on purpose, since each reply waits a randomised delay before submitting, and it gets its own status so the UI can say replying rather than leaving a stale polling on screen. If it were ordered first, every failure in the fancy feature would become a failure in the essential one.

One clock, kept in step

The last thing the cycle does is the least interesting and the most asked about by users:

// Keep every search's retry clock in step with its host. Only the search
// that was actually probed learns its own new wait during the cycle, and a
// countdown that has quietly expired without anything happening is worse
// than no countdown at all.
for (const site of enabledSites) {
  siteState.setNextRetry(site.url, _backoff.get(hostOf(site.url))?.nextRetryAt ?? null);
}
Enter fullscreen mode Exit fullscreen mode

The probe is the only search that learns the host's new wait first hand, so without this the other four would show a countdown that hit zero and then sat there. The live state the window reads is owned by the monitor precisely so that this kind of correction happens in one place.

What this looks like from outside

A user sees none of the above. What they see is a row per search with a state and, when a site is standing us down, a countdown that is the same on every search for that site.

The per site pages describe the behaviour that makes this necessary, because it differs a lot by portal: notifio.app/alerts/idealista, notifio.app/alerts/rightmove and notifio.app/alerts/funda each have their own notes. The cap of 15 saved searches comes out of the same cycle budget this post is about, since every search is checked in turn inside one cycle. Setup and troubleshooting are at notifio.app/help, and the app is at notifio.app/download.

Top comments (0)