DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our 30 second poll is measured start to start, and a cycle that overruns goes again in five seconds

Notifio is a desktop app that watches rental search pages and emails you when a listing appears that was not there on the previous check. The number on the marketing page is thirty seconds, and it is on the pricing page and on every one of the per site pages too, so it is a number we have to actually mean.

The interesting part is that "every thirty seconds" is ambiguous, and the obvious reading of it is the wrong one.

Two readings of an interval

A poll loop can mean either of these:

  1. Wait thirty seconds, then do the work.
  2. Start a cycle every thirty seconds.

Reading one is what you get for free. You write the cycle, you await it, you sleep at the end, you recurse. It is also a lie about the cadence, because the real gap between two checks is thirty seconds plus however long the cycle took. Notifio scrapes each saved search one after another in a real browser, which is somewhere between three and five seconds each. A user watching five searches is on a forty five second cadence under reading one, not thirty. A user watching fifteen is on a cadence of over a minute and nobody told them.

Reading two is what we ship. The cycle's own duration is subtracted from the interval:

const cycleStartedAt = Date.now();
await poll();
if (!_running) return;

// Subtract the cycle's own duration from the interval so the cadence is
// measured start-to-start. A cycle that already took longer than the
// interval goes again after MIN_GAP_MS instead of waiting a full interval.
const elapsed = Date.now() - cycleStartedAt;
const target = randomDelay(POLL_INTERVAL_MS, POLL_JITTER_MS);
const delay = Math.max(MIN_GAP_MS, target - elapsed);
Enter fullscreen mode Exit fullscreen mode

So a cycle that takes four seconds sleeps twenty six, and a cycle that takes two seconds sleeps twenty eight. The user gets the cadence on the box regardless of how many searches they added, right up to the point where the work no longer fits inside the interval.

When the work does not fit

Fifteen searches at four seconds each is sixty seconds of scraping inside a thirty second interval. target - elapsed goes negative, and this is the branch where the naive loop does the most damage: it would sleep a further thirty seconds on top of a cycle that already ran long, so the searches at the end of the list would be checked once every ninety seconds.

Math.max with a floor handles it, and the floor is deliberately tiny:

/**
 * Floor between cycles, even when one overran the interval. Back-to-back with
 * no gap at all would look mechanical to the sites we scrape and leaves no room
 * to hit Stop.
 */
const MIN_GAP_MS = 5_000;
Enter fullscreen mode Exit fullscreen mode

Five seconds, not zero. Two separate reasons are packed into that comment and only one of them is about the remote site. The other is that stop() clears a pending setTimeout, so if there is never a moment between cycles where the loop is parked in a timer, the Stop button in the UI has nothing to interrupt and the user watches a cycle grind to its end after they asked it to quit.

The overrun case also gets its own log line rather than being silently absorbed, because a user who added fourteen searches and quietly fell off the advertised cadence deserves to be able to see that in the activity log:

if (delay === MIN_GAP_MS && elapsed > target) {
  log(
    `Cycle took ${(elapsed / 1000).toFixed(1)}s, over the ${(target / 1000).toFixed(0)}s interval. Next poll in ${(delay / 1000).toFixed(1)}s`
  );
}
Enter fullscreen mode Exit fullscreen mode

That log line is the reason the search limit is fifteen and not fifty. I wrote about that number on its own in Our 15 search limit is a latency budget, not a pricing tier; this post is the other half of it, the part that decides what happens when you are at the limit rather than what the limit is.

The interval is not exactly thirty

const POLL_INTERVAL_MS = 30_000;
const POLL_JITTER_MS = 5_000;

function randomDelay(baseMs: number, jitterMs: number): number {
  return baseMs + Math.round((Math.random() * 2 - 1) * jitterMs);
}
Enter fullscreen mode Exit fullscreen mode

Twenty five to thirty five seconds, uniformly. The jitter is not a politeness gesture, it is load shaping. Without it, every copy of the app that happened to be started at the same moment, say by a login item after a reboot, stays phase locked forever and arrives at the same site in the same second for as long as both are running.

Note that the jitter is recomputed every cycle rather than once at startup. A fixed per-install offset would stop two machines colliding but would leave each one perfectly periodic, which is its own signature.

One blip is not a reason to stand down

The counterpart to a tight cadence is being reluctant to abandon it. Notifio has a per host backoff registry for sites that refuse us, but a scrape can also just fail: the renderer crashes, wifi drops for a second, DNS hiccups. Treating that as a refusal costs real minutes:

/**
 * Consecutive hard errors on a host before it is made to wait.
 *
 * One failed scrape is usually a blip (a renderer crash, a dropped wifi
 * connection) and the next cycle recovers, so backing off immediately would
 * turn a two-second hiccup into minutes of not looking.
 */
const ERROR_GRACE = 2;
Enter fullscreen mode Exit fullscreen mode

Two consecutive errors on the same host before any waiting starts. The first one is logged and ignored. Since the next cycle is roughly thirty seconds away, the cost of being wrong about a blip is one missed check, and the cost of being right is not thirty seconds of silence.

Check now is the same function

There is a "Check now" button in the app, and the temptation with a button like that is to write a second code path for it. Notifio does not, because a second path is a second set of bugs. start() assigns its own tick to a module level handle, and the button calls it:

let _tick: (() => Promise<void>) | null = null;
Enter fullscreen mode Exit fullscreen mode
/** Set by start(), so an on-demand check can run the same cycle the timer does. */
Enter fullscreen mode Exit fullscreen mode

The guard that makes this safe is one boolean:

async function tick(): Promise<void> {
  if (!_running || _polling) return;
  _polling = true;
  try {
    // ...
  } finally {
    _polling = false;
  }
}
Enter fullscreen mode Exit fullscreen mode

A manual check that lands while a cycle is in flight is a no-op rather than a second concurrent cycle, and the same flag is exposed so the UI can grey the button out instead of letting the user press something that silently does nothing:

/** True while a cycle is actually in flight, so the UI can disable "Check now". */
export function isPolling(): boolean {
  return _polling;
}
Enter fullscreen mode Exit fullscreen mode

A manual check also reschedules the timer, so pressing it does not mean the next automatic cycle arrives one second later.

Why any of this matters commercially

Every portal Notifio competes with offers alerts of its own, and several of them sell the fast tier as an upgrade. The structural thing to notice is that a portal alert is a chain of scheduled work: a listing is published, an indexer or matcher job picks it up, a notification job builds the message, an email service provider accepts it, a send queue drains, the recipient's mail server accepts it, the mail client polls. Every hop is batch work running at a volume of hundreds of thousands of messages a day. "Immediate" describes when the portal decided to notify you, not when you could read it.

Notifio has no chain. It loads the search results page from the user's own machine on the cadence above and the diff happens locally. That is the whole argument, and the reason the start-to-start measurement is worth the eight lines it costs.

You can see the shape of it per site: Kamernet and Funda batch their free alert to once a day, Rightmove fires on publication or a 2% reduction, and Idealista sells immediacy. Each of those pages cites the portal's own documentation so you can check it rather than take our word for it.

The app itself is on the download page for macOS and Windows, and the help page covers what to do when a session expires, which is the failure mode this cadence runs into most often.

Top comments (0)