DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Two files hold the same listings, and only one of them survives you editing a search

Notifio watches rental search pages from your own machine and tells you when a listing appears that was not there last time. For a long while a find existed in exactly one place: an email.

That is a strange thing for a desktop app to do. The app knew about the room first, and then had nothing to show for it. If you missed the email, or the send failed, or you just wanted to see what turned up overnight, the app's answer was a line in the activity log saying 3 new listing(s) on Kamernet - Leiden and the links were somewhere in a mailbox.

Fixing that looked like a small feature. It turned into a question about lifetimes, because the app already stored every listing it had ever seen.

Why the existing store could not be reused

Diffing needs a record of what a page held last time. Notifio keeps one snapshot per search, keyed by the search URL, and compares against it. It is the mechanism the entire product rests on, and the rules around it are stricter than you would guess: the first check after the app starts deliberately replaces the baseline rather than comparing with it, for reasons I wrote up in The first check after a restart is not allowed to tell you anything.

So the listings are already on disk. Reading history out of the snapshots would mean no new storage at all.

It is the wrong store, and the reason is in two lines of the tracker that manages those baselines:

/**
 * Forget a search, so its next check counts as the first again.
 *
 * Used when a search is removed or re-pointed at a different URL, which
 * throws its stored baseline away too.
 */
forget(url: string): void {
  this.checked.delete(url);
}
Enter fullscreen mode Exit fullscreen mode

A baseline is scoped to a search, and it is meaningless the moment that search changes. If you edit a search from Leiden to Utrecht, the old snapshot is not stale data to be migrated, it is an answer to a question nobody is asking any more, and keeping it would mean the next check diffs Utrecht listings against Leiden ones and reports the whole page as new.

Which means the moment a baseline gets deleted is a search being re-pointed or removed. And that is precisely the moment a user least wants their history of finds to disappear. You change a search because the old one was not working; the rooms it did find are the only thing you have to show for the week.

Two stores, then, with the duplication written down so nobody tidies it away:

/**
 * This is a small on-disk ring buffer of the most recent finds, written after
 * the alert goes out. It is separate from the per-search snapshots on purpose:
 * those are a diffing baseline keyed by URL and get deleted whenever a search
 * is re-pointed, which is exactly when the user least wants their history to
 * vanish.
 */
Enter fullscreen mode Exit fullscreen mode

The general shape, which I keep rediscovering: state that exists to answer "what should happen next" has a short life, and state that exists to answer "what happened" has a long one. Putting both in one store means one of the two questions gets the wrong answer.

The buffer

export interface Find {
  /** Listing id (origin + path), unique per listing. */
  id: string;
  siteName: string;
  /** The search URL this came from, so the UI can group by search. */
  siteUrl: string;
  title: string;
  url: string;
  foundAt: number;
}

/**
 * How many finds to keep. A busy hunt produces a few hundred a week; this is
 * enough to cover "what did I miss overnight" without the file growing forever.
 */
const MAX_FINDS = 200;
Enter fullscreen mode Exit fullscreen mode

siteUrl is kept even though the search it points at may since have been deleted. That is the opposite decision from the baseline store, and it is the whole reason this file exists: the history records where a listing came from as a fact about the past, not as a live reference to something that must still exist.

The write is the same pattern as everywhere else in the app:

function write(finds: Find[]): void {
  _cache = finds;
  try {
    const tmp = `${FINDS_PATH}.tmp`;
    fs.writeFileSync(tmp, JSON.stringify(finds), 'utf8');
    fs.renameSync(tmp, FINDS_PATH);
  } catch (err) {
    console.error('[finds] Failed to save:', err);
  }
}
Enter fullscreen mode Exit fullscreen mode

Temp file, then rename, so a crash mid-write leaves the previous history intact rather than a truncated JSON file. And the failure is logged rather than thrown, which pairs with how the read behaves:

function read(): Find[] {
  if (_cache) return _cache;
  try {
    const parsed = JSON.parse(fs.readFileSync(FINDS_PATH, 'utf8'));
    _cache = Array.isArray(parsed) ? (parsed as Find[]) : [];
  } catch {
    // Missing or corrupt: an empty history is a fine place to start and must
    // never stop the monitor.
    _cache = [];
  }
  return _cache;
}
Enter fullscreen mode Exit fullscreen mode

The Array.isArray check is there because a corrupt file can parse successfully into the wrong shape. JSON.parse on a truncated array throws and is caught; JSON.parse on a file that somehow contains {} returns an object that would then have .filter called on it. Catching the throw is the easy half.

Each of the app's on-disk files makes its own decision about what to do with garbage, which I went through file by file in Six files, one write pattern, and six different answers to "what if this is garbage?". This one's answer is the most permissive in the app, and that is deliberate: the history is a convenience, and an unreadable convenience must never be able to take down the thing it is a convenience for.

It also lives in its own file rather than in the app's config:

// The recent-finds history shown in the app. Kept out of config.json so a long
// history can never make the file the monitor reads every poll any bigger.
export const FINDS_PATH = path.join(USER_DATA_DIR, 'finds.json');
Enter fullscreen mode Exit fullscreen mode

Config is read on every poll cycle. Two hundred listings of history in that file would mean parsing two hundred listings of history every thirty seconds to find out which searches are enabled.

The deduplication does two jobs

/**
 * Add newly found listings, newest first. Returns the ones actually added, so a
 * caller can notify about exactly those and nothing twice.
 */
export function add(items: Omit<Find, 'foundAt'>[]): Find[] {
  if (items.length === 0) return [];
  const existing = read();
  const seen = new Set(existing.map((f) => f.id));
  const foundAt = Date.now();
  const fresh = items.filter((item) => !seen.has(item.id)).map((item) => ({ ...item, foundAt }));
  if (fresh.length === 0) return [];
  write([...fresh, ...existing].slice(0, MAX_FINDS));
  _listeners.forEach((fn) => {
    try {
      fn(fresh);
    } catch {
      /* a broken listener must not lose the find */
    }
  });
  return fresh;
}
Enter fullscreen mode Exit fullscreen mode

The return value is the interesting part. add does not return a boolean or the new length, it returns exactly the subset that was not already known, and the call site uses that rather than its own input:

const fresh = finds.add(
  newListings.map((listing) => ({
    id: listing.id,
    siteName: site.name,
    siteUrl: site.url,
    title: listing.title,
    url: listing.url,
  }))
);
if (fresh.length > 0 && config.notifications?.desktop !== false) {
  notifyFinds(site.name, fresh, log);
}

Enter fullscreen mode Exit fullscreen mode

So the history's own identity check is also the guard against notifying twice about one listing. The diff upstream should already have filtered those out, and sometimes it does not: a listing can leave a page and come back, a search can be re-pointed at a URL that overlaps the old one, a baseline can get replaced. Every one of those produces a listing that is genuinely new to the diff and not new to the user.

I like this more than a separate "already notified" set would have been. Two data structures with overlapping responsibility for what counts as the same listing is how you get a bug where one of them says yes and the other says no. Here there is one answer, it lives in the store that persists across restarts, and the notification follows it.

Omit<Find, 'foundAt'> on the parameter is a small thing that keeps that honest: a caller cannot pass a timestamp, so it cannot pass a wrong one. The store stamps the time, because the store is the thing that knows when the record was made.

Order of operations

The three things that happen when a search produces a find are ordered, and the order is a product decision:

// The local channel goes first. It costs no network round trip, so the
// banner is on screen before the email has left the building, and on a
// rental site the first few minutes are the whole difference.
Enter fullscreen mode Exit fullscreen mode

History, then desktop notification, then the email, then the auto-reply queue:

// Auto-reply, if this search is set to it. Queued after the alert list is
// built so a failure here can never stop the email going out.
if (site.replyMode === 'auto') {
  autoReplyQueue.push({ site, listings: newListings });
}
// Defer persisting this snapshot until the alert is actually delivered.
pendingSnapshots.push({ url: site.url, listings: current });
Enter fullscreen mode Exit fullscreen mode

And the baseline snapshot goes last, after delivery. The two stores this post is about therefore get written at opposite ends of the cycle, for opposite reasons. The history is written immediately because a find the user cannot see is a find that did not happen. The baseline is written only once the alert is out, because a baseline saved before delivery and a delivery that then fails means the listing is recorded as seen and never reported to anybody.

Into the window without a refresh

The app's window and its monitor are separate processes, and the window reloads itself every time you open it, which I wrote about in Our window reloads itself every time you open it, so the renderer may not remember anything. So the history has to be both fetchable and pushable.

Fetchable is two routes:

app.get('/api/finds', (req: Request, res: Response) => {
  res.json({ finds: finds.list() });
});

app.delete('/api/finds', (req: Request, res: Response) => {
  finds.clear();
  res.json({ ok: true, finds: [] });
});
Enter fullscreen mode Exit fullscreen mode

Pushable is one line:

finds.onAdd((fresh) => broadcast({ type: 'finds', finds: fresh }));
Enter fullscreen mode Exit fullscreen mode

The subscription exists so the listing appears in the window at the same moment the desktop banner shows, rather than on the next poll or the next time somebody opens the app. The store does not know what a server is; it knows it has listeners, and the server is one.

The try/catch around each listener is the bit I would defend hardest:

_listeners.forEach((fn) => {
  try {
    fn(fresh);
  } catch {
    /* a broken listener must not lose the find */
  }
});
Enter fullscreen mode Exit fullscreen mode

The write to disk has already happened by the time listeners run. If a listener throws, the find is safe, and the only consequence is a row that appears a few seconds later when the window next asks. Letting that exception out of add would propagate into the poll loop, where its only available interpretation is "this search failed", which is the opposite of the truth.

What is still wrong with it

Three things, in descending order of how much they bother me.

Every find in one batch shares a timestamp:

const foundAt = Date.now();
Enter fullscreen mode Exit fullscreen mode

Five rooms from one cycle are all stamped identically, so their order within the batch is whatever order the listings came off the page. For a relative time display reading "2m ago" that is invisible. It would stop being invisible the day anything wants to sort finds by exact time, and the fix is to stamp per item.

clear() is all or nothing. There is no "clear the finds from this one search", which is the thing somebody will want the first time they re-point a search and would like the old city's rooms out of the way. The data supports it, since every find carries its siteUrl; the API does not.

And the React key in the list is belt and braces:

key={`${find.id}-${find.foundAt}`}
Enter fullscreen mode Exit fullscreen mode

find.id is unique in the array by construction, because add filters on exactly that. Compounding it with the timestamp admits a small doubt about that guarantee. Either the invariant holds and the key should be find.id, or it does not hold and the dedupe is wrong. A compound key papers over the question instead of answering it.

If you want to see the thing this is all in service of, notifio.app/download is the app, and How fast do rental listings actually disappear is the page explaining why the ordering above has a desktop banner ahead of an email. The sites it can watch, with a page each, are at notifio.app/alerts.

Top comments (0)