Notifio is a desktop app that polls rental search pages on your machine and alerts you when a listing appears that was not there on the previous check. It keeps six separate pieces of state, and over time each one has ended up with a different rule about how long it is allowed to live.
That was not a plan. Each rule got decided on its own, usually because the previous arrangement produced a bad outcome a user could see. Written down together, though, they turn out to follow one question, and it is not a storage question: what is this state a statement about? Answer that and the lifetime is forced.
(A companion post, Six files, one write pattern, and six different answers to "what if this is garbage?", takes the same set of files along the other axis: what each one does when the bytes on disk are corrupt. This one is about what is allowed to survive.)
The six
| State | Lives in | Survives restart | Survives a search being re-pointed |
|---|---|---|---|
| Searches, email, licence | config.json |
yes | n/a, it is the thing being edited |
| Page baseline per search | data/<hash>.json |
on disk yes, in effect no | no, deleted |
| Recent finds | finds.json |
yes | yes, deliberately |
| Reply ledger | ledger/replies.ndjson |
yes, with compaction | yes |
| Upgrade entitlement | entitlement.json |
yes, with a staleness flag | n/a |
| Live per search run state | nothing | no | no |
The one with no file
The last row is the interesting one. While the monitor runs, each search has a live state: is it being checked right now, did the last check succeed, was it walled, how long did it take, how many listings were on the page, when is it next due. The app shows all of that against each search.
None of it is written down:
/**
* Deliberately in memory only: it describes the current run, and a stale
* "checking..." restored from disk after a restart would be a lie.
*/
A file here would be very easy to add and would make the app worse in a way that is hard to debug later. "Checking..." is a claim about the present tense. Persist it and the first thing the user sees after a crash is a search that is permanently mid check, and a lastCheckedAt from yesterday presented as the current state of a run that is not happening.
The same reasoning decides what the stop button does:
export function stop(): void {
...
siteState.reset();
emitStatus('stopped');
}
Clear it, because the run it describes has ended.
There is a second consequence of this store being the authority. Because the renderer cannot own any of it (the window reloads itself every time it is re-shown, which is its own post), the same snapshot has to serve both the initial fetch and the live stream. Changes arrive in bursts, because one poll touches every search in turn, so they are coalesced:
// Changes arrive in bursts (a poll touches every search in turn). Emitting on
// each one would push a stream of near-identical snapshots at the renderer, so
// they are coalesced into one frame.
const EMIT_DEBOUNCE_MS = 120;
_emitTimer = setTimeout(() => { ... }, EMIT_DEBOUNCE_MS);
// Don't hold the process open for a pending UI update.
_emitTimer.unref?.();
That unref() is a small thing that matters in a desktop app specifically. A pending UI update is never a reason for the process to stay alive at exit.
The baseline: on disk, but not trusted across runs
Each search's baseline is a file named for a hash of the search URL:
function snapshotPath(url: string): string {
const hash = crypto.createHash('sha1').update(url).digest('hex').slice(0, 16);
return path.join(DATA_DIR, `${hash}.json`);
}
Keyed by URL, not by the search's position in the user's list. That is a deliberate reaction to a real bug: the local REST API addresses searches by index, and an off by one there meant an operation landing on the wrong search. The worst version of that story is parseInt returned NaN, and splice(NaN, 1) deleted the first search. Positions change when a list is edited; a URL is the identity of the thing being watched.
The file persists across restarts, but the first check of each search after the app starts throws it away and writes a fresh one rather than diffing against it, because a baseline from a previous run describes what the page held when the app was last open. That rule has its own post too. So the honest entry in the table is "on disk yes, in effect no".
One detail I have come to like. The app tells the user how long it was closed for, and that number is not stored anywhere:
/**
* How long ago this search's baseline was last written, in ms.
*
* Taken from the file's timestamp rather than stored inside it: the baseline is
* rewritten on every successful check, so its mtime is exactly "when the app
* last looked at this search", which is the question being asked.
*/
function baselineAge(url: string): number {
try {
return Date.now() - fs.statSync(snapshotPath(url)).mtimeMs;
} catch {
return Infinity;
}
}
A lastCheckedAt field would have been a second copy of a fact the filesystem already holds, with the usual consequence: one day it is written and the file is not, or the other way round. Infinity for a missing file also falls out correctly, because a search with no baseline has been waiting forever by definition.
The two stores that outlive the thing they describe
Finds and the reply ledger are both histories, and the rule for a history is that it must not be deleted by an action that was about something else.
That sounds obvious until you see how it goes wrong. Finds used to be derivable only from the baseline files, and re-pointing a search deletes its baseline:
/**
* This is a small on-disk ring buffer of the most recent finds, written after
* the alert goes out. It is separate from the per-search snapshots on purpose:
* those are a diffing baseline keyed by URL and get deleted whenever a search
* is re-pointed, which is exactly when the user least wants their history to
* vanish.
*/
const MAX_FINDS = 200;
Editing a search's URL is a thing people do when a search is working well and they want to narrow it. Doing that and watching this morning's finds disappear is a bug, and no amount of explaining the data model makes it not one.
The ledger has the same property for a stronger reason. It exists to answer two questions across restarts: has this listing already been messaged, and how many replies have gone out recently. Both answers have to be right after a crash, so it is append only and folded on read:
/**
* NDJSON rather than sqlite on purpose: no native module, so the existing
* electron-builder pipeline stays untouched. Updates are appended as a new row
* with the same id and readers fold last-wins, which keeps writes atomic by
* construction.
*/
for (const line of raw.split('\n')) {
const trimmed = line.trim();
if (!trimmed) continue;
try {
rows.push(JSON.parse(trimmed) as T);
} catch {
// A torn final line (power loss mid-append) must not poison the read.
}
}
Append only gives you crash safety for free: a row either made it or it did not, and the only thing a power cut can damage is the last line, which is skipped. The price is unbounded growth, paid in compaction:
const REPLIES_MAX_ROWS = 8000;
const REPLIES_KEEP = 2000;
(The file format choice has its own post, and the reason was the build pipeline rather than anything about databases.)
The ledger also has the only deliberately pessimistic rule in the set. The statuses that mean "never touch this listing again" include failed:
/**
* Statuses that mean "this listing has been dealt with, never touch it again".
* `failed` is deliberately included: a submit that errored may still have gone
* through, and messaging a landlord twice is worse than missing one.
*/
That trade is the opposite of what you would choose for an internal job queue, where a failed attempt should obviously be retried. It is the right way round here because the recipient is a person who will form an opinion of the sender, and we cannot ask them whether the first message arrived. I wrote that one up separately in Treating "failed" as terminal: idempotency when you cannot ask the recipient.
The rule that falls out
For each piece of state, finish this sentence: "this is a statement about ___".
- A statement about the user's intent (their searches, their email) belongs in a file they effectively own, and it survives everything.
- A statement about what happened (we found this, we sent that) belongs in a history, and nothing except compaction may remove it.
- A statement about the outside world (what a page held, whether the server says the upgrade is paid for) belongs on disk with an explicit rule about when it stops being believed.
- A statement about right now belongs in memory, and writing it down is not caution, it is a future lie with a timestamp on it.
The app is at notifio.app/download, the data it keeps and where it keeps it is described on notifio.app/help and notifio.app/privacy, and notifio.app/alerts lists the sites it can watch.
Top comments (0)