I looked at what my monitoring product had actually been sending people.
78% of reports showed a delta of exactly zero. Nothing had changed since the previous run. The report was correct, on time, and empty.
My first reading was that my change-detection thresholds were too coarse — that real movement was happening below the granularity I was measuring, and a finer scale would surface it. I spent a while on that. I wrote about the alert side of it earlier in this series, and about the backtest that talked me out of the tuning; this is the part I got wrong underneath it, which is different and worse.
The thresholds weren't the problem. The artifact was the problem. A report whose content is "no change" is not a report with a sensitivity issue. It's a report with nothing in it.
Why "no change" is the expected state
Once you say it out loud it's almost embarrassing.
Someone signs up, runs an audit, gets a list of problems, fixes some of them, and stops. That's the success path. Their site is now static in the dimensions I measure, because sites are static in those dimensions. robots.txt doesn't drift. Schema markup doesn't decay. A well-formed llms.txt stays well-formed until somebody edits it, and nobody edits it.
So the population my weekly report addresses is mostly people for whom the honest answer is "nothing happened, and that's fine." Which is not a message worth an email, and definitely not worth a dashboard visit.
The uncomfortable implication: a monitoring product for a slow-moving signal has no natural weekly cadence. I had assumed one because monitoring products have weekly reports. That's an industry convention, not a property of my data.
What I'd actually measure now
The metric I was missing isn't a threshold. It's a property of the output:
What fraction of generated artifacts contain at least one thing the recipient didn't already know?
Call it content rate. Mine was 22%. Every product decision downstream looks different once you know that number, and none of them are threshold changes:
Change the cadence to match the signal. Not weekly because it's Tuesday. Send when there's something, and say so — "we watched, nothing moved" is a fine thing to put in a monthly digest and a bad thing to put in a weekly email.
Add signals that actually move. The things I was measuring are properties of a site the owner controls, so they change only when the owner acts. Competitor position, whether a citation appeared or disappeared, a new crawler user-agent showing up in the logs — those move without the user doing anything, which is exactly what makes them worth telling someone.
Make stability an explicit product claim, or drop it. "Nothing changed" is genuinely valuable information if the product is framed around confirming that. It's dead weight if the product is framed around finding problems. I had the second framing and the first output.
The measurement, if you want to run it on yours
If you generate a recurring artifact — report, digest, alert summary, weekly email — this is the query. Mine is scores, yours is whatever the artifact is about:
WITH ranked AS (
SELECT
domain_id,
score,
LAG(score) OVER (PARTITION BY domain_id ORDER BY created_at) AS prev
FROM reports
)
SELECT
count(*) AS reports,
count(*) FILTER (WHERE prev IS NOT DISTINCT FROM score) AS unchanged,
round(100.0 * count(*) FILTER (WHERE prev IS NOT DISTINCT FROM score) / count(*), 1)
AS pct_unchanged
FROM ranked
WHERE prev IS NOT NULL;
IS NOT DISTINCT FROM rather than = is the detail that matters: NULL = NULL is NULL, which is not true, so a plain = silently drops every pair where a score was missing — and missing scores cluster exactly where something odd is happening. IS NOT DISTINCT FROM treats two NULLs as equal and gives you the honest count.
Then read the number as a statement about your product rather than your thresholds. If most of your artifacts are empty, no amount of sensitivity tuning fills them.
The part I'd tell myself in advance
I spent weeks on threshold work because threshold work feels like the right kind of effort. It's measurable, it's in code, it produces diffs, and at the end of a day of it you can point at something.
Asking "should this report exist at all" produces no diff. It's also the question that was load-bearing, and I avoided it for a month by staying busy adjacent to it.
The tell, in hindsight: I was tuning the sensitivity of a detector without ever having measured how often the thing it detects occurs. That ordering is backwards, and it's a cheap check — one query, before the tuning, not after.
Top comments (0)