DEV Community

Cover image for My alert fired 7 times across 1963 pairs. I assumed the thresholds were too high. They were, and it didn't matter
Juan Camilo Auriti
Juan Camilo Auriti

Posted on AI-assisted

My alert fired 7 times across 1963 pairs. I assumed the thresholds were too high. They were, and it didn't matter

I shipped an alert that tells users when their score drops. Then I checked how often it had actually fired.

7 times, across 1963 consecutive-report pairs. A fire rate of 0.36%.

An alert that almost never fires and a system where almost nothing goes wrong produce exactly the same dashboard. There is no way to tell them apart by looking at the alert.

Where the thresholds came from

Two conditions, both had to hold:

total score dropped by >= 5 points
AND at least one category dropped by >= 3 points
Enter fullscreen mode Exit fullscreen mode

I would like to tell you those numbers came from an analysis. They came from me, in an afternoon, reasoning that a 5-point drop "feels significant" and that requiring a category drop too would "cut the noise."

Both halves of that sentence are assumptions, and I had shipped them as configuration.

The backtest

The check is cheap and I should have run it before shipping, not four months after. Replay every historical pair against candidate thresholds and look at the distribution:

WITH pairs AS (
  SELECT
    domain_id,
    score,
    LAG(score) OVER (PARTITION BY domain_id ORDER BY created_at) AS prev
  FROM reports
)
SELECT
  count(*)                                       AS pairs,
  count(*) FILTER (WHERE prev - score >= 1)      AS drop_1,
  count(*) FILTER (WHERE prev - score >= 3)      AS drop_3,
  count(*) FILTER (WHERE prev - score >= 5)      AS drop_5,
  round(100.0 * count(*) FILTER (WHERE prev - score >= 1) / count(*), 1) AS pct_any_drop
FROM pairs
WHERE prev IS NOT NULL;
Enter fullscreen mode Exit fullscreen mode

The LAG window is the whole trick: it puts each report next to its predecessor for the same domain, which is the unit the alert actually reasons about. Without PARTITION BY domain_id you compare one customer's report to another's, which produces a beautiful and completely fictional distribution. I did that first.

What it said

The thresholds were too strict. Dropping the AND to an OR, or the 5 to a 3, would have multiplied the fire count several times over. I was right that they were wrong.

Then I read the first column.

The score got worse in only 6.2% of all pairs.

That's the ceiling. Not 6.2% of pairs crossed my threshold — 6.2% of pairs moved in the direction the alert cares about at all, by any amount, including one point. An alert that fires on every single degradation, no matter how small, tops out at 6.2%.

So the tuning exercise I was about to do had a maximum payoff of going from 0.36% to 6.2%, and everything between those numbers is noise I'd be forwarding to users' inboxes.

The thresholds weren't the problem. The signal wasn't there.

Which is a product finding wearing engineering clothes

Once you say it plainly it's obvious: the thing I was monitoring mostly doesn't change. Sites don't degrade weekly. Someone who fixes their markup and leaves it alone will produce a flat line forever, and a flat line is the correct output.

That reframes what to build. Not a better threshold on a signal that's absent 93.8% of the time. Either:

  • Monitor something that moves. Competitor position, citation presence, the arrival of a new crawler in the logs — things with real week-to-week variance.
  • Change what the artifact says when nothing happened. "No change" is a legitimate and even reassuring answer, but only if the product is honest that stability is the expected state rather than shipping an empty-looking report and hoping.

Both of those are bigger than the config change I'd been about to make, which is exactly why the backtest was worth running. It talked me out of a week of threshold tuning that would have produced nothing.

The check, generalised

Before you tune an alert, measure two numbers:

  1. Fire rate — what fraction of evaluations fire today.
  2. Ceiling — what fraction of evaluations move in the monitored direction at all, at any magnitude.

The gap between them is your entire tuning budget. If the ceiling is small, tuning is not the work, and no amount of threshold adjustment will make a signal appear that isn't in the data.

The reason to compute the ceiling first is that it's the number that can tell you to stop. Fire rate alone always says "tune me" — it's low, it looks fixable, and it will absorb as much time as you give it.

Both numbers come out of one query. Mine took ten minutes to write and saved a week.

Earlier in this series: the trailing slash that ate 86.9% of this same database, and what a model does when your acronym turns out to be ambiguous. I wrote about the monitoring problem this alert was meant to solve in Why Your Brand Is Invisible to AI Search — this is the part where I checked whether the thing I built actually worked.

Top comments (0)