Static thresholds are the smoke detectors of data engineering: useful, but they only go off
when something's already on fire, and they miss the slow leaks entirely. "Alert if rows < 1000"
won't catch a table that quietly dropped 30%, or a feed that's complete in total but broken for
one segment. Anomaly detection asks a better question: is today normal for this thing?
Why "normal" needs a baseline, not a number
Data has shape — weekday/weekend rhythms, month-end spikes, seasonal trends. A good baseline
captures that: compare today against a rolling window of comparable days (same day-of-week),
not a flat constant. Then flag deviations relative to that baseline. Suddenly a Wednesday that's
35% below recent Wednesdays stands out, even though it clears your old static floor.
The three signals worth watching
- Volume — row counts per load vs baseline. Catches partial feeds and double loads.
- Freshness — is the newest record from the period you expect? Catches stalled upstreams.
- Distribution — did the shape of a key column shift (null rate, category mix, a metric's mean)? Catches silent logic bugs that keep volume normal but change meaning.
The lesson that matters most: check per segment, not just totals
This is the subtle one. A feed can match perfectly in aggregate while being badly wrong
underneath, because per-segment errors cancel out. One store is doubled, another is missing,
and the fleet total looks fine. If you only monitor the grand total, you'll certify broken data
as healthy. Run your anomaly checks at the grain that matters — per segment, per channel,
per role — and the offsetting errors stop hiding.
Keep it explainable and quiet
Fancy models are tempting, but an anomaly alert nobody understands gets muted, and a muted
alert is worthless. Favor methods whose output you can explain in a sentence ("28% below the
trailing 4-week median for this segment"). Tune aggressively against false positives — alert
fatigue kills these systems faster than missed anomalies do. Add seasonality and ML only where
a simple baseline provably isn't enough.
Takeaways
- Replace flat thresholds with baselines that respect weekday and seasonal shape.
- Watch volume, freshness, and distribution — the last one catches silent meaning changes.
- Detect at the real grain, not the total, so offsetting per-segment errors can't hide.
- Explainable and low-false-positive beats clever-but-muted every time.
Top comments (0)