Alerts that fire twice for the same event are worse than alerts that never fire at all. The first duplicate teaches you the system is noisy; the tenth teaches you to swipe notifications away without reading them. Once that habit forms, a real alert gets the same treatment.
The usual cause is not a broken alert rule. It is retries and overlapping schedulers. A cron job restarts mid-run, the previous process is still alive, both evaluate the same condition, both push. Or a task queue retries after a timeout while the original request actually succeeded. Or you run the same check from two machines during a deploy. Each path is individually correct, and together they produce two notifications for one event.
Deduplication belongs at the push boundary, not in the detection logic. Detection rules are where you want redundancy; the notification layer is where you want exactly-once behavior. This post covers a small, dependency-free dedup cache that enforces a five-minute window keyed on the notification title, with an explicit bypass for signals that must never be suppressed.
The dedup contract
Keep the rule short enough to state in one sentence:
Before pushing a notification, look up its normalized title in a JSON cache. If the same title was pushed within the last five minutes, skip the push and log the suppression. Otherwise, push and record the timestamp.
Two properties matter:
- Idempotency by title. Two calls with the same title inside the window produce one push.
- Fail-open. If the cache is unreadable or corrupt, push anyway. A duplicate is annoying; a swallowed alert is a bug you find out about too late.
The cache is deliberately tiny — one file, one dict, no database. This is not an event log. It is a short-lived suppression table, and anything older than the window is garbage.
Cache file layout
The file is a flat JSON object mapping normalized titles to Unix timestamps:
{
"disk_usage_high:/var": 1730541123.44,
"feed_stale:BTC-USD": 1730541401.02,
"worker_queue_backlog": 1730539807.91
}
A few decisions worth being explicit about:
- Title as key, timestamp as value. No arrays, no per-key history. You only ever need the most recent push time to answer "was this within the window?"
- Unix epoch floats. Timezone-free, trivially comparable, and easy to prune. If you store ISO strings you have to parse them back, and parsing is where bugs live.
- One file, not one file per title. A per-title layout turns a read into a directory scan and makes atomic updates harder. With a few hundred keys, a single file is microseconds.
- Prune on write. Entries older than the window are dropped when the file is rewritten. The file stays bounded by the number of distinct titles active in the last five minutes, which is small.
If you want a durable audit trail of what was suppressed, write that to your normal log sink. Do not grow the cache into a database — the moment it becomes stateful infrastructure, it becomes a thing that can fail and block pushes.
Atomic writes with tempfile and os.replace
The cache is read and written by processes that may run concurrently. A naive open(path, "w") truncates the file before writing, so a reader that lands in that gap sees an empty file. Worse, a crash mid-write leaves a half-written JSON document that fails to parse on the next run.
The fix is the standard write-to-temp-then-rename pattern:
import json
import os
import tempfile
def _write_cache(path: str, data: dict) -> None:
directory = os.path.dirname(path) or "."
fd, tmp_path = tempfile.mkstemp(dir=directory, prefix=".dedup-", suffix=".tmp")
try:
with os.fdopen(fd, "w", encoding="utf-8") as fh:
json.dump(data, fh)
fh.flush()
os.fsync(fh.fileno())
os.replace(tmp_path, path)
except BaseException:
# Clean up the temp file if anything went wrong before the rename.
try:
os.unlink(tmp_path)
except FileNotFoundError:
pass
raise
Why this works:
-
tempfile.mkstempcreates the temp file in the same directory as the target.os.replaceis atomic only within a filesystem, so a temp file in/tmprenamed onto a path in/var/libcan degrade to a copy — or fail outright withEXDEV. -
flushplusos.fsyncpushes bytes to disk before the rename. Without the fsync, a crash can leave you with a renamed file whose contents were never durable. -
os.replaceis atomic on POSIX and on Windows. Readers see either the old file or the new one, never a truncated intermediate. - On any exception before the rename, the temp file is removed so you do not accumulate
.dedup-*.tmplitter.
This pattern costs one extra file creation per write. For a cache that is written at most a few times per minute, that is irrelevant.
The dedup check
Putting it together: read, decide, push, record. The read path must tolerate a missing or corrupt file by returning an empty dict.
import json
import logging
import os
import time
log = logging.getLogger(__name__)
WINDOW_SECONDS = 300.0
CACHE_PATH = "/var/lib/alerts/dedup.json"
def _load_cache(path: str) -> dict:
try:
with open(path, "r", encoding="utf-8") as fh:
data = json.load(fh)
except FileNotFoundError:
return {}
except (json.JSONDecodeError, OSError) as exc:
# Fail open: a broken cache must not block notifications.
log.warning("dedup cache unreadable (%s); treating as empty", exc)
return {}
if not isinstance(data, dict):
log.warning("dedup cache has unexpected shape; treating as empty")
return {}
return data
def _prune(data: dict, now: float) -> dict:
cutoff = now - WINDOW_SECONDS
return {k: v for k, v in data.items()
if isinstance(v, (int, float)) and v >= cutoff}
def should_push(title: str, *, urgent: bool = False,
cache_path: str = CACHE_PATH,
window: float = WINDOW_SECONDS) -> bool:
"""Return True if the notification should be pushed."""
if urgent:
log.info("dedup bypass (urgent): %s", title)
return True
now = time.time()
data = _load_cache(cache_path)
last = data.get(title)
if isinstance(last, (int, float)) and (now - last) < window:
log.info("dedup suppressed: %s (last push %.1fs ago)", title, now - last)
return False
data = _prune(data, now)
data[title] = now
_write_cache(cache_path, data)
return True
Call it right before the transport:
def push(title: str, body: str, *, urgent: bool = False) -> None:
if not should_push(title, urgent=urgent):
return
send_to_device(title, body)
The suppression is logged with the age of the last push, which is what you want when someone asks why they did not get an alert. "Suppressed 42 seconds after the previous push of the same title" is a complete answer.
A note on the check-then-write race
Two processes can both read the cache, both see no recent entry, and both push. The window narrows the race but does not eliminate it. If you need strict exactly-once across processes, put an advisory lock around the read-modify-write:
import fcntl
def with_lock(path: str):
lock_path = path + ".lock"
fh = open(lock_path, "w")
fcntl.flock(fh, fcntl.LOCK_EX)
return fh # caller closes to release
On Windows, use msvcrt.locking or a named mutex. In practice, most alerting pipelines are single-process or serialize pushes through a queue, and the window alone is enough. Add the lock only when you have measured the duplicate rate.
Why the window keys on the normalized title
The single most important design choice is what you hash on. Keying on the full notification body is tempting — it is the thing you actually send — but it breaks dedup in exactly the cases you care about.
Consider a price or metric alert. The body almost always contains the current value:
Title: Disk usage high: /var
Body: /var is at 91.4% (threshold 90%)
The next check two minutes later produces 91.7%. Different body, same condition, second notification. The user sees two alerts for one ongoing problem. Keying on the body means dedup only catches byte-identical retries, which are the easy case you could handle with a request ID anyway.
The title, by contrast, is the identity of the condition. Normalize it and use that:
import re
def normalize_title(title: str) -> str:
t = title.strip().lower()
t = re.sub(r"\s+", " ", t)
return t
Normalization matters because titles are often assembled from templates and may differ in case or whitespace between code paths. "Disk usage high: /var" and "disk usage high: /var" are the same condition and must collapse to one key.
The rule that follows: put the stable identity in the title and the volatile detail in the body. If a title embeds a timestamp, a hostname that changes per replica, or a fluctuating value, dedup will not work — and the fix is to change the title, not to loosen the window.
If you genuinely need sub-condition granularity, make the condition part of the title deliberately: disk_usage_high:/var and disk_usage_high:/home are separate keys, which is correct. The key should be exactly as specific as the thing a human would consider "the same alert."
Bypassing the window for urgent signals
A five-minute window is right for most alerts and wrong for a few. If a process is crash-looping, or a safety condition trips, you do not want the second occurrence suppressed because the first happened ninety seconds ago.
The urgent flag above is the escape hatch. Two rules keep it from eroding the whole design:
- Urgency is a property of the alert rule, not of the caller's mood. Decide at definition time which titles are urgent, and keep that list short. If everything is urgent, nothing is deduplicated.
- Urgent pushes still update the cache. An urgent push writes its timestamp like any other, so a subsequent non-urgent push of the same title inside the window is still suppressed. You get the urgent signal immediately, and you do not get a redundant routine one a minute later.
A common refinement is a shorter window for urgent titles rather than a full bypass — say thirty seconds instead of five minutes — so that a crash loop produces at most two notifications per minute instead of one per restart. The implementation is the same; you just look up the window per title:
URGENT_WINDOW_SECONDS = 30.0
def window_for(title: str, urgent: bool) -> float:
return URGENT_WINDOW_SECONDS if urgent else WINDOW_SECONDS
Pass that value into should_push instead of the fixed window default.
Operating notes
-
Where to put the file. A local path with a stable directory, e.g.
/var/lib/alerts/dedup.json. If you run multiple hosts that each push independently, each needs its own cache, or you need a shared store. A local file per host is usually correct: duplicate suppression across hosts is a different problem, solved by routing pushes through one service. - Permissions. The file contains alert titles, which may leak hostnames or symbols. Restrict it to the service user.
- Monitoring the cache. Log the suppression count. A sudden jump in suppressions means either a retry storm or a window that is too wide. Both are worth knowing.
-
Testing. Inject
nowinstead of callingtime.time()inside the function so you can test window boundaries deterministically. The version above takestime.time()inline for brevity; in production, pass a clock.
The whole mechanism is under sixty lines, has no dependencies, and removes the most common source of alert fatigue. The window is a blunt instrument, but blunt and predictable beats clever and occasionally silent.
More notes like this ship every week on this site.
Top comments (0)