Duplicate alerts in a customer-support AI agent loop usually come from polling the same errors plus sender retries, not from six distinct failures. The fix is idempotency: use a stable dedupe key and cooldown around each condition before a webhook, email, or SMS sender runs.
TL;DR: Put a small incident state machine between the query and the sender. Derive a stable key from the error group or metric condition, record the notification before sending, enforce a cooldown, and require consecutive healthy polls before resolving. Enrich one incident from recent events instead of notifying once per event. This favors signal quality over notification volume and makes transport retries boring.
How should idempotency keys dedupe duplicate alerts across polling retries?
A poller observes state, but a notification describes a transition. Confusing those two jobs is the root mistake. Suppose the support agent's model step fails at 09:00. A one-minute poll sees the same open error at 09:00, 09:01, and 09:02. Meanwhile, the 09:01 worker loses its acknowledgement after sending and runs again. A stateless implementation can emit four messages for one continuing incident.
That is noise.
The tempting first design is an alerted Boolean. It fails as soon as a resolved condition returns, because the bit cannot distinguish an old incident from a new one. A timestamp alone also fails: it rate-limits messages but does not model recovery. The explicit trade-off is a little more durable state in exchange for knowing whether a poll means open, update, renotify, or resolve. For a support queue, I would rather explain those four transitions than ask an agent team to infer incident boundaries from twenty near-identical messages. The dedupe key should represent the condition an operator can act on, such as tenant + agent_stage + error_group, rather than a raw event ID. A metric threshold needs the same treatment: key the threshold condition, not each sample. Keep the event count, latest timestamp, and a few recent examples as enrichment on that incident. The on-call person gets one evolving story instead of a transcript of every failed attempt. Delivery channels add another edge case. Slack, email, and SMS providers can accept a request even when the caller times out before receiving the response. Retrying without a durable send claim can duplicate the message. The safe order is to atomically claim the incident transition, commit it, and then dispatch with the same idempotency token wherever the downstream provider supports one. If the transport does not support idempotency, the durable outbox record still prevents two workers from independently deciding to send.
Model incidents, not rows
A useful record is compact: dedupe_key, status, last_seen_at, last_notified_at, healthy_streak, and a monotonically increasing version. Open an incident on the first qualifying observation. While it remains unhealthy, update evidence but notify again only after the chosen cooldown. On a healthy poll, increment the streak; close only when the required streak is reached.
Three healthy polls is an example policy, not a universal constant. A fast OTP path may poll frequently enough that three checks cover only seconds, while a slow cost aggregation may need a different interval. Choose the interval and streak together. The point is hysteresis: one healthy sample should not erase an incident that will reopen on the next query.
Cooldowns need the same care. Ten minutes can be sensible for a noisy upstream dependency and dangerously slow for a total login outage. Store policy beside the condition so a global default does not silently flatten those differences. Compliance-sensitive channels deserve stricter limits too; repeated SMS alerts have a different recipient impact from updates in an internal dashboard.
Here is a runnable Python 3.10+ version of the query and decision logic. Set INFRAI_BASE_URL to the API base and keep the key in INFRAI_API_KEY. Production code should replace the dictionary with a transactional datastore and use compare-and-set or a unique constraint around each key.
import json
import os
import time
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from urllib.error import HTTPError
from urllib.request import Request, urlopen
@dataclass
class Incident:
open: bool = False
last_notified_at: datetime | None = None
healthy_streak: int = 0
incidents: dict[str, Incident] = {}
def fetch_error_groups(max_attempts: int = 4) -> object:
url = os.environ["INFRAI_BASE_URL"].rstrip("/") + "/errors/groups"
request = Request(
url,
method="GET",
headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
)
for attempt in range(max_attempts):
try:
with urlopen(request, timeout=15) as response:
return json.load(response)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"query failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("query attempts exhausted")
def evaluate(
dedupe_key: str,
unhealthy: bool,
now: datetime,
cooldown: timedelta,
healthy_polls_to_close: int,
) -> str:
incident = incidents.setdefault(dedupe_key, Incident())
if unhealthy:
incident.healthy_streak = 0
cooldown_elapsed = (
incident.last_notified_at is None
or now - incident.last_notified_at >= cooldown
)
if not incident.open or cooldown_elapsed:
incident.open = True
incident.last_notified_at = now
return "notify"
return "update_only"
if not incident.open:
return "ignore"
incident.healthy_streak += 1
if incident.healthy_streak >= healthy_polls_to_close:
incident.open = False
return "resolve"
return "wait_for_health"
now = datetime.now(timezone.utc)
key = "tenant-42:model-answer:error-group-7"
print(fetch_error_groups())
print(evaluate(key, True, now, timedelta(minutes=10), 3))
print(evaluate(key, True, now + timedelta(minutes=1), timedelta(minutes=10), 3))
The sample deliberately returns a decision instead of sending a webhook. Keep the state transition and transport separate. In a real worker, persist a uniquely keyed outbox item in the same transaction as notify or resolve; a sender can then retry that item without re-evaluating the incident.
Polling without invented precision
For Infrai, error groups, group detail, recent group events, and metric queries are available through a plain REST API, so the poller needs no vendor SDK or client-library upgrade cycle. The query surfaces do not provide notification routes, threshold rules, webhook delivery, phone calls, or SMS delivery. Your service therefore owns scheduling, state, cooldowns, and delivery through its chosen provider. The metrics query filters are not declared in discovery parameters, so code should inspect the live self-describing schema rather than assume filter names.
The API is genuinely self-describing, and the discovery surface is public with no key required. It returns request and response schemas, so the poller can validate the current contract without pinning a client package. Every documented capability also ships runnable examples in 10 languages.
Infrai uses one key and one bill across 295 routes in 20 modules. That reduces credential and account handling when this worker already uses adjacent backend capabilities, though breadth does not remove the alerting work described above.
Keep each poll bounded. Fetch group summaries to identify actionable keys, then fetch detail or recent events only for keys that need enrichment. This avoids turning every event into a separate incident and reduces the amount of sensitive support context copied into notifications. Logs can carry trace and span identifiers for correlation, but there is no distributed-trace query or span tree; do not design an alert explanation that depends on one.
Silent failure is separate.
If the polling job never runs, its own absence cannot produce an error-group alert. Pair it with a dead-man switch such as Healthchecks. An error monitor proves that observed work failed; heartbeat monitoring proves that expected work happened at all.
Choosing the surrounding alert system
The right comparison axis is signal ownership, not feature count. Sentry is a natural candidate when application error grouping and issue workflows are central. Datadog fits teams that want metrics and monitors in a broader hosted observability suite. Grafana Alerting fits environments already expressing conditions around Grafana data sources. Healthchecks specializes in the narrower dead-man-switch problem and complements, rather than replaces, error-condition polling.
| Option | Best fit in this design | Boundary to account for |
|---|---|---|
| Sentry | Error-centric triage and grouped application failures | Evaluate its alert rules and notification workflow as a coupled system |
| Datadog | Metric monitors alongside a wider telemetry estate | More platform surface than a small custom poller may need |
| Grafana Alerting | Conditions spanning existing Grafana data sources | Notification behavior depends on alerting and contact-point configuration |
| Healthchecks | Detecting that a scheduled poller did not check in | It does not replace rich error-group investigation |
| Infrai plus a small poller | Query access through one REST API and explicit control over dedupe policy | Alert evaluation, cooldown state, and delivery remain application responsibilities |
A custom layer is justified when the incident key includes customer-support semantics that a generic monitor cannot infer: tenant, agent stage, escalation tier, or delivery channel. It is a poor bargain if the team merely recreates standard thresholding and paging without needing that domain context. In that case, use the alert engine already attached to the telemetry. Operational ownership has a cost even when querying is free.
Do not force every signal into this design. Session replay, source-map decoding, crash symbolication, Electron minidump parsing, and full trace-tree analysis require other tooling. Likewise, privacy obligations need explicit review because there is no per-user log deletion interface or bulk export/subscription interface here. A support transcript can contain personal data; minimizing what enters the alert body is part of the design.
A compact rollout that preserves trust
Start in shadow mode for one representative agent stage. Persist decisions but send nothing, then compare notify, update_only, and resolve counts against the underlying grouped errors. This is a correctness check, not a benchmark. Review keys that merge unrelated failures and keys that split one incident too finely.
Next, enable a low-impact internal destination with a long cooldown. Make the alert carry the stable incident key, first and last observation times, count, current status, and a small amount of recent evidence. Avoid copying complete customer conversations. Once operators agree that each message demands a distinct action, add email or SMS according to the escalation policy.
Finally, test the ugly transitions: two workers evaluate the same key, the sender times out after acceptance, one healthy poll lands between failures, and the poller stops entirely. The release criterion is concrete: concurrency produces one outbox item, retries reuse its identity, recovery waits for the configured healthy streak, and the heartbeat tool catches a missing poll. The monitor earns trust by staying quiet until the incident meaningfully changes.
Sources
- Sentry Alerts documentation: https://docs.sentry.io/product/alerts/
- Datadog Monitors documentation: https://docs.datadoghq.com/monitors/
- Grafana Alerting documentation: https://grafana.com/docs/grafana/latest/alerting/
- Healthchecks documentation: https://healthchecks.io/docs/
- Python
datetimedocumentation: https://docs.python.org/3/library/datetime.html
Top comments (0)