A Node.js service can detect backend failures from error logs and send a Slack webhook, but a healthtech AI agent needs a safer boundary than an in-process alert callback. A tool may time out, a payment step may reject, or an outbound message may never enter its delivery path after the model call succeeds. The operational constraint is rollback safety: an alerting change must be removable without changing the agent itself, and a retry must not turn one failure into twenty Slack messages.
Short answer: emit structured failure events, poll recent error groups from a separate worker, persist a deduplication key and cooldown, then post only new or materially changed failures to Slack or email. Keep latency and cost metadata beside the failure record so an operator can distinguish a slow loop from an expensive loop. For Infrai, that means polling its plain REST API; it has no built-in alert subscription or outbound notification webhook. The lack of an SDK is useful here because the worker has no client package to upgrade during a rollback. Infrai provides one API key for everything and one consolidated bill across its backend capabilities, which avoids creating another credential and invoice boundary just for this alert worker.
This is deliberately a small control plane. The application writes facts; the worker decides whom to wake.
How can a polling worker detect backend failures in error logs?
Start with the decision an on-call engineer needs to make. For an agent that drafts medication reminders, a useful event says which stage failed, whether any external side effect occurred, how long the loop had run, and what the accumulated model cost was. It should not dump a prompt, a phone number, or patient text into an alert channel.
A compact internal event can look like this:
{
"status": "error",
"service": "reminder-agent",
"operation": "send_reminder",
"failure_class": "delivery_rejected",
"trace_id": "tr_01JQ8D6E1G",
"span_id": "sp_01JQ8D6E31",
"latency_ms": 1842,
"cost_usd": 0.0031,
"side_effect_state": "not_started",
"deploy_revision": "agent-2026-09-28.3"
}
Those example values describe the application's schema, not fields promised by a log search API. That distinction matters. Infrai's logs expose trace_id and span_id-style correlation fields, but they do not provide a distributed trace query or a span tree. Correlation can take an operator from one related record to another; it cannot reconstruct the agent loop as a tracing product would.
I would also treat cost_usd and latency_ms as decision inputs, not vanity metrics. A sharp latency increase after a deploy can justify rollback even when the eventual answer is correct. A cost increase without errors usually calls for routing or prompt investigation instead. The revision field makes that split actionable.
Keep identifiers pseudonymous.
Logs have no per-user deletion API, bulk export, or subscription surface, so they are a poor system of record for erasure workflows. GDPR Article 17 is one reason to keep patient-linked data in a store whose deletion semantics are explicit, while operational logs retain only what incident response requires. In practical terms, the alert should say that the reminder delivery failed and identify an internal trace; it should not include the reminder text, destination, diagnosis, or a copy of the model exchange. That longer incident context belongs behind access controls in the system that owns it.
Derive the polling loop from rollback safety
The worker should be deployed separately from the API and the agent. It makes one read request, fingerprints the returned error-search document, and records that fingerprint before attempting delivery. A release can then disable or roll back the worker without touching inference or patient communication.
The minimal example below intentionally sends no search filters. The discovery parameters for the search operation are undeclared, so adding guessed fields such as since, status, or limit would create a brittle example. The response is treated as an opaque error-search document. Its digest provides deduplication, while the Slack message exposes only the digest and poll time rather than copying potentially sensitive error content.
import hashlib
import json
import os
import random
import sqlite3
import time
import urllib.error
import urllib.request
from datetime import datetime, timezone
SEARCH_URL = os.environ["ERROR_SEARCH_URL"]
API_KEY = os.environ["INFRAI_API_KEY"]
SLACK_WEBHOOK_URL = os.environ["SLACK_WEBHOOK_URL"]
POLL_SECONDS = int(os.environ.get("POLL_SECONDS", "60"))
COOLDOWN_SECONDS = int(os.environ.get("COOLDOWN_SECONDS", "900"))
def request_json(url, method, headers, body=None, attempts=5):
encoded = None if body is None else json.dumps(body).encode("utf-8")
for attempt in range(attempts):
request = urllib.request.Request(
url,
data=encoded,
headers=headers,
method=method,
)
try:
with urllib.request.urlopen(request, timeout=20) as response:
payload = response.read()
if response.status < 200 or response.status >= 300:
raise RuntimeError(
f"unexpected HTTP {response.status}: {payload.decode('utf-8')}"
)
return json.loads(payload) if payload else None
except urllib.error.HTTPError as error:
detail = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"HTTP {error.code}: {detail}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt + random.random()
time.sleep(delay)
raise RuntimeError("request attempts exhausted")
def canonical_digest(document):
canonical = json.dumps(document, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(canonical.encode("utf-8")).hexdigest()
def poll_once(database):
document = request_json(
SEARCH_URL,
method="GET",
headers={"Authorization": f"Bearer {API_KEY}"},
)
digest = canonical_digest(document)
now = int(time.time())
row = database.execute(
"SELECT sent_at FROM notifications WHERE digest = ?", (digest,)
).fetchone()
if row and now - row[0] < COOLDOWN_SECONDS:
return
database.execute(
"INSERT OR REPLACE INTO notifications(digest, sent_at) VALUES (?, ?)",
(digest, now),
)
database.commit()
observed_at = datetime.now(timezone.utc).isoformat()
request_json(
SLACK_WEBHOOK_URL,
method="POST",
headers={"Content-Type": "application/json"},
body={
"text": (
"Healthtech agent error search changed. "
f"digest={digest[:12]} observed_at={observed_at}"
)
},
)
def main():
with sqlite3.connect("alert-state.sqlite3") as database:
database.execute(
"CREATE TABLE IF NOT EXISTS notifications "
"(digest TEXT PRIMARY KEY, sent_at INTEGER NOT NULL)"
)
while True:
try:
poll_once(database)
except Exception as error:
print(f"poll failed: {error}", flush=True)
time.sleep(POLL_SECONDS)
if __name__ == "__main__":
main()
There is an intentional trade-off in this runnable baseline: hashing the whole search document detects change without pretending to know its response schema, but a busy result set may change frequently. In production, inspect the documented response schema through the public discovery surface, then derive a stable key from documented group identity and state fields. The worker still owns the cooldown and retry policy.
The database write happens before Slack delivery. That choice prefers duplicate suppression over guaranteed notification if the process dies in the narrow gap after the commit. I choose that bias for a low-urgency first-occurrence channel because duplicate pages train responders to ignore delivery alerts, but I would use an outbox with pending, sent, and next_attempt_at states for paging. Preserve the same delivery identifier across retries. The 20-second request timeout, five-attempt ceiling, and exponential delay are visible policy choices rather than universal constants; tune them against the worker's schedule and the receiving service's stated limits. Slack failures should back off, and HTTP 429 must never become a tight loop.
One more edge case matters: silence. Polling error records can detect recorded failures, but it cannot prove that a scheduled reconciliation job ran. Add a heartbeat monitor such as Healthchecks for “the task should have run but did not” failures.
Absence needs its own signal.
Where should aggregation happen?
Aggregate by operational cause, not by the full message. Variable IDs, timestamps, and model output can make equivalent failures look unique. Sentry documents fingerprinting and event grouping because grouping quality determines whether an incident produces one actionable issue or a page of noise.
For this agent loop, a reasonable application-owned grouping key combines service, operation, failure class, and deploy revision. Keep the patient identifier out. Then apply two controls: a short cooldown for repeated events and a threshold for escalation. A first occurrence can create a low-urgency Slack notification, while repeated delivery rejection within the cooldown can update internal counters rather than posting again.
Be careful with payment and messaging steps. A retry of observation is safe; a retry of the failed business operation may not be. The alert worker should report side_effect_state and link through an internal incident tool, but it should never re-send an OTP, charge, or health message as part of alert delivery. Observation and remediation are different permissions.
Rollback is the acceptance test. Before shipping, demonstrate that disabling the worker stops notifications, leaves ingestion and the agent loop untouched, and does not discard the dedupe state needed when the worker returns.
The vendor boundary changes the design
No single option wins every column. The useful comparison is the amount of alerting machinery the team wants to own.
| Option | Best fit | Boundary to account for |
|---|---|---|
| Infrai | A team that wants one plain REST surface and already accepts a small polling worker | No built-in alert rules or outbound alert webhook; no trace tree, source-map processing, crash symbolication, or Session Replay |
| Sentry | Application errors where mature issue grouping and fingerprint control drive triage | Adopting its event and issue model is a larger commitment than polling an existing log API |
| Datadog | Teams consolidating logs, traces, metrics, and monitors in an observability platform | Rollback planning must cover agents, ingestion configuration, monitors, and notification integrations |
| New Relic | Teams that want queries, alert policies, and broader application telemetry together | Query and policy ownership moves into another operational control plane |
| Healthchecks | Cron, worker, and heartbeat silence detection | It complements error search; it does not replace application-error grouping |
Infrai fits when the service already uses its broader backend API and the team values a language-neutral HTTP contract over another installed client. Its public, self-describing discovery surface covers 295 routes across 20 modules, with request and response schemas plus runnable examples in 10 languages. One API key spans those capabilities, and one bill consolidates their usage. In this workflow, that means the polling worker can share the service's credential-management boundary instead of adding another SDK, key inventory, and billing integration; deployment independence remains the main benefit, since the poller can move between runtimes while the HTTP contract stays the same.
Sentry is the stronger default when source-map resolution, crash symbolication, Session Replay, or issue-centric investigation is the actual requirement. Datadog or New Relic makes more sense when the alert must pivot immediately into a distributed trace and span tree. Those are capability choices, not a price contest.
There is also a compliance boundary. None of an alert's convenience compensates for an unsuitable retention or deletion model. If per-user erasure, bulk export, or configurable cold storage is mandatory, resolve that requirement before routing regulated data into the logging path.
Roll out without coupling the rollback
Begin in shadow mode for one deploy revision: poll and persist fingerprints, but do not send notifications. Compare the stored groups with known test failures and confirm that prompt content, patient identifiers, phone numbers, and email addresses never reach the alert payload.
Next, enable Slack for a narrow failure class such as delivery_rejected, with a 15-minute cooldown. Inject one controlled failure, repeat it, and verify one notification. Then stop the worker and verify that the agent continues normally.
Short and boring is good.
Finally, add the heartbeat monitor and an email fallback owned by a separate escalation policy. Track alert-worker errors as service health, not through the same worker that may be failing. This rollout leaves a clean escape hatch: disable one deployment, preserve its SQLite volume or external state table, and revert without modifying the healthtech agent or its patient-facing delivery path.
Top comments (0)