Alert on a media experiment only when an unresolved error group crosses a cohort-specific decision boundary, and store enough state to distinguish a new failure from another observation of the same one. The deciding constraint is signal quality: a cron job that forwards every open group to email, Slack, or a generic webhook can be perfectly reliable and still make the experiment harder to interpret.
TL;DR: treat polling as a small stateful reconciliation loop. Read unresolved groups from the error API, normalize them into a vendor-neutral record, compare monotonic evidence with a durable checkpoint, and emit one idempotent notification per group transition. Keep delivery transport outside the decision function.
This architecture decision record assumes a Node.js media service running an experiment across tenant cohorts, while a Python job performs the scheduled reconciliation. It does not assume that an error event alone proves a cohort regression.
How should a Node.js cron poll unresolved error groups from an API?
Four invariants matter more than the scheduler or chat integration. First, the same remote snapshot must produce the same notification key. Second, a failed delivery must not advance the durable checkpoint. Third, one cohort's high traffic must not erase a lower-volume cohort's distinct signal. Fourth, resolution and recurrence must be represented as transitions rather than inferred from silence.
Silence is ambiguous.
The failure boundary deserves precision. The source API owns whether a group is unresolved and how events are grouped; the poller owns its checkpoint and decision policy; each destination owns final delivery. A timeout leaves the source state unknown. A 429 response means no decision should be made from that run. A malformed record is quarantined rather than converted into a misleading zero. Partial pagination is not a complete snapshot.
OpenTelemetry's logs model is useful at the ingestion boundary because a LogRecord can carry a timestamp, observed timestamp, severity, body, resource, and attributes. Those fields let the Node.js service attach stable experiment dimensions such as experiment.id and tenant.cohort before grouping occurs. They do not define an alerting policy, and a high-cardinality tenant identifier should not casually become a grouping dimension; that can turn one failure mode into thousands of tiny groups.
The checkpoint is part of the evidence. Put it in storage with conditional writes or an equivalent compare-and-set rule. A local file may be adequate for one non-overlapping process, but it stops being authoritative as soon as two scheduled runs overlap or the job moves between hosts.
Consider the concrete overlap. Run A reads source version 41, claims a notification, and stalls during delivery; 60 seconds later, run B reads version 42 for the same cohort and group. A single last_run value cannot describe both obligations, and updating it when B finishes can strand A's undelivered message. A ledger keyed by group, cohort, transition, and source version keeps the two observations separate. The storage operation must atomically move each key from absent to claimed and then to delivered, while an abandoned claim needs an expiry policy long enough to exceed the delivery deadline. This creates a trade-off: a short lease retries quickly but can duplicate a still-running request, while a long lease delays recovery after a crashed worker. Choose that lease from the measured request deadline and scheduler overlap, not from a convenient round number. The point is not to promise magical exactly-once delivery across an HTTP boundary. The point is to make retries explicit, bounded, and auditable.
One timestamp cannot do that job.
Decision record: compare the state models
The choice is between remembering deliveries, remembering observations, or remembering nothing. The table is intentionally about semantics, not brands or connector counts.
| State model | Duplicate control | Cohort fidelity | Primary failure mode | Valid boundary |
|---|---|---|---|---|
| Stateless snapshot forwarding | None across runs | Depends on each payload | Every run repeats every unresolved group | Manual, low-frequency audit |
| Last-seen timestamp per group | Suppresses unchanged observations | Good if cohort belongs to the key | Clock skew or reordered events can hide a change | Single writer with ordered source data |
| Monotonic source cursor plus delivery ledger | Stable across retries | Explicit in the notification key | Requires atomic checkpoint discipline | Scheduled production reconciliation |
| Event-count threshold only | Reduces small signals | Poor when cohort traffic differs sharply | Large cohorts dominate the threshold | Cohorts with comparable exposure |
A monotonic cursor is preferable when the source exposes one. If it does not, use a documented stable update token or tuple rather than pretending wall-clock time is ordered. The ledger key should bind the experiment, cohort, group, transition, and observed source version. This is more storage than a single last_run timestamp, but it answers the operational question that matters after a retry: was this exact decision already delivered?
There is also a data-retention trade-off. The ledger need not become a second error archive. Retain compact keys and delivery outcomes for the retry and audit window chosen by the team, while the source remains authoritative for event detail. Deleting ledger entries too early permits old snapshots to notify again; keeping raw payloads indefinitely expands both privacy scope and schema-coupling risk.
Critical path in Python
The following core is deliberately transport-agnostic. The API client is expected to return all pages or fail the run; the store must implement atomic claims; and the notifier may route to email, Slack, or another webhook without changing the decision rule. The Node.js application only needs to emit consistent experiment and cohort attributes into its error telemetry.
from dataclasses import dataclass
from hashlib import sha256
from typing import Iterable, Protocol
@dataclass(frozen=True)
class Group:
group_id: str
experiment_id: str
cohort: str
source_version: str
unresolved: bool
affected_sessions: int
exposed_sessions: int
class Ledger(Protocol):
def claim(self, key: str) -> bool: ...
def mark_delivered(self, key: str) -> None: ...
def release(self, key: str) -> None: ...
class Notifier(Protocol):
def send(self, subject: str, fields: dict[str, str]) -> None: ...
def notification_key(group: Group, transition: str) -> str:
material = "|".join((
group.experiment_id, group.cohort, group.group_id,
transition, group.source_version,
))
return sha256(material.encode("utf-8")).hexdigest()
def should_alert(group: Group, minimum_sessions: int) -> bool:
if not group.unresolved or group.exposed_sessions <= 0:
return False
return group.affected_sessions >= minimum_sessions
def reconcile(
groups: Iterable[Group],
ledger: Ledger,
notifier: Notifier,
minimum_sessions: int,
) -> None:
for group in groups:
if not should_alert(group, minimum_sessions):
continue
key = notification_key(group, "threshold-crossed")
if not ledger.claim(key):
continue
try:
notifier.send(
subject="Experiment cohort error signal",
fields={
"experiment": group.experiment_id,
"cohort": group.cohort,
"group": group.group_id,
"affected_sessions": str(group.affected_sessions),
"exposed_sessions": str(group.exposed_sessions),
},
)
except Exception:
ledger.release(key)
raise
else:
ledger.mark_delivered(key)
The integer threshold is a placeholder for an experiment policy, not a statistical claim. A real policy may require an exposure floor, a rate comparison, and a sustained interval; those values must come from the experiment design. The important property is that should_alert is deterministic over recorded inputs and can be tested separately from HTTP, authentication, and delivery.
No guesswork belongs here.
Do not catch an API failure and continue with an empty list. Empty means "the source reports no unresolved groups"; failure means "the poller does not know." Conflating them can manufacture a resolution transition. This is a nasty pitfall because the next successful run then appears to contain a fresh recurrence.
Retries need the same care. Use bounded exponential backoff for transient requests, honor server retry guidance where supplied, and place a deadline around the entire run so the next scheduled invocation cannot pile up behind it. Authentication errors, schema violations, and exhausted pagination limits should fail closed and produce an operational signal about the poller itself.
Testing the signal, not merely the connector
Unit tests should replay the same snapshot twice and assert one delivery, then inject a delivery exception and assert that the next run retries the same key. Add cases for two cohorts sharing a group fingerprint, a source version arriving out of order, zero exposure, a group becoming resolved, and an incomplete second page. Those cases exercise meaning. A test that merely receives a webhook proves very little.
At deployment time, ensure only one worker can own a given partition or make claim atomic enough that overlapping cron invocations are harmless. Track poll duration, source freshness, groups examined, decisions suppressed, deliveries attempted, and delivery failures. Avoid labeling those metrics with raw group or tenant identifiers unless their cardinality is strictly bounded.
Email and chat destinations should carry the same compact evidence: experiment, cohort, group identity, affected and exposed counts, observation time, and a link controlled by the operator. Keep credentials in the platform's secret store, redact payloads from routine logs, and sign outbound webhooks when the receiver supports request verification.
A useful acceptance test is blunt: disable the destination, cross the threshold, restore the destination, and verify exactly one eventual notification. Then run two pollers concurrently. If the result is two messages, the storage contract is unfinished.
I would accept extra ledger writes because they preserve the distinction between observation and delivery; I would not accept a smaller state model that silently merges them.
The rejected option still has a place
I reject stateless forwarding for experiment alerts because unresolved groups persist across polls and cohort traffic is uneven; repeated messages increase noise without adding evidence. It also couples schedule frequency directly to human interruption frequency, an architectural leak that gets worse when teams shorten the interval to reduce detection delay.
Stateless polling remains valid for an operator-triggered audit that writes a snapshot to a file or dashboard and never pages anyone. In that boundary, repetition is visible context rather than a notification defect, and there is no promise of exactly-once human attention.
Use durable reconciliation when a message changes a decision. The scheduler can run every minute or every hour without changing the semantics: unresolved state is read completely, cohort evidence is evaluated consistently, and delivery advances only after success. That separation preserves the signal the media experiment was built to measure.
Top comments (0)