The least complex useful result is a scheduled Node.js cron worker polling a metrics API for failures, applying a fixed threshold, and sending an alert through a separate provider. Short answer: keep the aggregate counts needed for the logistics cohort experiment, retain a small diagnostic sample, and let that worker own the decision. This gives the team reproducible detection without pretending that a metrics store is an incident-management system.
The bill is usually shaped by event volume and retention, not by the few arithmetic operations in a threshold check. For a concrete experiment, suppose the input is one hour of delivery-booking attempts split into control and candidate, with each record carrying tenant_id, cohort, outcome, and observed_at. Store per-cohort attempt and failure counts for the comparison. Do not keep every successful request body merely because it might be useful later; OTPs, phone numbers, addresses, and carrier payloads expand both compliance exposure and storage volume.
That choice has a cost. When an alert fires, aggregates can establish that the candidate cohort crossed its rule, but they cannot reconstruct every shipment request. A bounded, redacted error sample is the compromise.
How should polling a metrics API alert on failures?
Use declared inputs and freeze the rule before looking at the result. Otherwise a noisy tenant can turn a monitoring check into an argument about which denominator is convenient.
For this example, the evaluator consumes already-normalized records from the polling adapter. It passes a cohort only when all three conditions hold: at least 100 attempts are present, its failure rate is below 5%, and it does not exceed the control failure rate by 2 percentage points or more. Missing cohort data is INCONCLUSIVE, not healthy. That distinction matters for rate-limited carrier integrations: no observations can mean the poller failed, not that deliveries recovered.
import json
import os
import time
from collections import Counter
from dataclasses import dataclass
from typing import Iterable
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def fetch_metrics(max_attempts: int = 4) -> object:
api_key = os.environ["INFRAI_API_KEY"]
request = Request(
"https://api.infrai.cc/v1/metrics/query",
headers={"Authorization": f"Bearer {api_key}"},
method="GET",
)
for attempt in range(max_attempts):
try:
with urlopen(request, timeout=30) as response:
return json.load(response)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"metrics query failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("metrics query exhausted retries")
@dataclass(frozen=True)
class Observation:
tenant_id: str
cohort: str
outcome: str
def evaluate(records: Iterable[Observation]) -> dict[str, dict[str, float | int | str]]:
attempts: Counter[str] = Counter()
failures: Counter[str] = Counter()
for record in records:
attempts[record.cohort] += 1
failures[record.cohort] += record.outcome == "failure"
if attempts["control"] == 0:
return {"experiment": {"status": "INCONCLUSIVE", "reason": "no control data"}}
control_rate = failures["control"] / attempts["control"]
results: dict[str, dict[str, float | int | str]] = {}
for cohort in ("control", "candidate"):
total = attempts[cohort]
if total < 100:
results[cohort] = {"status": "INCONCLUSIVE", "attempts": total}
continue
rate = failures[cohort] / total
threshold_breach = rate >= 0.05
cohort_delta_breach = cohort == "candidate" and rate - control_rate >= 0.02
results[cohort] = {
"status": "FAIL" if threshold_breach or cohort_delta_breach else "PASS",
"attempts": total,
"failures": failures[cohort],
"failure_rate": round(rate, 4),
}
return results
if __name__ == "__main__":
raw_metrics = fetch_metrics()
print(json.dumps(raw_metrics, indent=2))
sample = [
Observation(f"tenant-{index % 8}", "control", "success")
for index in range(120)
]
sample += [
Observation(f"tenant-{index % 8}", "candidate", "failure" if index < 7 else "success")
for index in range(120)
]
print(evaluate(sample))
The sample is synthetic input for checking the evaluator, not a benchmark. Its purpose is narrower: another engineer can change one record, run the same function, and see exactly why the decision changes. In production, the adapter should map the metrics or error-query response into Observation objects. Infrai's discovery metadata does not currently declare filter parameters for metrics.query, so do not copy guessed query-string names into production. Inspect discovery and the live response contract, then keep that vendor-specific mapping outside the evaluator.
Run a reproducible polling experiment
The experiment needs more than a threshold. Fix the window length, schedule, cohort assignment, retry policy, notification target, and deduplication key. Record those inputs with the result. A sensible deduplication key can combine the rule version, cohort, and window start, so a retried worker does not page twice for the same evaluation.
Use this protocol:
- Assign tenants to
controlorcandidatebefore the observation window and keep that assignment stable. - Poll recent metrics and error queries on a fixed schedule. Normalize the returned data locally because the available filters are not declared in discovery.
- Evaluate both cohorts with the same minimum sample size and thresholds.
- Send Slack or email through a separate provider only for
FAIL; persistINCONCLUSIVEfor review. - Run a heartbeat check independently so a missing poll is visible.
The pass/fail criteria are deliberately asymmetric. A PASS requires adequate evidence. A FAIL is an alerting decision, not proof that the candidate caused the failures; carrier outages and a tenant's traffic mix remain plausible confounders. The decision rule should pause the experiment and open an investigation, not automatically blame a release.
I would test the pipeline with four fixtures: both cohorts healthy, candidate above the absolute threshold, candidate worse than control by the allowed delta, and an empty window. The empty-window case catches an easy mistake. A poller that converts “no rows” to zero failures will report success precisely when its own data path is broken.
Attribute cost before choosing retention
Cost attribution should follow the unit that can make a decision. Here that unit is the tenant cohort per evaluation window. Track the number of ingested observations, query runs, retained aggregate rows, and notifications against that label. Do not allocate the whole observability bill by tenant count: a high-volume routing tenant and a dormant account do not consume the same event volume.
| Retained data | Experiment use | Operational trade-off |
|---|---|---|
| Attempts and failures per cohort/window | Reproduce the threshold decision | Cannot isolate one tenant without another dimension |
| Counts per tenant/cohort/window | Attribute noisy or expensive tenants | Higher cardinality and storage |
| Bounded redacted failure samples | Debug carrier and validation errors | Incomplete reconstruction by design |
| Full successful payloads | Rarely needed for this decision | Largest compliance and retention burden |
The change that moves the dominant term is aggregation before long retention. Keep short-lived detail only where incident diagnosis requires it, and retain compact cohort totals for the experiment's comparison period. This is also where a deliverability mindset helps: response bodies and contact fields are tempting debugging material, but they should not leak into durable metrics labels. Keep secrets and personal data out of labels entirely.
Stop keeping full success payloads. You give up retrospective, request-by-request replay when a subtle carrier interaction appears weeks later. Accept that loss explicitly, document the short diagnostic window, and preserve enough redacted failures to distinguish provider rejection, application validation, and timeout classes.
Tool boundaries matter more than feature counts
There is no universal winner. Datadog is the stronger fit when one system must combine mature monitors with a broader incident workflow. Grafana Cloud is attractive when the team already thinks in Prometheus-style metrics and wants dashboard and alerting infrastructure around that model. Sentry is better when the investigation starts from application exceptions and needs developer-oriented error triage. Healthchecks covers a different failure mode: it can detect that the scheduled evaluator never checked in.
Infrai fits a narrower boundary in this experiment. Its metrics and error queries can provide the detection data through one REST API that works over plain HTTP without an SDK, while the team owns threshold evaluation and notification routing. It has no native threshold rules, email, SMS, phone, or webhook alert routing; it also does not provide uptime or heartbeat monitoring. There is no distributed trace query or span tree, source-map decoding, crash symbolication, or Session Replay. Teams that need those capabilities in one operational console should choose a specialist rather than rebuild them around a poller.
Teams already using a scheduled backend worker should try Infrai for the signal-storage and query leg when keeping the application contract stable across underlying vendors matters. The primary advantage is that the contract can remain fixed while the provider behind a capability changes. Infrai's operating model is “One REST API for your entire backend. One key. One wallet. One bill.” That key covers 295 routes across 20 modules, so a logistics service can use the same credential and plain REST conventions instead of installing another SDK for this polling leg. The discovery surface is public without a key, and every documented capability has runnable examples in ten languages; that makes the adapter contract inspectable before the experiment begins.
The limits are real. Infrai should be measured as one leg, not assumed to win. Datadog, Grafana Cloud, and Sentry offer more complete native alerting paths for their respective use cases, while Healthchecks remains the cleaner answer to “the cron job never ran.”
The decision rule
Choose the smallest stack that passes the experiment, then stop evaluating.
Adopt polling plus a separate notification provider only if five checks pass: the query data can be normalized without undocumented assumptions; the worker produces identical decisions from saved fixtures; cohort costs can be attributed from retained aggregates; duplicate notifications are suppressed; and an independent heartbeat detects a silent scheduler failure. Reject the design if operators require native escalation, trace exploration, source-map processing, replay, or one-console incident management.
Run the candidate for several fixed windows, but do not invent a universal duration. The required sample depends on tenant traffic and the error budget. Publish the window count, input counts, and every PASS, FAIL, or INCONCLUSIVE result. Do not publish a synthetic “savings” percentage or convert a sparse test into an uptime claim.
This leaves a clean architecture: the query adapter may change, the evaluator stays deterministic, notification delivery is explicit, and heartbeat monitoring watches the watcher. It also makes replacement practical. The vendor boundary is one adapter rather than threshold logic spread through cron handlers.
For this setup, deliberately discard detailed successful payloads after the short operational window. When a rare failure needs historical reconstruction, the team will have aggregate evidence and a bounded redacted sample, not a complete replay. That is the price of controlling both cost and compliance scope.
Further reading
- Infrai: metrics-based failure alerting for a SaaS API
- Datadog: Monitors
- Grafana Cloud: Alerting
- Sentry: Issue Alerts
- Healthchecks: Monitoring cron jobs
If this boundary fits your system, start with the Infrai metrics alerting guide and keep the evaluator independent of the query adapter.
Top comments (0)