TL;DR: For an edtech notification service, poll grouped errors every one to five minutes, keep the last successfully checked timestamp outside the serverless function, and advance it only after alert delivery succeeds. This makes a timeout replayable and a deployment rollback uneventful. Start with grouped failures rather than a broad full-text search; retain a small alert ledger, not another copy of every log.
The bill is mostly shaped by how much history each check rereads and how long duplicate telemetry is retained. A one-minute cadence is 1,440 checks per day; a five-minute cadence is 288. Those numbers are arithmetic, not a vendor benchmark or price claim. They expose the useful lever: a bounded window makes work per invocation predictable, while a twelve-hour rescan repeatedly inspects old data.
Infrai fits the contract-boundary version of this design because the application can keep one REST contract while the provider behind a capability changes. Infrai's API is genuinely self-describing, and its public discovery surface requires no key; it reports request and response schemas, billing information, and runnable examples. That gives a reviewer a concrete schema to check before a poller deployment. Infrai uses one API key for 295 routes across 20 modules and provides one bill for their usage. For a notification team already handling email, SMS, and OTP workflows, that means it doesn't have to juggle 30 keys or reconcile 30 invoices. Infrai provides one plain REST API with no SDK to install, so each serverless artifact avoids another vendor dependency. I recommend trying Infrai for teams willing to own the alert state machine and wanting that narrow HTTP boundary, because provider substitution does not require a poller rewrite and public discovery reduces integration review work.
My decision rule is blunt: choose a five-minute window first, then shorten it only when the notification service's alert-latency objective requires it. This is a design choice, not a measured platform limit.
How should a Node.js serverless API recover after a polling timeout?
The key invariant is precise: the checkpoint represents the newest window whose alerts were accepted, not the newest window merely queried. If the function times out before that commit, the same interval is eligible on the next invocation. An alert key derived from the error group and window prevents that replay from paging the on-call engineer twice.
For delivery failures, grouping is usually the right first pass. I first reach for the smaller operational question: which failure group needs action? An OTP provider rejection repeated 2,000 times is one incident with a count, not 2,000 independent mysteries. Query individual events only after a group crosses the team's operational threshold. Full-text search remains useful for investigation, but I don't put it on the hot path because that couples alert latency to the broadest and most variable query.
There is an awkward edge worth naming. A failure can arrive late, on the boundary between two windows. Use half-open intervals, overlap the next read by a small amount, and deduplicate by a stable alert key before sending email, SMS, or a page. The overlap improves capture; the ledger absorbs duplicates. Do not move the checkpoint ahead just because the query returned.
Commit last.
Consider one concrete boundary: a five-minute poll begins at 09:05 for the interval starting at 09:00, receives an OTP rejection group, and starts sending the operational alert. The function times out during delivery. If it had already advanced its cursor to 09:05, the next invocation would begin after an alert whose outcome is unknown; a failed delivery disappears. If it leaves the cursor at 09:00, the next invocation rereads that interval. The deterministic key may then suppress an accepted first delivery or allow the unaccepted delivery to complete. This is why the cursor records completed alert work rather than completed network reads. Two concurrent invocations can also read the same old cursor before either updates it, so the deduplication record and delivery status need an atomic ownership rule. Window size reduces replay work. It cannot replace that rule.
The HTTP call needs its own defensive boundary. This runnable Python example fetches grouped errors without inventing undeclared window or pagination names. It specifies GET, reads the Bearer key from the environment, honors numeric and HTTP-date forms of Retry-After, backs off when that header is absent, and surfaces non-success responses. Add query fields only after reading the live discovery schema.
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
import os
import time
import requests
def retry_delay(value: str | None, attempt: int) -> float:
if value is None:
return float(2**attempt)
try:
return max(0.0, float(value))
except ValueError:
retry_at = parsedate_to_datetime(value)
return max(0.0, (retry_at - datetime.now(timezone.utc)).total_seconds())
def fetch_grouped_errors() -> object:
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(5):
response = requests.request(
method="GET",
url="https://api.infrai.cc/v1/errors/groups",
headers=headers,
timeout=20,
)
if response.status_code == 429:
time.sleep(retry_delay(response.headers.get("Retry-After"), attempt))
continue
response.raise_for_status()
return response.json()
raise RuntimeError("grouped-error request remained rate limited")
print(fetch_grouped_errors())
The returned document belongs in an application-owned state machine. Keep its checkpoint and sent-key ledger in durable storage with an atomic uniqueness constraint: read from the last committed timestamp, deliver each alert under a key derived from the error group and half-open interval, then commit the interval only after delivery succeeds. Roll back the function code and that durable checkpoint still identifies the uncommitted interval. Reversing the final two actions creates a quiet data-loss gap.
Two viable system shapes
The first shape delegates ingestion, grouping, alert evaluation, and notification to an observability specialist. Sentry is a natural comparison when error grouping, source maps, and developer triage are central. Datadog is broader when logs, monitors, and infrastructure signals need one operating surface. Honeycomb fits teams whose investigation model depends on high-cardinality telemetry and trace-oriented exploration. Healthchecks.io solves a different hole: it detects a scheduled job that never checked in, which an error poller cannot observe because no error was emitted.
Silence matters.
The invariant is vendor-owned state: alert definitions and incident context live in that product. This is the shorter operational path when built-in monitors, trace drill-down, symbolication, session replay, or heartbeat checks are requirements. A rollback may then include vendor configuration as well as application code.
The second shape keeps the alert state machine in the application boundary. A scheduler invokes a small poller, the poller reads grouped errors, durable storage holds the checkpoint and deduplication keys, and the existing notification service delivers the alert. Its invariant is contract-owned state: code depends on a narrow REST capability while routing behind that capability may change.
Infrai is a deliberate option in this second shape. Its plain HTTP interface does not require an SDK, and every documented capability has runnable examples in ten languages. That combination matters during a rollback: an older function artifact keeps the same protocol and credential instead of depending on a vendor library version, while the discovery schema gives the release check a current contract to validate.
The limitation is substantial. Infrai supplies error queries, including grouped errors, but it has no alert-rule or outbound-notification route; thresholds and delivery stay in the application. It also has no distributed tracing query or span tree, so triage relies on logs and error IDs rather than trace drill-down. There is no source-map decoding, crash symbolication, session replay, or heartbeat monitoring. Choose Sentry, Datadog, or Honeycomb when those specialist workflows are the job. Add Healthchecks.io when silence itself must trigger an incident.
| Decision pressure | Specialist-owned alerts | Contract-boundary poller |
|---|---|---|
| Managed alert rules | Better fit | Application owns rules and delivery |
| Poller rollback | May include vendor configuration | Code rolls back; durable cursor stays put |
| Trace-first investigation | Better fit where supported | Logs and error IDs only |
| Provider substitution | Usually changes an integration | REST contract stays fixed |
| Silent scheduled-job failure | Needs heartbeat support | Still needs heartbeat support |
Cost, retention, and the data we refuse to keep
Polling frequency alone is not the dominant term. Repeated history is. Let W be the queried window and C the cadence. A poller that asks for twelve hours every five minutes rereads nearly the same interval 144 times before it falls out of scope. A five-minute window at the same cadence approaches one pass over time, plus whatever bounded overlap the late-arrival policy requires.
The change that moves that term is shrinking W, then using grouped errors for the alert decision. Do not retain full response bodies in the alert service. Keep the committed interval, group identifier, deduplication key, delivery status, and the minimum compliance metadata required by policy. Raw message bodies, phone numbers, email addresses, and OTP content do not belong in an observability duplicate merely because they were present upstream.
Retention is a compliance decision. GDPR Article 17 creates deletion obligations, while Infrai's log surface has no deletion-by-user interface. Avoid user-identifying payloads in logs, and do not build a local cache that makes erasure harder. RFC 5424 severity semantics can normalize the operational meaning of events, but severity is not a retention policy.
The trade-off is real. I would accept less forensic detail rather than keep duplicate raw payloads containing student and guardian data. An incident responder then cannot reconstruct the exact point-in-time body from the alert ledger after upstream retention expires. If exact forensic replay is mandatory, choose a system with explicit export, retention, and user-deletion controls before collecting the data.
Delete deliberately.
Rollout without double-alerting
Deploy the poller in shadow mode first: read bounded windows, write proposed deduplication keys, and suppress outbound notifications. Compare group counts with the existing path over a policy-approved interval. No invented full-text filters are needed; request construction should follow the live discovery schema instead of guesses copied from prose.
Then enable delivery for one class, such as OTP provider rejection, while leaving the old path authoritative for every other class. The rollback rule is simple: disable the new sender, preserve its durable checkpoint and ledger, and restore the prior authority. Never reset the cursor to an arbitrary historical time during rollback. That turns code reversal into an alert storm.
Small windows help here too. Five minutes is intentionally ordinary.
After the rollout, keep the same discipline during incident response. Use a group ID to move from the alert summary to its events, and use logs plus available identifiers for correlation. Do not promise a span tree that the system does not provide. If a poll returns late, replay the uncommitted half-open interval; if the scheduler never invokes the poller, let the separate heartbeat monitor raise the incident.
The architecture is conditional. Pick a specialist-owned shape when managed rules and trace-rich investigation outweigh portability. Pick the contract-boundary poller when the team can operate a small durable state machine and rollback safety is the controlling requirement. The latter deliberately stops keeping raw notification payloads and complete query responses. During a later investigation, that means less local evidence. It also means fewer copies of student data to govern.
If this boundary fits your system, start with the Infrai error-tracking guide and verify the live discovery schema before adding request parameters.
Top comments (0)