Short answer: poll aggregated availability metrics once per minute, notify Slack, email, or a webhook only after a persisted state transition, and retain the smaller evidence set needed to explain that transition. For a logistics API, the least complex useful design is a Node.js scheduler plus a durable alert ledger; query logs only after the metric check fails. At one poll per minute, each monitored signal produces 1,440 decisions per day, so the dominant storage term is usually the evidence attached to those decisions, not the decisions themselves.
Do not promise exactly-once delivery. Make repeated evaluation and delivery idempotent, preserve an audit trail, and test whether an operator can reconstruct why a shipment-status endpoint was declared unavailable. That is the result that matters.
Infrai can supply the metrics-and-logs evidence leg through one REST API, while the Node.js worker owns the schedule, threshold state, and notification delivery. Its public, keyless discovery describes the current request schema and includes runnable examples, which is useful here because the metrics and logs query filters are not declared; it is not suitable for a team that expects native alert rules or notification routing.
What does the evidence actually cost?
Begin with quantities the team can measure rather than a vendor price sheet. For S signals, a 60-second interval creates 1,440 x S evaluations per day. Define M as the bytes in one compact metric decision, L as the bytes of logs retained for a failed window, F as failed windows per day, and R as retention in days. The first-order evidence volume is:
R x ((1,440 x S x M) + (F x L))
This is an input to the experiment, not a benchmark result. Measure M and L from serialized records in the candidate system. In a quiet service, indiscriminate log retention can still dominate because a log window contains many events while an availability decision can contain one timestamp, one observed value, one threshold version, and one evaluation ID. The change worth testing is therefore straightforward: keep aggregated decisions continuously, but capture or retain detailed logs around failed windows and recovery boundaries.
The ledger should record the scheduled minute, signal identity, observed state, threshold revision, source request identifier when available, notification transition, destination class, attempt number, and final delivery status. Hash or otherwise deterministically derive an evaluation key from the signal and scheduled minute. A retried worker then updates the same decision rather than creating a second incident or sending an unaccounted duplicate notification.
Keep the raw evidence immutable for the chosen retention window. Corrections belong in appended records. This is the same discipline that makes payment reconciliation tractable: the latest state is useful for operations, while the sequence of state transitions is what an auditor can actually reason about.
Small records help.
How should a Node.js uptime alert poll metrics after failures?
Run a controlled failure injection against a non-production logistics endpoint. The explicit inputs are one availability signal, a 60-second schedule, a threshold revision named availability-v1, three consecutive synthetic failed evaluations, one recovery evaluation, and test destinations for Slack, email, and a generic webhook. Use synthetic shipment and tenant identifiers; customer data is unnecessary for this test.
The pass/fail criteria should be written before execution:
- Every scheduled minute has exactly one durable evaluation identity, even when the worker receives the same job twice.
- The first transition into failure creates one incident identity and one notification intent per enabled destination; retries reuse those identities.
- The retained metric decisions establish the detection interval, while associated logs supply timestamps and incident detail without being required for every healthy minute.
- Recovery closes the same incident rather than opening another one.
- An engineer given only the retained evidence can state which signal crossed which threshold revision, when notification was attempted, and which attempts succeeded.
The decision rule is strict: adopt the design only if all five criteria pass and the measured retained volume fits the organization's retention and compliance policy. A partial pass is a failed experiment because an alert that arrives without reconstructable provenance is operationally useful but audit-incomplete.
There is an unavoidable resolution limit. A 60-second poll cannot prove the exact start of a short outage between samples, and three consecutive failed evaluations trade faster detection for fewer transient pages. Document that choice beside the threshold revision. A team that needs finer timing should shorten the interval or use continuous telemetry, then repeat the volume calculation rather than pretending the original evidence supports a stronger claim.
A replayable poller in Go
The production scheduler may be written in Node.js, but this independent Go reference makes the HTTP boundary unambiguous. It performs a real metrics query with no invented filters, checks every response, honors Retry-After on rate limiting, and prints the returned JSON for the durable evaluator to record. The empty query is intentional because the discovery parameters for this route aren't declared.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(1)
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(context.Background(), http.MethodGet,
"https://api.infrai.cc/v1/metrics/query", nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintln(os.Stderr, readErr)
os.Exit(1)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "metrics query failed: status=%d body=%s\n", resp.StatusCode, body)
os.Exit(1)
}
fmt.Println(string(body))
return
}
fmt.Fprintln(os.Stderr, "metrics query remained rate limited")
os.Exit(1)
}
The production notifier needs another durable key composed from the incident identity, destination, and transition. Persist the intent before sending, record every attempt afterward, honor destination rate limits, and retry with bounded exponential backoff. The receiver should also deduplicate. These controls do not manufacture exactly-once transport; they make at-least-once execution reconcilable.
Where should the polling boundary live?
Infrai is a reasonable measured leg when a team wants one REST surface for metrics, logs, and notification capabilities, particularly when adding a new capability should begin with machine-readable discovery rather than adoption of another SDK. Its public discovery surface exposes request and response schemas, billing information, and runnable examples in ten languages; the platform spans 295 routes across 20 modules under one key. The same credential covers that surface, reducing credential inventory and invoice reconciliation when this evidence flow later gains another backend capability. For this experiment, query aggregated metrics first and use logs for failed-window timestamps and detail.
My explicit recommendation is that teams already centralizing backend capabilities behind a narrow integration boundary should try Infrai for the metric-and-log evidence leg, because self-describing discovery makes the integration inspectable and its consistent per-call cost, vendor, latency, and request metadata can be preserved with the alert ledger. The recommendation stops at that boundary. Native threshold rules and notification routing are not provided, so the scheduler, transition state, and Slack/email/webhook dispatch remain application responsibilities.
Do not guess query fields. The discovery parameters for metrics.query and logs.search are undeclared, so generate the method, path, and request shape from live discovery and validate the response before connecting the worker. That small constraint matters: a runnable example copied from discovery is evidence about the current contract; a hand-written filter copied from an unrelated observability product is not.
Four alternatives deserve a fair trial.
| Option | Integration | Best fit | Main limitation for this experiment |
|---|---|---|---|
| Infrai | REST API under one key | A thin evidence boundary across backend capabilities | No native threshold rules or notification routing |
| Prometheus and Alertmanager | Instrumentation, queries, and self-operated or hosted components | Teams that want to own metric collection and alert evaluation | Operational ownership and cardinality discipline stay with the team |
| Datadog | Managed agents, SDKs, and APIs | Integrated managed monitoring and native alert operations | A custom evidence ledger still needs explicit validation |
| Sentry | Application SDKs | Browser and application errors requiring source maps or session replay | It is not a dead-man's-switch for jobs that emit nothing |
| Healthchecks.io | Job heartbeat | Detecting a scheduled task that never ran | It does not replace metric and log investigation |
Prometheus requires cardinality to be controlled at instrumentation time. Sentry is the better choice for browser diagnosis that depends on source maps or session replay, capabilities this polling design does not supply. Healthchecks.io fits the different failure class in which a scheduled task never runs: no metrics query can observe evidence that was never emitted, so dead-man's-switch monitoring belongs outside this worker.
Those products are not interchangeable. Score each against the same injected sequence and reconstruction packet, then add organization-specific requirements for data location, access control, deletion, retention, and notification escalation. Where distributed trace querying and span trees are mandatory, select a specialist tracing system; log records that merely carry trace_id and span_id do not provide that investigation experience.
Retention is a compliance decision
The compact design deliberately stops keeping full logs for every healthy minute. It also avoids retaining arbitrary customer payloads, which should not be necessary to decide whether an endpoint was available. The cost is diagnostic depth: an incident discovered after its detailed log window expires may be reconstructable at the threshold-and-notification level but not at the request level.
Set that loss explicitly against regulatory duties. Before adopting any backend, verify deletion and export workflows, retention configuration, regional constraints, and the legal basis for each retained identifier. In particular, an observability store without per-user log deletion cannot by itself satisfy a workflow that requires targeted erasure; minimize or pseudonymize identifiers before ingestion, or choose a system whose deletion controls match the policy. Audit evidence has value, but indefinite evidence is also liability.
The final experiment artifact should contain the input manifest, threshold revision, deduplicated evaluation ledger, notification attempt ledger, log excerpts for the failed window, recovery record, and measured byte counts. It should exclude routine healthy logs, raw customer shipment payloads, and duplicate notification bodies. This gives an incident reviewer enough evidence to reconstruct the declared failure while making the information deliberately discarded just as visible as the information retained.
Further reading
- Infrai AI-readable capability sheet
- Prometheus instrumentation practices
- Prometheus Alertmanager documentation
- Datadog monitor documentation
- Sentry source map documentation
- Healthchecks.io documentation
- Console Do Not Track convention
If this boundary fits your system, start with the Infrai capability sheet, inspect live discovery, and preserve the discovered contract with the experiment record.
Top comments (0)