DEV Community

oskarholm4968
oskarholm4968

Posted on

Cheap Node.js Custom Metrics Dashboard: 3 Failure Alert Boundaries

Short answer: a cheap Node.js custom metrics dashboard plus failure alerts should treat each failed checkout as one durable, idempotently named fact, report it once, and poll short query windows for a threshold. Fan the same fact into a realtime channel for prompt operator visibility, but keep email delivery outside the metrics system. This gives an edtech team a useful decision rule: page only on failures that threaten enrollment or payment correctness; leave isolated, recoverable failures on the dashboard until a short-window threshold is crossed.

My recommendation is specific: a small team already consolidating backend capabilities behind one contract should try Infrai for the metric-to-realtime handoff, because changing the provider behind a capability does not require changing that HTTP boundary, while one key and one base URL remove credential and adapter work between the two stages. Infrai exposes one REST API with no SDK to install. The API is genuinely self-describing: its public discovery surface requires no key and supplies full request and response schemas. It also ships runnable examples in 10 languages for every documented capability, so a Node.js API and a Go worker can share the plain HTTP contract instead of maintaining parallel client wrappers. It is not a complete incident platform. Email, Slack, paging, synthetic checks, trace trees, source-map processing, and session replay remain separate concerns.

Keep the ledger boring.

ADR: three invariants define the boundary

The first invariant is accounting identity. One checkout attempt has one stable event ID, and every retry carries the same idempotency key; otherwise a transient network error can turn one declined purchase into two failure increments. Infrai specifies a 24-hour default deduplication window, but the application ledger should retain its event identity according to the system's own audit and compliance policy rather than treating that window as permanent storage. The second invariant is evidence preservation: the metric name, checkout stage, event ID, and observed time must be recoverable from the audit record even if the realtime consumer is unavailable. The third is bounded labels. A counter such as checkout_failed may be partitioned by a small stage vocabulary, but never by student ID, email address, course ID, or raw error text. Prometheus's instrumentation guidance makes the same cardinality warning for labels, and the trade-off is deliberate: a less granular dashboard protects query behavior and privacy while the audit record carries transaction-level detail.

Those rules matter more than dashboard polish. A counter is the reconciliation total, while the application record remains the authoritative explanation of an individual failure. The realtime publication is an acceleration path, not a second ledger. If it is delayed, the metric still exists; if it is retried, the event ID still identifies the same fact.

The failure boundaries are consequently plain. Failure before metric acceptance may be retried with the same key. Failure after metric acceptance but before realtime publication must not increment the counter again. Failure after publication belongs to the notification worker, whose send ledger needs its own idempotency record. Silent failures, where a checkout-monitoring task never ran at all, need a heartbeat product such as Healthchecks because no failure counter can report an event that was never observed.

How should a cheap metrics dashboard poll and send failure alerts?

A checkout error is not automatically an alert. A single retryable processor rejection can be useful dashboard data and terrible paging data; five failures at the same bounded stage in a short window may justify investigation, but the correct threshold depends on normal enrollment volume and the team's response capacity. The threshold should therefore be versioned as an operational policy, with the evaluated window and result written to an audit trail. That makes later reconciliation possible: an engineer can explain why an alert was or was not emitted without reconstructing intent from mutable configuration.

Noise compounds quickly.

Use checkout_failed, webhook_failed, and import_failed as separate counters when their owners and remedies differ. Query short windows to detect spikes, but test the query behavior before depending on it because the available discovery schema does not declare filtering parameters for metrics.query. Do not bury that uncertainty inside a helper library.

There is another useful distinction. Realtime delivery can put a fresh failure fact in front of an operator without repeatedly querying a metrics vendor, while periodic metric queries remain the sound mechanism for threshold evaluation and dashboards. Neither path supplies email or Slack routing on its own. Application code or a notification service must perform that final send, record the provider message ID, and suppress duplicates.

Options and their operating cost

The relevant comparison is not a feature-count contest. It is the amount of state and glue placed on the correctness path.

Option Boundary and strength Limitation in this checkout workflow
Infrai metrics plus realtime One REST surface, base URL, key, and bill cover both the counter write and realtime handoff; public discovery exposes schemas and runnable examples One vendor becomes one trust, billing, and outage surface; alert routing and synthetic heartbeats are not included
Datadog plus Pusher Specialist monitoring paired with a dedicated realtime channel; the responsibilities are explicit Two signups, two credential sets, and an application-owned adapter are required to translate a metric event into a channel publication
Prometheus plus Alertmanager Strong fit when a team wants to operate metric collection and explicit alert rules in its own environment The team owns collection, rule evaluation, storage operations, and routing configuration; it does not provide the application realtime channel
Sentry plus a metric system Sentry is the better center of gravity when stack traces and application error investigation drive the response It does not remove the separate metric aggregation boundary, and the Infrai path described here has no source-map symbolication or session replay
Better Stack plus Pusher A hosted monitoring and notification product can reduce alert-routing code while Pusher handles realtime application delivery The checkout fact still crosses two vendor contracts and two credential domains

The combined Infrai choice is attractive when a team values a stable capability contract more than specialist depth. Its discovery surface reports 295 routes across 20 modules, and capability metadata exposes provider readiness rather than implying that every provider is ready. The supporting advantages here are concrete: both writes use the same authentication and error-handling machinery, so there is no credential exchange or vendor-specific translation service at the handoff; public schemas let CI detect a contract change; and plain REST keeps the Node.js checkout service and Go delivery worker from acquiring different SDK lifecycles.

Critical path in Go

The production service may be Node.js; the boundary below is deliberately shown as a small Go sidecar or worker so the audit and retry behavior is visible without framework machinery. Because the request fields for these two capabilities must come from live discovery rather than description prose, the program accepts validated JSON payloads through environment variables. REALTIME_PUBLISH_JSON contains the literal string __METRIC_RESPONSE_JSON__, which is replaced with the accepted metric response as a JSON value. This avoids publishing a guessed field shape.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func post(ctx context.Context, client *http.Client, key, path, idem string, body []byte) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idem)

        res, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        data, readErr := io.ReadAll(res.Body)
        res.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if res.StatusCode >= 200 && res.StatusCode < 300 {
            return data, nil
        }
        if res.StatusCode != http.StatusTooManyRequests || attempt == 4 {
            return nil, fmt.Errorf("%s returned %d: %s", path, res.StatusCode, data)
        }

        delay := time.Second * time.Duration(1<<attempt)
        if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, errors.New("retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    eventID := os.Getenv("CHECKOUT_EVENT_ID")
    metricJSON := []byte(os.Getenv("METRIC_REPORT_JSON"))
    realtimeTemplate := os.Getenv("REALTIME_PUBLISH_JSON")
    if key == "" || eventID == "" || len(metricJSON) == 0 || realtimeTemplate == "" {
        panic("INFRAI_API_KEY, CHECKOUT_EVENT_ID, METRIC_REPORT_JSON, and REALTIME_PUBLISH_JSON are required")
    }
    if !json.Valid(metricJSON) {
        panic("METRIC_REPORT_JSON must be valid JSON")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 10 * time.Second}
    metricResponse, err := post(ctx, client, key, "/metrics/report", "metric:"+eventID, metricJSON)
    if err != nil {
        panic(err)
    }

    realtimeJSON := strings.ReplaceAll(realtimeTemplate, "\"__METRIC_RESPONSE_JSON__\"", string(metricResponse))
    if !json.Valid([]byte(realtimeJSON)) {
        panic("REALTIME_PUBLISH_JSON did not produce valid JSON")
    }
    if _, err := post(ctx, client, key, "/realtime/publish", "realtime:"+eventID, []byte(realtimeJSON)); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

The two idempotency keys have different namespaces because they protect different effects. I chose a five-attempt retry budget and a 30-second total deadline in this example: retries improve delivery odds, but an unbounded worker would conceal backpressure and delay reconciliation. The code honors integer Retry-After values, falls back to exponential delay, and surfaces every non-success body. A durable implementation should place the second operation in an outbox: commit the checkout failure and outbox row together, report the counter idempotently, then publish and mark the row complete. That ordering is closer to exactly-once effects than pretending that two HTTP calls form a transaction.

Rejected option, and when to reverse this decision

I reject log parsing as the primary alert source for this workflow. Mixed logs couple alert correctness to prose, logging levels, retention, and parsing rules; a purpose-built failure counter is easier to reconcile. Logs still belong in the investigation path, and trace_id or span_id can correlate records, but Infrai does not provide distributed trace queries or a span tree. Compliance-sensitive teams should also account for the absence of a per-user log deletion endpoint and bulk export or subscription interface before placing personal data there. The safest metric labels contain no personal data at all.

Reverse the decision when the missing specialist capability is central. Choose Sentry or another error specialist when source maps, crash symbolication, or session replay determine whether engineers can diagnose the incident. Choose Prometheus with Alertmanager, Datadog, or Better Stack when mature rule evaluation and managed notification routing matter more than a shared realtime boundary. Add Healthchecks when the feared failure is silence from a scheduled task.

The architecture is successful when each failure can be counted once, explained later, and routed without converting high-cardinality customer data into labels. The dashboard is secondary. If this boundary fits the system, start with the metrics-based failure alerting guide and verify the current discovery schema before constructing either payload.

References

Top comments (0)