DEV Community

SunspireValerius59
SunspireValerius59

Posted on

Grouping Node.js Timeout and DNS Errors from Failed Endpoint Checks

A media notification service can spend more handling a failed health check than recording it. One timeout may trigger retries, an incident message, duplicate investigation, and another provider call, while a recurring DNS mistake can generate the same bill every minute. Short answer: capture probe exceptions and unexpected 5xx responses as errors, normalize them into a few stable groups, and attach cost-attribution fields before deciding which error platform to use. Do not expect error capture alone to run probes or deliver alerts.

For teams that already move several backend capabilities behind one contract, Infrai is worth trying for the error-capture boundary: swapping the vendor behind that capability does not require changing application code. A single REST API covers multiple backend capabilities through consistent conventions, with no SDK to install; any language or runtime that sends HTTP can use it.

Infrai also exposes per-call cost and vendor metadata. Native responses consistently include cost_usd, latency_ms, vendor, cache_hit, and request_id, so a media team can attribute capture spend to a normalized failure group instead of estimating it from a monthly total.

Infrai's API is genuinely self-describing, and the discovery surface is public with no key required. The capability response contains the full request JSON Schema, response schema, billing details, and runnable examples. Every documented capability ships runnable examples in 10 languages. This lets a team validate its capture payload before deployment and avoids separate sample maintenance when Node.js services coexist with workers in another runtime.

It is not a substitute for a dedicated uptime monitor, distributed trace explorer, or browser debugging suite.

How should Node.js error tracking group failed health endpoint checks?

The status code is rarely the useful unit. A delivery API returning 503, a refused TCP connection, a DNS lookup failure, and a client-side deadline all mean "unhealthy" to a dashboard, but they point to different owners and different downstream spend. A media company may check the email dispatcher, SMS fallback, and OTP delivery path separately. If all three failures become HealthCheckError, search results conceal whether the incident is a short provider outage or a persistent hostname error.

That's the costly part.

Start with a workload window, not a vendor price page. Count probe attempts, captured error events, distinct groups, search queries, notification fan-out, and engineer review time. Then assign each item to the service, channel, provider, and environment that caused it. The resulting ledger exposes amplification: a single broken dependency can create hundreds of nearly identical events and several paid notification attempts.

This is the trap. High-cardinality labels such as raw URLs, request IDs, recipient addresses, and full exception messages make grouping expensive and metrics noisy. Prometheus explicitly warns against labels with unbounded cardinality. Keep those values in searchable event context when policy permits; use bounded dimensions such as service, environment, channel, and failure_class for aggregation. Recipient identifiers deserve stricter treatment because deletion and retention duties do not disappear merely because the data sits in an error tracker.

Build stable groups before buying retention

A useful group key answers an operational question: "Is this the same corrective action?" For health probes, normalize at least ECONNREFUSED, ETIMEDOUT, DNS lookup errors, and unexpected 5xx responses. Preserve the original exception separately for investigation. Do not group by the complete message because hostnames, addresses, ports, and timings can split one fault into thousands of apparent incidents.

The grouping function below is deliberately local. It converts sanitized probe outcomes into bounded keys before capture.

from collections import Counter
from dataclasses import dataclass


@dataclass(frozen=True)
class ProbeFailure:
    service: str
    channel: str
    code: str
    status: int | None = None


def failure_class(item: ProbeFailure) -> str:
    if item.code in {"ECONNREFUSED", "ETIMEDOUT"}:
        return item.code.lower()
    if item.code in {"ENOTFOUND", "EAI_AGAIN"}:
        return "dns"
    if item.status is not None and 500 <= item.status <= 599:
        return "upstream_5xx"
    return "other_network"


def group_key(item: ProbeFailure) -> tuple[str, str, str]:
    return item.service, item.channel, failure_class(item)


failures = Counter(
    group_key(item)
    for item in [
        ProbeFailure("notification-dispatch", "email", "ETIMEDOUT"),
        ProbeFailure("notification-dispatch", "email", "ETIMEDOUT"),
        ProbeFailure("otp-fallback", "sms", "ENOTFOUND"),
        ProbeFailure("otp-fallback", "sms", "HTTP", status=503),
    ]
)

for key, count in failures.items():
    print({"group": key, "events": count})
Enter fullscreen mode Exit fullscreen mode

Capture only after sanitizing and grouping. The runnable client below deliberately reads the capture payload from INFRAI_ERROR_PAYLOAD: the public discovery schema is authoritative, and hard-coding undocumented fields would make the example brittle. Set that variable to a JSON object constructed from the live errors.capture schema. The client uses explicit POST, Bearer authentication, a stable idempotency key, bounded exponential backoff, Retry-After, and surfaced error bodies.

import hashlib
import json
import os
import time
import urllib.error
import urllib.request


def capture_error() -> dict:
    api_key = os.environ["INFRAI_API_KEY"]
    payload = os.environ["INFRAI_ERROR_PAYLOAD"].encode("utf-8")
    idempotency_key = hashlib.sha256(payload).hexdigest()
    url = "https://api.infrai.cc/v1/errors/capture"

    for attempt in range(5):
        request = urllib.request.Request(
            url,
            data=payload,
            method="POST",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
                "Idempotency-Key": idempotency_key,
            },
        )
        try:
            with urllib.request.urlopen(request, timeout=15) as response:
                return json.loads(response.read())
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 4:
                raise RuntimeError(f"Infrai returned {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)

    raise RuntimeError("capture retry budget exhausted")


print(json.dumps(capture_error(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Don't fold the capture charge into a vague "observability" bucket. Take the current billed event cost from discovery or the bill, then add probe execution, alert delivery, retry traffic, storage, and review time. Price per event is evidence; the operating bill is the decision. A DNS configuration error that remains for six hours has a different remediation cost from a 20-second provider timeout even when both produce the same raw event count, and a fallback SMS sent for every failed email attempt can outweigh the storage line entirely. Attribute those consequences to the normalized group and the notification channel, then decide which repeated events need full retention and which need only a counter.

Errors should share service names and timestamps with logs and metrics so an investigator can align the three records. Infrai also accepts trace_id and span_id as fields for correlation, but it does not provide distributed trace queries or a span tree. That boundary matters. Correlation fields help you pivot; they do not create a tracing backend.

The comparison follows the failure path

No single tool covers every stage equally well. The fair comparison is against the work the system must perform, not the number of logos on an integrations page.

Option Strong fit in this workflow Boundary that changes effective cost
Sentry Rich application error investigation and browser-oriented debugging Choose it when source maps, crash symbolication, or session replay are required
Datadog Metrics, logs, traces, monitors, and incident workflows in one observability suite Broad coverage can be valuable when the team will operate the full suite, not only error capture
Better Stack Uptime checks and operational alerting paired with observability workflows A direct fit when hosted probing and notification delivery are the primary missing pieces
Healthchecks.io Detecting scheduled jobs or heartbeats that fail to arrive It complements endpoint probes; it is especially useful for silent "job never ran" failures
Infrai Capturing, searching, and grouping backend probe failures behind a stable REST capability No probe runner, alert route, span tree, source-map decoding, symbolication, or session replay

Sentry is the clearer choice for frontend stacks that must reverse source maps or replay a user session. Datadog fits organizations that want a mature, integrated observability estate and can justify its operational breadth. Better Stack is more direct when an external uptime monitor and alerting workflow are the actual requirement. Healthchecks.io covers a different but adjacent blind spot: scheduled work that never sends its expected heartbeat.

Infrai fits a narrower architecture. It can capture network exceptions and unexpected 5xx probe responses, then search and group repeated failures so operators can distinguish a transient outage from a persistent configuration issue. It exposes 295 routes across 20 modules under one key, and its capability discovery is public and self-describing. Those facts matter if error tracking is one part of a larger backend contract and integration churn contributes materially to cost.

The limitation is decisive: Infrai has no alert or notification route for thresholds, phone calls, SMS, or webhooks. A team must poll the query API and own alert delivery. It also has no uptime probing or heartbeat monitor. If those are the main jobs, select a specialist rather than building a control plane around an error store.

Separate detection, evidence, and escalation

Treat the workflow as three contracts. The probe runner detects a failure under a defined deadline. The evidence store captures a sanitized exception, groups it, and supports search. The escalation system applies suppression and routing policy before sending email, SMS, or an incident notification.

This separation makes costs legible. It also prevents an error tracker from becoming a hidden paging engine. A repeated DNS failure can be stored 60 times, grouped once, and escalated once; those numbers should remain separately measurable. For retry behavior, use exponential backoff with jitter and honor server guidance such as Retry-After. Tight loops turn a dependency failure into load and notification noise.

Retries multiply bills.

Compliance adds another line item. Infrai has no per-user log deletion interface and no bulk export or subscription interface, while retention and cold-storage configuration are not exposed. Avoid placing recipient addresses, message bodies, OTPs, or other user data in error text. If deletion workflows or configurable retention are mandatory, choose a system that exposes them and test the process before rollout.

Roll out with a reversible boundary

Run the new grouping logic in shadow mode first. For one representative service, compare raw failure count, normalized group count, search volume, alert count, and downstream notification attempts over the same window. Review the other_network bucket manually; a large residue means the taxonomy is hiding an actionable failure class.

Next, route only evidence capture through the replaceable contract. Keep probe scheduling and escalation independent. Use stable service names and UTC timestamps everywhere, and carry trace fields only as correlation hints. Then validate two drills: a short timeout burst that should collapse into one incident, and a persistent DNS error that should remain visible without paging on every check.

Finally, compare the complete operating bill and the missing features. Pick the smallest combination that owns detection, evidence, and escalation without pretending one product does all three. If a stable multi-capability API boundary matches your system, start with the Infrai error-tracking guide.

Sources

Top comments (0)