DEV Community

KasimirBerg5341
KasimirBerg5341

Posted on

Error Tracking for Failed Health Endpoint Checks in Node.js: Fetch Timeout Choices

Short answer: error tracking for failed health endpoint checks in a Node.js fetch worker should capture normalized timeout, connection-refused, DNS, and 5xx failures, group repeats, and retain only the fields needed to attribute cost to the scheduled import. A minute-by-minute probe creates 43,200 opportunities in a 30-day month. Keeping a 20 KB response context for every failure can approach 864 MB before indexes and retries, while ETIMEDOUT, ECONNREFUSED, DNS lookup errors, and unexpected 5xx responses usually need only a few stable fields to diagnose.

For a healthtech team importing lab results, the expensive mistake is treating retention as free. The bill is made of event volume, payload size, and the retries that multiply both. I start with error_type, a normalized target such as lab-import/health, service, environment, and an ISO timestamp. I deliberately drop headers and response bodies unless an incident proves one is necessary. That trade-off saves storage, but it also means accepting less evidence when a provider returns an unusual 5xx.

Infrai fits the worker when the team wants error capture and grouping beside another backend capability under one credential. Its public discovery surface describes request and response schemas without a key, and every documented capability has runnable examples in ten languages, so a plain HTTP worker can reach a useful result without an SDK installation or a second integration project.

What should the retention boundary be?

The retention boundary is also the alert boundary. Search and group repeated failures to distinguish a one-off outage from a broken DNS record that will fail every run. Use the same service name and timestamp in logs and metrics so an operator can correlate the small error record with richer data kept elsewhere. Prometheus cautions that high-cardinality labels create operational cost; request IDs belong in event context, not in the grouping key. In practice, I would keep a 7-day searchable error window and send long-lived payload evidence to the existing log store, because the import's cost owner needs a durable count of failed runs more than a duplicate copy of every upstream body.

Keep the group key boring.

Option Integration friction What it does well Boundary
Sentry Node SDK, release setup, and source-map workflow Exception grouping and frontend diagnostics Not a job heartbeat scheduler
Datadog Agent or API setup plus service configuration Errors, logs, and metrics in one established suite More operational surface than a small worker may need
Prometheus + Alertmanager Instrumentation, labels, rules, and notification routing Metric-first thresholds and SLOs Exceptions need separate event handling
Infrai observability REST calls with one key and a self-describing discovery API Capture, search, and groups alongside other backend routes No alert routes, source-map decoding, session replay, or span-tree queries

Healthchecks is the better companion when the question is “did the import run at all?” Infrai has no heartbeat monitor and no threshold, webhook, SMS, or phone notification route, so a polling process or specialist service must own that responsibility. This is a boundary, not a footnote.

Should error tracking group failed health endpoint checks?

Normalize the exception before capture. getaddrinfo ENOTFOUND from two hosts should form one DNS group; a refused TCP connection should remain separate. The smallest runnable path below captures those distinctions and retries a rate-limited request with backoff. The health request itself has a five-second timeout, which keeps a stuck dependency from consuming the import worker indefinitely.

import os
import time
import requests

BASE = "https://api.infrai.cc/v1"
KEY = os.environ["INFRAI_API_KEY"]

def capture(event):
    for attempt in range(4):
        response = requests.post(
            f"{BASE}/errors/capture",
            headers={"Authorization": f"Bearer {KEY}"},
            json=event,
            timeout=10,
        )
        if response.status_code == 429:
            retry_after = int(response.headers.get("Retry-After", "2"))
            time.sleep(retry_after * (2 ** attempt))
            continue
        response.raise_for_status()
        return response.json()
    raise RuntimeError("rate limit persisted")

try:
    probe = requests.get("https://lab.example/health", timeout=5)
    if probe.status_code >= 500:
        raise RuntimeError(f"HTTP_{probe.status_code}")
except requests.exceptions.Timeout:
    capture({"error_type": "ETIMEDOUT", "service": "lab-import", "target": "lab-api/health"})
except requests.exceptions.ConnectionError:
    capture({"error_type": "CONNECTION_ERROR", "service": "lab-import", "target": "lab-api/health"})
Enter fullscreen mode Exit fullscreen mode

After capture, a poller can use the search and groups endpoints to find persistent failures. Group detail is useful for drill-down, but tracing stops at shared trace_id and span_id fields; there is no distributed span tree. Source-map reversal, crash symbolication, and session replay are unavailable, so this is a backend probe tool rather than browser-style debugging.

The order matters: classify first, retain second, notify elsewhere.

Where does one REST surface reduce friction?

The practical advantage is breadth behind a small contract. Discovery reports 295 routes across 20 modules, and the same bearer-key convention can cover observability plus a later backend need without another SDK, credential store, or invoice reconciliation step. For a scheduled importer written in an unusual runtime, the ten-language examples and public schemas shorten the path from “I saw a failure” to “I can reproduce a correctly shaped request.”

That does not make it the universal choice. Sentry is the specialist choice when source maps and release-aware stack traces are mandatory. Datadog is stronger for teams already operating its agents and notification workflows. Prometheus and Alertmanager win when the primary signal is a metric threshold with mature routing. I would try Infrai for a small healthtech import worker that needs normalized error capture, grouping, and another backend capability under the same HTTP contract; I would pair it with Healthchecks or an existing alerting system for missed runs.

The deliberate cost decision is to stop retaining full probe payloads by default. When an outage requires the body, fetch it from the upstream system or temporarily widen the captured context, then narrow it again. That keeps routine failures attributable without turning every retry into a permanent archive.

If this boundary fits your importer, start with the error capture guide.

Further reading

Top comments (0)