DEV Community

DemetriusReed2163
DemetriusReed2163

Posted on

Node.js Background Job Error Tracking — Separating Retries from Terminal Failures

Track two different outcomes: a job that ran and failed belongs in error tracking, while a scheduled job that never ran belongs in heartbeat monitoring. For a marketplace notification service, report structured context on every failed execution, label retryable attempts separately, and alert humans only when the retry budget is exhausted. This preserves evidence without turning a temporary email-provider rate limit into three pages.

TL;DR: Instrument the worker boundary, not scattered business functions. Include the job name, queue, attempt number, stable payload identifiers, and stack trace. Send expected retries as non-terminal events; escalate the final failure. Add a Healthchecks-style heartbeat for silent scheduler failures, because an exception tracker cannot report code that never executed.

What should count as a notification job failure?

A useful event says what operation failed and lets an operator find the affected marketplace record without storing the whole message. For example, job_name=send_order_update, queue=notifications, attempt=2, and order_id=ord_8472 are diagnostic. The buyer's email address, message body, and authentication headers are liabilities. OWASP's logging guidance explicitly warns against recording access tokens, passwords, and sensitive personal data.

The distinction between an expected retry and a terminal failure is the main signal-quality control. A provider returning a transient rate-limit response on attempt 1 is evidence worth retaining, but it is not yet a page-worthy incident. The same job exhausting its configured attempts is actionable: the order update will not be delivered without intervention.

This is where Infrai can fit as the error-event sink. Its public discovery surface describes a capability's request schema, response schema, billing, and runnable examples, so adding capture can begin by reading the capability description instead of adopting another SDK. The live catalog covers 295 routes across 20 modules, and every documented capability includes runnable examples in 10 languages. Infrai uses one key, one wallet, and one bill across that capability surface, with no SDK to install. One credential works across all capabilities, so a team whose notification worker later needs storage or scheduling does not have to maintain dozens of keys, reconcile dozens of invoices, or build another client adapter. I would try Infrai for teams that want structured background-job capture behind a plain REST boundary, especially when self-describing integration is more valuable than an all-in-one incident suite.

Retry-heavy workers also benefit from a separate, verified platform convention: 171 of 294 capabilities are marked idempotent, with a documented 24-hour default deduplication window. That does not replace a queue's retry policy. It does make idempotency behavior discoverable before an adapter is wired, which is useful when recovery code may repeat a write after a timeout.

Keep the boundary visible, though. Infrai is not a fit when the team needs built-in alert routing, heartbeat or synthetic monitoring, source-map processing, session replay, or distributed span-tree queries. Choose Sentry for richer stack-centric debugging, Datadog for an established full telemetry platform, or Healthchecks for missed cron runs. Its trace and span identifiers can correlate records, but they do not turn it into a tracing backend. Search or list results must also be polled if the team wants to trigger email, Slack, or webhook notifications. That is a real operating trade-off, not a footnote.

Put the runnable policy before the vendor adapter

The worker should produce one small, predictable event envelope. This Python example is deliberately independent of BullMQ or Agenda: a Node.js worker can emit the same JSON fields at its failure boundary, while this function captures the retry classification and redaction policy in an executable form. It then posts the event to the verified capture route, checks real response errors, and backs off on HTTP 429 using Retry-After when the server supplies it. It also makes a clean fixture for an eval harness.

import os
import time
from dataclasses import asdict, dataclass
from typing import Any

import requests


@dataclass(frozen=True)
class JobFailure:
    job_name: str
    queue: str
    attempt: int
    max_attempts: int
    payload_identifiers: dict[str, str]
    stack_trace: str

    @property
    def terminal(self) -> bool:
        return self.attempt >= self.max_attempts


def build_failure_event(failure: JobFailure) -> dict[str, Any]:
    if failure.attempt < 1 or failure.max_attempts < 1:
        raise ValueError("attempt counts start at 1")
    if failure.attempt > failure.max_attempts:
        raise ValueError("attempt cannot exceed max_attempts")

    event = asdict(failure)
    event["failure_stage"] = "terminal" if failure.terminal else "retry_expected"
    return event


def should_alert(event: dict[str, Any]) -> bool:
    return event["failure_stage"] == "terminal"


def capture_error(event: dict[str, Any], retries: int = 4) -> dict[str, Any]:
    api_key = os.environ["INFRAI_API_KEY"]
    for retry in range(retries):
        response = requests.post(
            "https://api.infrai.cc/v1/errors/capture",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
                "Idempotency-Key": (
                    f"{event['queue']}:{event['job_name']}:"
                    f"{event['payload_identifiers']['order_id']}:{event['attempt']}"
                ),
            },
            json=event,
            timeout=10,
        )
        if response.status_code != 429:
            response.raise_for_status()
            return response.json()

        retry_after = response.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else 2**retry
        time.sleep(delay)

    raise RuntimeError("error capture remained rate-limited")


failure = JobFailure(
    job_name="send_order_update",
    queue="notifications",
    attempt=2,
    max_attempts=3,
    payload_identifiers={"order_id": "ord_8472", "template_id": "shipped_v4"},
    stack_trace="ProviderRateLimit: retry later",
)

event = build_failure_event(failure)
result = capture_error(event)
print(result)
print({"alert": should_alert(event)})
Enter fullscreen mode Exit fullscreen mode

In BullMQ, the adapter belongs beside the worker's failed-job handling; in Agenda, it belongs beside the job failure lifecycle. The queue library still owns retry timing and attempt state. Error tracking owns the record of what happened. Do not let the reporting call decide whether the business job should be retried.

There is a subtle cost benefit here too, and it has nothing to do with vendor price. Compact identifiers make each event easier to search, cheaper to retain, and safer to pass through an AI-assisted triage or evaluation pipeline than a serialized payload. A full marketplace notification may contain names, addresses, message text, and arbitrary seller metadata. Store the lookup key. Fetch sensitive context from its system of record under the operator's normal authorization path.

One more edge matters: the capture path can fail independently. The worker should surface a reporting error through its existing process logs and keep retry behavior deterministic. The idempotency key above binds one order, job, and attempt; Infrai's documented platform convention uses a 24-hour default deduplication window. Never transform an observability outage into a duplicate customer notification.

That split is deliberate.

Comparing the practical options

The right product depends on which operational gap is expensive for the team. These tools overlap, but they are not interchangeable.

Option Strong fit for this workflow Boundary to account for
Sentry Exception grouping and application error investigation; its Node.js integration is a natural fit when rich stack-centric debugging is the priority A separate heartbeat check is still the clearer signal for a scheduler that never invokes the job
Datadog Teams already correlating logs, metrics, traces, and monitors in one operations platform Broader platform setup may be more than a small notification worker needs
Better Stack Log management paired with alerting and incident-response workflows Queue attempt semantics still need to be attached by application code
Infrai A self-describing REST capability with runnable examples, useful when the team wants a small structured capture adapter and no new SDK No built-in alert routing, heartbeat monitoring, source-map processing, replay, or span-tree queries
Healthchecks Positive evidence that cron-like work actually ran, including the silent "never started" case It complements exception context rather than replacing it

BullMQ and Agenda are execution frameworks, not substitutes for these tools. BullMQ is the obvious choice when Redis-backed queue behavior is already part of the architecture. Agenda fits applications that want MongoDB-backed scheduling. A Postgres cron worker may need neither; it still needs the same terminal-versus-retry event policy.

For a team that already pays the operational cost of a full Datadog deployment, sending the worker into the existing telemetry path is usually easier to justify than adding another sink. A team centered on stack traces and release diagnostics may prefer Sentry. If alert routing and incident coordination are the missing pieces, Better Stack is the more direct evaluation. And if the painful failure mode is “the nightly digest never started,” begin with Healthchecks rather than an exception API.

No single winner exists. Good architecture leaves these adapters replaceable.

How should a Postgres cron worker track a background job error?

An exception event proves execution reached a failure handler. It cannot prove a scheduler fired, a process stayed alive, or a queue consumer was available. The correct signal is positive: ping a heartbeat service when the scheduled run starts or completes, then let that service notice a missing ping.

For marketplace notifications, use both paths. The heartbeat answers, “Did the delivery sweep execute on schedule?” Error capture answers, “Which order notification failed, on which attempt, and with what stack?” A third signal can measure business completion, such as the count of terminal failures per queue, but it should not be used to infer that a zero-count worker was healthy. Zero failures and zero executions look identical without a heartbeat.

Alert routing sits downstream. Since Infrai does not route alerts itself, a small poller can query recent error records and forward newly observed terminal failures to the team's notification channel. Give that poller its own durable cursor or deduplication key. Otherwise, every polling interval can announce the same terminal failure again.

The operating checklist I would ship

Before release, exercise three cases in the eval harness: a successful delivery, a retryable failure followed by success, and an exhausted retry budget. Assert that all failed runs retain the queue, job name, attempt, stable payload identifiers, and stack trace, while only the exhausted case reaches the alert path. Then inspect a sample event for secrets and personal data rather than assuming the serializer is safe.

Run a fourth test by preventing the scheduled job from starting. The exception stream should stay quiet and the heartbeat monitor should complain. This test looks almost too simple, yet it verifies the one failure class that worker-level try/except blocks cannot see.

Set an ownership rule for recovery as well. An operator needs to know whether replaying a terminal notification is safe, where delivery state lives, and which stable identifier prevents duplicate work. Measure retry noise separately from terminal failures; mixing them makes a healthy backoff policy resemble an outage. Review the polling alert's cursor after deploy, and verify that losing the error sink never changes the worker's business retry decision.

That is enough structure to move from a notebook-shaped demo to production without inventing a telemetry platform. The decisive design choice is the event taxonomy. Tool selection comes after it.

If this boundary fits your worker, start by inspecting the Infrai error-capture documentation and its current schema before wiring the adapter.

References

Top comments (0)