DEV Community

SolaceW31
SolaceW31

Posted on

Rollback-Safe Node.js Background Job Error Tracking for BullMQ and Postgres Workers

Short answer: report every failed execution with structured job context, but page only when the retry budget is exhausted. Monitor the schedule separately. Error tracking can explain why a BullMQ, Agenda, or Postgres-backed worker failed; it cannot prove that a job which produced no event ever ran.

For a B2B SaaS notification service, that split is the rollback-safe design. A bad deploy can be rolled back without erasing the evidence, expected retries do not flood the delivery team, and a missing OTP dispatch is detected even when there is no exception to capture. Keep the reporting adapter thin and asynchronous enough that observability trouble cannot turn a provider rejection into a worker outage.

How should a Postgres cron worker track background job errors?

The first invariant is mundane and important: acknowledging the queue item and reporting its failure are different decisions. BullMQ standard processing is at-least-once in practical failure conditions, so the notification operation needs its own idempotency key. A retry must not send the same email or SMS twice merely because a process lost its lock or restarted near acknowledgement.

The second invariant is that every attempted run has a stable execution identity and a stable business-operation identity. The execution identity distinguishes attempt 1 from attempt 3. The operation identity lets an operator connect those attempts to one notification without putting an email address, phone number, OTP, message body, or access token into the error record. OWASP's logging guidance is explicit about excluding secrets and sensitive personal data. Store opaque identifiers, then enforce deletion in the system of record; do not treat an error index as the authoritative customer record.

The third invariant is deploy independence. Capture the release identifier, job name, queue, attempt number, payload identifier, and stack trace. Preserve the schema across the old and new worker versions for at least one rollback window. Otherwise, rolling back the producer can leave the consumer unable to interpret queued work, which is a much larger incident than a noisy exception.

Finally, expected retries and terminal failures must be distinguishable. An SMTP timeout on attempt 1 of 4 is evidence, not necessarily an alert. The same failure after attempt 4 is actionable. Keep both for diagnosis, but give them different lifecycle states so the operations view is not a wall of transient delivery problems. The trade-off is extra policy in the worker adapter, and I prefer that explicit policy to a dashboard filter somebody has to remember during an incident.

Failure boundaries and ownership

There are four boundaries in this design. The scheduler proves that work was requested. The queue proves that work was accepted. The worker reports an attempted execution. The email or SMS provider reports downstream delivery state. Those signals answer different questions, and merging them under one generic failed counter destroys useful information.

The awkward case is silence. If a Postgres cron row was never selected, an Agenda lock was never acquired, or the scheduler process stopped before enqueueing, no exception reaches an error tracker. A Healthchecks-style heartbeat closes that gap: ping on schedule, and alert when the expected ping is absent. It complements error capture; it does not replace it.

No ping, no proof.

Alert routing is another boundary. If the error service has no threshold rules or email, Slack, phone, SMS, or webhook routing, a small control-plane process must poll its list or search surface, apply the terminal-failure policy, and notify through an independently operated channel. That process needs a cursor and deduplication key. Re-reading the same terminal event must not send the on-call message twice.

Comparing the realistic options

The product decision should follow the failure boundary, not brand familiarity. These options overlap, but they are not interchangeable.

Option Best fit in this design Rollback-safety trade-off
Sentry Exception grouping and worker error investigation A mature error-focused workflow is useful when stack traces are the primary evidence; missed schedules still need a heartbeat signal.
Datadog Error Tracking Teams already correlating application errors with Datadog telemetry Consolidation can reduce context switching, but rollout and rollback are coupled to the team's existing telemetry integration. Schedule silence remains a separate check.
Honeycomb Investigation centered on traces and high-cardinality event context Stronger fit when a job is one span in a distributed request; a span-oriented workflow is more instrumentation than a team needs for basic terminal-failure capture.
Healthchecks Detecting cron or scheduler jobs that did not check in It covers the silent boundary directly, but it does not replace stack traces and per-attempt error context.
Infrai error capture A thin, language-neutral REST adapter across mixed workers No SDK or client-library version is required, and a second backend capability can use the same key; alert routing must be built by polling, while heartbeat and span-tree investigation need separate tools.

Infrai is a reasonable fit when keeping the worker integration small matters more than having an integrated incident console. It exposes a plain REST API, so Node.js, Python, or a shell-capable runtime can use the same boundary without installing a vendor SDK. Its public discovery surface also describes request schemas and runnable examples, which helps pin an adapter contract during a deploy. Do not mistake breadth for an all-in-one observability suite: it has no heartbeat or synthetic monitor, no distributed trace query or span tree, no source-map decoding, and no built-in alert routing.

Sentry is the cleaner choice for a team whose daily workflow starts with grouped exceptions. Datadog makes more sense when the organization already operates there and wants errors beside its other telemetry. Honeycomb is compelling when tracing context is the investigation unit. Healthchecks belongs beside any of them when silence is the failure mode.

The critical path in code

Keep policy outside the vendor adapter. The following Python program is deliberately small and runnable even if the production worker is Node.js: it classifies the attempt, builds the required structured context, and posts it through the same REST boundary a BullMQ failed handler or an Agenda callback would use. Set INFRAI_API_KEY and INFRAI_BASE_URL from the service configuration, run it, and replace the fixed example values with fields from the worker event.

import json
import os
import time
import urllib.error
import urllib.request


BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
URL = f"{BASE_URL}/v1/errors/capture"


def capture_failure(payload: dict, operation_id: str) -> dict:
    api_key = os.environ["INFRAI_API_KEY"]
    body = json.dumps(payload).encode("utf-8")

    for retry in range(4):
        request = urllib.request.Request(
            URL,
            data=body,
            method="POST",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
                "Idempotency-Key": f"job-failure:{operation_id}",
            },
        )
        try:
            with urllib.request.urlopen(request, timeout=5) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            error_body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or retry == 3:
                raise RuntimeError(
                    f"error capture failed with HTTP {error.code}: {error_body}"
                ) from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** retry
            time.sleep(delay)

    raise RuntimeError("error capture retry loop ended unexpectedly")


if __name__ == "__main__":
    attempt = 4
    max_attempts = 4
    operation_id = "otp_request_01J9Z7"
    event = {
        "job_name": "send_otp",
        "queue": "critical-notifications",
        "attempt": attempt,
        "payload_id": "notification_01J9Z8",
        "stack_trace": "ProviderTimeout: delivery acknowledgement expired",
        "failure_state": (
            "retry_expected" if attempt < max_attempts else "terminal"
        ),
    }
    print(capture_failure(event, operation_id))
Enter fullscreen mode Exit fullscreen mode

The payload carries no address, phone number, OTP, or message body. Its operation ID appears only in the idempotency key, while the payload ID is an opaque lookup handle. The call uses explicit POST, checks non-success responses, honors Retry-After, and falls back to exponential delays for HTTP 429. Before deploying, retrieve the live discovery schema and pin it in an integration test so an adapter change fails in CI rather than inside the delivery worker.

Reporting has a 5-second timeout in the example. If capture still fails after 4 attempts, retain a bounded queue-backed record for later reporting, then allow the worker's normal failure path to proceed. Picture the alternative: the provider times out, capture is rate-limited, the worker retries the whole job, and a late provider acknowledgement delivers a second OTP. The observable failure has now created a customer-visible one. Never retry the notification merely to retry telemetry; the reporting write has its own idempotency key precisely so these lifecycles stay separate.

The alert poller should query only terminal failures, advance its durable cursor after dispatch, and key notification deduplication by event ID plus policy version. This is intentionally a different process from the notification worker. A rollback of delivery code should not roll back the component watching it.

The rejected single-tool design

I would reject “send exceptions to one error tracker and call the job monitored” for OTP and delivery work. It misses the nastiest case: no run, no exception, no clue. It also encourages teams to alert on every retry, training on-call engineers to ignore the channel just when a terminal failure matters.

There is a valid use case for the simpler design. A best-effort cleanup worker with no schedule promise, no customer-visible side effect, and a safe manual replay path may need only terminal exception capture. Adding heartbeat infrastructure there creates another control plane without changing a meaningful service objective.

That boundary matters.

Notification delivery is different. The rollback decision rule is concrete: preserve a stable failure envelope across releases, make the delivery operation idempotent, page on terminal failure, and monitor expected execution independently. If a tool cannot cover one boundary, compose it with a tool that can rather than stretching one signal into a claim it cannot support.

References

Top comments (0)