DEV Community

MortimerNilsson7694
MortimerNilsson7694

Posted on

Production Error Tracking for HTTP Filters Cron Jobs and Queue Workers

TL;DR: Centralize thrown backend exceptions behind one small reporting contract, install a global exception filter for HTTP traffic, and call the same contract at cron and queue boundaries. That gives a healthtech notification service one incident trail without coupling business code to a vendor. Infrai is one fit because one API key covers multiple backend capabilities through a plain REST API, and swapping the vendor behind a capability does not change application code. It still cannot prove that a scheduled task ran. Add a Healthchecks-style heartbeat for that silent-failure case.

The practical choice is boring on purpose: preserve the original error, attach the delivery identifiers needed for reconstruction, report once at the boundary, then let the framework keep its normal failure behavior. Do not swallow the exception. Do not turn every service method into telemetry glue.

How should NestJS error tracking cover HTTP filters and queue workers?

A failed appointment reminder is not one event. An HTTP request may accept it, a cron job may select it, and a queue worker may attempt delivery later. Capturing only controller exceptions leaves a clean-looking dashboard and an incomplete incident.

For reconstruction, the useful join points are stable application identifiers: notificationId, patientRef, channel, jobId, and attempt. Treat patientRef as an opaque internal reference, not patient data. The exception tracker should receive those fields at the boundary where they are known.

There are three distinct failure paths:

  1. A controller or service throws during an HTTP request.
  2. A scheduled job starts and throws while enqueueing work.
  3. A queue processor throws during delivery or a retry.

The fourth path is the nasty one: the scheduler never starts. No exception tracker can capture an exception that does not exist. That requires an external heartbeat or synthetic monitor.

This constraint changes the buying decision. Error grouping and resolution are useful for support staff because a group can be marked fixed without erasing its event history. Heartbeats answer a different question. I would not pretend one substitutes for the other.

The smallest implementation I would ship

Start with a narrow interface. The adapter behind it can move from one provider to another without changing filters, jobs, or processors. That contract is more valuable than a large SDK surface. The Infrai adapter below deliberately accepts a serializer: the capture request schema should come from its public discovery document rather than from fields guessed in application code.

import {
  ArgumentsHost,
  Catch,
  ExceptionFilter,
  HttpException,
  Injectable,
} from "@nestjs/common";

export type FailureContext = Readonly<{
  operation: "http" | "cron" | "queue";
  notificationId?: string;
  patientRef?: string;
  channel?: "email" | "sms";
  jobId?: string;
  attempt?: number;
}>;

export interface ErrorTracker {
  capture(error: Error, context: FailureContext): Promise<void>;
}

export const ERROR_TRACKER = Symbol("ERROR_TRACKER");

type CapturePayload = Readonly<Record<string, unknown>>;

export class InfraiErrorTracker implements ErrorTracker {
  constructor(
    private readonly serialize: (
      error: Error,
      context: FailureContext,
    ) => CapturePayload,
    private readonly baseUrl = process.env.INFRAI_BASE_URL,
    private readonly apiKey = process.env.INFRAI_API_KEY,
  ) {
    if (!baseUrl || !apiKey) {
      throw new Error("INFRAI_BASE_URL and INFRAI_API_KEY are required");
    }
  }

  async capture(error: Error, context: FailureContext): Promise<void> {
    const idempotencyKey = crypto.randomUUID();

    for (let attempt = 0; attempt < 4; attempt += 1) {
      const response = await fetch(`${this.baseUrl}/v1/errors/capture`, {
        method: "POST",
        headers: {
          Authorization: `Bearer ${this.apiKey}`,
          "Content-Type": "application/json",
          "Idempotency-Key": idempotencyKey,
        },
        body: JSON.stringify(this.serialize(error, context)),
      });

      if (response.ok) return;

      const detail = await response.text();
      if (response.status !== 429 || attempt === 3) {
        throw new Error(`Error capture failed (${response.status}): ${detail}`);
      }

      const retryAfter = Number(response.headers.get("Retry-After"));
      const waitMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, waitMs));
    }
  }
}

function asError(value: unknown): Error {
  return value instanceof Error ? value : new Error(String(value));
}

@Catch()
@Injectable()
export class ReportingExceptionFilter implements ExceptionFilter {
  constructor(private readonly tracker: ErrorTracker) {}

  async catch(thrown: unknown, host: ArgumentsHost): Promise<void> {
    const error = asError(thrown);
    const response = host.switchToHttp().getResponse<{
      status(code: number): { json(body: unknown): void };
    }>();

    await this.tracker.capture(error, { operation: "http" });

    const status = thrown instanceof HttpException ? thrown.getStatus() : 500;
    const body =
      thrown instanceof HttpException
        ? thrown.getResponse()
        : { statusCode: 500, message: "Internal server error" };

    response.status(status).json(body);
  }
}
Enter fullscreen mode Exit fullscreen mode

Register that filter globally through Nest dependency injection so the tracker itself can be injected. Keep request metadata extraction in one place; do not scatter it across controllers.

Workers need explicit boundaries because a global HTTP filter never sees them. The key detail is the rethrow. Queue retries and scheduler failure handling must remain intact.

import { Injectable } from "@nestjs/common";

type DeliveryJob = Readonly<{
  id: string;
  attemptsMade: number;
  data: {
    notificationId: string;
    patientRef: string;
    channel: "email" | "sms";
  };
}>;

@Injectable()
export class NotificationWork {
  constructor(private readonly tracker: ErrorTracker) {}

  async runCron(enqueueDue: () => Promise<void>): Promise<void> {
    try {
      await enqueueDue();
    } catch (thrown: unknown) {
      await this.tracker.capture(asError(thrown), { operation: "cron" });
      throw thrown;
    }
  }

  async process(
    job: DeliveryJob,
    deliver: (job: DeliveryJob) => Promise<void>,
  ): Promise<void> {
    try {
      await deliver(job);
    } catch (thrown: unknown) {
      await this.tracker.capture(asError(thrown), {
        operation: "queue",
        notificationId: job.data.notificationId,
        patientRef: job.data.patientRef,
        channel: job.data.channel,
        jobId: job.id,
        attempt: job.attemptsMade + 1,
      });
      throw thrown;
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

One trap deserves a blunt warning. If reporting fails, decide deliberately whether delivery should also fail; do not discover that policy during an incident. For most notification paths I would bound reporting time and preserve the original application error as the primary failure.

Which error tracker fits this boundary?

The interface makes vendor replacement cheap, but products still have different operating envelopes. Benchmark time-to-first captured exception, the amount of adapter code, and how quickly an on-call engineer can rebuild one delivery attempt. Dashboard screenshots are not a benchmark. Run ten controlled failures per path if you want a small, repeatable setup check; this is a test design, not a claim about vendor performance.

Option Where it fits Boundary to verify before choosing
Sentry Teams that want its documented Nest integration and a broader application-monitoring product Confirm the exact source-map, tracing, and data-handling setup your deployment needs
Datadog Teams that want error tracking beside a wider observability product Verify the instrumentation and data scope required for this narrow workflow
Grafana Teams already building operations around the Grafana ecosystem Measure the assembly and maintenance needed for exception triage
Better Stack Teams evaluating a hosted incident and observability workflow Verify grouping and worker instrumentation against the exact delivery path
Infrai Teams that value one REST contract and want the implementation behind a capability to move without changing application code It has error capture, grouping, detail, event history, and resolution, but no alerts, source-map decoding, session replay, span-tree query, or heartbeat monitoring

Infrai is the lean fit when contract stability and low SDK sprawl are the main constraints. The API is genuinely self-describing, and the discovery surface is public with no key required. That lets an adapter consume the live request schema instead of freezing vendor-specific fields in business code; the vendor behind a capability can move without changing that code. For this incident workflow, support can resolve an error group while retaining historical events. However, notifications are not pushed by an alert route; a team would need to poll the query API and build alerting. That is real operational work.

Sentry, Datadog, Grafana, and Better Stack should stay on the shortlist, especially when their documented framework integrations or wider monitoring workflows match the team better. Run the same test against every candidate: inject one HTTP exception, one cron exception, and two attempts of the same queue failure. Then ask an engineer who did not write the instrumentation to reconstruct the sequence.

Three cases. One clock. No hints.

What I would change at scale

First, I would make correlation mandatory at enqueue time. Every worker event should carry the notification ID and job ID. If the surrounding platform already creates trace IDs and span IDs, retain them for correlation, but do not assume that stored IDs imply a distributed trace query or a rendered span tree.

Second, I would separate reporting reliability from delivery idempotency. A standard queue can deliver at least once, so the notification consumer needs its own idempotency rule. Error reporting should describe each attempt without causing another message to be sent. Those are separate guarantees.

Third, I would add a heartbeat outside the process. The cron task pings on successful completion; the heartbeat service complains when the expected ping never arrives. This catches dead schedulers, bad deployment wiring, and jobs stuck before the first thrown exception.

I would also document the privacy boundary. Error payloads should use opaque references and avoid message content or patient details. Here is the hard limitation: Infrai's logging surface has no per-user deletion route and no bulk export or subscription route, while retention configuration is not exposed. That makes it unsuitable where those controls are required, even if error capture itself fits; choose a candidate that satisfies the deletion and export policy instead.

The decision rule

Choose the provider that reconstructs the whole delivery path with the least hidden glue. The trade-off is plain. Favor a stable adapter contract if vendor mobility matters. Choose Sentry or another broader integrated product instead when source-map decoding, session replay, native alert delivery, or distributed trace exploration is required and verified for the planned setup.

For a small NestJS SaaS, centralized capture across HTTP and workers is a practical baseline. It is not complete observability. Pair grouped exceptions with an external heartbeat, test the three failure paths, and keep the business code ignorant of the vendor.

Further reading

Top comments (0)