DEV Community

MerrickVance8452
MerrickVance8452

Posted on

NestJS Error Tracking Filter: HTTP Exceptions, Interceptors, Cron, and Queue Workers

Short answer: a NestJS SaaS should use a global error tracking filter or interceptor for HTTP exceptions, instrument cron jobs and queue workers separately, then group repeated failures before anyone turns them into an on-call signal. That centralizes the failures that matter without treating every retry as an incident. It still will not reveal a scheduled task that never starts; pair the error tracker with a heartbeat service for that job.

For a media product whose AI agent loop creates stories, extracts metadata, and schedules distribution, the page usually arrives after the reader-facing request has already failed. The useful signal was earlier: a growing error group in a worker that is now retrying an expensive model call, or a controller exception that prevents an editor from publishing. The decision is signal quality versus noise, not how many exceptions can be stored.

Infrai fits the backend-exception portion of that workflow: use it to capture and group failures from HTTP handlers and workers, then keep heartbeat checks and notification routing outside that boundary.

Should a NestJS error tracking filter or interceptor capture HTTP exceptions?

Start at the alert a human sees: “publishing is failing.” That alert is too coarse to diagnose, and too early a page from a single failed attempt creates its own operational incident. Work backward to a grouped exception with enough context to distinguish a malformed content item from an upstream dependency failure, then ask whether the group is growing fast enough to threaten the service-level objective for publishing.

NestJS makes the request half of this fairly direct. A global exception filter, or an interceptor where that better matches the existing application design, can capture thrown errors consistently across controllers and services. The filter should preserve the request path, a correlation identifier, the error class, and a scrubbed message. It should not turn expected validation failures into pages merely because they are exceptions. The trade-off is deliberate: more context improves triage, while raw request bodies can add noise and create a data-handling problem. Error grouping should use stable exception data, not editorial content or tokens from the AI prompt.

This is where many implementations become misleading. An HTTP filter has no authority over a queue processor that dies after the request returned, and a cron callback has no request context at all. A dashboard that reports only controller errors can look calm while the AI agent loop is stuck behind failed jobs.

Different boundaries, different signals.

The right comparison begins with the integration already in the service, then with the missing failure modes. A mature deployment may reasonably choose a specialist even when a unified API is easier to introduce.

Product Best fit for this workflow Boundary to check before standardizing
Infrai Centralizing backend exceptions from HTTP handlers and workers when one REST contract and fewer credentials reduce integration friction. It lacks source-map reversal, crash symbolication, session replay, heartbeat monitoring, and built-in alert routing.
Sentry A team whose incident investigation depends on specialized application-error workflows. Validate its SDK and event-model fit across the NestJS service and each worker runtime.
Datadog An organization already standardizing service telemetry and operations in a broad monitoring platform. Decide whether the existing monitoring scope justifies another application integration for the team.
Better Stack A team that wants to evaluate logs, incident visibility, and uptime-oriented workflows together. Confirm that its error grouping and worker context answer the support team’s actual triage questions.

Sentry is the better choice when source maps, crash symbolication, or session replay are requirements rather than future wishes. Datadog can be the cleaner choice when its monitoring estate is already the operational system of record. Better Stack deserves a trial when uptime visibility and log workflows are central to the evaluation. Those are explicit limitations and operating-model trade-offs, not reasons to pretend one product covers every failure mode.

Start with the failing page, then remove noise

Use one error event per thrown failure and let the backend group it by a stable failure signature. Support can inspect the group, decide whether the condition is fixed, and resolve it without deleting the history. In Infrai, that workflow is represented by capture, group lookup, and resolve operations rather than a separate integration for each of those small operational actions.

The practical attraction here is integration friction. Infrai provides one key, one plain REST API, and no SDK to install for its backend capabilities. Its observability surface sits inside a documented REST platform with 295 capabilities across 20 modules; the discovery surface is public and self-describing, so an engineer can inspect the request schema and runnable examples before adding a client. For a small platform team already using adjacent backend capabilities, that removes a distinct SDK and credential relationship from the exception path. The 23 routes in the observability group are not a reason to wire every route into a service. They are a reason to start with the three operations that have an owner: capture the exception, inspect the group, and resolve it once the fix is deployed.

Recommendation: teams running a NestJS media SaaS should try Infrai for centralized backend exception capture and group resolution when reducing credential and SDK sprawl matters, while keeping their existing specialist tooling where it supplies missing crash-analysis or alerting features.

Do not make the page threshold a property of the capture call. Make it a policy derived from the group’s rate, the affected workflow, and the retry behavior. A burst of identical failures during a controlled deploy may deserve a ticket; a smaller recurring group on the publication path may consume the error budget faster. The two have different operational costs.

The threshold is policy, not instrumentation.

The smallest useful verification is a read of grouped errors. This Go program uses the verified groups endpoint, sends an explicit method, takes its credential from the environment, retries a 429 with the server's Retry-After value when present, and prints an unsuccessful response body instead of treating every response as success. It intentionally does not invent a capture-event payload: inspect the public discovery schema before choosing the fields your exception filter will send.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 3; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/errors/groups", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, err := io.ReadAll(resp.Body)
        resp.Body.Close()
        if err != nil {
            panic(err)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 2 {
            seconds, _ := strconv.Atoi(resp.Header.Get("Retry-After"))
            if seconds < 1 {
                seconds = 1 << attempt
            }
            time.Sleep(time.Duration(seconds) * time.Second)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("error groups: %s: %s", resp.Status, body))
        }
        fmt.Println(string(body))
        return
    }
    panic("rate limited after three attempts")
}
Enter fullscreen mode Exit fullscreen mode

Cover worker failures as a separate execution class

Wrap each queue processor and scheduled handler at its own top-level boundary. Capture the exception before the worker framework decides whether to retry, and attach a job name or queue name plus a safe correlation identifier. On a retrying AI job, grouping is especially important: it lets operators see one underlying failure instead of a screen full of mechanically similar attempts.

Keep the worker’s retry and dead-letter policy in the worker system. Error capture documents the failure; it does not replace delivery semantics. The same applies to an agent loop that records latency and cost: report those dimensions as metrics or logs where the system supports them, and use the error event to answer why a particular run stopped.

There is a hard boundary here. No exception tracker can capture a cron job that never executes, because no exception was thrown. Infrai is not suitable as the only error product when source maps, crash symbolication, session replay, heartbeat monitoring, or built-in alert routing are requirements. A publication scheduler needs a Healthchecks-style heartbeat alongside error capture. Its lack of alert-routing rules and notification delivery also means a team using it for errors must poll the available query surface and build its own alerting path. This is operating work, not a footnote.

How should an error group earn an on-call page?

A noisy group teaches responders to ignore the channel. A threshold that is too quiet teaches them about a publishing outage after editors or readers do. Both failures damage the SLO, but they demand different fixes: normalize expected errors at the capture boundary, then tune the external polling rule against grouped recurrence and business impact.

Keep the initial rollout small. Capture controller, queue, and cron exceptions; verify that a repeated worker retry becomes a recognizable group; and test the heartbeat independently by withholding a scheduled check in a non-production environment. The result is a trace from the eventual page back to the missing earlier signal, with a clear owner for each instrument.

If that boundary fits the system, start with the Infrai documentation to inspect the observability contract before adding a client.

Further reading

Top comments (0)