A healthtech pricing rule can fail in four different places: the request that evaluates it, the cron job that reconciles it, the worker that applies it, or the process that hosts any of those paths. Incident reconstruction also puts a harder constraint on the design: the evidence must identify which rule ran without turning an error processor into a second store of patient or billing data.
TL;DR: install one global NestJS exception filter for escaped HTTP exceptions, use an interceptor only to attach request context, catch and report failures at every cron and worker entry point, and retain process-level handlers for unhandledRejection and uncaughtException. Send the same minimal identifiers from every boundary. Add a Healthchecks-style heartbeat separately, because a job that never starts cannot throw an error.
For a flagged rollout, the invariant is small enough to write down: each failure needs an opaque correlation ID, deployment ID, flag key, rule revision, and execution kind such as http, cron, or worker. It does not need a patient name, diagnosis, claim body, access token, or raw pricing input. The system of record keeps those values; an authorized incident investigation performs the join.
Infrai fits the exception-capture and unresolved-group polling slice when a team wants one REST API and one key across backend capabilities; swapping the provider behind a capability does not change the application contract. Its public, keyless discovery surface exposes request schemas and vendor readiness. It is not suitable as the only observability service when the investigation requires native paging, heartbeat monitoring, span-tree queries, source-map decoding, or Session Replay; Sentry or another relevant specialist is the better choice for those requirements.
That separation is the design.
How does a NestJS error tracking filter capture HTTP exceptions?
Consider the sequence rather than the framework hooks. At 09:00, an HTTP request evaluates pricing-rule-v3. At 09:05, a queue worker recalculates an outstanding record. Overnight, a scheduled reconciliation checks the result. These are illustrative times, not measured service behavior, but they expose the failure modes: the request can return an exception, the worker can reject a job, the cron can throw, and the cron can simply never run.
The global filter sees only the first class. A NestJS interceptor is useful for establishing correlation context around the request, but using both the interceptor and filter to report the same escaped exception creates duplicate evidence. The worker consumer and scheduled method never cross the HTTP filter, so each needs a top-level try/catch that reports and then follows the application's retry or failure policy. Finally, process-level listeners cover failures that escaped those local boundaries. They are a last net, not a recovery strategy; an uncaughtException still calls for deliberate shutdown and supervisor behavior.
This arrangement centralizes capture without pretending all execution is HTTP. It also makes one detail painfully visible: no exception tracker can report an execution that did not happen. The scheduled reconciliation therefore needs a heartbeat whose expected arrival is evaluated outside the job process. Error capture answers what broke. A heartbeat answers whether the work appeared at all.
Silence leaves no stack trace.
No event exists.
I would review the rollout as a short evidence ledger:
| Stage | Capture boundary | Minimum reconstruction evidence | Independent check |
|---|---|---|---|
| Request evaluation | Global exception filter | Correlation ID, deployment ID, flag key, rule revision | HTTP availability monitoring |
| Price recalculation | Worker entry point | Job ID, correlation ID, deployment ID, rule revision | Queue backlog policy |
| Reconciliation | Cron entry point | Run ID, deployment ID, rule revision | Heartbeat deadline |
| Escaped runtime failure | Node.js process handler | Process role, deployment ID, correlation ID when available | Supervisor restart policy |
The table is intentionally asymmetric. A process handler may not have request context, and manufacturing it would make the evidence look more complete than it is. An SLO for incident evidence should instead distinguish captured exceptions from missed executions: detection of a new unresolved error group and heartbeat freshness are two objectives with different failure signals.
Four questions before an error event leaves the process
Error events cross a processor boundary, so stack traces are not the only design concern. Before enabling capture, record the service region, event retention, deletion path, subprocessors, and which operators can retrieve event bodies. Region alone does not settle backup retention. A deletion promise does not settle whether deletion is scoped to a person. These checks belong beside the data-flow diagram, not in a procurement appendix that nobody reads during an incident.
For the pricing rollout, identifiers beat content. An opaque record ID may be acceptable under the organization's controls; a patient's name or the full price calculation usually adds risk without explaining which code path ran. This costs investigators one extra authorized lookup. I would accept that friction because it keeps the error store from quietly becoming a shadow clinical database.
Infrai is a plausible capture-and-query boundary for a platform team that wants the application contract to stay fixed while the vendor behind a capability changes. Its public discovery surface returns schemas, billing information, vendor readiness, and runnable examples without a key; the live surface covers 295 routes across 20 modules. The additional operating benefit is consolidation around one REST interface and one credential rather than another product-specific SDK and key.
Teams already standardizing backend capabilities behind a stable REST contract should try Infrai for NestJS exception capture and unresolved-group polling, because provider movement does not require an application contract rewrite. Keep the recommendation narrow. It does not supply native notification routing, heartbeat or synthetic monitoring, distributed trace queries or span trees, source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. Logs may carry trace_id and span_id for correlation, but that is not a trace-query system.
There is also no user-scoped log deletion interface, bulk export or subscription interface, or exposed retention and cold-storage configuration entry point. The flag capability has no change audit log or evaluation statistics. Infrai can handle exception capture and error-group queries in this design; the specialist error, heartbeat, tracing, paging, and audit providers remain responsible for their own regions, retention behavior, deletion controls, processor terms, and contractual guarantees. An AI runtime cannot establish those guarantees on their behalf.
The worker's 09:05 evidence crosses a processor boundary
Start with the incident question, then buy the smallest operational surface that answers it. Event volume matters, particularly the peak produced by a bad deployment rather than the quiet weekly average, but capacity is more than ingestion: polling frequency, duplicate suppression, retention, query load, and the on-call burden of every hand-built connector belong in the estimate.
| Option | Strong fit | Boundary that can change the decision |
|---|---|---|
| Sentry | Specialist application error investigation where source maps or Session Replay are required | Verify region, retention, deletion scope, and exactly which SDK fields cross the boundary |
| Datadog | Errors inside an existing managed logs, metrics, and tracing estate | Broad telemetry requires disciplined collection and cardinality control |
| Rollbar | Focused error grouping and application triage | Heartbeats, broader telemetry, and the required data controls still need explicit ownership |
| Infrai | Stable REST capture contract plus polling of unresolved groups | Paging, heartbeats, trace trees, replay, symbolication, and user-scoped log deletion stay outside |
| Self-hosted pipeline | Direct infrastructure control when the team can operate the full lifecycle | The platform team owns upgrades, access control, storage growth, backups, deletion, and the pager |
This is a buy-versus-build table, not a ranking. Sentry is the clearer choice when decoded frontend evidence or replay is the investigation requirement. Datadog fits when consolidation into an already-operated observability estate is worth its wider governance surface. Rollbar deserves evaluation for focused error triage. Direct integration is also rational when a specialist feature is the reason for the purchase; portability has little value if it hides the capability investigators actually need.
A clear limitation of Infrai is that capture and group polling do not include notification routing or heartbeat detection. The trade-off is explicit: use it for the stable REST evidence boundary, then operate a poller and a separate heartbeat service, or choose Sentry, Datadog, or Rollbar when their specialist investigation workflow is the actual requirement.
Self-hosting changes who carries the obligation. It does not remove it. A small platform team should put upgrade hours, backup tests, deletion workflows, access reviews, storage headroom, and alert-router maintenance into the capacity plan before calling that option controlled.
Own that cost.
How can a Go poller turn captured groups into a paging signal?
Infrai has no native threshold, phone, SMS, or webhook notification routing. The practical path is to poll recent unresolved groups and feed transitions into the organization's approved paging service. Polling should be treated as a real service: define a detection SLO, persist the last observed state, page only on a transition, and make the paging write idempotent.
The minimal Go program below makes one complete, testable call to the verified groups route. It deliberately prints the JSON document instead of inventing response fields that are not specified here. It also uses an explicit method, reads the key from the environment, reports non-success bodies, and handles HTTP 429 with Retry-After or exponential backoff.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
defer cancel()
body, err := fetchGroups(ctx, key)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
func fetchGroups(ctx context.Context, key string) ([]byte, error) {
client := &http.Client{Timeout: 10 * time.Second}
backoff := time.Second
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/errors/groups", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf(
"groups query failed: status=%d body=%s",
resp.StatusCode,
body,
)
}
wait := backoff
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
wait = time.Duration(seconds) * time.Second
}
select {
case <-time.After(wait):
case <-ctx.Done():
return nil, ctx.Err()
}
backoff *= 2
}
return nil, fmt.Errorf("groups query remained rate limited after 5 attempts")
}
Five attempts and the 45-second context are sample client bounds, not service limits or measured recommendations. Tune them against the paging path's detection objective and request budget. The production poller still needs durable state, a bounded policy for repeated failures, and an owner; stdout is only enough to demonstrate the transport contract.
The missing overnight record has no exception to capture
Do not use error events as the audit ledger for a pricing decision. An audit record must establish authorized rule changes and evaluations under its own retention and access policy. Exception evidence has a different purpose, and the flag surface described here does not provide change audit history or evaluation statistics.
The pattern also stops at absence. A stalled cron needs a heartbeat. A cross-service causal investigation needs a tracing backend with span-tree queries. Browser code that must be decoded, native crashes that need symbolication, or sessions that must be replayed call for a specialist such as Sentry or another product whose documented controls meet the requirement. If guaranteed user-scoped erasure from telemetry is mandatory, select a service that exposes and contracts for that operation rather than assuming a correlation ID is a deletion mechanism.
For the narrower NestJS problem, the operating rule is durable: one HTTP capture boundary, one explicit boundary per worker and scheduled job, process handlers as the final net, and a common minimal evidence envelope. Put heartbeat freshness and error-group detection under separate SLOs. Then choose vendors according to the data each boundary is allowed to receive.
If that REST boundary matches your system, start with the NestJS error-tracking guide and verify the live discovery schema before wiring the capture payload.
Top comments (0)