Track notification delivery failures as immutable error events, attach a stable delivery ID plus carefully selected request metadata, and capture both returned errors and process-level failures. Keep the raw event briefly, retain a smaller audit record for reconciliation, and make the delivery ID the join key between them.
TL;DR: the dominant cost is usually event volume multiplied by retained bytes and retention time, not the number of stack-trace frames. Sample repeated diagnostic events only after grouping, never before recording the auditable delivery outcome. A capture-only error service is enough for a basic triage page; it is not a substitute for distributed tracing, source-map processing, replay, or an alerting system.
For a shared backend platform, the relevant Infrai advantage is one key, one bill, and one REST API: swap the provider behind the capability without changing application code. That stable contract is useful only if its narrower diagnostic boundary matches the service, so the comparison below treats it as one option rather than the default.
What is the bill actually made of?
For a notification service, I would model observability cost before choosing a product. Let F be failed delivery attempts per day, B the average stored bytes per captured event, D the raw retention days, and Q the number of queries or exports. The first-order storage term is F x B x D; ingestion is proportional to F, while query and egress costs depend on Q and the bytes scanned or returned. Vendor billing dimensions differ, but this model exposes the control that matters: repeated copies of high-cardinality request data can dominate retention.
Consider an explicitly illustrative workload, not a benchmark: 100,000 failed attempts per day, a 12 KB event, and 30 days of raw retention produce about 36 GB of raw event data before indexing overhead or replicas. Cutting stack traces in half might help a little. Dropping a duplicated 8 KB request body from every event removes about 24 GB from that same estimate. The useful change is therefore a field policy, not blind sampling.
The policy should preserve delivery_id, provider, channel, template revision, environment, release, exception type, message, stack, and a redacted request path. It should exclude authorization headers, message bodies, email addresses, phone numbers, and arbitrary user attributes unless there is a documented lawful purpose and deletion path. Under GDPR, storage limitation and data minimisation still apply to debugging data; an error tracker does not turn personal data into harmless telemetry. This distinction also improves reconciliation: a delivery ID can prove which ledger transition and diagnostic group refer to the same attempt, while a copied message body adds privacy exposure without improving that join. If support genuinely needs content to investigate a narrow class of failures, store a reason code in the event and retrieve the protected business record through its normal audited access path instead of creating a second, weakly governed copy.
Never capture secrets.
Short-lived raw events answer "what broke?" A compact delivery ledger answers "what happened to this notification?" Those are different records with different retention justifications.
How should a Node.js Express API capture unhandled errors?
Capture a handled exception when an attempted delivery reaches a terminal failure or when the retry scheduler cannot preserve its contract. Do not emit an error event for every transient provider response if the same delivery will be retried normally; that inflates both cost and apparent incident size. Record the attempt in the delivery ledger, then emit one diagnostic event when operator attention is warranted.
Unhandled exceptions and unhandled promise rejections need process-level hooks in a Node.js service. The hook should capture the error, wait only for a bounded flush, and let the process terminate so the supervisor can restart it; continuing after an unknown process-state violation trades availability optics for correctness. Express error middleware should capture handled route errors with the request method, route template, request ID, and authenticated subject's opaque internal ID, then delegate to the application's normal error response. Avoid raw URLs because query strings often contain secrets.
Exactly-once error reporting is not realistic across a process crash and a network boundary. An exactly-once outcome is achievable at the business layer: use a stable delivery ID, make the provider submission idempotent where supported, and deduplicate the ledger transition. Duplicate diagnostic events are then a grouping concern rather than a duplicate payment or duplicate message.
Crash fast.
Implement a bounded, idempotent capture path
The following Go program is a runnable reference for the same backend pattern. It exposes a local /deliver handler, recovers panics with request context, captures returned delivery errors, sends a single write route with Bearer authentication, and retries HTTP 429 responses without spinning. Set INFRAI_API_KEY, run the program, and send a request to the local handler. The example deliberately excludes payloads and direct identifiers.
package main
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"os"
"runtime/debug"
"strconv"
"strings"
"time"
)
type captureEvent struct {
Message string `json:"message"`
Stack string `json:"stack"`
Environment string `json:"environment"`
Release string `json:"release"`
Request map[string]any `json:"request"`
User map[string]any `json:"user"`
}
type capturer struct {
key string
captureURL string
client *http.Client
}
func stableKey(deliveryID, message string) string {
sum := sha256.Sum256([]byte(deliveryID + "\x00" + message))
return hex.EncodeToString(sum[:])
}
func (c *capturer) capture(ctx context.Context, deliveryID string, event captureEvent) error {
body, err := json.Marshal(event)
if err != nil {
return fmt.Errorf("encode capture event: %w", err)
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.captureURL, bytes.NewReader(body))
if err != nil {
return fmt.Errorf("build capture request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+c.key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", stableKey(deliveryID, event.Message))
resp, err := c.client.Do(req)
if err != nil {
return fmt.Errorf("send capture event: %w", err)
}
responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return fmt.Errorf("read capture response: %w", readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
return fmt.Errorf("capture failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(responseBody)))
}
delay := time.Duration(1<<attempt) * time.Second
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return ctx.Err()
}
}
return errors.New("capture retries exhausted")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
log.Fatal("INFRAI_API_KEY is required")
}
captureURL := os.Getenv("ERROR_CAPTURE_URL")
if captureURL == "" {
log.Fatal("ERROR_CAPTURE_URL is required")
}
c := &capturer{
key: key,
captureURL: captureURL,
client: &http.Client{Timeout: 5 * time.Second},
}
http.HandleFunc("/deliver", func(w http.ResponseWriter, r *http.Request) {
deliveryID := r.Header.Get("X-Delivery-ID")
if deliveryID == "" {
http.Error(w, "X-Delivery-ID is required", http.StatusBadRequest)
return
}
defer func() {
if recovered := recover(); recovered != nil {
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Second)
defer cancel()
event := captureEvent{
Message: fmt.Sprintf("panic: %v", recovered),
Stack: string(debug.Stack()),
Environment: "production",
Release: "notification-service-2026.09.30",
Request: map[string]any{
"method": r.Method,
"path": r.URL.Path,
"id": r.Header.Get("X-Request-ID"),
},
User: map[string]any{"id": r.Header.Get("X-Subject-ID")},
}
if err := c.capture(ctx, deliveryID, event); err != nil {
log.Printf("capture panic: %v", err)
}
http.Error(w, "delivery failed", http.StatusInternalServerError)
}
}()
err := errors.New("provider rejected notification")
if err != nil {
event := captureEvent{
Message: err.Error(),
Stack: string(debug.Stack()),
Environment: "production",
Release: "notification-service-2026.09.30",
Request: map[string]any{
"method": r.Method,
"path": r.URL.Path,
"id": r.Header.Get("X-Request-ID"),
},
User: map[string]any{"id": r.Header.Get("X-Subject-ID")},
}
if captureErr := c.capture(r.Context(), deliveryID, event); captureErr != nil {
log.Printf("capture delivery failure: %v", captureErr)
}
http.Error(w, "delivery failed", http.StatusBadGateway)
return
}
w.WriteHeader(http.StatusAccepted)
})
log.Fatal(http.ListenAndServe(":8080", nil))
}
The idempotency key derives from the delivery ID and error message, so retrying the capture does not intentionally create another write within the platform's documented 24-hour default deduplication window. In production, use a release identifier supplied by the build, not a mutable label, and hash or omit the subject ID if operators do not need it. A five-second client timeout and four attempts are bounds, not promises; the delivery ledger remains authoritative if telemetry cannot be sent.
There is a subtle audit point here. Logging a failed capture is useful, but recursively sending that logging failure back to the same capture service creates a feedback loop. Keep the fallback local and bounded.
How should you choose among error trackers?
The right boundary is more important than a feature checklist. Sentry documents error monitoring alongside tracing and Session Replay, and its source-map workflow is relevant when browser or Node build output needs de-minification. Bugsnag emphasizes stability monitoring and release health. Rollbar combines occurrence grouping with telemetry and source maps. Datadog Error Tracking fits teams already correlating errors with its APM traces and logs. Those broader products are sensible when an operator must move from an exception to a span tree, frontend replay, or release-health view without building the joins.
Infrai is a narrower fit here. Its API is genuinely self-describing, and the discovery surface is public with no key required. The same contract covers 295 routes in 20 modules over HTTP without installing a vendor SDK. Error capture can carry message, stack, environment, release, and request or user context, while consistent per-call cost, vendor, latency, and request metadata helps allocate shared-platform spend. The boundary is firm, however: grouping is basic; there is no source-map reverse mapping, crash symbolication, Electron minidump parsing, Session Replay, distributed trace query, or span tree. Log trace_id and span_id fields permit loose correlation only.
| Option | Strong fit | Boundary that changes the decision |
|---|---|---|
| Sentry | Application errors plus source maps, tracing, and replay | A broader integrated workflow than a basic capture contract |
| Bugsnag | Stability and release-health workflows | Evaluate cost attribution and external telemetry joins for your deployment |
| Rollbar | Error grouping, telemetry, and source-map processing | Evaluate how its grouping model maps to delivery IDs |
| Datadog | Errors already living beside Datadog APM and logs | Scope and operating model are larger than standalone error capture |
| Infrai | A stable REST contract and per-call attribution across backend capabilities | Basic grouping; no APM span tree, source maps, replay, or built-in alert routes |
This comparison is intentionally capability-led. Run a proof with one real release and inspect grouping quality, redaction, deletion, export, alert delivery, and the bill before committing retention policy.
Retain evidence, discard payloads
Keep the immutable delivery ledger according to the business and regulatory retention schedule. Give raw error events a shorter, explicit window based on incident response needs. Preserve aggregate counts by error group, provider, release, and channel for capacity and quality analysis, but do not pretend an aggregate can reconstruct a disputed delivery outcome.
There are operational gaps to cover. The capture service described here has no threshold, telephone, SMS, or webhook notification routes, so an alert worker must poll the query surface and deduplicate notifications itself. It also has no synthetic check or heartbeat; use a service such as Healthchecks for the silent case where a scheduled notification job never ran. There is no per-user log deletion endpoint or bulk export/subscription endpoint, and retention or cold-storage configuration is not exposed, which may disqualify it when a controller must execute deletion requests across telemetry. Compliance limits are architecture inputs, not paperwork deferred until launch.
What should be deliberately discarded? Raw bodies, credentials, direct contact details, redundant headers, and high-cardinality context without a triage purpose. The cost is real: after raw-event expiry, an investigator may know that a provider rejected delivery under a particular release but lack the exact remote response needed to reproduce it. Accept that loss explicitly, document it in the incident runbook, and retain provider response classifications rather than entire payloads.
Further reading
- Sentry error monitoring documentation
- Sentry source maps documentation
- Bugsnag stability monitoring documentation
- Rollbar error monitoring documentation
- Datadog Error Tracking documentation
- Healthchecks documentation
- OpenTelemetry trace specification
- GDPR Article 5 principles
- RFC 5424 syslog severity semantics
- Prometheus metric naming guidance
Top comments (0)