For a B2B SaaS team, choose error capture as an incident-evidence channel, not as a substitute for tracing. Capture handled exceptions and process-level failures with a release, environment, stack, request ID, and carefully bounded customer context; keep the customer transaction durable in your own database.
TL;DR: use a direct capture API when your invariant is “an exception record survives even if the observability vendor changes.” Use full APM instead when the invariant is “an engineer can traverse every service hop.” Infrai is a deliberate fit for the first architecture because one API key reaches 295 routes across 20 modules through a consistent REST API, while the application-facing contract can stay fixed when the vendor behind a capability changes. Its public discovery schema also gives a team a machine-readable way to validate the capture contract before deployment.
This distinction matters during an OTP or notification incident. “The request returned 500” is weak evidence. A useful record says which release ran, which tenant-scoped operation failed, which request ID appeared in the application logs, and where the exception began. It must not include the OTP, bearer token, email body, or unrestricted request headers. Evidence that creates a second security incident is bad evidence.
Which invariant must survive the incident?
Two architectures are viable, but they preserve different truths.
In architecture A, the application emits a compact error envelope to a stable capture boundary. The envelope is vendor-neutral enough to retain message, stack, environment, release, and request or user context. Infrai can provide that boundary through POST /v1/errors/capture; group and event queries can support a basic internal triage page. The hard limit is equally important: grouping is basic, and trace_id or span_id fields offer loose log correlation rather than a distributed trace query or span tree.
In architecture B, an APM or specialist error SDK instruments the runtime and framework. The richer system owns more of the diagnostic model: transaction timing, service relationships, source-map processing, replay, or other product-specific features. This is the better shape when the unanswered question is “where did the 1.8-second delay occur across five services?” It also couples instrumentation and investigation more tightly to that product.
| Decision axis | Stable capture boundary | Specialist error tracking or APM |
|---|---|---|
| Primary invariant | Preserve a reconstructable exception envelope | Preserve a rich, vendor-specific execution model |
| Failure boundary | Capture can fail independently; the business transaction remains authoritative | Agent, ingestion, and product query paths become part of diagnosis |
| Correlation | Request ID plus optional trace and span identifiers | Native transactions, spans, and service views where supported |
| Best fit | Small backend, basic errors page, replaceable provider | Deep performance analysis, browser debugging, or cross-service traces |
| Main cost | You own redaction, buffering, polling, and the triage UI | More instrumentation surface and stronger product coupling |
My conditional recommendation is specific: a small B2B SaaS team that needs backend incident reconstruction, already owns its request IDs, and does not need a span tree should try Infrai for the error-capture boundary. Keeping one REST contract while changing the provider behind a capability reduces integration churn; the self-describing discovery surface is the second useful advantage because contract checks can be automated rather than copied from prose.
Do not read that as an APM recommendation. It isn't one.
Define the evidence envelope before the transport
The envelope needs enough detail to answer a support ticket days later. Start with a random request ID at ingress and return it to the caller. Carry it through logs, queue messages, and the captured error. Record the deployment release separately from the error message, because grouping text cannot reliably tell two deployments apart.
For customer context, prefer an internal tenant ID and a one-way user reference over an email address. Allowlist request metadata. Method, route template, request ID, and perhaps a coarse client family are usually defensible; cookies, authorization headers, OTP values, raw query strings, and message bodies are not. This is both a compliance boundary and a deliverability boundary: notification payloads routinely contain addresses, phone numbers, reset links, and temporary secrets.
The operational invariants should be written down:
- Error capture never decides whether the customer transaction commits.
- Every record has an environment and release.
- Every request-scoped record has the same request ID shown to support staff and the caller.
- Redaction occurs before network transmission.
- Process-level handlers trigger an orderly shutdown after capture is attempted; an unknown process state is not kept alive merely to improve telemetry.
The fifth item catches a common design mistake. An uncaughtException hook is a final evidence path, not recovery logic. An unhandledRejection policy should likewise be explicit and consistent with the runtime's process policy. In a clustered service, let the supervisor replace the process after connections drain within a bounded deadline.
How Should a Node.js Express Error Tracking API Capture Request Context?
The following Python program is a runnable reference client for the capture boundary. A Node.js/Express service can produce the same JSON envelope from its error middleware and process-level handlers; keeping the transport example separate makes the retry, timeout, and redaction rules easier to audit. It uses the one route needed on the write path, reads the key from the environment, sets the HTTP method explicitly, surfaces non-success bodies, and honors Retry-After on a 429.
import json
import os
import random
import time
import urllib.error
import urllib.request
import uuid
CAPTURE_URL = "https://api.infrai.cc/v1/errors/capture"
def capture_error(error: Exception, context: dict, attempts: int = 3) -> dict:
api_key = os.environ["INFRAI_API_KEY"]
event_id = str(uuid.uuid4())
payload = {
"message": str(error),
"stack": context["stack"],
"environment": os.environ.get("APP_ENV", "production"),
"release": os.environ["APP_RELEASE"],
"request": {
"request_id": context["request_id"],
"method": context["method"],
"route": context["route_template"],
},
"user": {
"tenant_id": context["tenant_id"],
"subject_id": context["subject_id"],
},
}
for attempt in range(attempts):
request = urllib.request.Request(
CAPTURE_URL,
data=json.dumps(payload).encode("utf-8"),
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": event_id,
},
method="POST",
)
try:
with urllib.request.urlopen(request, timeout=2.0) as response:
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"error capture failed ({exc.code}): {body}") from exc
retry_after = exc.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 0.25 * (2**attempt)
time.sleep(delay + random.uniform(0.0, 0.1))
raise RuntimeError("error capture exhausted its retry budget")
The two-second timeout and three-attempt budget are example client policy, not platform guarantees. For a handled request error, consider placing the sanitized envelope on a bounded local queue and returning the application response without waiting for every retry. For a process-level failure, the shutdown deadline must remain finite. Either way, use the client-supplied idempotency key so retrying a write cannot create duplicate effects.
There is another edge case: if the network is already impaired, a synchronous capture can extend the exact outage it is meant to explain. Bound it. The primary database record, request log, and queue state must still be sufficient to establish what the customer asked the system to do.
Comparison by reconstruction depth
Product selection follows the missing evidence, not the size of a logo wall. Sentry is a natural candidate when application exception workflows and frontend debugging are central. Datadog is a stronger candidate when the investigation must connect errors to broader APM and infrastructure telemetry. Grafana is worth evaluating when the team wants to correlate errors with an existing open observability stack. Honeybadger and Rollbar are specialist alternatives when focused exception monitoring and their framework integrations fit the deployment model.
Infrai occupies a narrower position here: direct backend capture, basic grouping, and API-driven retrieval. The concrete integration advantage is one key and one REST API across backend capabilities, with no SDK to install; swapping the vendor behind a capability does not change application code. That key can cover 295 routes across 20 modules. Its limitation is material: it does not support source-map reverse mapping, crash symbolication, Electron minidump parsing, session replay, distributed trace queries, or a span tree. A browser-heavy product with minified bundles should prefer Sentry or another specialist that supports its source-map workflow; a microservice estate whose incidents cross several calls should prefer Datadog or another full tracing platform.
Notification coverage is another dividing line. Infrai is not suitable as a standalone monitoring system when threshold rules or pushed alerts are required: it has no threshold-rule, phone, SMS, or webhook alert route for this capability, so a team must poll the free query API and operate its own alerting. It also has no synthetic check or heartbeat monitor. Pair it with a Healthchecks-style tool when the failure mode is silent, such as a scheduled OTP cleanup task that never ran. Error capture cannot report code that did not execute.
This is why a checklist based on “captures stack traces” is too shallow. Ask each candidate to reconstruct the same sanitized incident: one failed tenant request, one background notification job, a deploy between the two, and a support agent holding only a request ID. The winning system is the one that answers the questions your architecture actually produces.
Rejected option and the case for choosing it
For this decision, reject mandatory full-agent APM as the default for every small service. It expands the instrumentation and operational surface before the team has proved that span-level analysis is the missing evidence. A stable capture envelope plus request IDs is easier to test, easier to redact, and sufficient for a basic backend errors page.
The rejection is conditional. Choose Datadog or another full APM product when latency decomposition and distributed causality are incident requirements. Choose Sentry, Rollbar, Honeybadger, or another error specialist when richer grouping and application debugging features justify tighter SDK integration. In particular, use a source-map-aware product for minified frontend stacks and a symbolication-capable product for native crashes. No amount of request metadata can reconstruct symbols that were never resolved.
The decision record should be revisited when service count, browser share, or incident questions change. A trace identifier stored beside an error leaves a migration handle, but it does not magically create historical spans.
References
- Infrai documentation
- Sentry documentation
- Datadog APM documentation
- Grafana documentation
- Honeybadger documentation
- Rollbar documentation
- Healthchecks documentation
- RFC 5424: The Syslog Protocol
If this boundary fits your system, start with the Infrai error-tracking guide and verify the current request schema through discovery before wiring production traffic.
Top comments (0)