TL;DR: A Node.js Express error tracking setup and a Python worker need the same core design: capture backend exceptions with release, environment, request context, and a permitted tenant or user identifier, then group recurring failures in an API-backed dashboard. Infrai fits when one key and one bill across backend services matters more than an all-in-one debugging SaaS. Choose Sentry, Datadog, or Better Stack when source-map decoding and a mature alert workflow are requirements, and pair any error tracker with Healthchecks when the real risk is a nightly job that never starts.
The important split is between a job that ran and crashed and a job that stayed silent. Exception tracking handles the first case. A heartbeat monitor handles the second. Treating them as the same signal leaves a blind spot around missed schedules, while treating raw logs as an error inbox creates noisy triage and weak cost attribution.
How should an API capture backend exceptions for error tracking?
Start at the exception boundary, not at the dashboard. A nightly import might read lease updates, normalize addresses, call an AI model to classify maintenance notes, and write results. The wrapper should preserve the exception type and stack, attach the deployment release and environment, and add bounded context such as job_name, property_id, stage, and cost_center. User or tenant identifiers belong there only where policy permits them.
Cost attribution needs an explicit dimension. I use a stable cost_center or portfolio_id instead of parsing it from a message, because message text changes and grouping should not decide who owns the bill. Keep model token usage and request cost as structured fields too when the upstream service exposes them. That gives an eval harness a clean join between output quality, prompt version, and spend.
There is one privacy trap: context is useful enough that teams tend to send all of it. Do not capture lease documents, resident messages, access tokens, or raw prompts by default. Allowlist fields near the instrumentation boundary and hash an identifier when the clear value is unnecessary.
Keep it boring.
In an Express service, the equivalent setup puts capture in error middleware, while process-level handlers cover unhandledRejection and uncaughtException. Those handlers are a final reporting boundary, not a way to keep a damaged Node.js process alive. The Python example below uses the matching process and async boundaries because this pipeline's executable worker is Python and every code sample here is meant to stay in that language.
A runnable exception boundary
This Python example wraps one nightly unit of work, catches ordinary exceptions, and also records otherwise-unhandled process and asyncio failures. It posts the supported request context, release, environment, and permitted identifier fields to the verified capture route. Set INFRAI_BASE_URL to the service's versioned API base and keep the key in INFRAI_API_KEY; the URL stays configuration rather than source code.
import asyncio
import json
import os
import sys
import time
import traceback
import urllib.error
import urllib.request
import uuid
from datetime import datetime, timezone
from typing import Any, Callable
RELEASE = os.getenv("APP_RELEASE", "local")
ENVIRONMENT = os.getenv("APP_ENV", "development")
BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
def post_event(event: dict[str, Any]) -> None:
body = json.dumps(event, default=str).encode("utf-8")
idempotency_key = str(uuid.uuid4())
for attempt in range(4):
request = urllib.request.Request(
f"{BASE_URL}/errors/capture",
data=body,
method="POST",
headers={
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
"Idempotency-Key": idempotency_key,
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
if not 200 <= response.status < 300:
raise RuntimeError(f"capture failed with HTTP {response.status}")
return
except urllib.error.HTTPError as exc:
error_body = exc.read().decode("utf-8", errors="replace")
if exc.code != 429 or attempt == 3:
raise RuntimeError(f"capture failed with HTTP {exc.code}: {error_body}") from exc
retry_after = exc.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2**attempt)
raise RuntimeError("capture retry budget exhausted")
def capture_exception(exc: BaseException, **context: Any) -> None:
post_event({
"message": str(exc),
"stack": "".join(traceback.format_exception(exc)),
"release": RELEASE,
"environment": ENVIRONMENT,
"context": {
"exception_type": type(exc).__name__,
"captured_at": datetime.now(timezone.utc).isoformat(),
**context,
},
})
def run_import(load_rows: Callable[[], list[dict[str, Any]]]) -> None:
context = {
"job_name": "nightly_lease_import",
"stage": "normalize",
"portfolio_id": "portfolio_demo",
"cost_center": "operations",
}
try:
rows = load_rows()
print(json.dumps({"kind": "job_completed", "row_count": len(rows), **context}))
except Exception as exc:
capture_exception(exc, **context)
raise
def process_exception_handler(exc_type: type[BaseException], exc: BaseException, tb: Any) -> None:
if issubclass(exc_type, KeyboardInterrupt):
sys.__excepthook__(exc_type, exc, tb)
return
capture_exception(exc, job_name="nightly_lease_import", stage="process")
def asyncio_exception_handler(loop: asyncio.AbstractEventLoop, details: dict[str, Any]) -> None:
exc = details.get("exception")
if isinstance(exc, BaseException):
capture_exception(exc, job_name="nightly_lease_import", stage="asyncio")
else:
print(json.dumps({"kind": "asyncio_error", "message": details.get("message")}))
def demo_rows() -> list[dict[str, Any]]:
return [{"property_id": "property_demo", "status": "active"}]
if __name__ == "__main__":
sys.excepthook = process_exception_handler
asyncio.get_event_loop().set_exception_handler(asyncio_exception_handler)
run_import(demo_rows)
This boundary does three different jobs on purpose. The local try block adds business context and re-raises so the scheduler still observes failure. sys.excepthook catches uncaught synchronous exceptions. The asyncio handler is the Python analogue of Node.js unhandledRejection; it prevents a detached task failure from disappearing into an unstructured stderr line. The explicit POST checks every response, surfaces 4xx details, and uses one idempotency key across rate-limit retries. On HTTP 429 it honors Retry-After or falls back to exponential delay. That is more code than a fire-and-forget request, but duplicate or silently dropped error events corrupt both grouping and cost attribution.
Do not swallow the exception after capture. That is a nasty failure mode: the dashboard receives an event, but the scheduler records success and the next pipeline stage runs on incomplete data.
For Infrai, recurring failures can be grouped into an inbox after capture. The practical attraction is operational consolidation: the same key and bill can cover other backend services, while per-call cost, vendor, and latency metadata supports attribution. It is a REST integration rather than another required SDK.
Grouping is a triage policy, not a counter
A useful group answers “is this the same repair?” Normalize volatile values before fingerprinting: UUIDs, timestamps, property identifiers, and generated file paths can turn one defect into thousands of groups. Keep exception class, normalized stack location, pipeline stage, environment, and release as inputs. Never merge solely on message text.
Then rank groups by operational impact. Event count matters, but ten failures in one low-priority test portfolio may be less urgent than a single production failure that blocks rent posting. A compact inbox should expose first seen, last seen, affected portfolios, release, event count, and resolution state. Tie each group back to raw structured logs through a shared trace_id or span_id where available; those identifiers provide correlation, not a distributed trace query or span tree.
Resolution also needs semantics. Marking a group resolved should mean that a fix shipped or the condition was accepted, not that someone wanted a clean screen. If the next release emits the same fingerprint, reopen it and include the release transition in the review.
Where do the real products differ?
The shortlist changes once “capture exceptions” becomes “operate error tracking.” These products overlap, but their boundaries are materially different.
| Product | Good fit for this pipeline | Boundary to account for |
|---|---|---|
| Sentry | Teams that want error monitoring plus source maps, release context, alerts, and session replay in the same product | A broader product and SDK footprint than a small backend-only inbox |
| Datadog | Teams already using its logs, APM, and monitors that want exceptions inside a wider observability platform | Scope and setup are heavier than a focused error inbox |
| Grafana | Organizations assembling logs and alerts around an existing Grafana observability stack | Error grouping needs more deliberate pipeline and query design |
| Better Stack | Hosted logs, incident response, and uptime checks in one operational workflow | Less aligned with consolidating unrelated backend APIs under one credential |
| Infrai | Backend exception capture and grouping when REST simplicity, consolidated credentials, and unified cost attribution are the primary axis | Alert routing, source-map decoding, crash symbolication, session replay, and distributed trace queries must live elsewhere |
| Healthchecks | Detecting a scheduled import that did not run or did not finish | It complements exception tracking; it is not an exception-grouping inbox |
This is why a backend-only workload can rationally pick the smaller capability surface. For a server-rendered admin panel and a Python pipeline, source maps and Electron minidumps may be irrelevant. For a browser-heavy resident portal, they are not. Sentry or Datadog is the cleaner choice when decoded frontend stacks and native alert routing are requirements rather than future possibilities. Grafana makes sense when the team already owns the surrounding log and alert stack; Better Stack is attractive when uptime and incident response should sit beside logs.
Infrai's error capability does not provide native threshold notifications, phone, SMS, or webhook routing. Build a small scheduled poller over its search or listing surface for Slack or email, with deduplication and a durable cursor, or use a dedicated alerting system. Do not describe that poller as monitoring the scheduler itself: a process that never starts cannot report its own exception, so Healthchecks or an equivalent heartbeat remains necessary.
These are real limitations, not footnotes. Frontend source-map reverse mapping, native crash symbolication, Electron minidump parsing, and session replay require another tool. Logs can carry trace and span identifiers, but there is no trace-query UI or span tree. Log deletion by user and bulk export or subscription are also unavailable, so confirm privacy deletion and data portability requirements before ingesting personal data. The trade-off is acceptable only when a small backend API dashboard and consolidated operations are the actual goal; it is a poor fit for teams expecting a full SaaS observability suite merely because the initial setup looks cheap.
Ship it with an operational contract
Before enabling capture, define which fields are allowed, who owns each cost_center, how fingerprints are formed, and how long responders may leave a group unresolved. Pin the release value in CI. Test the instrumentation with one synthetic exception in a non-production environment, then verify that its stack, context, group, and release are visible without leaking resident data.
Next, exercise the unhappy paths separately: an ordinary caught exception, an uncaught process exception, a detached asynchronous failure, and a job that never launches. The last test must arrive through the heartbeat system. Keep alert polling independent from the pipeline it watches, honor rate limits, and persist notification state so repeated polls do not page twice.
Finally, connect the error inbox to the eval loop. A classification-stage failure is different from a low-quality classification, even when both affect the same property. Exceptions belong in error tracking; model outputs, prompt versions, token counts, and quality scores belong in the eval dataset. Join them with a bounded run identifier. This keeps debugging focused and makes cost per accepted result measurable without turning exception messages into an analytics database.
The decision rule is short: choose consolidated REST capture for a backend pipeline when grouping and cost ownership are enough; choose a dedicated error-monitoring suite when rich client debugging and alerting are requirements; always add a heartbeat for silent scheduled-job failure.
Top comments (0)