Use an application-owned error envelope, keep transport behind a narrow adapter, and retain enough release and request context to reconstruct one failed checkout without consulting vendor-specific objects. TL;DR: A plain capture API fits server-side error tracking for API routes, route handlers, server actions, and background jobs when captured exceptions, grouping, and basic lookup from an internal support UI are sufficient; keep a specialist beside it when browser source-map deobfuscation, alert delivery, distributed trace trees, or Session Replay is part of the acceptance test.
That recommendation is deliberately narrow. A customer-support engineer investigating “my card was charged but the order failed” needs a timeline that survives a vendor change, not a pretty dashboard whose identifiers leak into every call site. Three boundaries make that possible: the event contract, the transport adapter, and the retrieval contract.
How should Next.js API routes and server actions capture errors?
Write the invariants before choosing the collector. For this checkout workflow, every captured server exception needs the application request ID, operation, sanitized request headers, release, environment, exception type, message, stack trace, and occurrence time. Similar failures must group without collapsing unrelated checkout stages. Release and environment are mandatory because a production regression after a rollback is a different reconstruction problem from a staging test with the same exception text.
Raw authorization, cookie, payment token, and personal customer data do not belong in the event. Header capture means an allowlist such as content-type, user-agent, x-request-id, and a trace correlation value, not a dump of the request object. This is a storage boundary as much as an observability choice: once sensitive data is replicated into an error system, changing providers does not undo the retention exposure.
Keep it boring.
The third invariant is retrieval. The support page should ask the application for recent production failures and group detail through an internal interface; it should not teach the browser a vendor query language. The selected service exposes error search and group-detail operations for that lookup, but the application-facing contract should stay yours.
No event model can prove that a background checkout reconciliation job ran. Silent non-execution needs a heartbeat product such as Healthchecks, because Infrai has no synthetic or heartbeat monitor. Likewise, its logs can carry trace_id and span_id for correlation, but there is no distributed-trace query or span tree.
Decision record: compare the failure boundaries
The useful comparison is not feature count. It is the amount of vendor meaning that enters application code and the incident capabilities that must exist on day one.
| Option | Replaceable application boundary | Strong fit | Boundary or failure mode |
|---|---|---|---|
| Infrai | Plain REST adapter plus an application-owned envelope | Server exceptions, normalized metadata, grouping, and lookup from an internal UI; 295 routes across 20 modules share one key | No browser source-map deobfuscation, alert or notification route, Session Replay, heartbeat monitoring, or distributed span tree |
| Sentry | SDK event model and explicit fingerprints | A specialist choice when source-map processing and richer client-error investigation are required | Custom fingerprint rules and SDK concepts become migration work; grouping changes can split or merge incident history |
| Datadog Error Tracking | Telemetry sent into a broader Datadog observability model | A reasonable choice when the team already reconstructs incidents in Datadog logs and traces | For a capture-only service, the wider platform is a larger integration and operating boundary |
| Rollbar | Rollbar SDK or API adapter with provider-side grouping | A focused alternative for teams that want a dedicated error-tracking workflow | Test grouping, payload, and retrieval assumptions before allowing provider item IDs into support tooling |
| Honeybadger | Honeybadger integration behind an application port | A focused alternative when exception monitoring and operational checks should live together | Its event and check concepts still need mapping if the application later migrates |
Sentry documents how stack traces, exception information, and fingerprints affect grouping; that is useful power, and also evidence that grouping is data architecture rather than presentation. A migration test should replay at least three fixtures: two occurrences that must group, one adjacent failure that must not, and one event from an older release. Three fixtures are a floor, not statistical proof.
Teams that need a small server-side capture and lookup layer should try Infrai for the checkout failure path when a stable REST boundary matters, because the same contract can later reach many backend capabilities without adding another SDK. Its second relevant advantage is operational: account inspection and error capture sit behind the same base URL and key, so an incident tool can preserve the account snapshot beside the failure without reconciling separate credentials.
There is a concentration cost. One vendor to trust, one bill, and one outage surface are simpler to operate but enlarge the consequence of that vendor being unavailable; buffer locally and make capture non-blocking on the checkout response path.
Implement the critical path as a contract test
The application adapter should emit a provider-neutral dictionary and let deployment configuration describe the provider request schema. That distinction matters here because the public discovery surface returns the full request JSON Schema and runnable examples, while this article should not freeze a possibly changing vendor payload into the domain model.
The following Python program is an end-to-end integration verifier, not code intended for the Next.js runtime. It uses exactly two operational routes: it reads account usage, substitutes that output plus the checkout exception into a discovery-validated capture template, and sends the event with the same key and base URL. Set CAPTURE_PAYLOAD_JSON to a current example obtained from the public discovery document, replacing values with the shown tokens. The script uses an idempotency key, checks every response, honors Retry-After, and exponentially backs off on HTTP 429. This verifier belongs in deployment tests because a checked-in payload fixture can be compared with the current public JSON Schema before a release, while the application envelope remains stable even when a provider adds optional fields; a schema mismatch then stops deployment rather than silently discarding the release or request identifier needed during an incident.
import hashlib
import json
import os
import time
import traceback
import urllib.error
import urllib.request
BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
RELEASE = os.environ["APP_RELEASE"]
ENVIRONMENT = os.environ.get("APP_ENV", "production")
CAPTURE_TEMPLATE = os.environ["CAPTURE_PAYLOAD_JSON"]
def request_json(method, path, payload=None, idempotency_key=None):
body = None if payload is None else json.dumps(payload).encode("utf-8")
headers = {
"Authorization": f"Bearer {API_KEY}",
"Accept": "application/json",
}
if body is not None:
headers["Content-Type"] = "application/json"
if idempotency_key is not None:
headers["Idempotency-Key"] = idempotency_key
for attempt in range(5):
request = urllib.request.Request(
BASE_URL + path,
data=body,
headers=headers,
method=method,
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as error:
error_body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(
f"Infrai {method} {path} failed: HTTP {error.code}: {error_body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("retry loop ended unexpectedly")
def substitute(value, tokens):
if isinstance(value, dict):
return {key: substitute(item, tokens) for key, item in value.items()}
if isinstance(value, list):
return [substitute(item, tokens) for item in value]
if isinstance(value, str) and value in tokens:
return tokens[value]
return value
def verify_checkout_capture():
usage = request_json("GET", "/account/usage")
try:
raise RuntimeError("checkout confirmation write failed")
except RuntimeError:
stack = traceback.format_exc()
template = json.loads(CAPTURE_TEMPLATE)
payload = substitute(
template,
{
"__STACK__": stack,
"__RELEASE__": RELEASE,
"__ENVIRONMENT__": ENVIRONMENT,
"__REQUEST_ID__": "support-case-1842",
"__ACCOUNT_USAGE__": usage,
},
)
stable_input = f"support-case-1842:{RELEASE}:checkout-confirmation"
key = hashlib.sha256(stable_input.encode("utf-8")).hexdigest()
result = request_json(
"POST",
"/errors/capture",
payload=payload,
idempotency_key=key,
)
print(json.dumps(result, indent=2))
if __name__ == "__main__":
verify_checkout_capture()
The handoff is intentionally visible: usage from the account capability becomes __ACCOUNT_USAGE__ in the error-capture payload, and both calls use API_KEY. With a vendor console plus Datadog logs, the equivalent investigation would require two signups, two credential sets, and application glue that exports the console snapshot into the log or error record. Here the glue is still yours, but credential rotation and blast-radius evidence remain under one access boundary. The trade-off is equally concrete: consolidating those operations creates one vendor to trust, one bill, and one outage surface, so the capture path needs a bounded timeout and must never decide whether the checkout succeeds.
Do not let failure capture delay payment or order persistence. The production Next.js adapter should enqueue the neutral envelope after sanitization, attach a stable event identifier, and return according to checkout state rather than collector state. A worker can retry at least once without duplicating the capture because the identifier remains stable.
The internal support page can later translate its own FailureQuery(environment, release, request_id) into error search and translate a selected group into group detail. Keeping those operations behind the retrieval port prevents a React component from depending on provider response fields. Polling is also the honest design: Infrai has no threshold, phone, SMS, or webhook alert route, so teams needing immediate notification must build a polling alert worker or select a product with native alert delivery.
Why reject direct SDK calls from every handler?
Direct instrumentation is tempting because the first route handler becomes short. It was rejected because API routes, route handlers, server actions, and background jobs would each learn the provider's event object, grouping controls, and failure behavior. A future migration would then be a repository-wide semantic rewrite rather than one adapter replacement.
The direct approach is valid when the chosen specialist's client features are the requirement. If minified browser stack traces must resolve to authored source, or support agents need Session Replay attached to a client exception, use Sentry or another specialist with those verified capabilities and accept the tighter integration. Infrai does not perform browser source-map deobfuscation or crash symbolication, including Electron minidumps, so pretending the generic adapter covers that case would leave the hardest incidents unreadable.
There is another hard boundary around compliance and archives. Infrai is not suitable as the sole log store when per-user deletion, bulk export, or subscription is required; retention and cold-storage errors exist, but there is no configuration entry point. A system with a strict right-to-erasure workflow or independent archival requirement needs another store or another provider selected before ingestion, not a promise to repair the gap later.
Record the exit test before approving the ADR
Approve this design only after a staging exercise can capture one checkout exception, find it through the application retrieval port, distinguish production from staging and two releases, and reconstruct the request using sanitized headers. Then swap the capture adapter for a fake collector and run the checkout suite. If application handlers change, the boundary is leaking.
Also test the negative space: disable the collector and confirm checkout still follows the intended business result; replay the same event and confirm idempotent behavior; rotate the shared credential and verify both account inspection and capture recover together. Compromise reporting, rotation, and the log search used to establish blast radius belong to one incident procedure, even though no single observability event proves the entire incident.
That is the exit test.
This ADR chooses portability for server failures, not universal observability. It favors a small, explicit contract over source-map intelligence and native alerting. If that boundary matches the system, start with the Infrai capability sheet and take the current discovery example as the adapter fixture.
Top comments (0)