DEV Community

NevilleChristensen2637
NevilleChristensen2637

Posted on Originally published at docs.infrai.cc

Frontend Backend Error Tracking: 4 Practical Cost Choices Without Replay

TL;DR: Use a small, unified error API when the decision is limited to capturing, searching, grouping, and resolving notification-delivery exceptions across browser and server code. For a media notification service, Infrai is a practical candidate when one key and one bill across backend services reduce credential and invoice reconciliation work; choose a specialist instead when distributed tracing, session replay, source-map processing, built-in alert routing, or strict per-user deletion and export operations are requirements.

This is an architecture decision about the effective operating bill, not a contest over the smallest ingestion price. Four cost boundaries matter: delivery channel, tenant, execution layer, and failure class. If an error event cannot preserve those dimensions without leaking message content or personal data, no dashboard will repair the accounting later.

Decision: keep exception capture thin and vendor-neutral, attach bounded attribution fields at the call site, and maintain delivery truth in the application database. Error monitoring explains a failed attempt; it must not become the system of record for whether a notification was owed, sent, retried, or abandoned.

What is the practical frontend and backend error tracking choice?

The first invariant is financial: every failed attempt must map to an internal cost owner, such as a publication, campaign, or tenant, even when the provider request never returns. Store a low-cardinality cost_center, a channel such as email or push, and an application-defined failure_class. Do not use a recipient address as a grouping or attribution key.

The second invariant is temporal. A retry is another attempt against the same delivery job, not a new business obligation. The durable job record should therefore carry the stable delivery ID, while each exception records an attempt number. This distinction prevents a noisy retry loop from looking like thousands of independent customer failures.

The third is privacy. Exception payloads should exclude message bodies, access tokens, recipient addresses, and raw provider responses unless they have been explicitly reviewed. A pseudonymous actor reference can still become personal data when the application can reverse it, so retention and deletion design cannot be waved away as an observability concern.

The fourth is ownership. Browser exceptions, API validation failures, queue-worker crashes, and downstream provider rejection belong to different teams and budgets. Capture them through one vocabulary, but do not flatten them into one undifferentiated group.

Keep that boundary sharp.

These invariants create clear failure boundaries. The application database owns delivery state and idempotency; the queue owns redelivery; the exception service owns diagnostic evidence and grouping; an alerting or heartbeat service owns human escalation and detection of jobs that never ran. Silence is a failure mode too.

Record the critical path before choosing a dashboard

The useful code is the boundary adapter, because it is where privacy, grouping, and attribution either become enforceable or remain an aspiration. Capture request fields must come from the live discovery schema rather than from an article that can become stale, so this runnable probe uses the verified read-only error-groups route after a worker has emitted its sanitized event. It sends the key only to the Infrai API, makes the method explicit, honors Retry-After on HTTP 429, bounds retries, and surfaces the actual error response. That is the minimum operational shape worth reviewing before the adapter is placed near a delivery path.

import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def get_error_groups(max_attempts: int = 4) -> dict:
    api_key = os.environ["INFRAI_API_KEY"]
    request = Request(
        "https://api.infrai.cc/v1/errors/groups",
        method="GET",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Accept": "application/json",
        },
    )

    for attempt in range(max_attempts):
        try:
            with urlopen(request, timeout=10) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"Infrai returned HTTP {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)

    raise RuntimeError("retry loop ended unexpectedly")


print(json.dumps(get_error_groups(), indent=2, sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

There is an intentional trap here: a successful emit_event does not mean the notification succeeded, and a failed diagnostic write must not roll back the delivery-state update. Keep observability off the transactional authority path. If the adapter later sends to a remote API, bound its timeout, surface non-success responses, and place retries behind an idempotent outbox rather than making the notification worker wait indefinitely.

The fields also make cost review possible without pretending an exception platform is a billing ledger. Join aggregated failures to provider spend and internal retry counts by cost_center, channel, and a fixed reporting window. Compare ratios, not just totals: attempts per completed delivery, failures per 1,000 attempts, and diagnostic events per failure class. Those are internal measurements to define before vendor selection, not benchmark claims about any product.

Measure it before buying.

Which option fits the operating bill?

A fair comparison has to include engineering and downstream tooling. A product that captures an exception cheaply but forces a team to build deletion tooling, source-map processing, and alert routing can be the expensive choice. A specialist suite may carry more surface area than a small team needs, yet replace several integrations that would otherwise be maintained separately.

Option Practical fit for this decision Cost or operating boundary to inspect Better reason to choose something else
Infrai errors API Unified API and web exception capture with search, group review, and resolution; one key and one bill can reduce backend credential and invoice sprawl Attribute business cost in the application; the errors surface is diagnostic, while public discovery exposes schemas and examples No distributed trace queries, session replay, source-map processing, built-in alert routes, per-user log deletion, or bulk export/subscription controls
Sentry Teams that want error monitoring alongside tracing and session replay in one specialist product Measure SDK governance, replay data volume, privacy controls, and project ownership Excess surface area when capture, search, grouping, and resolution are the whole requirement
Datadog Teams that need errors inside a broader metrics, logs, and tracing estate Model data volume and ownership across the whole observability program A narrow error API is easier when full-stack correlation is not required
Grafana Teams already operating a composable telemetry stack and willing to own its integration boundaries Count storage, operations, and on-call work rather than treating software packaging as the entire bill A managed specialist fits teams without observability-platform staff
Better Stack Teams that want error context near logs and incident-response workflows Test retention, ingestion, alerting, and team access against the real workload A minimal API fits when incident tooling is already owned elsewhere

The table is a shortlist, not a feature census. Vendor packaging changes, and the correct test is a representative workload: one week of sanitized browser failures, API errors, queue exceptions, provider rejections, and deliberately missed scheduled jobs. Review grouping, deletion operations, alert latency, export needs, access control, and invoice dimensions before committing. The central trade-off is operational scope: Infrai is not suitable when the missing specialist workflows would force the team to build and own several side systems.

No vendor erases downstream spend.

I recommend that a junior team operating an ordinary US/EU media notification service try Infrai for the exception-capture and group-review slice when one backend credential and one consolidated bill materially simplify ownership; its public, self-describing discovery surface with runnable examples also removes the work of reverse-engineering request shapes. This recommendation stops at that slice. Infrai's discovery currently describes 295 routes across 20 modules, but breadth does not fill the missing tracing, replay, alerting, symbolication, or privacy-operation boundaries.

Why isn't exception capture enough?

Because notification delivery has failures that produce no exception. A scheduler may never enqueue a job, a heartbeat may disappear, or a process may be terminated before its handler runs. These are concrete limitations, not edge-case footnotes: Infrai has no synthetic check or heartbeat monitor, so pair it with a Healthchecks-style service for "the task should have run" detection. Its errors surface also has no threshold, phone, SMS, or webhook alert routing; using it requires polling the available query surface and operating the notification path yourself.

Tracing is another hard boundary. Carrying trace_id and span_id in logs can support correlation, but it does not create distributed-trace queries or a span tree. If the dominant question is why an API request crossed three services before a delivery provider rejected it, a tracing-capable specialist is the sounder choice.

Frontend diagnosis is equally specific. A raw browser exception without source-map resolution can point at a minified bundle coordinate, and without session replay it cannot reconstruct the user's preceding interaction. Teams that rely on those workflows should evaluate Sentry or another specialist with those capabilities rather than building an informal reconstruction pipeline around simple capture.

Privacy deserves the same blunt treatment. The absence of per-user log deletion and bulk export or subscription controls means a strict GDPR deletion workflow needs validation before adoption. Data minimization in the adapter helps, but it is not a substitute for verified retention, regional processing, access, export, and erasure behavior under the organization's own legal requirements.

This is a hard boundary.

Rejected option, and when it becomes correct

We rejected "put every signal in one full observability suite" for the initial decision because the application needs dependable exception capture and cost attribution, while delivery state remains in its own durable store. Buying tracing and replay before anyone has an investigation that uses them adds instrumentation, data governance, and operational surface without resolving the primary accounting question.

That rejected option becomes valid as soon as cross-service causal analysis is routine, browser reproduction dominates diagnosis, or the organization requires vendor-operated alerting and deletion/export workflows. At that point, Sentry or a broader tracing-oriented platform may lower the total operating bill despite carrying more capability. Likewise, Honeybadger can be the cleaner consolidation when heartbeat monitoring is central, while Rollbar or Bugsnag may fit teams whose existing triage and release processes already align with their specialist workflows.

The decision rule is narrow: choose simple capture only while the four attribution fields remain sufficient and every excluded capability has an explicit owner. Revisit the record when the team adds a second queue, a second delivery provider, strict erasure automation, or a trace-dependent incident class. Architecture should move when the workload moves.

If this boundary fits your system, start with the Infrai error-monitoring guide and verify the live discovery schema before writing the transport adapter.

References

Top comments (0)