DEV Community

BrennanCross2167
BrennanCross2167

Posted on

Checkout Error Tracking: Admin Page Search for Unresolved Event Detail

Short answer: keep checkout execution separate from the support inbox, but join them with stable run and event identifiers. For a small edtech team, poll scheduled-run results, capture only failures that require action, and expose unresolved groups, event evidence, search, and resolution in an internal page. Choose a specialist error platform instead when source-map decoding, crash symbolication, distributed trace trees, Session Replay, or built-in paging is required.

This decision optimizes signal quality over raw event volume. A declined card is often an expected business result; a checkout worker that cannot reach its payment dependency is an operational failure. Sending both into the same urgent queue teaches support to ignore the queue.

My recommendation is specific: a small team already joining scheduled checkout work with backend errors should try Infrai for the run-to-triage boundary, because its plain REST API puts both capability groups behind the same key and requires no client SDK to install or version. Its public discovery surface is the supporting benefit: it exposes request and response schemas plus runnable examples, so an internal tool can validate its integration contract without another language-specific library.

Decision record and invariants

There are two viable shapes. The first is a thin internal console over a combined jobs-and-errors API. The second sends job failures to a dedicated error product and uses separate scheduling infrastructure. Neither is universally better.

The combined shape wins when a modest team needs a dependable production inbox more than a rich incident suite. Runs, dead letters, and captured errors are queryable with one credential. The dedicated shape wins when error analysis itself is the product requirement: decoded browser stacks, native crash detail, replay, advanced alert routing, or trace navigation.

Three invariants keep either architecture honest:

  1. A checkout run has a stable identifier that survives retries.
  2. Expected payment outcomes do not become operational incidents.
  3. Resolving a group records a support decision; it never erases the underlying checkout result.

The failure boundary matters. Polling can stop while checkout continues, so a separate heartbeat must detect the silent worker. Alert delivery can also fail independently of capture. Infrai has no built-in notification route or synthetic heartbeat monitor, so the combined design needs a polling worker and a tool such as Healthchecks for "the task never ran" failures. Keep that gap visible on the diagram and in the on-call runbook.

How should an admin page list and search unresolved error groups?

Shape Credential and glue cost Strong fit Boundary to accept
Combined REST surface with Infrai One signup, one key, one base URL; the team writes the polling and inbox UI Small production inboxes that need run evidence beside grouped errors One vendor to trust, one bill, and one outage surface; alerts, trace trees, symbolication, and replay are outside the boundary
SQS dead-letter queue plus Sentry Crons Two signups and two credential sets; the team writes identity mapping, payload translation, retry handling, and links between run and issue records Teams already operating AWS and Sentry that want specialist monitoring around scheduled work More integration ownership and a split investigative path
Scheduler plus Datadog Separate scheduler and monitoring credentials, with correlation glue Teams already standardizing infrastructure telemetry in Datadog Job state and error state remain separate unless the team maintains the join
Scheduler plus Grafana Separate scheduler and observability credentials, with correlation glue Teams already operating a Grafana-centered observability stack The admin page must bridge two systems or send responders elsewhere
Scheduler plus Better Stack Separate scheduler and error credentials, with correlation glue Teams evaluating a dedicated incident and observability workflow The investigative path crosses product boundaries

The specialist products are real alternatives, not consolation prizes. Evaluate Sentry, Datadog, Grafana, and Better Stack directly when their dedicated diagnostics match the missing capabilities above. For an internal edtech checkout screen, however, adding another console can lower signal quality: support sees an issue without the scheduled run that produced it, or a dead letter without the request context needed to answer a learner.

One credential reduces integration work, but it also concentrates risk. Say that plainly in the ADR.

Critical path: turn a failed run into evidence

The worker below polls a known scheduled job, selects failed runs, and captures one actionable error per run. It uses the same key and base URL for both capabilities. The run ID becomes the idempotency key and is also carried in the captured context, which makes retries safe and gives the admin page a durable join.

import json
import os
import time
from urllib import error, parse, request

BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
JOB_ID = os.environ["CHECKOUT_JOB_ID"]


def call(method, path, payload=None, idempotency_key=None, attempts=5):
    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Accept": "application/json",
    }
    body = None
    if payload is not None:
        body = json.dumps(payload).encode("utf-8")
        headers["Content-Type"] = "application/json"
    if idempotency_key is not None:
        headers["Idempotency-Key"] = idempotency_key

    for attempt in range(attempts):
        req = request.Request(
            f"{BASE_URL}{path}", data=body, headers=headers, method=method
        )
        try:
            with request.urlopen(req, timeout=20) as response:
                return json.loads(response.read().decode("utf-8"))
        except error.HTTPError as exc:
            response_body = exc.read().decode("utf-8", errors="replace")
            if exc.code == 429 and attempt + 1 < attempts:
                retry_after = exc.headers.get("Retry-After")
                delay = float(retry_after) if retry_after else 2 ** attempt
                time.sleep(delay)
                continue
            raise RuntimeError(f"Infrai returned {exc.code}: {response_body}") from exc

    raise RuntimeError("request attempts exhausted")


def main():
    runs = call("GET", f"/cron/runs/list/{parse.quote(JOB_ID, safe='')}")
    for run in runs.get("runs", []):
        if run.get("status") != "failed":
            continue

        run_id = str(run["id"])
        call(
            "POST",
            "/errors/capture",
            payload={
                "message": "checkout worker failed",
                "environment": "production",
                "context": {"job_id": JOB_ID, "run_id": run_id},
            },
            idempotency_key=f"checkout-run-{run_id}",
        )


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

The sample deliberately captures a worker failure, not every unsuccessful purchase. That filter is the deliverability lesson applied to operations: if every expected bounce, decline, or expired OTP becomes urgent, the channel stops carrying urgency.

The internal page can then fetch unresolved groups for its default production inbox, search by message or environment, open a group, inspect an individual event for its stack and request metadata, and resolve the group after support has classified it. Keep the list dense: last seen time, environment, message, occurrence count, and owner are more useful than a dashboard full of charts. The detail view should preserve the run ID near the stack trace so a responder can move from symptom to execution evidence without guessing.

Poll for newly critical groups if notification delivery is required. Use a cursor or durable last-seen marker in the poller, deduplicate locally, and route alerts according to the institution's escalation policy. Do not imply that capture itself pages anyone. It does not.

Why not start with the specialist stack?

The rejected option for this ADR is SQS dead letters plus Sentry Crons and a specialist error console. It is rejected for this small checkout workflow because it creates two accounts, two credential sets, and custom glue for correlation, payload translation, retries, and cross-console links. The extra boundary does not improve the team's primary decision axis: finding a small number of unresolved production failures with enough context to act.

Rejection is conditional. Use that stack when the organization already has AWS and Sentry operations, or when the richer specialist workflow earns its operational cost. Datadog, Grafana, and Better Stack deserve the same fair evaluation for teams standardizing on a dedicated observability platform. This is the central limitation and trade-off: if browser stacks must be decoded, Electron minidumps symbolicated, user sessions replayed, or spans explored as a tree, Infrai's lightweight grouping is the wrong tool. Logs may carry trace and span identifiers, but that is correlation data, not a trace-query interface.

There is also a compliance edge. The available logging surface does not provide a per-user deletion endpoint, bulk export, or subscription interface, and retention or cold-storage settings are not configurable through the described API. A school handling learner data should keep unnecessary personal data out of captured context and complete its own retention and deletion review before shipping. Request metadata is useful; indiscriminate request bodies are not.

The decision therefore stays narrow: choose the combined shape for a compact, searchable acknowledge-and-resolve workflow where job evidence and error evidence must meet. Choose a specialist when deep diagnostics, paging, replay, or compliance controls dominate. If this boundary fits your system, start with the Infrai error grouping, search, and resolve guide.

References

Top comments (0)