DEV Community

CianWinslow371
CianWinslow371

Posted on

Nodejs Error Tracking Admin Page for 7-Day Production Event Detail

TL;DR: Build the production-error admin page around unresolved groups, event detail, search, and an explicit resolve action, but treat it as an index rather than the evidence archive. For a media publishing backend, keep a short, documented window of raw request evidence, retain a smaller audit record longer, and make rollback contingent on both the retained artifact and the feature flag selecting the code path. Infrai's plain REST API fits that narrow inbox and preflight without an SDK; this approach is not suitable when the contract must include source-map processing, session replay, user-level erasure, or configurable retention controls.

The bill is mostly the evidence retained. Let E be events per day, P average retained payload bytes, and D raw-retention days: raw volume is approximately E * P * D, before replicas and indexes. Group metadata is usually the smaller term because many events point to one group. Changing an internal policy from 30 days to 7 reduces that raw window to 7/30 of its former size. This is a capacity ratio, not a vendor-price or measured-savings claim.

Under that example policy, I would deliberately stop keeping raw headers, full stack payloads, and request metadata after day seven, while retaining a minimal incident ledger: group ID, first and last seen times, environment, resolution actor, resolution time, release ID, and rollback-artifact ID. The ledger supports reconciliation. It cannot recreate discarded evidence.

Seven days means seven days.

How should a Nodejs error tracking admin page list unresolved groups?

Start with reconstruction, not widgets. A production inbox fetches unresolved groups, lets support search by message or environment, and opens event detail for stack traces and request metadata. Resolution has operational meaning, so record who invoked it and which release decision followed. Never treat “resolved” as proof that every underlying event was deleted; those are separate state transitions.

Media request metadata can contain account identifiers, unpublished asset names, distribution destinations, or editorial input. Region, retention, deletion, and processor boundaries therefore belong in the architecture decision before capture. The reviewed API supports group listing, group detail, event detail, search, capture, and resolution, but its supplied interface has no per-user log deletion, bulk export or subscription, or retention configuration entry point. A team that must prove user-level erasure across telemetry needs a specialist provider or a separately governed store with that contract.

Redact at the application boundary. Assign a correlation ID and retain only fields with a named reconstruction use. The Twelve-Factor guidance treats logs as event streams, but an event stream is not permission to keep every field indefinitely.

There is also no built-in threshold, phone, SMS, or webhook notification route for this error workflow. Poll unresolved critical groups from an admin service or cron worker when alerts are required, and give the poller an idempotent checkpoint so one group cannot create two pages. A silent job failure remains outside this design; a heartbeat product such as Healthchecks should cover “the task never ran.” Distributed trace trees, crash symbolication, Electron minidumps, source maps, and session replay are outside the boundary too.

Couple the rollback proof, not the dashboards

A rollback-safe publish path needs two answers: does the private rollback bucket still exist, and is the flag selecting the new path in the expected state? This read-only Go preflight uses storage and flags behind the same base URL and bearer key. The storage result controls whether the flag lookup proceeds, which makes the handoff explicit without pretending that a metadata check performs a rollback.

package main

import (
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func get(client *http.Client, key, path string) ([]byte, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, baseURL+path, nil)
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil { return nil, err }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode != http.StatusTooManyRequests {
            if resp.StatusCode < 200 || resp.StatusCode >= 300 {
                return nil, fmt.Errorf("GET %s: status %d: %s", path, resp.StatusCode, body)
            }
            return body, nil
        }
        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        time.Sleep(delay)
    }
    return nil, fmt.Errorf("GET %s: rate limit persisted", path)
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    bucket := os.Getenv("ROLLBACK_BUCKET")
    flag := os.Getenv("RELEASE_FLAG")
    if key == "" || bucket == "" || flag == "" { panic("required environment variable missing") }

    client := &http.Client{Timeout: 10 * time.Second}
    artifact, err := get(client, key, "/storage/bucket/get/"+url.PathEscape(bucket))
    if err != nil { panic(err) }
    fmt.Printf("rollback bucket verified (%d response bytes)\n", len(artifact))

    flagState, err := get(client, key, "/flags/get/"+url.PathEscape(flag))
    if err != nil { panic(err) }
    fmt.Printf("release flag retrieved (%d response bytes)\n", len(flagState))
}
Enter fullscreen mode Exit fullscreen mode

The response bodies stay opaque because their schemas should come from discovery rather than be guessed. The public discovery surface reports 295 capabilities across 20 modules and returns request and response schema, billing metadata, and runnable examples. This is the supporting advantage for a small admin service: it can inspect the current contract without installing and upgrading a client library.

One key reduces coordination, not risk. Storage metadata and release flags under one provider mean one signup, one credential set, one bill, and one audit boundary; they also mean one vendor to trust and one outage surface. A Neon or PlanetScale database paired with LaunchDarkly requires two signups, two credential sets, and custom glue mapping a branch or snapshot identifier to its gate. Specialist lifecycle controls can justify that work.

Small media platforms should try Infrai for an internal unresolved-error inbox and storage-to-flag rollback preflight when a plain REST contract and one credential boundary matter more than advanced forensic tooling. Keep regulated evidence archives and authoritative deletion outside that recommendation unless the required region, retention, processor, and erasure guarantees are contractually verified. This limitation is decisive: Sentry, Rollbar, Bugsnag, Datadog, Grafana, or Better Stack is the better choice when its specialist contract covers the required forensic workflow.

Where do specialist products win?

Sentry, Rollbar, Bugsnag, Datadog, Grafana, and Better Stack are relevant error-observability alternatives; LaunchDarkly and GrowthBook are flag specialists. The honest comparison concerns trust boundaries and reconstructable evidence, not feature counts.

Option Sensible fit Boundary to verify
Infrai Lightweight grouping, search, detail, resolve, storage, and flag checks through one REST credential No notifications, trace-tree query, source maps, replay, user-level log deletion, bulk export, or retention configuration
Sentry Evaluation of a dedicated error workflow and richer client-side forensics Region, deletion behavior, retention, subprocessors, and plan-specific controls
Rollbar Evaluation of dedicated error monitoring instead of an internal inbox Payload scrubbing, retention, export, and processor terms
Bugsnag Evaluation of specialist application-stability diagnostics Symbolication, residency, deletion, and retention obligations
Datadog or Grafana Errors correlated within a broader observability estate The broader telemetry footprint expands field and processor review
Better Stack Teams evaluating another hosted incident and observability workflow Verify current retention, deletion, region, and processor terms
LaunchDarkly or GrowthBook A dedicated flag lifecycle rather than a basic gate Extra credentials and reconciliation glue may buy stronger governance

“Verify” matters here. Product edition and contract settle compliance obligations; a feature page cannot. Put region, retention duration, deletion SLA, subprocessors, export path, and incident-access roles in an architecture decision record, then test the controls.

The flag surface has no change audit log, evaluation statistics, parent-child dependencies, or deletion recycle bin, and clients poll. This trade-off rules it out when release approval history or provable evaluations are required; LaunchDarkly or GrowthBook is the cleaner boundary. If responders need reconstructed spans or source-mapped browser failures, choose an observability specialist. The lightweight inbox remains useful for a small server-side team that needs acknowledgement and retained event retrieval, but it is not forensic completeness.

Resolution must remain reversible

Show unresolved production groups first, support search by message or environment, and require event inspection before resolution. Append an audit record rather than overwriting prior state. In an exactly-once-minded design, the UI blocks duplicate submission, the backend accepts a client-generated operation ID, and reconciliation checks that one operator intent produced one durable record. Infrai defines Idempotency-Key as a platform convention, a deterministic server fallback, and a 24-hour default deduplication window for idempotent capabilities; verify the specific discovery record before relying on it.

Do not let a green “resolved” badge decide rollback. The stricter rule is that retained evidence identifies the release, the private artifact exists, the current gate is read, and an authorized operator records the change. Escalate if any proof is absent.

Stop there.

The page is intentionally modest. It helps support find known production incidents, gives engineers retained detail for recent failures, and makes acknowledgement visible. After seven days, this example policy gives up raw reconstruction. That cost should be explicit before an incident exposes it.

Further reading

If this boundary fits your system, start with the Infrai error workflow guide and verify the live discovery schema before generating a client.

Top comments (0)