DEV Community

FluxH91
FluxH91

Posted on

Error Tracking Admin Page: Retained Event List for Rollback-Safe Triage

Keep one searchable record per failure group, a bounded sample of its events, and immutable state transitions for resolution and rollback. That is the least complex admin page that can answer which production notification failures remain open, show enough event detail to diagnose them, and reverse a mistaken resolution without retaining every payload forever.

TL;DR: storage cost is driven mainly by retained event volume: delivery attempts multiplied by failure rate, bytes per normalized event, replicas, indexes, and retention time. Group first, sample deliberately, redact before persistence, and treat “resolved” as an append-only transition rather than deletion. The screen then needs four operations: filter unresolved groups, search stable fields, inspect representative events, and resolve or reopen with an actor and reason.

For a logistics notification service, the expensive object is not the row that says “carrier webhook timed out.” It is the repeated evidence behind that row: destination identifiers, attempt history, stack traces, request context, indexes, and replicas. Consider an illustrative planning case, not a benchmark: 12 million delivery-notification attempts per day, a 0.5% failed-attempt rate, and a 4 KiB normalized event produce about 234 MiB of raw failure events daily before indexing or replication. Keeping 30 days instead of 7 multiplies that dominant term by roughly 4.3. A prettier list page does nothing to change it.

The useful change is to separate group metadata from event evidence. Retain compact group records and their state history longer; keep a capped, time-distributed sample of raw events for a shorter window. Deliberately stop keeping every duplicate stack trace and every original payload. The cost is real: after the raw-event window expires, an engineer can establish frequency and lifecycle from group counters, but may no longer reconstruct a rare payload-specific failure. That loss must be an explicit retention decision, not an accidental side effect of a cleanup job.

What should the storage model preserve?

Start with three records: a normalized event, a group, and a group-state transition. The event carries a generated event ID, observed time, deployment identifier, environment, exception class, normalized stack fingerprint, transport channel, carrier or route identifier, and a redacted context envelope. The group carries the fingerprint, first-seen and last-seen times, occurrence count, latest deployment, and current state. The transition carries group ID, previous state, next state, actor, reason, and time.

Do not make the raw error message the group key. Shipment IDs, phone numbers, retry counters, and upstream request IDs turn one defect into thousands of groups. Normalize volatile values, then hash the exception class plus selected application frames and a schema version. Keep that schema version in the fingerprint input: changing normalization rules without versioning makes old and new groups appear comparable when they are not.

import hashlib
import json
import re


VOLATILE = re.compile(r"\b(?:[0-9a-f]{16,}|\d{6,})\b", re.IGNORECASE)


def group_fingerprint(exception_class: str, frames: list[str], version: int = 1) -> str:
    stable_frames = [VOLATILE.sub("<id>", frame) for frame in frames[:8]]
    material = {
        "version": version,
        "exception_class": exception_class,
        "frames": stable_frames,
    }
    encoded = json.dumps(material, sort_keys=True, separators=(",", ":")).encode()
    return hashlib.sha256(encoded).hexdigest()
Enter fullscreen mode Exit fullscreen mode

This Python example describes the normalization boundary; a Node.js producer can emit the same canonical fields. Cross-service agreement on the input matters more than the language used to compute the digest.

The Twelve-Factor guidance treats logs as event streams and says applications should not manage log-file routing or storage. That is a sound boundary for emission. An error-tracking store, however, is a derived index with lifecycle state, so it should consume the stream rather than become the only copy of application output. If ingestion or grouping fails, the source stream remains available according to its own retention policy.

How can resolution remain safe to roll back?

Never implement resolve as DELETE, and do not represent it only as a mutable Boolean. A mistaken click, an automation bug, or a late recurrence then destroys either evidence or intent. Append a transition with optimistic concurrency: the command supplies the state version the operator saw, and the write succeeds only if that version is still current.

The invariant is small but valuable. A group can move from unresolved to resolved, and a new matching event can either leave it resolved while recording recurrence or append a transition back to unresolved, according to a documented policy. Pick one policy and expose it in the event timeline. Silent reopening is confusing; silent suppression is worse.

from dataclasses import dataclass


@dataclass(frozen=True)
class ResolveCommand:
    group_id: str
    expected_version: int
    actor_id: str
    reason: str


def resolve(repository, command: ResolveCommand):
    group = repository.get_group(command.group_id)
    if group.version != command.expected_version:
        raise RuntimeError("group changed; refresh before resolving")
    if group.state == "resolved":
        return group
    return repository.append_transition(
        group_id=group.id,
        from_state="unresolved",
        to_state="resolved",
        actor_id=command.actor_id,
        reason=command.reason,
        expected_version=group.version,
    )
Enter fullscreen mode Exit fullscreen mode

Reopen is the inverse transition, not a database restore. This also makes deployment rollback safer: an older application version may ignore a newer optional field, but it must not reinterpret or erase prior transitions. Test mixed-version readers before release, and use expand-then-contract schema changes. A destructive column change and a UI release should never share one rollback boundary.

What should a Node.js error tracking admin page list?

The default list should be boring: production environment, unresolved state, newest recurrence first. Each row needs the group title, first and last seen, count, affected delivery channel, latest deployment, and an explicit state version. Search should cover bounded, stable fields such as group ID, exception class, deployment ID, route ID, and normalized message. Searching arbitrary retained payloads expands both the privacy surface and the index bill. Event detail should distinguish observed facts from derived labels, showing timestamps with timezone, the normalized stack, release or deployment identity, retry attempt, and redacted structured context. Link representative events across the retention window rather than presenting only the newest event, because the newest event can conceal a payload shape that appeared once near the beginning. The page must also survive partial failure: if event samples have expired, render the group counters and transition history with an “event detail expired” state; if a group changes after the page loads, reject the stale resolve command and ask for a refresh; if search indexing lags, display the projection watermark so an operator knows which ingestion time is covered.

Evidence expires.

Here is the decision surface I would insist on before choosing storage or index technology:

Concern Safer default Failure mode it contains Limit
Group identity Versioned normalized fingerprint Cardinality explosion from volatile IDs Normalization can merge genuinely distinct defects
Raw evidence Time-distributed capped samples Duplicate events dominate retention Rare payload variants can age out
Resolution Append-only transition with version check Lost updates and irreversible clicks Transition history grows continuously
Search Allowlisted structured fields Payload leakage and unbounded indexing Free-form forensic search is narrower
Deployment Expand, migrate, then contract Rollback reads an incompatible schema Requires a longer migration window

No storage engine removes those trade-offs. It merely changes where they surface.

Retention, testing, and rollout

Set separate retention classes for group summaries, transition history, sampled events, and source logs. The exact durations depend on investigation latency, legal obligations, delivery volume, and the time it takes a defect to recur; a universal number would be false precision. Measure bytes written per class, index bytes per searchable field, sample acceptance rate, groups created per thousand failures, and the age of the oldest unresolved group. Those measurements reveal whether cardinality or raw evidence is controlling the bill.

Test grouping with fixtures that vary shipment IDs, destination identifiers, retry counts, and line numbers while preserving the underlying call path. Then add negative fixtures whose exception class or meaningful application frame differs. The first set should converge; the second should not. Test redaction before persistence, because deleting a sensitive field from the UI does not remove it from snapshots, replicas, or backups.

Roll out a new fingerprint version in shadow mode. Compute both keys, compare split and merge rates, and only then make the new key authoritative. A feature flag can separate computation from read-path activation, but the flag is a control mechanism rather than a substitute for compatible data. Keep the old projection readable through the rollback window, and record which version created every group.

One subtle trap remains: resolving all groups associated with a deployment immediately before rollback can make the rollback look clean while older failures continue. Resolution describes triage state, not software health. Deployment markers and recurrence policy must remain independent.

A practical acceptance rule

Ship the admin page when an operator can find unresolved production notification failures by stable identifiers, open representative redacted evidence, see exactly why and when state changed, reject a stale update, and reverse a resolution without restoring deleted data. Also prove that expired detail degrades into an honest retention message and that an application rollback can still read every record written during the rollout window.

The durable design is the one that preserves decisions longer than bulky evidence while making the loss of evidence visible. It gives logistics operators a useful production queue, keeps storage growth bounded, and treats rollback as a normal state transition rather than an emergency reconstruction exercise.

Further reading

Top comments (0)