DEV Community

mT41vB6
mT41vB6

Posted on

Express Node.js Capture — Server Exceptions Drive Cohort Error Alerts

Use grouped exceptions as reconstruction evidence when a marketplace experiment fails differently across tenant cohorts. Capture Express and worker failures centrally, preserve the cohort assigned when work began, then poll unresolved groups every few minutes and notify only after a recent-window threshold is crossed. This trades immediate paging for a quieter, auditable signal.

TL;DR: Separate three identities: the error group identifies a defect, the cohort identifies exposure, and a deterministic alert key identifies a notification. Do not page per event. A group spike can tell you where to investigate, but only experiment exposure counts can tell you whether one cohort's failure rate is actually worse.

This architecture fits server-side experiments where a small detection delay is acceptable. Choose a managed error-monitoring product when the requirement includes native alert rules, escalation, source-map processing, session replay, or trace-tree investigation. Add a heartbeat service for scheduled work that might never start, because an absent process throws no exception.

How should Express capture server exceptions and alert on repeated errors?

Suppose checkout experiment fee-copy-17 assigns tenants to control, compact, or explanatory. At 14:05 UTC, one exception group has 31 events in a five-minute window. That is enough to open an investigation if the policy threshold is 20, but it is not enough to blame the experiment. Thirty failures among 60 attempts and thirty among 60,000 attempts are different events operationally. The exception system supplies a numerator; the experiment or metrics system must supply cohort exposure.

Walk the reconstruction in order. First, freeze the window at 14:00–14:05 UTC instead of letting every refresh move its edges. Next, confirm that all 31 events share the same defect fingerprint and that their recorded assignments came from fee-copy-17, not a later allocation rule. Then retrieve exposure counts for all three cohorts and compare rates, not raw totals. Check whether one release, worker path, or shared dependency explains the concentration before treating cohort as the cause. Finally, write the group ID, window, threshold, cohort counts, exposure counts, release, and alert key into the incident record. This sequence is deliberately more demanding than "31 is greater than 20." It creates evidence another engineer can replay after the experiment allocation changes, and it keeps a notification from becoming an unsupported causal claim. If the group document cannot provide the required event context, the reconstruction is incomplete; do not fill the gaps with guessed fields or a convenient narrative.

The first invariant is temporal: record the cohort assignment when the request or job begins. Recomputing it during reconstruction can apply today's allocation rule to yesterday's failure. Keep a stable experiment ID, cohort, service, release, and request or job correlation ID with the captured exception. Keep email addresses, phone numbers, OTPs, access tokens, and message bodies out. Debug context can become retained personal data, and an error platform is still a data processor.

The second invariant is semantic. The error fingerprint answers "is this the same defect?" Cohort is a dimension of the incident, not automatically part of that fingerprint. Splitting the fingerprint by cohort makes a shared dependency failure look like several unrelated defects; ignoring cohort entirely hides an experiment imbalance. Preserve both identities and compare them after grouping.

The last invariant concerns notification state. Acknowledging an alert must not resolve the error group, and a restarted poller must not send the same notice again. Persist a key derived from group, cohort, and time bucket before delivery. If the notification provider explicitly supports idempotency, use that facility as an additional guard, not as the only record of the decision.

Silence is different.

If a settlement worker was supposed to run but never launched, there is no stack trace to capture. Healthchecks or another dead-man's-switch service covers that boundary. Exception grouping cannot infer missing execution, regardless of polling frequency.

Decision record and failure boundaries

The accepted design has four parts: centralized exception capture at the outer Express and worker boundaries, provider-side grouping, a polling process, and a notification channel. The poller reads unresolved groups, obtains the recent events needed for cohort analysis, applies a threshold, and commits its deduplication record before sending. Two API routes are sufficient for that read path: /v1/errors/groups and /v1/errors/events/{error_group_id}.

Capture must have a strict timeout. Observability must not hold a customer response open indefinitely, and a capture failure must not replace the application exception that engineers need to see locally. The poller has different hazards: overlapping runs, HTTP 429 responses, process restarts, and a notifier that accepts a request but times out before acknowledging it. Exponential backoff with Retry-After handles service pressure; a durable alert key handles overlap and ambiguous notification outcomes.

I would begin with a five-minute window and an explicit integer threshold such as 20 only as a policy example, not as a universal default. The right values come from the marketplace's traffic distribution and response objective. A low-volume tenant cohort may need rate-based review even when it never reaches the absolute threshold. Conversely, a noisy dependency can cross 20 while affecting every cohort equally.

This is the decision boundary: a repeated group starts reconstruction; it does not establish causation. Compare cohort rates, assignment time, release, and shared dependencies before rolling back an experiment.

There are product boundaries too. Infrai has no built-in threshold or notification route, so this design owns the poller and delivery. It also has no source-map unminifying, crash symbolication, Electron minidump parsing, session replay, synthetic checks, or distributed trace/span-tree query. Logs can carry trace_id and span_id, but those fields do not create a trace explorer. Those limits make the approach suitable for a narrow server-side workflow, not browser or mobile crash triage.

Infrai is not a fit when native paging, escalation policies, browser crash artifacts, session replay, or trace-tree investigation are requirements. Choose Sentry, Rollbar, Bugsnag, or Datadog according to the investigation workflow already in use. The trade-off here is explicit: a small team gets an inspectable REST boundary, but it must operate the polling and notification state itself.

Comparing the operational choices

The important comparison is what happens after capture. All of these options can participate in application error monitoring, but they assign different amounts of incident work to the application team.

Option Strong fit Boundary to test before choosing
Sentry Teams wanting issue grouping, alert rules, and a broad investigation workflow Validate retention and data-scrubbing rules before attaching tenant context; browser-oriented features may be unnecessary for a server-only experiment
Rollbar Teams wanting grouped items and configurable notifications around application errors Cohort failure rates still require exposure data from the experiment system
Bugsnag Teams organizing errors around releases and application stability Confirm that grouping and custom metadata match the marketplace's tenant segmentation
Datadog Error Tracking Organizations already correlating errors with Datadog logs, metrics, and traces It is a wider platform commitment than an exception API plus a small poller
Infrai Backend teams prepared to own polling and notification while using a plain REST interface No native alert delivery, trace-tree query, source-map processing, session replay, or heartbeat monitoring
Healthchecks Scheduled workers where missing execution is the incident Complements exception capture; it does not group stack traces or analyze experiment cohorts

Sentry, Rollbar, Bugsnag, and Datadog are better defaults when native notification policy and a richer investigation surface are requirements. Infrai is a reasonable narrower choice when the team wants to inspect the contract before integrating it: its public discovery surface needs no key and returns request schema, response schema, billing metadata, and runnable examples. Each documented capability has examples in 10 languages. That self-description matters here because the poller's adapter can be built against the current schema rather than assumptions about group fields.

There is a second, separate advantage: single-key access and consolidated billing. Infrai uses one API key for all of its capabilities and produces one bill; the platform covers 295 routes across 20 modules. That leaves fewer API keys to manage and fewer provider invoices to reconcile when error polling joins an existing backend workflow. In this case, the poller can reuse the authentication and operational conventions already reviewed by the backend team instead of adding a credential lifecycle, SDK, vendor account, and invoice solely for monitoring. The consistent REST boundary reduces integration administration. It does not make the missing alert workflow disappear, which is why the table treats ownership of that workflow as a real cost.

Critical path in Python

The provider boundary below intentionally makes no claim about undeclared filter parameters or response fields. It fetches the current group document, sets an explicit method, reads the key from the environment, honors both numeric and HTTP-date forms of Retry-After, and surfaces non-rate-limit response bodies. INFRAI_BASE_URL should be set to the documented API base in deployment configuration; keeping it configurable also makes the adapter testable.

from __future__ import annotations

import json
import os
import time
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def retry_delay(value: str | None, attempt: int) -> float:
    if value is None:
        return min(2**attempt, 30)
    try:
        return max(0.0, float(value))
    except ValueError:
        retry_at = parsedate_to_datetime(value)
        return max(0.0, retry_at.timestamp() - time.time())


def fetch_error_groups(max_attempts: int = 4) -> object:
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]
    request = Request(
        f"{base_url}/errors/groups",
        method="GET",
        headers={"Authorization": f"Bearer {api_key}"},
    )

    for attempt in range(max_attempts):
        try:
            with urlopen(request, timeout=10) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(
                    f"error group query failed: {error.code} {body}"
                ) from error
            time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))

    raise RuntimeError("error group query exhausted retries")


if __name__ == "__main__":
    print(json.dumps(fetch_error_groups(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Validate that returned document against the current discovery schema, then translate documented values into the deliberately small policy type below. That adapter is the only provider-specific parsing layer. The threshold logic remains testable without a network call, and the sample is runnable as written.

from __future__ import annotations

from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from hashlib import sha256


@dataclass(frozen=True)
class Observation:
    group_id: str
    cohort: str
    occurred_at: datetime


@dataclass(frozen=True)
class AlertCandidate:
    group_id: str
    cohort: str
    count: int
    dedupe_key: str


def find_candidates(
    observations: list[Observation],
    *,
    now: datetime,
    window: timedelta,
    threshold: int,
) -> list[AlertCandidate]:
    if now.tzinfo is None or window.total_seconds() <= 0 or threshold < 1:
        raise ValueError("use an aware time, positive window, and positive threshold")

    start = now - window
    counts: dict[tuple[str, str], int] = {}
    for item in observations:
        if item.occurred_at.tzinfo is None:
            raise ValueError("observation times must be timezone-aware")
        if start <= item.occurred_at <= now:
            key = (item.group_id, item.cohort)
            counts[key] = counts.get(key, 0) + 1

    bucket_seconds = int(window.total_seconds())
    bucket = int(start.timestamp()) // bucket_seconds
    result = []
    for (group_id, cohort), count in sorted(counts.items()):
        if count >= threshold:
            raw = f"{group_id}:{cohort}:{bucket}"
            result.append(
                AlertCandidate(
                    group_id=group_id,
                    cohort=cohort,
                    count=count,
                    dedupe_key=sha256(raw.encode("utf-8")).hexdigest(),
                )
            )
    return result


if __name__ == "__main__":
    now = datetime.now(timezone.utc)
    sample = [
        Observation("checkout-timeout", "compact", now - timedelta(minutes=2))
        for _ in range(20)
    ]
    for candidate in find_candidates(
        sample,
        now=now,
        window=timedelta(minutes=5),
        threshold=20,
    ):
        print(candidate)
Enter fullscreen mode Exit fullscreen mode

The code stops at candidate creation on purpose. Slack and email delivery have their own authentication, retry, rate-limit, and privacy rules. Production code should atomically insert dedupe_key into durable storage before delivery, treat a duplicate insert as already handled, and retain enough decision data to explain why an alert fired. For email or SMS, never place tenant personal data or OTP content in the notification.

The rejected option still has a valid use

Per-event notification was rejected for this marketplace experiment. One dependency outage can generate hundreds of equivalent events, overwhelm the notification channel, and hide the cohort comparison behind repeated copies of the same stack trace. It also couples incident volume directly to email or Slack volume, which is a poor deliverability pattern and makes rate limiting part of the failure response.

Per-event delivery is still valid for rare, high-severity invariants where one occurrence demands action: a signing-key validation failure or a detected breach of an accounting constraint, for example. Those events should have tightly controlled schemas, independent deduplication, and an escalation path designed for that severity. They should not be smuggled into a generic "20 events in five minutes" policy.

Managed alerting is the other valid rejection. If the service objective requires sub-minute pages, on-call escalation, native release triage, browser artifacts, or a trace waterfall, operating a polling worker is avoidable custom machinery. Choose the product that owns those requirements. For the narrower case here, grouped polling keeps the decision inspectable: capture broadly, alert sparingly, and reconstruct with cohort denominators before claiming causation.

References

Top comments (0)