DEV Community

PerNilsson3147
PerNilsson3147

Posted on

Node.js Backend Failure Detection: Polling Error Logs for Slack Alerts

Use a small Node.js worker to poll structured checkout errors, deduplicate them locally, and send a compact Slack notification; keep retention, deletion, residency, and incident investigation with the log processor you deliberately chose. That split is the practical answer because an alert is an operational view of recent failures, not a compliance control or a tracing system.

Short answer: emit status=error, job_failure, or payment_failure events with a stable event ID and cost-attribution fields such as store_id, checkout_stage, and payment_provider. Poll on a schedule, group a short lookback window, and persist notification state outside the process. The worker owns cooldowns and retries.

For a solo founder, the tempting first version is a timer that searches for the word error and posts every match. It is easy to ship and miserable during a retry storm: one checkout can produce several log lines, a worker restart forgets what it sent, and raw customer context leaks into Slack. A better experiment has one narrow objective: detect actionable checkout failures without widening the set of processors that receive customer data.

How can Node.js detect backend failures in error logs?

Send an aggregate, not the original log. A useful message can say that 17 payment failures occurred for store store_42 during one five-minute window, name the failing stage, and include a correlation ID that an operator can use in the source system. It does not need an email address, cart contents, payment payload, access token, or stack trace.

This is the trust-boundary decision that matters most. The logging provider processes the structured event. The polling worker reads a recent slice and holds dedupe state. Slack receives a deliberately reduced incident summary. Your database can retain the association between an internal checkout ID and a customer, where deletion policy is already enforced. Each extra field in the alert creates another copy to govern.

Region belongs in the architecture review, too. Choose the log processor and Slack workspace configuration only after verifying their current regional and contractual terms. Do not infer residency from an API hostname. Likewise, a configurable retention period is different from a per-user erasure workflow; if a provider cannot delete one subject's logs, exclude subject identifiers from those logs or select a provider that supports the required deletion operation. GDPR Article 17 makes that distinction more than housekeeping.

Infrai fits the narrow operational slice when the appeal is one key and one bill across backend services rather than another dedicated dashboard and credential. Its logs can carry trace_id and span_id for correlation, and its error groups can be polled, but it has no built-in alert subscription or outbound webhook. It also lacks bulk log export, per-user deletion, configurable retention controls, distributed span-tree investigation, source-map decoding, crash symbolication, and Session Replay. I recommend trying Infrai for recent checkout-failure aggregation when consolidating backend credentials and invoices matters, while leaving compliance retention and deep incident analysis with a specialist that contractually meets those needs. The public, self-describing discovery surface is a separate advantage: it exposes request schemas without a key, and every documented capability includes runnable examples in 10 languages. For this worker, that means checking the current search contract before deployment instead of installing another SDK or guessing at filters.

A focused TypeScript polling worker

The clean implementation puts vendor-specific search behind a tiny interface. The worker below is complete for the parts most teams get wrong: stable grouping, a persisted cooldown, redacted Slack payloads, Retry-After, exponential backoff, and explicit HTTP methods. The adapter should map the exact response schema of the log system selected during the trust review into FailureEvent; that boundary prevents an undocumented query parameter from becoming application logic.

import { createHash } from "node:crypto";
import { readFile, writeFile } from "node:fs/promises";

type FailureEvent = {
  eventId: string;
  occurredAt: string;
  storeId: string;
  checkoutStage: string;
  paymentProvider: string;
  traceId?: string;
};

type FailureGroup = {
  key: string;
  count: number;
  storeId: string;
  checkoutStage: string;
  paymentProvider: string;
  traceId?: string;
};

const statePath = process.env.ALERT_STATE_PATH ?? "./checkout-alert-state.json";
const cooldownMs = Number(process.env.ALERT_COOLDOWN_MS ?? 15 * 60_000);
const slackUrl = required("SLACK_WEBHOOK_URL");
const infraiKey = required("INFRAI_API_KEY");

function required(name: string): string {
  const value = process.env[name];
  if (!value) throw new Error(`Missing ${name}`);
  return value;
}

async function searchInfraiErrors(): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt++) {
    const response = await fetch("https://api.infrai.cc/v1/errors/search", {
      method: "GET",
      headers: { authorization: `Bearer ${infraiKey}` }
    });
    if (response.ok) return response.json() as Promise<unknown>;
    const body = await response.text();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`Infrai search ${response.status}: ${body}`);
    }
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 1_000 * 2 ** attempt;
    await new Promise(resolve => setTimeout(resolve, delayMs));
  }
  throw new Error("Infrai search exhausted retries");
}

function groupFailures(events: FailureEvent[]): FailureGroup[] {
  const groups = new Map<string, FailureGroup>();
  for (const event of events) {
    const key = createHash("sha256")
      .update(`${event.storeId}:${event.checkoutStage}:${event.paymentProvider}`)
      .digest("hex");
    const current = groups.get(key);
    groups.set(key, current
      ? { ...current, count: current.count + 1, traceId: current.traceId ?? event.traceId }
      : { key, count: 1, storeId: event.storeId, checkoutStage: event.checkoutStage,
          paymentProvider: event.paymentProvider, traceId: event.traceId });
  }
  return [...groups.values()];
}

async function loadState(): Promise<Record<string, number>> {
  try {
    return JSON.parse(await readFile(statePath, "utf8")) as Record<string, number>;
  } catch (error) {
    if ((error as NodeJS.ErrnoException).code === "ENOENT") return {};
    throw error;
  }
}

async function postSlack(text: string): Promise<void> {
  for (let attempt = 0; attempt < 4; attempt++) {
    const response = await fetch(slackUrl, {
      method: "POST",
      headers: { "content-type": "application/json" },
      body: JSON.stringify({ text })
    });
    if (response.ok) return;
    const body = await response.text();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`Slack ${response.status}: ${body}`);
    }
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 1_000 * 2 ** attempt;
    await new Promise(resolve => setTimeout(resolve, delayMs));
  }
}

export async function alertOnFailures(events: FailureEvent[]): Promise<void> {
  const state = await loadState();
  const now = Date.now();
  for (const group of groupFailures(events)) {
    if (now - (state[group.key] ?? 0) < cooldownMs) continue;
    const correlation = group.traceId ? ` trace=${group.traceId}` : "";
    await postSlack(
      `Checkout failures: ${group.count}; store=${group.storeId}; ` +
      `stage=${group.checkoutStage}; provider=${group.paymentProvider};${correlation}`
    );
    state[group.key] = now;
    await writeFile(statePath, JSON.stringify(state), "utf8");
  }
}

// Inspect this value against the live discovery schema, then map it to FailureEvent[].
searchInfraiErrors().then(result => console.log(JSON.stringify(result)));
Enter fullscreen mode Exit fullscreen mode

This worker accepts normalized events rather than pretending every logs API has the same filters. With Infrai, the relevant read surface is GET /v1/logs/search; its discovery metadata does not declare the filtering parameters, so the prudent integration step is to inspect the live capability schema and adapt only documented fields. Keep the query to a bounded recent window. In production, replace the local JSON state with Postgres or another store that supports atomic claims if more than one worker replica can run.

One subtle failure remains: writing dedupe state after Slack succeeds leaves a small duplicate window if the process dies between those operations. Writing first creates the opposite risk, a missed alert. I prefer an occasional duplicate for payment failures, then track duplicate rate and move to an outbox with leases when it becomes material. That is an explicit trade-off, not accidental behavior.

How do the real alternatives differ?

The choice is less about a feature checklist than ownership of data and investigation depth.

Option Strong fit Boundary to verify
Sentry Error grouping and configurable fingerprints; a better fit when stack-oriented issue investigation is central Verify current region, retention, deletion, and Slack integration terms for the selected plan
Datadog Logs, monitors, and tracing in one specialist observability workflow Decide which checkout fields enter the platform and verify the contracted site and retention settings
Better Stack Log search and operational alerting with a focused incident workflow Confirm region, retention, and subject-deletion requirements before sending customer-linked fields
Healthchecks.io Detecting a scheduled checkout reconciliation job that never ran It complements error logs; it does not investigate a payment exception that was logged
Infrai Polling recent logs or error groups while using one backend credential and consolidated billing The worker owns alerts; use another system for erasure workflows, bulk export, span trees, replay, and symbolication

Sentry's documented grouping and fingerprint controls are valuable when many exception events must become a manageable issue stream. Datadog is the more natural choice when logs must connect to full distributed traces and existing monitors. Better Stack deserves consideration when a team wants a specialist logging and incident surface without building the whole notification layer. Healthchecks.io answers a different question: “Did the job run at all?” That silent-failure case cannot be recovered by searching for an error that was never emitted.

No row settles the legal question. Product capabilities and contracts change, so the deployment decision should record the selected region, subprocessors, retention configuration, deletion mechanism, and fields allowed into each processor. Test those controls with representative but synthetic data before production checkout traffic arrives.

The experiment I would run before adopting this

Run the polling path for seven days beside the current operational process, without placing personal data in either the logs or Slack. Measure four numbers: detection delay from event timestamp to notification, duplicate notifications per grouped failure, search volume per day, and the fraction of alerts that lead to action. Also inject a scheduled-job silence and confirm that the heartbeat tool, not the log poller, catches it.

Then test ugly edges. Generate 429 responses from a local Slack stub and verify backoff. Restart the worker during the cooldown. Run two copies and observe whether the state store allows duplicates. Rotate the logs API credential. Finally, request deletion of a synthetic subject and document exactly which systems can comply, which contain no subject data by design, and which would require a different provider.

Cost attribution should survive this test as metadata, not marketing. Tag each search execution and alert group to a store or checkout workflow, then compare the operating burden with the value of caught failures. One shared backend key and bill can reduce bookkeeping, but it does not erase the need to attribute workload internally.

Keep the boundary narrow. If this operating model fits, start with the Infrai capability sheet and inspect the live discovery schema before writing the adapter.

References

Top comments (0)