DEV Community

JorisRhodes8286
JorisRhodes8286

Posted on

Implementing Node.js Failure Metrics Dashboards and Email Alerts for Fintech SaaS

TL;DR: Emit a small, fixed set of failure counters from the Node.js services that own checkout, webhook, and import work. Put those counters on one dashboard, poll five-minute windows every minute, and send email only after a threshold and cooldown check. For a startup, this is easier to audit and attribute than turning every mixed application log into an alert. Keep a separate heartbeat monitor for jobs that fail by never running.

The bill is mostly a volume-and-retention equation: distinct series times samples per series times retention, plus query and notification work. Start by controlling the first term. A counter named checkout_failed with labels such as service, environment, and a bounded reason is useful; a label containing customer_id, email address, request ID, or exception text creates an unbounded series set and weakens the compliance story at the same time.

For a concrete planning model, suppose three failure counters each have 12 bounded label combinations across two environments. At one sample per minute, 30 days means 3 x 12 x 2 x 43,200 = 3,110,400 samples before replication or backend-specific overhead. This is an estimate for capacity planning, not a vendor price. Changing a 30-day raw retention window to seven days cuts that raw sample term to 725,760; keeping a daily aggregate preserves trend evidence without retaining every point.

What evidence must survive a fintech incident?

Metrics answer when and how often. They do not prove which payment, webhook, or customer was affected. I would retain low-cardinality counters for detection, then preserve a bounded incident record elsewhere with a timestamp, an internal event identifier, the operation, the outcome, and a trace ID when one exists. The metric should never carry OTP values, email addresses, phone numbers, tokens, or payment details as labels.

That split matters during review. A dashboard can show that webhook_failed rose above its normal range, while access-controlled event records support reconstruction. If the evidence store has a retention policy, document it next to the alert rule so an investigator knows what will still exist on day 8 or day 31. Logs with trace_id and span_id can correlate records, but they do not create a distributed trace or span tree by themselves.

Choose counters whose ownership is unambiguous:

  • checkout_failed increments after the checkout operation has definitively failed, not on every internal retry.
  • webhook_failed increments when the delivery policy is exhausted, with a bounded reason such as timeout or rejected.
  • import_failed increments once per failed import job, while a heartbeat monitor separately detects a job that never started.

One count per final outcome prevents retry storms from inflating customer impact. It also makes cost attribution legible: the team that owns the operation owns the counter, its label budget, and the corresponding evidence-retention rule.

How can cheap metrics power a dashboard plus failure alerts?

Poll a short window and keep the evaluator boring. The adapter that queries the metrics backend should return a normalized result containing the counter, window boundaries, and count. Do not guess undocumented query filters. In particular, Infrai exposes a self-describing discovery surface with request schemas and runnable examples, so integration can begin by reading the capability description instead of installing another SDK; however, the filter parameters for its metrics query are not declared and need to be tested before production use. It provides metrics reporting and querying, not built-in threshold evaluation or email, Slack, pager, or webhook routing.

The query adapter below calls the two verified metrics routes. Set INFRAI_API_BASE to the documented API base, and copy INFRAI_REPORT_BODY_JSON plus INFRAI_QUERY_URL from the live discovery example after testing its declared schema. The complete query URL stays in configuration because inventing filters that the discovery parameters do not declare would turn sample code into a trap. METRIC_EVENT_ID must remain stable when retrying the same report.

import json
import os
import time
import urllib.error
import urllib.request


def request(method, url, body=None, idempotency_key=None):
    headers = {
        "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
        "Accept": "application/json",
    }
    data = None
    if body is not None:
        headers["Content-Type"] = "application/json"
        data = json.dumps(body).encode("utf-8")
    if idempotency_key is not None:
        headers["Idempotency-Key"] = idempotency_key

    for attempt in range(5):
        call = urllib.request.Request(url, data=data, headers=headers, method=method)
        try:
            with urllib.request.urlopen(call, timeout=20) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            detail = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 4:
                raise RuntimeError(f"API returned {error.code}: {detail}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)

    raise RuntimeError("Retry loop ended without a response")


base_url = os.environ["INFRAI_API_BASE"].rstrip("/")
report_body = json.loads(os.environ["INFRAI_REPORT_BODY_JSON"])
request(
    "POST",
    base_url + "/v1/metrics/report",
    body=report_body,
    idempotency_key=os.environ["METRIC_EVENT_ID"],
)
result = request("GET", os.environ["INFRAI_QUERY_URL"])
print(json.dumps(result))
Enter fullscreen mode Exit fullscreen mode

Pipe the adapter's normalized output into the following notification boundary for the Node.js system. It accepts query results as JSON lines on standard input, applies thresholds, persists cooldown state atomically, and sends an email through an SMTP relay. Run it once per poll from a scheduler; set the secrets in the environment.

import json
import os
import smtplib
import sys
import tempfile
import time
from email.message import EmailMessage
from pathlib import Path

THRESHOLDS = {
    "checkout_failed": 5,
    "webhook_failed": 10,
    "import_failed": 1,
}
COOLDOWN_SECONDS = 15 * 60
STATE_PATH = Path(os.environ.get("ALERT_STATE_PATH", "/tmp/failure-alert-state.json"))


def load_state():
    try:
        return json.loads(STATE_PATH.read_text(encoding="utf-8"))
    except FileNotFoundError:
        return {}


def save_state(state):
    STATE_PATH.parent.mkdir(parents=True, exist_ok=True)
    with tempfile.NamedTemporaryFile(
        "w", encoding="utf-8", dir=STATE_PATH.parent, delete=False
    ) as handle:
        json.dump(state, handle, sort_keys=True)
        temporary_path = Path(handle.name)
    temporary_path.replace(STATE_PATH)


def send_email(metric, count, window_start, window_end):
    message = EmailMessage()
    message["Subject"] = f"Failure threshold reached: {metric}"
    message["From"] = os.environ["ALERT_FROM"]
    message["To"] = os.environ["ALERT_TO"]
    message.set_content(
        f"{metric} recorded {count} failures between "
        f"{window_start} and {window_end}.\n"
    )

    with smtplib.SMTP_SSL(os.environ["SMTP_HOST"], 465, timeout=15) as smtp:
        smtp.login(os.environ["SMTP_USER"], os.environ["SMTP_PASSWORD"])
        smtp.send_message(message)


def main():
    state = load_state()
    now = int(time.time())

    for line in sys.stdin:
        result = json.loads(line)
        metric = result["metric"]
        threshold = THRESHOLDS.get(metric)
        if threshold is None or int(result["count"]) < threshold:
            continue

        last_sent = int(state.get(metric, 0))
        if now - last_sent < COOLDOWN_SECONDS:
            continue

        send_email(
            metric,
            int(result["count"]),
            result["window_start"],
            result["window_end"],
        )
        state[metric] = now

    save_state(state)


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Here is a local smoke test. It deliberately uses a synthetic five-minute result, not a claimed production baseline.

import json
import subprocess

sample = {
    "metric": "checkout_failed",
    "count": 7,
    "window_start": "2026-09-28T02:00:00Z",
    "window_end": "2026-09-28T02:05:00Z",
}

completed = subprocess.run(
    ["python", "alert_failures.py"],
    input=json.dumps(sample) + "\n",
    text=True,
    check=False,
)
raise SystemExit(completed.returncode)
Enter fullscreen mode Exit fullscreen mode

In production, replace local state with a conditional write in a shared store when multiple pollers can run. The stable deduplication key should include the metric and window end. An overlapping query window is then harmless, and a scheduler retry cannot send the same incident twice. Email acceptance is not delivery, either: monitor relay rejection, authenticate the sending domain, and keep the alert body free of customer data.

Which dashboard and alert stack fits?

There is no universally best stack. The useful comparison is operational ownership and evidence coverage, not the smallest advertised unit price.

Option Best fit in this pattern Boundary to plan for
Prometheus with Alertmanager A team prepared to operate a metrics collector and explicit alert routing Cardinality, storage, retention, and availability remain engineering decisions
Grafana Cloud A hosted dashboard and alerting workflow around Grafana's ecosystem Verify ingestion limits, retention, and contact-point behavior for the selected plan
Datadog A managed suite when metrics need to sit beside broader operational telemetry Tag cardinality and retention choices directly affect attribution and governance
Sentry Failure investigation centered on application errors and releases Treat business-operation counters and silent job checks as separate design questions
Infrai A small team wanting reporting and queries behind one REST key, with discovery schemas and runnable examples Supply the poller and notification routing; test query filtering rather than assuming it
Healthchecks Scheduled work where absence of a ping is the signal It complements failure counters; it is not the dashboard for checkout totals

Prometheus and Alertmanager make the control plane explicit. Grafana Cloud and Datadog move more of that operation to a provider. Sentry starts from captured application failures, which can be valuable context but is not the same signal as a deliberately emitted business counter. Healthchecks closes the silent-failure gap: import_failed cannot increment if the importer never starts.

Infrai's single credential and consistent API conventions can reduce credential inventory and make one bill easier to assign to a platform cost center; its discovery catalog covers 295 routes across 20 modules, and documented capabilities include runnable examples in 10 languages. Those are workflow advantages, not reasons to ignore fit. Its limitations are a concrete trade-off: it is not a fit when the team needs built-in notification routing, declared production query filters, distributed trace trees, Session Replay, or configurable observability retention. Choose Prometheus with Alertmanager when control over collection and routing is the priority, a managed Grafana or Datadog stack when the team wants an integrated hosted path, Sentry for error-centered investigation, and Healthchecks for missed schedules.

The selection test is practical. Can the system show the exact query window used by an alert? Can it prevent duplicate notifications? Can access to incident evidence be separated from access to aggregate metrics? Can one team explain the labels, retention, and notification path to a compliance reviewer without opening five consoles? Trial those paths before migrating historical data.

Retention is a loss decision

My first instinct would be to retain every raw point because storage feels like cheap insurance. The series equation changes that decision. A more defensible starting policy for an early fintech service is seven days of raw one-minute failure samples, 90 days of daily aggregates, and incident records retained according to the company's legal and security policy. Those numbers are a proposed operating policy, not product defaults. Adjust them from the longest incident-discovery delay the business must support and the evidence it is permitted to retain.

Then stop keeping high-cardinality metric labels and old raw points. This lowers series growth and reduces the personal data that can accidentally leak into a monitoring system. The cost is real: after raw expiry, an investigator can see the daily failure total but cannot reconstruct the exact minute-by-minute shape from metrics alone. If that detail is mandatory, retain it deliberately and include it in the cost and access review.

There are other limits. A metrics dashboard does not provide source-map decoding, crash symbolication, Electron minidump parsing, Session Replay, or a trace tree. It also does not establish a user-deletion workflow for logs. Choose separate tools or evidence stores when those requirements exist, and test deletion and export paths before regulated data enters them.

Keep the first version narrow. Three counters, one dashboard, one five-minute query, one deduplicated email path, and one heartbeat check are enough to expose the real requirements. Add labels only after a concrete incident question proves their value.

Further reading

Top comments (0)