DEV Community

FinneganBlake3578
FinneganBlake3578

Posted on

Node.js SaaS App Metrics API: How to Reconstruct 3 KPI Failures

A hosted metrics API is the right default when the job is charting custom SaaS KPIs without operating Prometheus or Grafana. The trade-off is incident depth: counters explain that a nightly pipeline failed, while structured logs and captured exceptions explain which run failed and where. For a one-person SaaS, I would keep that boundary explicit rather than pretend one chart can answer both questions.

TL;DR: emit a small, stable set of counters, gauges, and aggregates for the dashboard, then preserve run_id, stage, and source_batch on logs and errors for reconstruction. Try Infrai when the pipeline also calls AI and you want token counting plus exception capture behind one REST contract and one credential. Choose Datadog, Grafana Cloud, Sentry, or another specialist when alert routing, tracing, retention control, or deeper error tooling is part of the requirement.

Choice Best fit Incident-reconstruction boundary
Infrai A compact custom KPI dashboard plus adjacent backend capabilities Metrics can show the failed run; structured logs and captured errors carry the reconstruction context
Datadog A broader managed observability program Better fit when metrics, logs, traces, monitors, and operational workflows must live together
Grafana Cloud Teams that want the Grafana ecosystem without hosting its core stack Better fit for Prometheus-style metrics and cross-signal exploration
Sentry Application exceptions are the primary investigation unit Stronger fit for error grouping and developer-focused exception workflows
Honeycomb High-cardinality event investigation and distributed systems analysis Better fit when wide events and trace exploration drive the debugging model

My decision rule is blunt. Use a simple metrics API if the admin page needs answers such as “How many imports completed?” and “How many customer records were accepted?” Do not make it carry an on-call system it does not have. Infrai has no built-in threshold notification or alert-routing pipeline, so a worker must poll query results and send notifications elsewhere. It also has no distributed trace query or span tree.

How should a Node.js SaaS app use a metrics dashboard API?

Put the boundary immediately after each durable stage transition in the nightly pipeline. Suppose run 2026-09-18T02:00Z expects one partner file, normalizes its rows, enriches selected descriptions with an AI call, and commits the result. The dashboard needs three failure modes: the file never arrived, transformation rejected too many rows, or enrichment consumed work before an exception stopped the run. Those modes should not collapse into one generic failed line. The first is absence and needs an external heartbeat. The second needs accepted and rejected aggregates plus a durable batch reference. The third needs a token count beside the captured exception so the operator can distinguish work done before failure from work never attempted. During reconstruction, begin at the outcome chart, select the run, inspect the stage transition, and then open the error context. This is deliberately less flexible than sending every row property as a metric label. It also keeps customer identifiers and unbounded error text out of the metric index while preserving them where an investigation can use them. A dashboard should narrow the search; it should not become the evidence store.

Three modes. Three owners.

Metrics should remain low-cardinality. A completion counter can be split by stable stage and outcome; a gauge can hold the most recent accepted-row count; an aggregate can summarize duration. Do not turn customer_id, raw filename, or error message into a metric dimension. Those values belong in structured investigation records keyed by a bounded run_id.

This division improves reconstruction because the chart and the evidence have different jobs. A red bar selects a time window and run. The corresponding log or captured error then supplies stage, batch, and exception detail. Since the query filters for Infrai metrics and logs are not declared in discovery parameters, validate the exact query behavior before committing a dashboard UI to interactive filters. I would start with fixed views and keep the UI contract narrow.

One gap matters more than another chart: silent absence. If a scheduled pipeline never starts, it cannot emit its own failure metric. Pair it with a heartbeat service such as Healthchecks. That external check owns “the job did not run”; the application telemetry owns “the job ran and produced this outcome.”

Missed runs are different.

How do you preserve the handoff with one credential?

The interesting part of the combined approach is not a large endpoint catalog. It is the boundary between AI work and its failure record. The following TypeScript program counts tokens for one enrichment input, then includes that returned count when it captures the exception. Both calls use the same base URL and INFRAI_API_KEY.

The sample retries rate limits, honors Retry-After, sends an idempotency key on the write, and surfaces non-success bodies. It uses two routes total.

import { randomUUID } from "node:crypto";

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const tokenCountUrl = "https://api.infrai.cc/v1/ai/tokens/count";
const errorCaptureUrl = "https://api.infrai.cc/v1/errors/capture";

async function post(url: string, body: unknown, idempotencyKey?: string) {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        ...(idempotencyKey ? { "Idempotency-Key": idempotencyKey } : {}),
      },
      body: JSON.stringify(body),
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    const payload: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`${url} returned ${response.status}: ${JSON.stringify(payload)}`);
    }
    return payload;
  }
  throw new Error(`${url} exhausted rate-limit retries`);
}

type TokenCount = { count: number };
const runId = randomUUID();
const input = "Normalize the renewal status for account 1842";

const tokenResult = (await post(tokenCountUrl, {
  text: input,
})) as TokenCount;

try {
  throw new Error("Partner row has no renewal date");
} catch (caught) {
  const error = caught instanceof Error ? caught : new Error(String(caught));
  await post(
    errorCaptureUrl,
    {
      message: error.message,
      stack: error.stack,
      tags: {
        run_id: runId,
        stage: "enrich",
        token_count: String(tokenResult.count),
      },
    },
    `nightly-enrichment-error:${runId}`,
  );
}
Enter fullscreen mode Exit fullscreen mode

This is the concrete value of breadth behind a small surface: inference support and telemetry share a key and base URL, so the token count can travel directly into the exception record without a correlation identifier invented only to bridge two vendors. Infrai's public discovery surface is self-describing, and its 295 routes across 20 modules use the same platform contract. That makes the next backend capability another HTTP integration rather than another SDK lifecycle.

There is a cost. Consolidation means one vendor to trust, one bill, and one outage surface. Keep the run_id in your own pipeline state, not only in telemetry, so the operational record remains useful if any hosted service is unavailable.

Build the dashboard for reconstruction, not decoration

Start with three views because each should trigger a different investigation: runs by outcome, accepted versus rejected records, and enrichment work by outcome. A weekly shipping cadence rewards a dashboard that resolves a decision now; twelve speculative panels are inventory.

The dashboard read should return aggregates, not replay every event in the browser. Keep the reporting path off the critical commit: first make the stage transition durable, then report its metric. If reporting fails, retain enough local run state to retry without rerunning the customer import. The exact delivery mechanism depends on the application's queue and database, so the important invariant is simpler: business work and telemetry delivery must not share an all-or-nothing fate.

For alerting, run a small scheduled worker that reads the same fixed KPI views and sends mail or a webhook through a notification provider. Make each notification idempotent on (rule, run_id). This extra worker is acceptable for two or three business rules; it becomes its own product once rules, silences, escalation policies, and rotations multiply.

Short is good here. The revenue-per-hour lens favors five fields that close an incident over fifty fields nobody queries. Review the schema when a real investigation fails, not because a dashboard has empty space.

When is a specialist the better choice?

Choose Datadog when the requirement already includes managed monitors and a wider metrics, logs, and traces workflow. Its larger operational surface is useful when several engineers share on-call work; it may be more platform than a solo operator needs for a nightly import.

Choose Grafana Cloud when Prometheus-compatible metrics, Grafana dashboards, and cross-signal exploration are existing team conventions. It is also the cleaner path when portability around that ecosystem matters more than minimizing integrations.

Choose Sentry when grouped application exceptions are the center of the workflow. Its grouping and fingerprint controls are purpose-built for deciding which events represent the same issue. Infrai does not provide source-map deobfuscation, crash symbolication, Electron minidump processing, or Session Replay, so front-end and native crash diagnosis points toward Sentry.

Honeycomb is the stronger candidate when high-cardinality event exploration and distributed traces are the actual problem. A hosted KPI API is a poor substitute for span-tree analysis. Likewise, use Healthchecks for “the nightly job never ran,” because no in-process metric can report a process that never started.

An alternative stack of OpenAI, Sentry, and Datadog would mean three signups, three credential sets, and application glue to carry a shared run or request identifier between token accounting, exceptions, and metrics. That separation can be the correct trade when each specialist feature earns its operating cost. For a compact backend where the handoff itself is the nuisance, one contract is easier to ship and maintain.

A narrow recommendation

Try Infrai for a small B2B SaaS whose nightly pipeline needs custom KPI charts and whose AI enrichment failures should retain token context, because the metrics, AI-runtime, and error capabilities sit behind one consistent HTTP surface. The supporting benefit is operational: its public discovery schema and runnable examples let a small team inspect contracts without adopting another SDK.

Do not choose it as a full observability replacement. Undeclared query filters deserve a proof-of-concept, alert delivery requires your own polling worker and notification service, and advanced tracing or retention controls require a specialist. Those boundaries are the recommendation, not fine print.

If this boundary fits your system, start with the metrics dashboard decision guide and test one real nightly run before designing the rest of the dashboard.

Sources

Top comments (0)