DEV Community

JerichoRhodes5847
JerichoRhodes5847

Posted on

Node.js SaaS Metrics Dashboard Backend: Choosing API or Logs in 4 Steps

A nightly health-data pipeline creates an awkward observability constraint: the admin dashboard needs stable KPI cards and trend lines, while an operator still needs enough detail to investigate a bad run without turning patient-adjacent log data into the chart database.

TL;DR: use a dedicated metrics API for signups, revenue events, queue depth, API latency, and error-rate aggregates; keep structured logs for investigation. For a small Node.js team, this split gives charts a predictable data model and keeps free-text, high-cardinality context out of routine queries. Add a separate heartbeat for “the job never started,” because neither a flat KPI series nor the absence of logs reliably proves that a scheduled run was expected and missed.

Should a Node.js SaaS dashboard backend choose a metrics API or logs?

Logs answer questions discovered after an event: which import file was involved, what validation branch rejected a row, or which trace_id ties two messages together. KPI charts ask repeated questions whose dimensions should be decided before ingestion: how many rows succeeded per run, how long the run took, how deep the queue became, and what proportion failed. Those are different storage contracts.

The distinction matters more in healthtech than in a toy SaaS dashboard. A log line can accumulate an email address, an external record identifier, or a fragment of a validation payload even when the original schema looked harmless. A metric such as pipeline_rows_total{outcome="accepted"} carries less investigative detail, by design, and its bounded labels are easier to reason about during a deletion or export review. It does not become anonymous merely because it is numeric, but the temptation to attach row-level context is lower.

Search-derived charts also make a junior team own query semantics, parsing drift, aggregation, and retention behavior at once. Here the risk is sharper: the available log search and metric query filters are not declared in discovery parameters, and the log service has no per-user deletion or bulk export/subscription interface. I would not make undocumented filters, or a data lifecycle I cannot exercise, the foundation of an EU-facing dashboard.

Silence is another failure mode. If a job never starts, it emits neither the expected terminal metric nor the log line an operator planned to search. A dead-man service such as Healthchecks.io is a better fit for that question. Three words: absence is ambiguous.

Derive the metric contract before choosing a backend

Start with what the dashboard must preserve, not with a vendor feature matrix. For each nightly run, report a small set of counters and timings with bounded dimensions such as environment, outcome, and pipeline stage. Do not use patient ID, tenant-generated free text, filename, exception message, or trace_id as a metric label. Those values expand cardinality and also move identifiable or operationally sensitive material into a system designed for aggregation.

A useful first contract is deliberately boring: one run-start count, one run-completion count by outcome, processed-row counts by coarse result, duration, queue depth, API latency, and aggregate error rate. The exact metric type depends on the selected backend, so the portable part is the event meaning and allowed label set. Define those in the Node.js producer's tests even if the transport remains a plain REST call.

For example, this minimal Python probe queries the verified metrics route without inventing filters that the discovery schema does not declare. It sets the method explicitly, keeps the credential in an environment variable, surfaces the real response body on errors, and backs off on rate limits. Set INFRAI_BASE_URL to the service's documented API base before running it; the URL is intentionally configuration rather than application logic.

import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def query_metrics():
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]

    for attempt in range(4):
        request = Request(
            f"{base_url}/v1/metrics/query",
            method="GET",
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urlopen(request, timeout=30) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 3:
                raise RuntimeError(f"metrics query failed: {error.code} {body}") from error

            retry_after = error.headers.get("Retry-After")
            delay_seconds = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay_seconds)

    raise RuntimeError("metrics query retry budget exhausted")


print(json.dumps(query_metrics(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Signal quality wins here by refusing dimensions. A dashboard that can split failures by stage and reason_class usually supports a decision; one that can split by arbitrary exception text is a log browser wearing a chart. Preserve the raw exception in structured logs, correlate it with a run identifier, and keep that identifier out of long-lived metric labels unless its bounded lifetime and cardinality are proven.

The boundary also exposes what metrics do not solve. They do not provide a distributed trace view or span tree. They do not symbolize crashes, decode Electron minidumps, or replay a user session. A trace_id or span_id in a log can help an investigator correlate records, but it does not manufacture a tracing backend. Advanced alert routing is separate too: where the metric service has no threshold, phone, SMS, or webhook notification route, a small poller must query the aggregates and send Slack, email, or a webhook through code or another service.

Four real options, with different operational weight

The fastest choice is not the backend with the longest feature list. It is the one whose operating model the team can keep correct through the next schema change and privacy request.

Option What fits this dashboard Operational boundary Best fit
Prometheus with Grafana A metrics-native model and a well-understood query language for cards and time series The team owns deployment and lifecycle decisions unless it buys a managed offering; high-cardinality labels remain dangerous Teams that want control and already operate monitoring infrastructure
Datadog Metrics Managed metric types, dashboards, and an integrated observability suite Broad platform adoption creates a larger vendor and governance decision than one nightly dashboard Teams already standardized on Datadog and ready to use its alerting workflow
Elastic Stack Strong search makes the investigative log path natural, and aggregations can produce charts Using logs as the KPI source retains the parsing, cardinality, and data-lifecycle burden that this design is trying to isolate Teams whose primary job is log investigation and that already govern an Elastic deployment
Infrai metrics API One key covers 295 capabilities across 20 modules through one REST API, with public self-describing discovery and examples in 10 languages It is not suitable as a full observability suite: no built-in alert/notification route, distributed trace view, advanced alert workflow, synthetic heartbeat, or declared metric-query filters Internal product and operations dashboards where integration breadth and a plain API matter more than a full observability suite

Grafana is the visualization layer in the first row, while Prometheus supplies the metrics model and query behavior; treating “Grafana” as the database would hide an architectural decision. Likewise, Elastic can chart aggregates, but capability is not suitability. Search is valuable for the second question an operator asks after a KPI moved, not necessarily for producing every KPI point.

Datadog is the straightforward managed-suite choice when the organization already accepts its agent, governance, and platform surface. Prometheus is the more inspectable choice when a team has the operational capacity to own it. The REST option is compelling when adding metrics should look like adding one capability to an existing backend contract, but its missing alerting and tracing surfaces are firm boundaries, not footnotes.

No option makes label design optional.

The trade-off is concrete. A small team can prefer one key and one REST integration because that removes credential and SDK sprawl, while a team that needs native paging, span trees, sophisticated alert policy, or synthetic checks should choose Datadog or assemble Prometheus, Grafana, and dedicated alerting components instead. The verified breadth numbers, 295 capabilities across 20 modules with examples in 10 languages, describe integration scope; they do not prove that the narrower observability features match a mature suite.

Treat alerting, heartbeats, and investigation as separate paths

For production incident notification, run a small polling job against the chosen metric queries, compare a short list of explicit thresholds, and send through the team's Slack, email, or webhook provider. Make notifications idempotent around a stable tuple such as rule, evaluation window, and state transition; otherwise a retry can turn one pipeline failure into a page storm. The polling interval and evaluation window should tolerate the nightly job's normal completion variance, which must come from the team's schedule rather than an invented universal value. That poller still cannot detect every silent miss by looking for a metric that was never emitted. Register the schedule with a heartbeat service and signal success only after the pipeline reaches its terminal checkpoint; Healthchecks.io documents this dead-man-switch pattern directly, and the heartbeat payload should remain free of health data. When an alert fires, pivot to logs using a coarse run correlation value and a tight time range. This is where structured logs earn their storage cost: they retain failure context that would be destructive as metric dimensions. For EU GDPR work, document which identifiers can enter logs, how deletion requests are executed, and how data is exported before committing to the log backend. If those operations cannot be demonstrated, the architecture is incomplete even if the search screen looks excellent.

Keep those paths separate.

Roll out the split without losing evidence

First, inventory the existing dashboard queries and classify each output as a bounded aggregate, investigative detail, or liveness assertion. Move only the bounded aggregates to metrics. Leave diagnostic context in logs and send the liveness assertion to a heartbeat monitor.

Next, dual-write the agreed aggregates from the nightly pipeline while the old charts remain available. Compare totals over complete pipeline runs, paying particular attention to late events, retries, and the definition of “failed”; those semantic mismatches cause more damage than chart rendering. Do not claim equivalence from a single successful run.

Then switch the internal dashboard cards and trend lines to the metric source, retain a direct link or run identifier for log investigation, and exercise one failed-run alert plus one never-started heartbeat. Only after the team can perform those drills should it retire the log-derived KPI queries. The result is modest: metrics for repeated decisions, logs for evidence, and a heartbeat for silence. That modesty is the point.

Sources

References used for the storage and operating-model distinctions:

Top comments (1)

Collapse
 
docify profile image
Docify •

Hey, I'm building a small open-source CLI that analyzes a codebase and generates architecture/structure documentation. I'm looking for a few developers willing to run it against a real project and tell me where it gets things wrong.
You don't need to upload your code anywhere just run:
npx @autodocify/autodocs analyze .
Requires Node 20+.
If you try it, I'd especially like to know what it missed or misunderstood.