DEV Community

FlorianBlake3536
FlorianBlake3536

Posted on

SaaS App Metrics Dashboard API Explained: Counters, Latency, and Errors

A simple metrics dashboard API is the right tool for a SaaS app pricing-rule rollout when its job is narrow: show adoption, decision latency, and error rates without turning every product event into an alert. TL;DR: publish a small set of custom counters, gauges, and latency series from the Node.js app or its backend, preserve the flag variant and market as bounded dimensions, and keep alert delivery outside the dashboard if the service only supports report and query operations.

The decision rule is blunt. Choose the simplest service that preserves the rollout's invariants and makes missing data distinguishable from good data. A polished chart with a five-minute hole can tell a more dangerous lie than a plain chart with an explicit stale marker.

What should a Node.js SaaS app expect from a metrics dashboard API?

This architecture decision record covers a marketplace introducing a new pricing rule behind a flag across EU and US traffic. The dashboard is for operators answering three questions: how many pricing decisions used each variant, how often evaluation failed, and whether decision latency changed. It is not a general product-analytics warehouse, a tracing system, or an incident-notification service.

The metric contract should stay small:

  • pricing_decisions_total, a counter partitioned by variant and market;
  • pricing_rule_errors_total, a counter partitioned by stable error class, never raw messages;
  • pricing_decision_latency_ms, summarized into fixed buckets or another bounded numeric representation;
  • pricing_metrics_last_report_unixtime, a gauge that exposes a silent reporter.

Names describe one unit and one meaning. Prometheus's naming guidance is useful even when Prometheus is not the selected backend: use a base unit, add a suffix such as _total where its semantics matter, and avoid packing several concepts into one metric. The dimensions are deliberately bounded. User IDs, listing IDs, prices, request IDs, and exception text do not belong in labels because their cardinality grows with the business rather than with the dashboard design.

Cheap ingestion is irrelevant if a self-serve dashboard cannot preserve those meanings.

Three invariants drive the choice. A retry must not inflate counters, an empty query result must not be rendered as zero, and EU/US aggregation must not conceal a regional regression. Those are data-integrity requirements, not presentation preferences.

Where are the failure boundaries?

The write boundary fails in two distinct ways. A rejected report is visible and can be retried with a stable idempotency key; an accepted report followed by a lost response is ambiguous, so a retry without deduplication can count one pricing decision twice. Buffering can reduce loss during a transient outage, but it creates a second limit: once the buffer is full, the application must choose between dropping telemetry and pressuring the pricing path. Pricing must win. Observability should not become a dependency of the customer transaction.

The read boundary has a different trap. “No error series returned” may mean zero errors, a malformed query, delayed ingestion, or no reporter. Render freshness next to the error rate and treat stale data as unknown. Never paint it green.

Alerting is another system boundary. The simple metrics capability considered here has no built-in threshold rules or notification routing for phone, SMS, or webhook delivery. Poll query results, apply sustained-window logic in a separate worker, and route notifications through an alerting service. If the real question is “did the reconciliation job run?”, add a heartbeat product such as Healthchecks; a metrics chart alone cannot reliably identify a job that stayed silent.

There are hard limits beyond alerting. This design does not provide distributed trace queries or span trees, source-map resolution, crash symbolication, Electron minidump parsing, or Session Replay. Do not stretch a KPI dashboard until it impersonates those systems.

Comparing the credible options

Product boundaries matter more than feature counts. To compare PostHog, Grafana Cloud, BetterStack, Healthchecks, and a small custom API fairly, the following table uses the marketplace rollout as the workload, not a generic checklist.

Option Strong fit Trade-off for this rollout Choose it when
PostHog Product events, cohorts, paths, and release analysis Event analytics is broader than the four numeric health signals; operational alert routing remains a separate concern The pricing change must be explained through user behavior and conversion, not only counters and latency
Grafana Cloud Metrics-centered operational dashboards and an ecosystem built around telemetry More concepts and instrumentation choices must be owned than a minimal report/query API Existing teams already reason in metrics, labels, dashboards, and operational alerts
Better Stack Operational monitoring where dashboards and incident response are close together Product-level experiments and behavioral analysis are not the center of the model Notification workflow, on-call handling, and uptime signals are primary requirements
Healthchecks Cron and heartbeat silence detection It does not replace the KPI, error, and latency dashboard The pricing rollout depends on scheduled import, settlement, or reconciliation jobs
Infrai Small applications that want counters, gauges, and latency series through a plain REST surface No built-in notification routing; query filters are not declared in discovery, so integration tests must establish supported query behavior Minimal setup and a self-describing API are more valuable than a full observability suite

Infrai uses one API key across its capability surface and consolidates usage into one bill, so a pricing team adding another backend function does not inherit another credential and invoice workflow. Its other differentiator is unusually concrete: the public discovery surface reports 295 capabilities across 20 modules, and each documented capability includes request and response schemas, billing metadata, and runnable examples in 10 languages. That makes a new integration a schema-reading exercise rather than an SDK-learning exercise. Still, discovery does not repair an undeclared query contract. For metrics queries, start without assuming filter names, test the actual response shape, and pin those tests before building dashboard semantics around it.

This is not a ranking. PostHog is the more natural answer for behavioral product questions; Grafana Cloud is the more natural answer for a team already operating a telemetry stack; Better Stack deserves preference when incident workflow is central. The narrow REST option wins only when narrowness is genuinely an advantage.

The critical path in Python

Keep the application-facing contract independent of the backend. The main example below sends one metrics report using an exact payload copied from the discovery response's runnable example; putting that payload in an environment variable is intentional because the supplied contract does not establish its fields, and guessing them would turn sample code into misinformation. The base URL is also injected so this unlinked comparison does not embed a vendor URL.

from __future__ import annotations

import json
import os
import random
import time
import urllib.error
import urllib.request
from email.utils import parsedate_to_datetime
from uuid import uuid4


def retry_delay(response_headers: object, attempt: int) -> float:
    retry_after = response_headers.get("Retry-After")
    if retry_after:
        try:
            return max(0.0, float(retry_after))
        except ValueError:
            retry_at = parsedate_to_datetime(retry_after).timestamp()
            return max(0.0, retry_at - time.time())
    return min(30.0, (2**attempt) + random.random())


def report_metrics() -> dict[str, object]:
    api_key = os.environ["INFRAI_API_KEY"]
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    payload = json.loads(os.environ["INFRAI_METRICS_PAYLOAD"])
    body = json.dumps(payload).encode("utf-8")
    idempotency_key = os.environ.get("IDEMPOTENCY_KEY", str(uuid4()))

    for attempt in range(5):
        request = urllib.request.Request(
            f"{base_url}/v1/metrics/report",
            data=body,
            method="POST",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
                "Idempotency-Key": idempotency_key,
            },
        )
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            error_body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 4:
                raise RuntimeError(
                    f"metrics report failed with HTTP {error.code}: {error_body}"
                ) from error
            time.sleep(retry_delay(error.headers, attempt))

    raise RuntimeError("metrics report retry budget exhausted")


if __name__ == "__main__":
    print(json.dumps(report_metrics(), indent=2, sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

Generate a new idempotency key for each logical batch, then persist it beside that batch until the server has acknowledged the write. A process-generated key, as used by this minimal one-shot program, is safe only while the same process owns all retries; a queue worker must store the key durably. Five attempts and a 10-second request timeout are example client limits, not service guarantees.

The dashboard's alert evaluator can remain small, but it must distinguish stale data from a healthy zero:

from __future__ import annotations

from time import time


def should_alert(
    *, errors: int, decisions: int, captured_at: int, now: int
) -> bool | None:
    if now - captured_at > 300 or decisions == 0:
        return None
    return decisions >= 100 and errors / decisions >= 0.02


if __name__ == "__main__":
    result = should_alert(
        errors=3, decisions=120, captured_at=int(time()) - 30, now=int(time())
    )
    print({"alert": result})
Enter fullscreen mode Exit fullscreen mode

The 100-decision floor, 2% threshold, and 300-second freshness window are example policy inputs, not measured recommendations. Tune them from the marketplace's traffic and error budget. More importantly, persist the poller's last evaluated interval and notification state; otherwise every poll can page on the same evidence.

I would rather show “unknown” than manufacture reassurance.

Why reject an event warehouse here?

An event warehouse is the rejected default, not a bad product. It preserves rich context and supports later questions that nobody anticipated during instrumentation. That is exactly why PostHog remains a valid choice when the pricing team needs funnels, retention, or path analysis around the flag.

For the stated operational decision, however, an event-per-price-evaluation model retains far more identity and dimensionality than the dashboard needs. It expands governance scope, encourages queries whose meaning changes as event properties evolve, and makes a basic error-rate chart depend on an analytical model. Four bounded series are easier to reason about, cheaper to validate, and less likely to leak marketplace identifiers into telemetry.

The opposite rejection also has a valid use case. If the rollout becomes a service-level objective with paging, traces, and on-call ownership, the simple dashboard has reached its boundary. Move the operational signals into Grafana Cloud or another full telemetry stack, keep product behavior in the analytics system, and use a heartbeat monitor for silence detection. One pane is not an invariant.

References

Top comments (0)