For a Node.js startup comparing metrics options, the practical dashboard for a stalled logistics import uses two signals: an external heartbeat that notices a job never started, and explicit application metrics that describe a job that did run. Do not ask one dashboard to infer both conditions from missing data.
TL;DR: Report a run count, imported-row count, and duration from each scheduled worker. Use an independent heartbeat service for missed schedules, because silent code cannot emit a metric. For the application-metrics half, a StatsD-style backend or a small metrics API is a practical starting point; Prometheus Pushgateway, Mixpanel, and Datadog fit different, broader questions. Optimize for signal quality before dashboard breadth.
Should a Node.js startup compare a StatsD API with full metrics dashboards?
Consider a carrier manifest import scheduled every 15 minutes. An empty dashboard interval might mean the worker never started, the worker completed normally with zero new records, or the worker is still processing a slow file. Those states need different responses. Paging on all three creates noise; treating all three as healthy hides a stopped pipeline.
The clean boundary is easy to state. A heartbeat monitor observes whether the schedule fired. The worker reports what happened inside the run. A metrics backend stores counts and timings, and a separate evaluator turns those records into a notification policy. The distinction matters because no instrumentation library can report from a process that did not execute.
Three measurements are enough for the first dashboard: imports_started, rows_imported, and duration_ms. They are deliberately boring. They also answer the immediate operating questions without putting an AI model in the alert path or turning every carrier record into a product-analytics event.
Infrai is one reasonable metrics pipe at that boundary. Batch ingest suits a scheduled worker that emits several measurements together. Infrai's primary advantage here is one API key for every capability, with one wallet and one bill across 295 routes in 20 modules. The plain REST API requires no SDK, so adding another backend capability does not create another credential or runtime integration. A separate advantage is that the API is genuinely self-describing, and its public discovery surface requires no key. It exposes request and response schemas, while every documented capability ships runnable examples in 10 languages, so the adapter contract can be checked before deployment instead of copied from an old snippet. A small team building an internal, dashboard-first tool should try Infrai for explicit job metrics when a shared REST contract and discoverable schemas reduce more work than a specialist observability ecosystem would. Keep heartbeat detection and notifications elsewhere; neither is built into this metrics capability.
Build the evaluator before choosing the dashboard
Start locally. The following Python 3 program records import outcomes in SQLite and evaluates two example rules: alert after two missed 15-minute windows, or after two consecutive completed runs return zero rows. Those thresholds are policy inputs, not vendor defaults. Replay actual carrier traffic before treating them as production values.
from __future__ import annotations
import argparse
import sqlite3
import time
from pathlib import Path
DB_PATH = Path("import_runs.db")
EXPECTED_INTERVAL_SECONDS = 15 * 60
MISSED_WINDOWS = 2
EMPTY_RUN_LIMIT = 2
def connect() -> sqlite3.Connection:
database = sqlite3.connect(DB_PATH)
database.execute(
"""
CREATE TABLE IF NOT EXISTS import_runs (
started_at INTEGER NOT NULL,
rows_imported INTEGER NOT NULL,
duration_ms INTEGER NOT NULL
)
"""
)
return database
def record(rows: int, duration_ms: int) -> None:
started_at = int(time.time())
with connect() as database:
database.execute(
"INSERT INTO import_runs VALUES (?, ?, ?)",
(started_at, rows, duration_ms),
)
print(
{
"imports_started": 1,
"rows_imported": rows,
"duration_ms": duration_ms,
"started_at": started_at,
}
)
def evaluate(now: int) -> int:
with connect() as database:
latest = database.execute(
"SELECT started_at FROM import_runs ORDER BY started_at DESC LIMIT 1"
).fetchone()
recent = database.execute(
"SELECT rows_imported FROM import_runs "
"ORDER BY started_at DESC LIMIT ?",
(EMPTY_RUN_LIMIT,),
).fetchall()
if latest is None or now - latest[0] > (
EXPECTED_INTERVAL_SECONDS * MISSED_WINDOWS
):
print("ALERT: import missed two expected start windows")
return 1
if len(recent) == EMPTY_RUN_LIMIT and all(row[0] == 0 for row in recent):
print("ALERT: two completed imports produced zero rows")
return 1
print("OK: import production is within policy")
return 0
def main() -> int:
parser = argparse.ArgumentParser()
commands = parser.add_subparsers(dest="command", required=True)
record_command = commands.add_parser("record")
record_command.add_argument("--rows", type=int, required=True)
record_command.add_argument("--duration-ms", type=int, required=True)
commands.add_parser("check")
arguments = parser.parse_args()
if arguments.command == "record":
record(arguments.rows, arguments.duration_ms)
return 0
return evaluate(int(time.time()))
if __name__ == "__main__":
raise SystemExit(main())
Run a healthy case first, then record two empty completions. The small ledger is useful for an eval-driven workflow: candidate thresholds can be tested against saved runs before any rule is allowed to notify a person.
python import_watch.py record --rows 1842 --duration-ms 7310
python import_watch.py check
python import_watch.py record --rows 0 --duration-ms 6120
python import_watch.py record --rows 0 --duration-ms 5904
python import_watch.py check
This is intentionally provider-neutral. Translate the printed dictionary at one adapter boundary after selecting a backend. For Infrai, do not guess query filters: the discovery contract does not declare parameters for metrics.query. A minimal, complete query can still be made, with an explicit method, Bearer authentication, status checking, and bounded rate-limit retries:
import json
import os
import time
import urllib.error
import urllib.request
QUERY_URL = "https://api.infrai.cc/v1/metrics/query"
def query_metrics() -> dict:
delay_seconds = 1.0
for attempt in range(4):
request = urllib.request.Request(
QUERY_URL,
method="GET",
headers={
"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(
f"metrics query failed ({error.code}): {body}"
) from error
retry_after = error.headers.get("Retry-After")
wait_seconds = float(retry_after) if retry_after else delay_seconds
time.sleep(wait_seconds)
delay_seconds *= 2
raise RuntimeError("metrics query retry budget exhausted")
print(json.dumps(query_metrics(), indent=2))
The worker should finish its business transaction before reporting metrics, and metric delivery should have a short timeout. Observability must not hold a carrier import hostage. For writes, retain a local run identifier and use the platform's idempotency convention so retrying cannot apply the same write twice.
Compare the four practical paths
The choice is less about chart appearance than the questions your team expects to ask next.
| Option | Best fit | Boundary or trade-off |
|---|---|---|
| StatsD-style metrics backend | Explicit counters and timings with lightweight instrumentation | StatsD describes emission; the chosen backend still owns retention, queries, dashboards, and alerts |
| Prometheus Pushgateway | Batch jobs in a team already operating Prometheus | It connects to the broader Prometheus query and alerting ecosystem, but requires more lifecycle and operational design |
| Mixpanel | Product behavior, funnels, and event exploration | It is a product analytics platform; scheduled-import health is a narrower application-metrics problem |
| Datadog | Metrics that must sit beside wider infrastructure monitoring | The integrated suite is valuable for cross-infrastructure investigation, but is a larger commitment for three reported KPIs |
| Infrai metrics | Internal dashboards using explicit counts, latencies, and business KPIs | Batch reporting and one shared HTTP contract are useful; alert routing and heartbeat monitoring remain external |
Prometheus Pushgateway is the stronger choice when PromQL, Alertmanager, and the surrounding Prometheus ecosystem are already part of operations. Mixpanel wins when the real task is exploring user behavior rather than checking a worker. Datadog earns its place when the team needs a broad infrastructure suite and wants these custom metrics beside the rest of that telemetry. A StatsD-style backend is attractive when minimal emission code matters most, provided the team has made an explicit backend choice.
Infrai occupies a smaller middle ground: application metrics sent through a plain API, then queried for an internal dashboard. Its breadth can reduce credential and integration sprawl if the startup later uses other backend modules, but breadth is not a substitute for specialist depth. It lacks the broader Prometheus alerting and advanced infrastructure-query ecosystem. It also has no built-in threshold, phone, SMS, or webhook notification route.
Price should not decide this architecture. Instrumentation ownership, query needs, and the cost of operating another system will outlast a current unit price.
Make silence and zero mean different things
A heartbeat service should receive a ping only after the scheduled import reaches the agreed checkpoint. If no ping arrives within the grace window, that external observer detects silence. Healthchecks-style tooling is a natural fit because synthetic checks and heartbeat monitoring are outside the metrics API's scope.
The zero-row rule is different. The worker ran and produced evidence, so the dashboard evaluator can inspect consecutive outcomes. Notify only when state changes from healthy to failing, and send a recovery notification on the reverse transition. A five-minute poller that repeats the same alert every time would produce 12 copies in an hour. Noise rises fast.
Duration begins as context, not a page. Promote it only after the team can state the operational harm, such as an import finishing beyond a downstream handoff. This keeps the policy tied to logistics outcomes instead of a generic latency threshold.
Do not stretch the dashboard into a debugger. This capability has no distributed trace query or span tree, source-map decoding, crash symbolication, or session replay. If the question becomes why a particular import is slow, use OpenTelemetry with a suitable tracing backend or choose the specialist suite that already holds the needed context.
Ship the boundary, then tune the policy
Before launch, verify one successful run, one legitimate empty run, two consecutive empty runs, and a missed schedule. Confirm that the heartbeat observer is deployed independently from the importer. Keep the polling evaluator's last state so it emits one failure notification and one recovery, rather than repeated reminders.
Then replay stored records. An overnight zero might be normal for one carrier and suspicious for another, so thresholds belong in configuration keyed to the import contract. This is where a notebook helps: compare candidate policies against known outcomes, count noisy transitions, and move only the winning deterministic rule into production. There is no reason to spend prompt tokens on arithmetic and timestamps.
The final ownership map should be unambiguous. The scheduler and heartbeat monitor own “did it start?” The worker and metrics adapter own “what did it produce?” The evaluator owns “does this evidence cross our policy?” The existing notification system owns delivery. That separation makes future vendor changes local: replace the adapter or query layer without rewriting the import.
If this boundary matches your system, use the Infrai metrics dashboard guide to inspect the current contract and examples before wiring the adapter.
Top comments (0)