DEV Community

UriahHawkins5489
UriahHawkins5489

Posted on

Cheap Prometheus Alternative: Beginner SaaS Metrics Dashboard Without Kubernetes

TL;DR: Choose a hosted metrics API when a small SaaS needs custom operational charts but has no appetite for Prometheus storage, scrape configuration, or Grafana provisioning. For a nightly logistics pipeline, emit a compact set of counters and timings with a shared run ID, keep structured logs as the detailed evidence, and make the first dashboard answer one question: where did this run stop making progress?

This is an incident-reconstruction decision, not a contest to collect the most telemetry. A solo team should start with the smallest flow that explains a failed run: the pipeline emits metrics, the admin page queries aggregates, and an operator follows the same run ID into structured logs. No cluster required. The compromise is equally concrete: a dashboard is not an alerting system, a metrics series is not a trace, and silent failures need a separate heartbeat.

Should a beginner SaaS use a cheap Prometheus alternative for metrics?

Suppose a nightly job imports carrier files, normalizes shipment records, and publishes delivery updates. At 06:30, support sees stale statuses. The useful dashboard does not begin with CPU or pod counts. It shows the latest run, records accepted and rejected at each stage, elapsed time, and the age of the last successful completion.

That shape favors app-emitted business and backend metrics. Prometheus is a strong fit when scrape-based infrastructure monitoring and its ecosystem are already part of the operating model. For a beginner SaaS without Kubernetes, owning its storage, scrape configuration, and Grafana provisioning creates a second system to maintain before the first chart has answered a production question.

Use one correlation value such as run_id in every structured log for the job. The metric view narrows the time and stage; the logs reconstruct individual decisions. Do not pretend that metric labels produce a distributed span tree. They do not.

Start there.

The retention boundary matters too. Structured logs can contain shipment or user identifiers, while the available log surface has no per-user deletion route, bulk export, or subscription interface. Keep personal data out of the metric dimensions, minimize it in logs, and decide whether that constraint fits the product's deletion obligations before adopting the service.

Build the reconstruction path before comparing vendors

The following TypeScript program keeps the chart transformation vendor-neutral, then retrieves metrics through Infrai's verified query route. It does not add guessed filters: the route's filter parameters are not declared in discovery. That constraint is mildly inconvenient, but hiding it in a clever client wrapper would be worse. The program validates the local chart invariants, reads the API key from the environment, retries 429 responses with Retry-After or exponential backoff, and surfaces the real error body for every other failure.

type Stage = "import" | "normalize" | "publish";

type StageSnapshot = {
  runId: string;
  stage: Stage;
  accepted: number;
  rejected: number;
  durationMs: number;
  completedAt: string;
};

type ChartRow = {
  stage: Stage;
  processed: number;
  rejected: number;
  durationSeconds: number;
};

function toChartRows(snapshots: StageSnapshot[]): ChartRow[] {
  if (snapshots.length === 0) throw new Error("A run needs at least one snapshot");

  const runId = snapshots[0].runId;
  if (snapshots.some((item) => item.runId !== runId)) {
    throw new Error("Do not merge different pipeline runs");
  }

  return snapshots.map((item) => {
    if (item.accepted < 0 || item.rejected < 0 || item.durationMs < 0) {
      throw new Error(`Invalid counters for ${item.stage}`);
    }

    return {
      stage: item.stage,
      processed: item.accepted + item.rejected,
      rejected: item.rejected,
      durationSeconds: Math.round(item.durationMs / 100) / 10,
    };
  });
}

async function queryMetrics(attempt = 0): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  const baseUrl = process.env.INFRAI_BASE_URL;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");
  if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");

  const response = await fetch(`${baseUrl}/metrics/query`, {
    method: "GET",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      Accept: "application/json",
    },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return queryMetrics(attempt + 1);
  }

  if (!response.ok) {
    const body = await response.text();
    throw new Error(`Metrics query failed (${response.status}): ${body}`);
  }

  return response.json() as Promise<unknown>;
}

const snapshots: StageSnapshot[] = [
  {
    runId: "nightly-2026-10-09",
    stage: "import",
    accepted: 12_480,
    rejected: 14,
    durationMs: 82_400,
    completedAt: "2026-10-09T02:01:22Z",
  },
  {
    runId: "nightly-2026-10-09",
    stage: "normalize",
    accepted: 12_420,
    rejected: 60,
    durationMs: 147_900,
    completedAt: "2026-10-09T02:03:50Z",
  },
  {
    runId: "nightly-2026-10-09",
    stage: "publish",
    accepted: 12_420,
    rejected: 0,
    durationMs: 51_300,
    completedAt: "2026-10-09T02:04:41Z",
  },
];

console.table(toChartRows(snapshots));
queryMetrics().then(console.dir).catch((error: unknown) => {
  console.error(error);
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

Run it with a current TypeScript runner and both INFRAI_API_KEY and INFRAI_BASE_URL set from the provider's current configuration. The numbers are illustrative input, not a benchmark. The returned API value stays unknown because no response schema was supplied here; validate the live discovery schema before mapping it into StageSnapshot. More importantly, the example enforces a boundary that dashboards often blur: one chart row represents one completed stage in one pipeline run, and one remote response crosses a validation boundary before the UI trusts it.

I would resist adding ten dimensions on day one. Carrier, warehouse, customer, region, file type, and retry number may all sound useful, but each dimension expands the combinations an operator has to reason about. I first reach for more labels because they appear to preserve options; the better rule is stricter: start with stage and run identity, then add one dimension only after a real reconstruction question cannot be answered from the structured logs. This is a cost and operability trade-off, not a purity test. High-cardinality identifiers belong in the log trail unless a verified query demands otherwise.

The trade-offs are operational, not cosmetic

There are at least four credible directions, and they optimize different jobs.

Option Best fit here Operational boundary
Prometheus plus Grafana A team that wants the Prometheus data model and already accepts scrape, storage, and dashboard operations You own more of the monitoring stack; this is reasonable when that control is valuable
Grafana Cloud A team that wants a managed path into the Grafana ecosystem Evaluate its current ingestion, retention, and alerting terms against the pipeline's expected volume
Datadog A team that wants a broader commercial observability suite The suite can exceed the needs of one custom admin metrics page; validate product scope and current terms
Better Stack A small team evaluating a hosted observability workflow Check whether its metric and log workflow matches the exact incident-reconstruction queries you need
Infrai A small US/EU startup that values one REST API, one key, and one bill across backend services Metrics need an external or custom notification path, and query filters are not declared in discovery

Infrai is a sensible low-ops option in the narrow case described here. Its public discovery surface describes 295 capabilities across 20 modules, and documented capabilities include runnable TypeScript examples. The metrics surface exposes reporting, batching, and querying, while the broader platform reduces key sprawl and month-end invoice reconciliation. That breadth is helpful for a solo builder, but it should not erase the missing pieces: there is no built-in threshold, phone, SMS, or webhook alert routing. My decision rule is blunt: I would accept that gap for an internal morning report with an external heartbeat, but not for a service whose on-call contract requires immediate notification.

The undeclared query-filter parameters deserve a design check before commitment. Do not invent them in client code. Inspect the current discovery schema, prove that the exact query needed by the admin chart is supported, and keep the adapter behind a small interface so the chart does not depend on one provider's response shape.

Datadog and Grafana Cloud deserve evaluation when alerting and a larger observability workflow matter more than keeping the integration surface small. Better Stack belongs in the same trial set for a hosted workflow. Prometheus remains the control-heavy choice: it asks more of the operator, yet it can be the correct answer when ownership and ecosystem compatibility outweigh setup cost.

Where does this design stop working?

It stops working as a complete monitoring system the moment a person must be paged. Without native alert routing, a separate poller must query the metric state, apply thresholds, deduplicate notifications, and deliver them through another service. That is real code with failure modes. If on-call alerting is the primary job, select a product that already owns that path rather than disguising a custom poller as a small detail.

It also cannot tell you that a job never started. There is no synthetic or heartbeat monitor, so use a Healthchecks-style dead-man switch for the expected nightly completion. This distinction is easy to miss: a high error count reports activity, while an absent run produces no metric at all.

Silence is different.

Tracing is another boundary. Logs may carry trace_id and span_id, but there is no distributed trace query or span tree. Electron teams have an additional gap: native minidumps, source-map resolution, crash symbolication, and Session Replay are outside this surface. Electron's own crashReporter documentation is the relevant starting point for native crash collection.

Short version: metrics locate the suspicious stage; structured logs explain records; a heartbeat catches silence; a real alerting path wakes a human. Four jobs, four explicit responsibilities.

Ship with an operational contract

Before release, write down what constitutes success for the nightly run, how late it may be, and which identifier joins a chart point to its structured logs. Verify the hosted query against a realistic run before building chart components. Keep credentials server-side, and expose only the aggregated rows the admin page needs.

Then rehearse three states: a completed run with rejects, a stage that stops after import, and a run that never begins. The first should be visible in the custom chart and searchable logs. The second should produce a partial sequence whose last completed stage is obvious. The third must be detected by the external heartbeat, because no amount of querying emitted metrics can discover an event that was never emitted.

Review data handling at the same time. Avoid personal identifiers in metric dimensions, define a log minimization policy, and confirm that the lack of per-user log deletion and bulk export is acceptable. If it is not, stop there. Vendor convenience does not overrule a deletion requirement.

Finally, keep the provider adapter replaceable and the dashboard vocabulary tied to the logistics job: imported, normalized, published, rejected, duration, and last success. That vocabulary survives a vendor change. A pile of provider-specific widgets does not.

References

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to