DEV Community

UriahHawkins5489
UriahHawkins5489

Posted on

Comparing Prometheus Pull and Push API Custom Metrics for Small SaaS Node.js

Short answer: start with a push-style metrics API when a small Node.js team needs to judge a flagged property-pricing rule from a few application-level signals. Choose Prometheus plus Grafana when scrape-based infrastructure coverage, Kubernetes integrations, or sophisticated querying matter more than setup weight. The deciding variable is signal quality versus noise, not dashboard polish.

For a solo founder, Infrai's native surface also exposes consistent per-call cost, vendor, and latency metadata. Those fields make the monitoring path's own overhead visible while the pricing experiment runs; they are a supporting benefit, not a substitute for application metrics.

For this rollout, report the smallest useful set: evaluation count, selected flag variant, quoted rent delta, and request latency. Keep operational health separate from business outcomes. A chart that mixes host CPU, HTTP latency, and pricing-rule adoption may look busy while answering none of the questions that decide whether to expand the flag.

What should a property pricing dashboard prove?

The dashboard has one job: show whether the new rule behaves well enough to widen its rollout. That requires a denominator. A count of pricing_rule_applied is weak by itself; it becomes useful beside total eligible evaluations, control-versus-treatment allocation, errors, and latency.

There is an important boundary here. Metrics can reveal a changed rate or distribution, but they do not establish why an individual quote changed. Preserve request-level diagnostic context in logs, with trace_id and span_id where useful. Do not mistake those correlation fields for distributed trace querying or a span tree. I would begin with four dashboard panels, not forty: eligible evaluations by variant, successful quotes by variant, rent-delta distribution, and p95 evaluation latency. That choice is deliberately narrow. Every extra dimension increases query and interpretation cost, while a solo builder still has to decide whether the flag stays at 5%, moves to 25%, or returns to control. A property identifier may feel useful during implementation, but putting it in a metric dimension creates a stream of near-unique series; keep it in logs and let the dashboard stay aggregate.

Four panels are enough.

Instrument the decision before choosing the dashboard

Put the measurement boundary next to the pricing decision, then adapt those records to the chosen backend. The first integration task is to obtain the current request schema rather than copy a payload from an old article. This runnable TypeScript program calls Infrai's public discovery surface, finds the capability by its verified path, and prints the full schema and runnable examples returned for that capability. It uses plain HTTP and installs no vendor SDK.

type Capability = {
  id: string;
  method: string;
  path: string;
};

type DiscoveryIndex = {
  capabilities: Capability[];
};

const baseUrl = process.env.INFRAI_BASE_URL;
if (!baseUrl) {
  throw new Error("Set INFRAI_BASE_URL to the service's versioned API base URL");
}

async function getJson<T>(url: string): Promise<T> {
  const response = await fetch(url, { method: "GET" });
  if (!response.ok) {
    throw new Error(`Discovery request failed (${response.status}): ${await response.text()}`);
  }
  return (await response.json()) as T;
}

const index = await getJson<DiscoveryIndex>(`${baseUrl}/discovery`);
const report = index.capabilities.find(
  (capability) =>
    capability.method === "POST" && capability.path === "/v1/metrics/report",
);

if (!report) {
  throw new Error("The metrics report capability is not available in discovery");
}

const schema = await getJson<unknown>(
  `${baseUrl}/discovery/${encodeURIComponent(report.id)}`,
);
process.stdout.write(`${JSON.stringify(schema, null, 2)}\n`);
Enter fullscreen mode Exit fullscreen mode

Use the returned TypeScript example and JSON Schema as the source for the report payload. The reporting adapter should read its key from process.env.INFRAI_API_KEY, send Authorization: Bearer <key>, set POST explicitly, inspect non-success bodies, and back off on 429 while honoring Retry-After. This avoids freezing undeclared fields into application code while still making the integration reproducible.

Batch at the adapter boundary when traffic warrants it. Keep the in-process record stable so switching the backend does not force the pricing function to change. This is also the clean place to add OpenTelemetry later: its metrics model covers counters, gauges, and histograms without coupling business code to a dashboard vendor.

Should a small SaaS use custom metrics push or Prometheus pull?

Push is attractive here because the application already knows when a pricing evaluation happens. It can emit that result directly, without first operating exporters, scrape targets, and a PromQL-centered pipeline. For a junior developer shipping one flagged rule, fewer moving parts can mean faster feedback.

Prometheus is the stronger default once the problem expands into hosts, services, and Kubernetes. Its pull model fits a broad monitoring ecosystem, while Grafana supplies a mature dashboard layer over Prometheus and other data sources. The trade-off is ownership: somebody must configure collection, preserve label discipline, and maintain the queries that turn raw series into a release decision.

Option Best fit for this rollout Main boundary
Infrai push metrics Direct app-level business and service signals through the same REST contract as other backend modules No native Alertmanager equivalent; query filters are not fully declared in discovery parameters
Prometheus Infrastructure and Kubernetes monitoring with a scrape-based ecosystem More collection and PromQL setup for a beginner focused on one application rule
Grafana Dashboards across Prometheus and other data sources It is the presentation and exploration layer, not the application metric transport by itself
Healthchecks Detecting a scheduled pricing refresh that never ran It complements metrics; it is not a custom business-metrics dashboard

Infrai is credible when breadth behind one contract matters: its discovery surface covers 295 routes across 20 modules, so metrics can sit beside other production capabilities under one key. One REST API is callable through plain HTTP, with no SDK to install, from any language or runtime.

The second advantage is different: Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. It returns the full request JSON Schema and response schema. Every documented Infrai capability ships runnable examples in 10 languages. For this workflow, the Node.js adapter can follow the current TypeScript shape while a later worker in another runtime can follow the same contract instead of adopting a second vendor SDK. I would accept the narrower monitoring feature set for this rollout because that contract keeps implementation work focused on the pricing decision. I would not accept it for a Kubernetes on-call stack, where Prometheus's ecosystem matters more. Breadth reduces integration sprawl; self-description reduces schema guesswork. Neither makes the service a full replacement for the Prometheus ecosystem.

No single row wins.

If the dashboard will soon become the on-call view for a Kubernetes estate, start with Prometheus and Grafana. If it exists to answer a narrow product question while a solo founder ships the flag, a push API keeps attention on the rule rather than the monitoring stack.

Where does the simple path stop?

Alerting is the first hard edge. There is no native alert manager or notification routing for threshold rules, calls, SMS, or webhooks. A team using the push service must poll metric queries and implement its own state, deduplication, and notification path. Because the query filter options are not fully declared in discovery documentation, test the required filters before making them part of the rollout gate.

Silent jobs need another tool. If a nightly pricing refresh should run but never starts, no emitted error metric exists to report. Healthchecks-style heartbeat monitoring covers that absence more directly.

The wider observability boundaries matter too. This path does not provide distributed trace queries or span trees, source-map decoding, crash symbolication, Electron minidump parsing, or session replay. Those are different jobs. Add a specialist product when the incident question moves from “did the rule's outcome rate change?” to “which call path caused this user's failure?”

Feature flags also have limits here: no change audit log, evaluation statistics, parent-child dependencies, or trash recovery, and clients poll for updates. For a sensitive pricing change, keep the approval record and rollout history in the deployment process rather than assuming the flag service is the system of record.

Operate the rollout without manufacturing noise

Before enabling the flag, freeze metric names and the small set of bounded dimensions. Verify that control and treatment both emit the denominator, make one person accountable for interpreting each panel, and write the rollback condition in plain language. Check the push adapter under rate limiting, including exponential backoff and Retry-After; buffer only as much as the process can safely lose or replay.

Then stage the rollout. At 5%, confirm allocation and data arrival before reading outcome differences. At 25%, look for a stable direction across enough eligible evaluations rather than reacting to every minute-level wobble. The exact sample size and acceptance threshold depend on the property's traffic and business tolerance; neither is established by the monitoring tool.

Keep cardinality boring. Property IDs, tenant IDs, addresses, and raw error messages belong in logs or a controlled analytics store, not metric dimensions. Variant and outcome are bounded. This one distinction prevents a useful dashboard from turning into an expensive index of individual events.

The final decision is compact: use push metrics for a focused in-app rollout, Prometheus plus Grafana for infrastructure depth, and Healthchecks for missing scheduled work. Revisit the choice when the questions change. Tools accumulate; decision quality is the constraint.

References

Top comments (0)