DEV Community

Keria
Keria

Posted on

Hosted KPI Dashboard Backend API: Rollback-Safe Batch Metrics Ingestion

TL;DR: For an internal e-commerce dashboard that tracks an AI agent loop, send periodic metric batches from a worker and keep a small provider-neutral record at the application boundary. Batch ingestion cuts request overhead for snapshots such as order counts, active users, queue depth, agent latency, and token cost. The boundary is the important part: switching the hosted metrics backend should change one adapter, not the dashboard or the business code.

Choose this pattern when the dashboard can tolerate polling and the main job is charting operational KPIs. Do not mistake it for a complete observability stack. Native alert routing, distributed trace exploration, synthetic checks, session replay, and explicit long-term retention controls are separate buying decisions.

Should a cheap hosted KPI dashboard use a batch backend API?

The tempting implementation sends one HTTP request every time the agent calls a model, invokes a tool, or finishes an order-support task. It's easy to understand and expensive in the wrong currency: connections, retries, rate-limit pressure, and code paths that can interfere with the customer request. A five-step agent loop can turn one business action into several telemetry writes before the order workflow is done.

That's the wrong trade.

The better constraint is boring: business traffic must not wait for the KPI dashboard. Record measurements locally, aggregate them in a queue consumer or scheduled worker, and flush a bounded batch. For a small internal panel, a 60-second interval may be a sensible starting hypothesis, but it is not a universal default. Test it against the freshness your operators need and the amount of data you are willing to lose if a process exits before flushing.

Rollback safety changes the design more than vendor feature count does. A release should be reversible even if the metrics provider is slow or unavailable. Keep the previous adapter deployable, make each batch safe to retry, and avoid putting provider-specific query objects into React components. Then a provider change is plumbing rather than a dashboard rewrite.

I choose batching here because request isolation and a reversible adapter matter more than second-by-second freshness. I would measure six values for this e-commerce loop: completed runs, failed runs, end-to-end duration, model duration, input and output tokens, and an outcome such as order lookup completed. Cost belongs beside latency, but neither number explains value without that outcome. Keep dimensions low-cardinality: model family, agent version, tool name, and result class are useful; raw customer IDs, prompt text, order IDs, and trace IDs are poor metric labels. That choice has a boundary: if an operator must react within seconds, a minute-level batch is the wrong mechanism.

A narrow contract that survives a provider swap

The contract below keeps the hosted API at one boundary. The caller supplies the provider payload after validating it against the provider's public discovery schema; the transport owns authentication, retries, and error handling. This matters because the verified discovery surface supplies the full request JSON Schema, while the metric query filters aren't declared and shouldn't be guessed in application code.

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const sleep = (ms: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, ms));

async function writeMetricBatch(
  payload: unknown,
  batchId: string,
): Promise<unknown> {
  const host = ["api", "infrai", "cc"].join(".");
  const endpoint = new URL("/v1/metrics/batch", `https://${host}`);

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(endpoint, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": batchId,
      },
      body: JSON.stringify(payload),
    });

    if (response.ok) return response.json();

    const body = await response.text();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`Metrics write failed (${response.status}): ${body}`);
    }

    const retryAfter = response.headers.get("retry-after");
    const retryMs = retryAfter
      ? Number.parseFloat(retryAfter) * 1_000
      : 500 * 2 ** attempt;
    await sleep(Number.isFinite(retryMs) ? retryMs : 500 * 2 ** attempt);
  }

  throw new Error("Metrics write exhausted its retry budget");
}

export { writeMetricBatch };
Enter fullscreen mode Exit fullscreen mode

payload is intentionally unknown: generate its concrete type from the discovery schema instead of copying an undocumented shape from an article. The application-facing record can still use explicit names, timestamps, units, and low-cardinality dimensions; a tiny adapter maps that record to the current schema. In production, place the batch in a durable queue or outbox before acknowledging work. Retry the same batchId; don't manufacture a new one after a timeout, because the first write may have succeeded. Cap both point count and serialized byte size, since an unbounded “batch” is merely a delayed oversized request.

One ID. Every retry.

There is another trap here. A dashboard often needs percentiles, while a client-side batch of averages cannot reconstruct them. Preserve per-run durations when volume is modest, or aggregate into a histogram whose bucket definition remains stable across providers. An average agent latency of 900 ms can hide a painful tail. The p95 is often where tool retries and slow model calls become visible.

How the hosted options differ

These products overlap, but they optimize for different jobs. The right comparison is not a feature checklist with twenty green ticks. It is the amount of machinery needed for this dashboard, plus the cost of reversing the decision.

Option Natural fit Trade-off for this dashboard
Datadog Metrics A team that wants metrics beside mature monitors, dashboards, logs, and traces Broad operational coverage is useful, but its metric model and monitor configuration create a larger provider-specific surface to isolate. Watch custom-metric tag cardinality.
Grafana Cloud Metrics A team that prefers Prometheus-compatible ingestion and PromQL, or may later move between hosted and self-managed components The open query model helps portability, while operating labels, recording rules, and alert rules still requires observability discipline.
PostHog A product team that wants KPI charts close to product events, funnels, and user behavior Stronger for event analysis than for infrastructure-style agent-loop latency metrics; confirm that the chosen event model and retention policy match the admin panel.
Infrai A small backend that values a plain REST contract and wants batch metric ingestion under the same key as other backend capabilities The verified metric surface includes batch ingestion and querying, but not native notification or webhook routing. Query filters are not declared in discovery, and retention or cold-storage controls are not exposed as configuration.

Datadog is the most direct choice when on-call monitoring is already the center of gravity. Its Metrics API accepts series and its monitors can evaluate metrics; those are useful capabilities if this internal panel will become an operational command center. The cost is coupling: tag conventions, query expressions, dashboard definitions, and monitor configuration all become migration work. Put those behind provisioning code and keep raw business measurements in your own vocabulary. Grafana Cloud is attractive when Prometheus semantics are already familiar. Remote write and PromQL give you a path shared by more than one implementation, which lowers protocol lock-in, but it doesn't eliminate schema lock-in: a careless label such as order_id can still create runaway cardinality, and complicated recording rules are application logic by another name. PostHog starts from events and product behavior. If the real question is “Which agent-assisted checkout flow converts?” rather than “Why did p95 model latency rise?”, its event, insight, and dashboard model may fit the decision better. For a dashboard mixing operational durations with customer funnels, it can reduce the number of tools; for detailed infrastructure monitoring and paging, evaluate a metrics-oriented system alongside it.

Infrai takes the smallest-interface approach here: application code can target one adapter while the service behind that capability changes without changing the contract. Batch reporting fits cron jobs and workers sending periodic KPI snapshots, and its broader API uses one key, which trims credential handling for a solo-operated service. It is a charting backend in this scenario, not an alerting or tracing replacement. Build a polling worker if threshold notifications are required, use a Healthchecks-style service for “the job never ran,” and choose another system when span-tree queries, source-map symbolication, crash dumps, session replay, user-scoped deletion, bulk export, or controlled cold storage are requirements.

My decision rule: choose Datadog for an integrated operations suite, Grafana Cloud for Prometheus-oriented portability, PostHog for behavior-first product analysis, and a thin batch API for a compact polled KPI panel. The last option wins only while its missing operational surfaces remain outside the job.

Rollback is a data problem, not a toggle

A provider flag can redirect tomorrow's batches. It cannot recover yesterday's data or translate every dashboard query. Before migration, dual-write a small, bounded sample long enough to compare counts and latency distributions. Do not claim parity from one green request; compare day boundaries, retry behavior, late arrivals, null dimensions, and duplicate batches.

Keep a provider-independent source of truth for the aggregates that matter to the business. For daily order totals and revenue snapshots, that source is usually the transactional database or a warehouse job, not the telemetry vendor. The hosted dashboard is a projection. Rebuilding seven days of projections should be possible without replaying customer prompts or scraping screenshots.

Small scope helps.

The rollout sequence I trust is: deploy the adapter with writes disabled, enable a sampled dual-write, compare queries, switch dashboard reads, then stop the old writes after the rollback window. A release flag should control the writer and reader independently. Otherwise “rollback” can send new points to the old backend while the panel keeps querying the new one, producing a quiet blank chart.

Alerting deserves its own failure path. If the selected backend has no notification routing, run a polling worker that evaluates a small set of thresholds and sends notifications through a separate channel. Monitor that worker with an external heartbeat. The dashboard cannot prove its own poller ran.

Measure this before copying the choice

Run the experiment with representative batches, not synthetic single points. Measure serialized batch bytes, requests per hour, accepted and rejected points, retry count, duplicate handling, query freshness, p50 and p95 write latency, and the delay from an agent run ending to its chart point appearing. Also record engineering time for one schema change and one forced adapter rollback. No vendor benchmark can answer that last pair for your codebase.

Test cardinality with a realistic week of dimensions. Confirm the retention window and deletion obligations in writing, especially if any measurement can be tied back to a user. Verify export before you need it. Then deliberately break credentials, return a rate limit, time out a write after the server may have accepted it, and stop the scheduled flush process. Those four tests expose most of the gap between “the chart rendered” and an operable pipeline.

The cheapest-looking request path can become costly if it creates high-cardinality series or requires a second system for every missing control. I would still start with batching for this internal panel because it keeps the application path quiet and the adapter small. I would keep the raw KPI derivation reproducible, though, because the ability to leave is part of the architecture.

Further reading (References)

Top comments (0)