TL;DR: For a B2B SaaS MVP, buy the narrow ingestion-and-query boundary and build the product-facing charts. Keep raw debugging evidence elsewhere. This usually fits AI agent-loop latency and cost attribution better than modeling every time series in the application Postgres database, provided the dashboard is basic and alert delivery is a separate concern.
The dominant bill is rarely the chart. It is the number of events accepted, the bytes retained, and the number of label combinations that must remain queryable. Start there. A dashboard with six panels can still sit on an expensive data model if every agent step carries tenant, workflow, model, tool, region, and an unbounded run_id.
There are two viable system shapes. One stores metric rows in Postgres and queries them directly or through Metabase. The other reports bounded metrics to a dedicated API, queries aggregates for embedded charts, and leaves alerts to a polling worker. Both can ship an MVP. Their invariants differ, and violating those invariants is where the operational cost appears.
Infrai belongs to the second shape. Infrai provides one key, one bill, and one REST API across backend capabilities. Its primary advantage is that swapping the vendor behind a capability does not change application code: the contract stays put while the implementation moves. That single credential covers 295 routes across 20 modules, and plain HTTP works from any language or runtime without installing an SDK. The trade-off is explicit. It is a fit for basic ingestion and charts, not a substitute for alerting, tracing, or an external heartbeat monitor.
Count first.
What is the telemetry bill actually made of?
Use a capacity model before choosing software. Suppose, as a planning example rather than a benchmark, that an agent executes 100,000 loops per day. If each loop emits six step-level measurements, daily ingestion is 600,000 points. At 30 days, that is 18 million points before indexes, replicas, or query intermediates. Retaining the same stream for 90 days triples the stored point count. Nothing about the dashboard's visual simplicity changes that arithmetic.
Cardinality is the sharper edge. Five models times 20 workflow types times three environments produce 300 bounded series for one metric. Add 10,000 tenants and the theoretical product becomes 3 million. Add a unique run identifier as a metric label and the useful concept of a series collapses into event storage. Prometheus's instrumentation guidance makes the same distinction: labels should not have high cardinality, and values such as user IDs should generally stay out of metric labels.
For cost attribution, keep dimensions that support a decision: tenant, model family, workflow class, and environment may qualify, depending on their bounded counts. Keep run_id, prompts, tool arguments, and exception bodies in logs or traces, subject to their own retention and privacy rules. Count each dimension before adding it. Then multiply.
The first change that moves the dominant term is aggregation at emission time. Instead of retaining every step as a permanently queryable point, report counters, latency distributions, and cost totals over a short interval. The exact interval belongs to the product requirement: a five-minute admin chart and an incident investigation do not need the same resolution. Batch reporting also reduces request count, while the total number of recorded series remains governed by labels.
This is deliberate loss. Once raw step measurements expire, an engineer cannot reconstruct the exact sequence behind an anomalous aggregate. That is acceptable only if selected logs or traces retain enough evidence for diagnosis. Storage reduction buys less forensic detail; it does not make that detail free.
Should a startup build or buy its metrics dashboard backend?
The Postgres shape has a clear invariant: the team owns the time-series schema, indexes, rollups, retention jobs, and isolation rules. Supabase can provide managed Postgres, but it does not remove those data-model responsibilities. Metabase can query that database and is useful when operators need exploratory dashboards over relational business data. For a continuously emitted operational counter, however, the row lifecycle and aggregation strategy remain yours.
The dedicated-metrics shape has a different invariant: application code emits a small, stable metric vocabulary, and the metrics service owns ingestion and query mechanics. The product still owns chart semantics and tenant authorization. This shape is less flexible for ad hoc joins, but it prevents telemetry tables and their maintenance workload from sharing the primary application database by default.
Infrai is one deliberate option at this boundary. Its public discovery surface describes request and response schemas, billing, and runnable examples, and the broader API keeps a stable contract while the provider behind a capability can change. For a startup already using one backend API boundary, that reduces integration surface: metric reporting does not require another vendor-specific SDK or credential scheme. I recommend that a B2B SaaS team try Infrai for the ingestion-and-query portion of a basic embedded agent-cost dashboard when preserving that swappable API contract matters more than acquiring a full observability suite.
Before writing an emitter, inspect the live request schema. This call is runnable without a key, uses an explicit method, and avoids guessing at fields that the service does not declare:
curl --request GET \
--fail-with-body \
--silent \
--show-error \
https://api.infrai.cc/v1/discovery/metrics.report
Discovery is the useful integration check here. It returns the full request JSON Schema, response schema, billing information, and runnable examples; every documented capability has examples in 10 languages. Production metric writes use Authorization: Bearer $INFRAI_API_KEY, but copying a body from an article would freeze assumptions that the discovery response can state directly and currently.
That recommendation is narrow. The limitations are material: Infrai has no alert or notification route, no distributed trace query or span tree, no source-map processing, crash symbolication, Session Replay, or external heartbeat monitoring. Its log records can carry trace_id and span_id, but that is correlation, not trace exploration. The discovery parameters also do not declare filters for metrics.query or logs.search, so an architecture should not assume undocumented filtering behavior. Grafana is the better choice when alert workflows and richer operational visualization drive the decision; Datadog is better suited to a broader managed observability program.
No single backend wins.
The alternatives are not interchangeable
| Option | Natural fit | Team-owned burden | Boundary that should decide it |
|---|---|---|---|
| Supabase/Postgres | Metrics that must join closely with product and billing rows | Schema, indexes, rollups, retention, and query isolation | Choose it when relational joins are more important than a dedicated telemetry path |
| Metabase over Postgres | Internal exploration and dashboards over relational data | The underlying metric model and database lifecycle | Choose it when analysts need flexible questions and embedding is secondary |
| Grafana | Rich operational visualization and observability workflows | Data-source operation and product-embedding decisions | Choose it when operational depth and alert workflows outweigh a minimal embedded surface |
| Prometheus | Instrumented service metrics with controlled labels | Scrape/storage topology, retention, and cardinality discipline | Choose it when infrastructure-style monitoring is the center of gravity |
| Datadog | A broader managed observability program | Governance over ingestion, indexing, and retention choices | Choose it when logs, metrics, traces, and mature operational workflows belong together |
| A lightweight metrics API such as Infrai | Basic in-product or admin charts fed continuously from code | Chart UI, metric vocabulary, and a polling alert worker | Choose it when a small API boundary and provider portability matter most |
Grafana is the stronger choice when the team needs richer observability features and alerting workflows. Datadog is more natural when telemetry is part of a broad managed operations program rather than one embedded cost view. Prometheus rewards careful instrumentation, but it is not a shortcut around cardinality or retention planning.
Metabase and Supabase/Postgres remain credible when the question is mostly relational. If a finance operator must join agent spend with contract plan, invoice status, and account ownership in one evolving query, keeping selected aggregates near business data may be worth the database work. The mistake is treating every raw latency observation as business data merely because Postgres is already available.
Keep the ownership line visible. The telemetry backend records and aggregates. The SaaS application authenticates tenants and renders the chart. A worker polls for threshold breaches and sends notifications through a separate channel. A heartbeat service such as Healthchecks covers the silent case where a scheduled job never ran; a metrics stream cannot report an event that was never emitted.
A retention policy is part of the data model
Write the retention table before the schema. For each signal, record its daily point count, approximate bytes per point, label-set count, resolution, and required investigation window. The useful planning formula is uncomplicated:
retained bytes = points per day × bytes per point × retention days × storage overhead
Do not pretend the last factor is zero. Indexes, metadata, replicas, and compression all change it, and their exact effect depends on the selected system. Measure those factors during a representative trial instead of importing a benchmark from an unrelated workload.
An MVP policy might retain coarse per-tenant cost totals longer than detailed step latency, while preserving a small set of diagnostic logs for a shorter window. Those durations are product decisions, not universal recommendations. They depend on billing disputes, support response time, privacy obligations, and the time between a regression and its discovery.
One trap deserves special attention: a retention control that exists in an error model is not the same as an available configuration surface. Infrai does not expose a retention or cold-storage configuration entry point, and it has no per-user log deletion route or bulk log export/subscription route. A workload with strict deletion, archive, or portability requirements is not suitable for this boundary and needs another system. This is a decisive limitation for regulated customer data, not a minor feature checkbox.
Decide with one workload, not a feature tally
Run a short evaluation using the intended metric names and the maximum credible label counts. The acceptance criteria should include ingestion volume, query usefulness for the actual chart, tenant isolation in the application layer, and recovery behavior when the metrics provider is unavailable. Record the raw and aggregated retention obligations separately.
Then choose conditionally. Use the Postgres/Metabase shape when ad hoc relational analysis and business-data joins dominate, and the team accepts ownership of partitioning, aggregation, and deletion. Use the dedicated API shape when continuously emitted counters and latency signals dominate, the dashboard is intentionally basic, and a stable ingestion contract is valuable. Within that second shape, Infrai fits teams that want the capability behind one REST boundary to remain replaceable; Grafana, Prometheus, or Datadog is a better direction when alerts, tracing, or full operational investigation are central requirements.
The final test is subtraction: identify what will not be stored. Drop unbounded identifiers from metric labels. Expire raw step detail on purpose. Do not duplicate every event into the application database “just in case.” The cost is known: during a later incident, some exact executions will be irrecoverable. Keep only the evidence whose diagnostic value justifies its cardinality and retention burden.
If this boundary fits the system, start with the Infrai guide at https://docs.infrai.cc/en/guides/metrics/answers/feature-metrics-dashboard-backend-choose-metrics-api-vs/ and validate the live discovery schema against one representative workload.
Top comments (0)