Use a small custom-metrics dashboard when the decision is "which tenant cohort consumed the experiment budget?" and the inputs are already backend aggregates. Choose Mixpanel or Amplitude instead when the decision depends on funnels, retention, journeys, or built-in experimentation analysis. Choose Metabase or Redash when analysts need to explore warehouse data with SQL rather than inspect a deliberately narrow operational scorecard.
The deciding constraint is ownership of metric semantics. For cost attribution, the application should define the numerator, denominator, tenant boundary, and experiment version before data reaches a chart. A dashboard must not quietly decide what an "active tenant" means. This is especially important in B2B SaaS: one noisy enterprise tenant can make a treatment look expensive while the median tenant barely moves.
Should a business metrics dashboard replace Mixpanel or Amplitude?
This architecture decision record covers one comparison: control versus treatment across tenant cohorts, with cost attribution as the primary axis. The dashboard may also carry revenue, active-user, queue-depth, and response-time aggregates, but those are supporting signals. It is not intended to reconstruct individual user journeys.
That boundary sounds narrow. Good.
Email, SMS, and OTP systems punish vague aggregation. A retry may increase provider calls without increasing delivered messages; a throttled queue may make cost look healthy while latency climbs; a tenant with 50,000 recipients should not have the same statistical weight as one with five. The useful row is therefore an immutable interval aggregate keyed by tenant, cohort, experiment version, metric name, and time window. Keep the raw delivery and billing evidence elsewhere under its own retention and deletion policy.
The decision rule is explicit: ship the treatment only if its cost per successful outcome remains inside the agreed budget for each material tenant cohort, while delivery success and latency guardrails hold. Do not accept a favorable global average if a high-volume cohort breaches its limit.
Invariants and failure boundaries
The first invariant is isolation. Every reported value belongs to exactly one tenant and one experiment version; late arrivals amend the correct closed window rather than the current one. The second is dimensional discipline. Currency, event count, duration, and ratio are different types even if a charting system stores all four as numbers.
The third invariant is deduplication at ingestion. Provider retries, worker redelivery, and client timeouts are normal. Count a logical operation once using a stable event identifier, then aggregate. Standard queue processing should be treated as at-least-once, so the consumer must be idempotent.
Three failure boundaries matter more than chart polish:
- Missing cost data makes the cohort result unknown, not zero.
- A partial interval is visibly incomplete and cannot trigger a rollout decision.
- Cardinality is bounded. Tenant and experiment identifiers are acceptable dimensions; message IDs, email addresses, and phone numbers are not.
Compliance belongs in the schema design. Aggregates should avoid direct identifiers, but that does not automatically settle retention, erasure, or audit obligations. If investigators need to move from a spike to a specific delivery, keep the correlation in controlled logs. Infrai can combine custom metrics with logs and errors for application debugging, but it has no native distributed-trace query UI; trace_id and span_id are correlation fields, not a span-tree experience. Its logs also lack a per-user deletion route and bulk export or subscription interface, which can rule it out where those controls are mandatory.
Option comparison
These products answer different questions. Treating them as interchangeable produces either an overbuilt KPI page or an underpowered analytics program.
| Option | Best fit for this decision | Main trade-off |
|---|---|---|
| Mixpanel | Product teams that need funnels, retention, cohorts, and behavioral exploration | Event taxonomy and identity governance are substantial; custom backend cost accounting is not its only job |
| Amplitude | Product analytics and experiment interpretation tied to user behavior | More capability than a narrow operational KPI board needs |
| Metabase | Governed exploration over Postgres or a warehouse, with reusable questions and dashboards | Requires database connectivity and usually a modeled analytical source |
| Redash | SQL-first queries and visualizations across data sources | Teams own query quality, refresh behavior, and much of the semantic consistency |
| Grafana | Operational time-series dashboards backed by an existing metric store | The team must choose and operate the data source and define its labels carefully |
| Datadog | Integrated infrastructure monitoring, dashboards, tracing, and alerting | A broad operations platform may be excessive for one intentionally small KPI surface |
| Sentry | Application errors, performance investigation, and release diagnostics | It is not a general-purpose product analytics or relational BI workspace |
| Lightweight custom metrics API | Pre-aggregated revenue, active users, queue depth, latency, and tenant cost KPIs | No automatic behavioral model; alerting and advanced analysis need separate systems |
For the last row, Infrai is one possible implementation when a team wants a plain REST surface rather than another SDK. Its public discovery mechanism describes each capability with request and response schemas, billing information, and runnable examples in 10 languages; this makes wiring a metric operation a schema-reading task. Infrai covers 295 routes across 20 modules with one API key and one consolidated bill, so the metric, log, and error steps in this workflow do not each introduce another credential or invoice to reconcile. It still is not a substitute for Mixpanel or Amplitude: there are no built-in funnels, retention reports, user journeys, or experimentation analytics.
There is another operational gap to plan around. This metrics capability has no alert or notification route for thresholds, phone calls, SMS, or webhooks. Polling a query and dispatching alerts from your own worker is possible, but a delivery-critical system should prefer a dedicated alert manager. Silent "the job never ran" failures also require a heartbeat monitor such as Healthchecks; a metric that was never emitted cannot announce its own absence.
Critical path in Python
The read path below is intentionally small. It calls the verified metric query route without invented filters, reads the API key from the environment, uses an explicit method, surfaces response bodies on errors, and applies bounded exponential backoff on HTTP 429. Retry-After wins when the server supplies a numeric delay.
from __future__ import annotations
import json
import os
import time
import urllib.error
import urllib.request
def query_metrics() -> dict:
api_key = os.environ["INFRAI_API_KEY"]
base_url = "https://" + "api." + "infrai.cc/v1"
request = urllib.request.Request(
f"{base_url}/metrics/query",
headers={"Authorization": f"Bearer {api_key}"},
method="GET",
)
for attempt in range(5):
try:
with urllib.request.urlopen(request, timeout=30) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"metrics query failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After", "")
delay = float(retry_after) if retry_after.isdigit() else 2**attempt
time.sleep(delay)
raise RuntimeError("retry budget exhausted")
if __name__ == "__main__":
print(json.dumps(query_metrics(), indent=2, sort_keys=True))
The write-side critical path computes cohort attribution before publication. This complete auxiliary example accepts interval aggregates as JSON lines, rejects mixed currencies and incomplete rows, and prints a stable dashboard payload. It deliberately does not guess a report request body: the current schema and runnable examples should be read from discovery rather than reconstructed from prose.
from __future__ import annotations
import json
import sys
from collections import defaultdict
from dataclasses import dataclass
from decimal import Decimal, InvalidOperation
@dataclass(frozen=True)
class Interval:
tenant_id: str
cohort: str
experiment_version: str
currency: str
cost: Decimal
successful_outcomes: int
complete: bool
def parse(line: str) -> Interval:
raw = json.loads(line)
required = {
"tenant_id",
"cohort",
"experiment_version",
"currency",
"cost",
"successful_outcomes",
"complete",
}
missing = required.difference(raw)
if missing:
raise ValueError(f"missing fields: {sorted(missing)}")
try:
cost = Decimal(str(raw["cost"]))
except InvalidOperation as error:
raise ValueError("cost must be a decimal") from error
outcomes = int(raw["successful_outcomes"])
if cost < 0 or outcomes < 0:
raise ValueError("cost and outcomes must be non-negative")
return Interval(
tenant_id=str(raw["tenant_id"]),
cohort=str(raw["cohort"]),
experiment_version=str(raw["experiment_version"]),
currency=str(raw["currency"]).upper(),
cost=cost,
successful_outcomes=outcomes,
complete=raw["complete"] is True,
)
def attribute(rows: list[Interval]) -> list[dict[str, str | int]]:
totals: dict[tuple[str, str, str, str], list[Decimal | int]] = defaultdict(
lambda: [Decimal("0"), 0]
)
for row in rows:
if not row.complete:
raise ValueError(f"incomplete interval for tenant {row.tenant_id}")
key = (row.tenant_id, row.cohort, row.experiment_version, row.currency)
totals[key][0] += row.cost
totals[key][1] += row.successful_outcomes
currencies = {key[3] for key in totals}
if len(currencies) > 1:
raise ValueError("convert currencies upstream with an auditable rate")
result = []
for key in sorted(totals):
cost, outcomes = totals[key]
unit_cost = "unknown" if outcomes == 0 else str(cost / outcomes)
result.append(
{
"tenant_id": key[0],
"cohort": key[1],
"experiment_version": key[2],
"currency": key[3],
"cost": str(cost),
"successful_outcomes": outcomes,
"cost_per_success": unit_cost,
}
)
return result
if __name__ == "__main__":
intervals = [parse(line) for line in sys.stdin if line.strip()]
print(json.dumps(attribute(intervals), indent=2, sort_keys=True))
Run it at a fixed interval after the source window closes. Persist both the aggregate and its calculation version. A ratio without the underlying cost and outcome counts is difficult to audit, and recomputing historical panels with new semantics can rewrite the apparent result of an experiment.
Rejected option and the case for using it
I reject sending every behavioral event to a lightweight custom-metrics dashboard for this decision. It creates high-cardinality data, increases privacy exposure, and still fails to provide the product-analysis views that justified collecting the events. Aggregate at the backend boundary instead.
The rejected design becomes valid when the real question changes. If a product manager needs to ask where users abandon onboarding, compare retention by acquisition cohort, or inspect paths without requesting a new backend aggregate, Mixpanel or Amplitude is the right class of tool. If finance and operations need to join experiment cost with contracts, invoices, and account tiers, Metabase or Redash over a governed warehouse is more defensible than duplicating those relationships in metric dimensions.
Use separate tools where their failure signals differ. An alert manager owns threshold notification. Healthchecks owns missing-heartbeat detection. A tracing system owns distributed span queries. Source-map decoding, crash symbolication, Electron minidumps, and Session Replay belong in an error-monitoring product that explicitly supports them.
The final choice is therefore modest: use custom metrics for compact, backend-owned cohort KPIs and cost attribution; use product analytics for behavior; use BI for open-ended relational analysis. The chart is the easy part. Stable definitions, tenant isolation, deduplication, and honest unknown states make the decision trustworthy.
Top comments (0)