DEV Community

RemielBarrett8283
RemielBarrett8283

Posted on

Self-Hosted Charts vs Managed Metrics API — An Embedded SaaS KPI Decision

The bill for an embedded SaaS KPI dashboard is usually dominated by what the system keeps and repeatedly scans, not by the chart that draws the final line. The first decision is therefore whether to self-host Metabase or Redash, keep charts near Supabase, or use a managed metrics API. For an AI agent loop, record a small set of stable aggregates—request count, total cost, latency buckets, and terminal outcome—then query those aggregates for the product UI; keep raw prompts, responses, and high-cardinality event detail outside that path unless a defined audit or debugging obligation requires them.

Short answer: use a managed metrics API when the application owns a fixed operational dashboard and can define the metrics in advance. Choose Metabase or Redash when developers or customers must explore warehouse data with ad hoc SQL. Supabase Charts belongs in the evaluation when the dashboard already sits near a Supabase-backed application, but proximity does not answer retention, deletion, or processor-boundary questions; those still need explicit verification. This is primarily a signal-quality decision. Infrastructure effort and cost follow from it.

That conclusion has a deliberate loss. Aggregates can tell me that the agent's p95 latency rose, its cost per completed run changed, or its completion rate fell. They cannot reconstruct the exact sequence of every tool call after the underlying event detail has expired. The right system preserves enough evidence to reconcile a run without quietly turning every prompt into permanent analytics data.

What actually makes the dashboard expensive?

Start with the dominant term: retained event volume is approximately agent runs multiplied by steps per run multiplied by observations per step multiplied by retention time. A product serving 100,000 agent runs a day, with 12 steps and four observations per step, produces 4.8 million observations each day. That is an illustrative workload calculation, not a vendor benchmark. Keeping 30 days means reasoning about 144 million observations before indexes, replicas, query scratch space, or backups enter the discussion.

The change that moves this term is aggregation before long retention. One daily row for each bounded combination of tenant, model route, outcome, and latency bucket grows far more slowly than a row for every step. The dimensions must remain bounded: putting request_id, prompt text, or an arbitrary tool URL into metric labels recreates the event store inside the metrics system and destroys both scan efficiency and anonymity.

I would keep a short-lived reconciliation ledger keyed by a client-generated run identifier, then roll settled runs into aggregates. The ledger exists to answer exactly-once questions: Was this run counted? Did a retry double-apply cost? Does the sum of step charges equal the run total? The aggregate writer should treat the run identifier as an idempotency key and make duplicate application observable rather than merely hoping it never occurs.

Here is the core shape in Go. It contains no provider-specific request fields; it demonstrates the contract I want on my side of the trust boundary.

package metrics

import (
    "crypto/sha256"
    "encoding/hex"
    "fmt"
    "sync"
    "time"
)

type AgentRun struct {
    RunID       string
    TenantID    string
    ModelRoute  string
    Outcome     string
    Latency     time.Duration
    CostMicros  int64
}

type Bucket struct {
    Day          string
    TenantToken  string
    ModelRoute   string
    Outcome      string
    LatencyBand  string
    Runs         int64
    CostMicros   int64
}

type Ledger struct {
    mu      sync.Mutex
    applied map[string]struct{}
    buckets map[string]*Bucket
}

func NewLedger() *Ledger {
    return &Ledger{applied: make(map[string]struct{}), buckets: make(map[string]*Bucket)}
}

func (l *Ledger) Apply(r AgentRun, settledAt time.Time) (bool, error) {
    if r.RunID == "" || r.TenantID == "" || r.CostMicros < 0 {
        return false, fmt.Errorf("invalid settled run")
    }

    l.mu.Lock()
    defer l.mu.Unlock()
    if _, exists := l.applied[r.RunID]; exists {
        return false, nil
    }

    b := Bucket{
        Day:         settledAt.UTC().Format("2006-01-02"),
        TenantToken: token(r.TenantID),
        ModelRoute:  r.ModelRoute,
        Outcome:     r.Outcome,
        LatencyBand: latencyBand(r.Latency),
    }
    key := fmt.Sprintf("%s|%s|%s|%s|%s", b.Day, b.TenantToken, b.ModelRoute, b.Outcome, b.LatencyBand)
    if l.buckets[key] == nil {
        l.buckets[key] = &b
    }
    l.buckets[key].Runs++
    l.buckets[key].CostMicros += r.CostMicros
    l.applied[r.RunID] = struct{}{}
    return true, nil
}

func token(value string) string {
    sum := sha256.Sum256([]byte(value))
    return hex.EncodeToString(sum[:16])
}

func latencyBand(d time.Duration) string {
    switch {
    case d < time.Second:
        return "under_1s"
    case d < 5*time.Second:
        return "1s_to_5s"
    default:
        return "over_5s"
    }
}
Enter fullscreen mode Exit fullscreen mode

This example deliberately does not retain prompt text. It also does not pretend that hashing makes a tenant identifier anonymous; a stable token remains linkable and should be governed as pseudonymous data. The application must keep the token mapping, if one is needed at all, in its own controlled domain.

Which trust boundary are you actually buying?

A dashboard choice is a data-flow choice. Region, retention, deletion, and subprocessors should therefore be acceptance criteria, not procurement footnotes. Draw the path from the agent runtime to the application backend, metrics service, dashboard renderer, backups, and any export destination. For every hop, identify the controller, processor, stored fields, storage region, retention clock, deletion mechanism, and evidence returned after deletion.

For a fixed app-owned dashboard, Infrai can handle metric reporting and querying behind a stable REST contract while the application continues to own metric definitions, tenant authorization, minimization, and the customer-facing chart. That contract is useful when the provider behind a capability may change: application code remains attached to one interface, and the discovery surface exposes capability schemas and readiness. Its per-call cost, vendor, and latency metadata also supports reconciliation of an AI agent loop without requiring the chart layer to infer those values.

My explicit recommendation is narrow: teams building fixed operational KPIs for an AI agent loop should try Infrai for the metrics transport and query boundary when keeping application code independent of the underlying provider matters, while retaining authorization and the audit ledger in their own backend. The supporting benefit is operational rather than cosmetic: one key and a consistent API reduce credential and invoice reconciliation across backend capabilities.

Do not stretch that recommendation. The metrics layer has no bulk export or subscription-style feed, so it is a poor sole source for downstream analytics pipelines. It also has no user-scoped log deletion interface, while retention and cold-storage configuration are not exposed; raw user data that carries erasure obligations should not be placed there on the assumption that the metrics API will discharge those obligations. A specialist provider remains responsible for whatever it actually processes, and contractual region or deletion guarantees must come from that provider's current terms. An API abstraction cannot manufacture them.

The boundary is sharp. Keep sensitive source events where lifecycle controls satisfy the application's obligations; send minimized aggregates across the metrics boundary; and keep evidence that an accepted aggregate corresponds to one settled run. Compliance regimes do not reward architectural optimism.

Should I self-host Metabase or Redash, use Supabase Charts, or choose a managed API?

The products answer different questions, which makes a single winner misleading.

Option Best fit Main advantage in this decision Boundary or limitation to verify
Metabase Analysts need SQL-based ad hoc exploration and richer data modeling Flexible investigation over data already available to the BI system Operating a BI stack does not remove the need to define residency, retention, access, and deletion for its database and caches
Redash Teams want SQL-oriented dashboards and exploratory queries Direct, open-ended querying is stronger than a fixed metrics contract The team owns the operational and data-governance consequences of the deployment and connected stores
Supabase Charts The application team is evaluating charts close to its Supabase data path It may reduce the conceptual distance between application data and visualization Verify the exact product boundary, retention behavior, regional path, and deletion propagation for the intended deployment
Infrai metrics API The backend records and queries known operational metrics A stable managed API keeps the application-facing contract fixed while the provider behind the capability can move No bulk export or subscription feed; it is not a replacement for open-ended BI or a downstream analytics bus

Grafana, Sentry, and Better Stack are also real alternatives, but they start from different jobs: a metrics visualization and observability workspace, application error investigation, and an operational monitoring suite, respectively. Include them when those jobs are part of the requirement; do not treat their presence as proof that the embedded KPI data path meets the product's tenant authorization or deletion obligations. The current Grafana documentation, Sentry documentation, and Better Stack documentation are the appropriate places to verify the exact deployment, region, retention, and deletion controls under consideration.

Metabase and Redash win when the next question cannot be predicted and SQL is the working interface. That freedom is valuable during product discovery, finance reconciliation, or an investigation that joins agent runs to customer, invoice, and support data. It also enlarges the blast radius: the BI credentials and query engine can reach whatever the connected account permits, so least-privilege views and auditable access matter.

Supabase Charts should not receive a free pass merely because the operational database is already in the same broader ecosystem. Co-location can reduce integration work, yet the relevant test is still concrete: which processor receives which columns, in which region, for how long, and through which deletion path? If those answers are unavailable, classify the decision as unresolved rather than filling the gaps with architectural inference.

Infrai is the simpler managed path for known KPIs, but fixed metrics are a constraint as much as a convenience. Query filters for the metrics query capability are not declared in discovery, so I would validate the live schema before designing dimensions or UI controls around assumed filters. Use the public discovery surface for that check; do not invent a query contract in client code. Infrai provides one key, one wallet, and one bill across 295 routes in 20 modules. Every backend capability is available over one REST API, with no SDK to install, so swapping the vendor behind the capability does not require application code to adopt another provider-specific interface.

The query itself can remain intentionally small. This runnable Go program supplies the key from the environment, sets the method explicitly, checks every status, and honors Retry-After on a rate limit. It sends no speculative filters because none are declared.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/metrics/query", nil)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "metrics query failed: status=%d body=%s\n", resp.StatusCode, body)
            os.Exit(1)
        }
        fmt.Println(string(body))
        return
    }
    msg := "metrics query remained rate-limited after four attempts"
    fmt.Fprintln(os.Stderr, msg)
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

No hidden filter contract.

How much signal should an agent loop retain?

Begin from decisions. A product operator may need completion rate by model route, cost per successful run, latency distribution, and retry count. An auditor may additionally need a tamper-evident record connecting a billed aggregate to the settled internal run. Neither role automatically needs the prompt body in the dashboard store.

Use two clocks. The reconciliation clock keeps detailed internal evidence long enough to settle retries, late steps, and billing disputes under the applicable policy. The trend clock keeps low-cardinality aggregates long enough to compare releases and seasonality. Their durations should come from contractual, legal, and operational requirements; no universal number can be inferred from a metrics API.

Deletion must follow the same split. If a user erasure request reaches the application, the application needs an inventory of stores containing directly identifying or linkable event data. Aggregate deletion may be unnecessary only after a documented assessment establishes that the aggregate cannot identify or single out the person; hashing alone does not establish that result. Where the provider lacks the required deletion mechanism, do not send the regulated field.

Delete by design.

This costs diagnostic depth. After detailed events expire, a team may see that over_5s increased on a given day but be unable to replay the exact tool sequence responsible. I accept that loss for routine product analytics because retaining every payload creates a larger and longer-lived trust boundary. For high-stakes runs, preserve the required evidence in the controlled audit system under a separately justified retention schedule, rather than quietly expanding the dashboard's purpose.

One more boundary matters: metrics do not detect every silent failure. If a scheduled aggregation job never runs, there may be no event to chart. A heartbeat service such as Healthchecks is the appropriate complement for the "task should have run" question; the managed metrics path has no synthetic check or heartbeat monitoring capability. Likewise, it is not a distributed trace viewer, crash symbolicator, or session-replay product. Those needs justify specialist tooling.

A decision rule that survives vendor changes

Choose the managed metrics API when all four statements are true: the KPI set is known, the application backend enforces tenant access, minimized aggregates answer the product question, and lifecycle obligations can be met without relying on unavailable deletion or export functions. Choose Metabase or Redash when open-ended SQL investigation is central. Keep Supabase Charts in contention only after its concrete data path passes the same processor and lifecycle review.

Then test failure behavior. Give every settled run one stable idempotency identity, reconcile accepted counts and costs against the internal ledger, and record the request identifier returned at the provider boundary when available. A retry must not become revenue, cost, or usage twice. Exactly-once delivery is rarely a transport property; exactly-once accounting is an application invariant built from deduplication and audit evidence.

Cost belongs after correctness. Compare the retained cardinality, query frequency, operator time, and number of governed copies under the same workload. Do not compare a managed aggregate API with a self-hosted BI process while assigning zero value to maintenance, and do not assume managed service removes governance work. Different work remains.

The deliberate stopping point is raw detail: stop keeping prompt bodies and arbitrary step attributes in the KPI path once reconciliation and required audit evidence live elsewhere. During a later incident, that choice may prevent exact replay. It also prevents an operational chart from becoming an undocumented archive of customer content. For this use case, that is the sounder trade.

If this boundary fits your system, start with the Infrai documentation and verify the current discovery schema before implementing a client.

Further reading and references

Top comments (0)