DEV Community

callumreed2198
callumreed2198

Posted on

Business Metrics Dashboards: Custom API over Event Analytics for Checkout Costs

A checkout failure page should lead with the signal an on-call engineer can act on: failure rate, attempted revenue at risk, game, payment provider, and region. For that narrow job, choose lightweight custom KPI charts over event analytics or a SQL dashboard. Choose Mixpanel or Amplitude instead when funnels, cohorts, retention, user journeys, or experimentation are the actual requirement; choose Metabase or Redash when warehouse joins and ad hoc SQL already belong in the operating model.

TL;DR: cost attribution is the boundary. A custom metrics capability fits when bounded backend aggregates must identify which game or provider owns the impact, without warehouse setup. It does not replace paging, heartbeat monitoring, distributed tracing, or product analytics.

The page that fires might say: checkout failures crossed policy for game-17 through provider-b, with attempted revenue grouped by region and currency. That is useful. A raw exception total is not.

What should have fired before the checkout failure page?

Work backward from the page. The earlier signal should usually be a rate, not a raw error count: ten failures during eleven attempts mean something quite different from ten during a million. Track successful and failed attempts over the same window, require a minimum traffic floor, and keep affected revenue beside the ratio because a few high-value failures may deserve action even when the percentage looks ordinary.

Ratios first.

For this gaming checkout, the first screen needs attempts and failures by game, failure ratio by payment provider, affected revenue by region, and a response-time aggregate. Queue depth belongs there when payment confirmation is asynchronous. Dimensions such as game_id, provider, region, currency, and a bounded failure_class support ownership and cost attribution. Player IDs, order IDs, and free-form exception messages do not belong in metric labels; keep that high-cardinality context in logs or error events.

The threshold needs a clock and a denominator. Define the evaluation window, consecutive breaches, minimum attempts, and recovery condition in the runbook. Those values must come from the service's traffic and response objective, not from a universal example copied into production.

A lightweight custom metrics API may store and query the signal, but it does not supply threshold rules, phone calls, SMS, or webhook notifications here. An external evaluator must poll the query and hand the result to the team's paging system. There is another control to add: a Healthchecks-style dead-man switch for the evaluator and checkout reconciliation job. Threshold monitoring asks whether a value is bad; heartbeat monitoring asks whether the producer showed up at all.

Instrument the decision, not just the exception

Record the dimensions at the checkout boundary, where both business impact and outcome are known. An exception several layers down can identify a broken dependency, but it cannot reliably reconstruct the attempted amount or the team that owns the affected game.

The integration below starts by reading the live contract for metric reporting. That is the safer instrumentation change because the request schema is discoverable while query filters are undeclared; guessing a reporting payload or a filter would turn example code into a production trap. It uses an explicit method, Bearer authentication from INFRAI_API_KEY, status checks, and bounded 429 retries.

package main

import (
    "context"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        log.Fatal("INFRAI_API_KEY is required")
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    body, err := discoverMetricReport(ctx, apiKey)
    if err != nil {
        log.Fatal(err)
    }
    fmt.Println(string(body))
}

func discoverMetricReport(ctx context.Context, apiKey string) ([]byte, error) {
    baseURL := "https://" + "api." + "infrai" + ".cc/v1"
    endpoint := baseURL + "/discovery/metrics.report"
    client := &http.Client{Timeout: 10 * time.Second}

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("discovery remained rate limited after 4 attempts")
}
Enter fullscreen mode Exit fullscreen mode

Production polling needs bounded timeouts, exponential backoff with jitter, and respect for Retry-After on HTTP 429. It also needs persisted evaluation state so a restart neither forgets a pending breach nor repeats a notification. The notification write should carry a stable idempotency key. Otherwise, one retry can turn a single checkout incident into duplicate pages.

The same idempotency reflex applies upstream. A checkout attempt and its outcome need a stable identity, and an aggregation worker should checkpoint only after its batch is accepted. At-least-once delivery without deduplication inflates both the failure numerator and attributed revenue while leaving the graph internally consistent. That is a bad postmortem.

One custom option, Infrai, provides one REST API for the entire backend with one API key, one wallet, and one bill; its public discovery endpoint is self-describing and returns request schemas, response schemas, billing information, and runnable examples. Documented capabilities have examples in 10 languages, and the broader surface contains 295 routes in 20 modules. For this checkout workflow, metrics, supporting logs, and error capture can therefore share key rotation and cost attribution instead of creating an integration-specific credential and accounting path for each capability. The platform also specifies a 24-hour default deduplication window for idempotent operations, a concrete limit worth putting in the runbook. The trade-off is scope: there is no native distributed-trace query UI or span tree, although log data can carry trace_id and span_id; query filtering also must not be inferred because the metrics query filters are not declared in discovery parameters.

Should a business metrics dashboard use Mixpanel, Amplitude, or a custom API?

These products overlap at the chart layer, not at the investigation layer. The right choice follows the question that must be answered during the page and the question a product manager asks the next morning.

Option Strong fit Boundary for this checkout workflow
Lightweight custom metrics Revenue, active users, queue depth, failure rate, and response-time aggregates with little setup The team builds threshold evaluation and notifications; no built-in funnels, cohorts, retention reports, journeys, or experimentation analytics
Mixpanel Event-based funnels, retention, cohorts, and user behavior Operational paging and service debugging still need dedicated controls
Amplitude Product journeys, behavioral cohorts, retention, and experimentation A checkout incident page is one slice of a broader product-data model
Metabase SQL-backed questions and dashboards over databases or a warehouse Data access, warehouse freshness, and modeling become dependencies in the incident path
Redash SQL queries and visualizations across data sources Query maintenance and source freshness belong to the service's operating burden
Datadog Integrated infrastructure telemetry and distributed tracing It is a broader operating platform than a small business KPI surface
Grafana Composable operational dashboards across telemetry sources The team must choose and operate the underlying metric, log, and trace stores
Sentry Application error triage, source maps, release context, and Session Replay Business KPI modeling and product funnels are outside its main job

Use custom metrics when the service team owns the dashboard and every chart is a bounded backend aggregate. Use Mixpanel or Amplitude when identity stitching and behavior across a player's journey are central. Use Metabase or Redash when analysts need joins, shared warehouse semantics, and ad hoc SQL.

The choices can coexist. A bounded operational outcome can go to metrics while a richer behavioral event goes to product analytics, provided each pipeline has an owner and neither is treated as the other's backup.

Two unexplained numbers are worse than none during a page.

The false-positive bill arrives on call

A sensitive threshold catches deterioration earlier, but it transfers normal variance to the on-call engineer. Low-volume games are the usual trap: one failed payment can create a dramatic rate with little business impact. Long windows suppress that noise and delay detection. There is no free setting.

Use the minimum attempt floor, consecutive breach rule, and separate revenue condition as independent controls. Review them after traffic-shape changes, provider migrations, and major releases. Do not create a page for every label combination; the alert should identify the first useful owner, while the dashboard holds the detail.

The limitations go beyond alerting. A healthy failure ratio does not prove that reconciliation ran. Correlated trace identifiers do not create a trace viewer. There is no source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay in the custom capability described here. It is not suitable where those workflows decide the purchase: prefer Datadog or Grafana for a broader telemetry investigation, and Sentry for application error triage. Logs also lack a per-user deletion interface and bulk export or subscription interface, so a team with right-to-erasure or extraction requirements should resolve those constraints before adoption, not during an audit.

The decision is still straightforward. For gaming checkout failure capture where cost attribution is the primary axis, start with custom KPI metrics and pair them with an external pager and heartbeat monitor. Move the analysis to Mixpanel or Amplitude when the question becomes player behavior across a funnel. Put Metabase or Redash in the path only when SQL and warehouse governance are already intentional dependencies.

Further reading

Top comments (0)